ledger-mcp
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@ledger-mcpRun the month-end review and tell me which vendors I need to contact."
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
ledger-mcp
An MCP server that lets LLM agents (Claude, ChatGPT, Cursor, any MCP client) query and reconcile an accounts-payable ledger, plus an eval suite that measures how well the server does the job.
It ships with a synthetic ledger (8 vendors, 131 invoices, 125 payments) where real-world problems are planted on purpose: underpayments, duplicate payments, payments in the wrong currency, overdue invoices, payments without a reference, and a payment that could belong to two invoices. The eval suite checks the server against that ground truth.
Why
Reconciling invoices against payments is a classic job for an agent: lots of rows, fuzzy references, and a few exceptions that matter. Letting a model write arbitrary SQL against financial data is risky, so this server gives it:
Safe access: the database is opened read-only, and free-form SQL is limited to a single
SELECTwith a row cap.Deterministic tools for the hard parts: matching logic lives in tested Python, not in the model's head.
Measurable quality: precision and recall per exception type, a hostile-SQL safety check, and a contract check, all runnable in CI.
Related MCP server: ledgerkit-mcp
What it exposes
Type | Name | What it does |
Tool |
| Read-only |
Tool |
| Billed, paid and open balance per vendor |
Tool |
| Matches invoices to payments and lists exceptions by type |
Tool |
| Raw invoice, payment and vendor rows behind an exception |
Resource |
| Table definitions, so the agent writes valid SQL |
Prompt |
| Step-by-step month-end close workflow |
Matching rules
Reference match: the payment reference contains the invoice number (case-insensitive).
Fuzzy match: same vendor and currency, amount within tolerance, paid within N days of issue, and exactly one candidate invoice. If more than one invoice fits, the payment is flagged as
ambiguous_paymentinstead of guessed.
Exception types: underpaid, overpaid, duplicate_payment, currency_mismatch, overdue_unpaid, unmatched_payment, ambiguous_payment.
Quick start
pip install -e ".[dev]"
ledger-seed # creates data/ledger.db and evals/golden.json
ledger-mcp # stdio transport, for Claude Desktop / Cursor
ledger-mcp --transport streamable-http --port 8000 # HTTP transportInspect it interactively with the MCP Inspector:
mcp dev src/ledger_mcp/server.pyUse it from Claude Desktop
Add this to claude_desktop_config.json:
{
"mcpServers": {
"ledger": {
"command": "ledger-mcp",
"env": { "LEDGER_DB": "/absolute/path/to/data/ledger.db" }
}
}
}Then ask: "Run the month-end review and tell me which vendors I need to contact."
Docker
docker build -t ledger-mcp .
docker run -p 8000:8000 ledger-mcp # streamable HTTP on :8000/mcpTests and evals
pytest -q # unit tests + protocol tests through an in-memory MCP client
python evals/run_evals.py # golden-set eval, exits 1 if a quality gate failsSample eval output:
kind exp pred precision recall
underpaid 3 3 1.00 1.00
duplicate_payment 2 2 1.00 1.00
currency_mismatch 2 2 1.00 1.00
overdue_unpaid 4 4 1.00 1.00
unmatched_payment 3 3 1.00 1.00
ambiguous_payment 1 1 1.00 1.00
fuzzy_matched 2 2 1.00 1.00
hostile SQL rejected: 100% data intact: True
PASSThe golden file is written by the seed script when it plants each anomaly, so the eval never grades the code against its own output. CI runs lint, tests, evals and a Docker build on every push.
Project layout
src/ledger_mcp/
server.py MCP tools, resource and prompt (FastMCP)
reconcile.py matching logic, pure functions
db.py read-only SQLite access and SQL guard
seed.py synthetic data with planted anomalies + golden file
tests/ unit and protocol tests
evals/ end-to-end eval suite over MCPIdeas for next steps
Migrate to MCP Python SDK 2.x (
MCPServer), currently pinned to 1.x.Postgres backend behind the same tools.
An agent-level eval: run a model through
month_end_reviewand score its final table against the golden file.Multi-currency matching with FX rates.
License
MIT
Available Tools
4 toolsexplain_exceptionC
Return the raw invoice and payment rows behind an exception, plus the vendor.
| Name | Required | Description | Default |
|---|---|---|---|
| invoice_id | No | ||
| payment_id | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description must carry the full behavioral burden. It implies a read operation by saying 'Return,' but does not state read-only safety, authentication needs, rate limits, or side effects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence with no wasted words. It is appropriately sized for a tool whose output schema already covers return values.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The output schema covers return values, but the description leaves out parameter usage, exception definition, and when-to-use guidance. Given two optional parameters and no annotations, it is not complete enough for reliable invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description does not mention invoice_id or payment_id at all. An agent must infer their filtering behavior from the schema names alone, with no explanation of optionality or how they select an exception.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: it returns raw invoice and payment rows behind an exception plus the vendor. This distinguishes it from generic siblings like run_sql, though it does not explicitly name alternatives or define what 'exception' means in this domain.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit guidance on when to use this tool versus run_sql, vendor_balances, or reconcile_ledger. The phrase 'behind an exception' implies a context, but no alternatives or exclusions are provided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
reconcile_ledgerA
Match invoices to payments and list every exception.
Exception kinds: underpaid, overpaid, duplicate_payment, currency_mismatch, overdue_unpaid, unmatched_payment, ambiguous_payment.
Args: as_of: ISO date used to decide which unpaid invoices are overdue. amount_tolerance: max difference (same currency) still treated as equal. date_window_days: how far after issue date a reference-less payment can land. include_matches: also return every matched invoice (large output).
| Name | Required | Description | Default |
|---|---|---|---|
| as_of | No | 2026-10-01 | |
| include_matches | No | ||
| amount_tolerance | No | ||
| date_window_days | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does well: it enumerates the seven exception kinds, explains what as_of decides, and warns that include_matches produces a 'large output'. It is silent on whether the operation is purely read-only, on required permissions, and on failure modes, which keeps it short of a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose is front-loaded, followed by the exception taxonomy and then the args, so the most selection-relevant information comes first. The 'Args:' block is a fragment list rather than prose but every line is informative and non-redundant.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists so return values need not be re-explained, and the description covers purpose, exception vocabulary, and all four parameters. Against zero annotation coverage the only real gap is the safety/mutation profile, which is not stated.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, and it does: all four parameters get real semantics (as_of as the overdue cutoff, amount_tolerance as same-currency equality bound, date_window_days as the landing window for reference-less payments, include_matches as output expansion). This meaning could not be recovered from the bare parameter titles.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a concrete verb+resource pair: 'Match invoices to payments and list every exception', naming the resource (invoices/payments) and the output (exceptions). It is distinguishable from siblings like vendor_balances (balances) and explain_exception (explains one), but it never names or contrasts a sibling explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to invoke this versus run_sql, vendor_balances, or explain_exception, and no stated prerequisites or exclusions. Usage must be inferred entirely from the purpose sentence.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
run_sqlA
Run a read-only SELECT (or WITH ... SELECT) against the ledger.
Returns columns, rows and whether the result was truncated. Write statements, multiple statements and PRAGMAs are rejected.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| query | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does well: it declares the operation read-only and enumerates what is rejected (write statements, multiple statements, PRAGMAs), which is exactly the safety profile an agent needs before invoking. It also notes truncation is reported. It stops short of timeout, rate-limit, or cost behavior, so it is strong but not exhaustive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences, with the core operation front-loaded and the constraint/rejection list following. No filler, no restatement of the name.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained, and the query constraints are covered. However, for a two-parameter tool with zero schema description coverage, the omission of the 'limit' parameter leaves a real gap an agent must guess at.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It partially covers 'query' by constraining the accepted statement forms, but the second parameter 'limit' (default 50) is never mentioned, so an agent gets no guidance on row capping or how limit interacts with truncation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Run a read-only SELECT ... against the ledger') and pins down the exact accepted grammar (SELECT or WITH ... SELECT). This is clearly distinguishable from the narrower siblings vendor_balances, reconcile_ledger, and explain_exception, which cover specific analytical tasks rather than arbitrary queries.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The 'read-only SELECT' framing implies this is the ad-hoc query escape hatch, but the description never says when to prefer it over vendor_balances, reconcile_ledger, or explain_exception. Usage is inferable rather than stated, and no exclusions or alternatives are given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
vendor_balancesA
Billed, paid and open balance per vendor, optionally filtered by currency.
Payments are counted in the vendor's own currency only, so a payment sent in the wrong currency shows up as an open balance (and as an exception in reconcile_ledger).
| Name | Required | Description | Default |
|---|---|---|---|
| currency | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does disclose a non-obvious behavioral rule: payments are matched only in the vendor's own currency, so a mismatched-currency payment surfaces as an open balance. This is exactly the kind of semantic caveat structured fields cannot convey, though it does not state read-only nature or any rate/latency characteristics.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences with the core output front-loaded and the caveat second; no filler, and the exception cross-reference to reconcile_ledger is placed where it is most useful.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return formatting is covered elsewhere, and the description supplies the key semantic gotcha an agent needs. Only the read-only nature and the precise meaning of the currency filter remain unstated.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% and the single currency parameter's enum is nested inside an anyOf, so the schema adds little prose. The description says the filter is optional, implying null means all currencies, but does not clarify what the filter actually constrains (vendor rows vs payments) given the currency-matching rule above.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states the resource (vendor) and the exact measures returned (billed, paid, open balance), which is specific enough to distinguish it from a generic query tool. It does not, however, explicitly differentiate itself from run_sql, the closest sibling for ad-hoc queries.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Optionally filtered by currency' and the note about wrong-currency payments imply the reconciliation use case, but there is no explicit statement of when to choose this over run_sql or reconcile_ledger. Usage is inferable but not spelled out.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
4 tool updates
v0.1.0- First observed
explain_exception - First observed
reconcile_ledger - First observed
run_sql - First observed
vendor_balances
TDQS
Scored across 4 tools
run_sql is a generic read-only query escape hatch that could overlap with vendor_balances, but vendor_balances provides currency-aware aggregation that SQL alone wouldn't easily replicate. reconcile_ledger and explain_exception form a clear drill-down pair, so boundaries are mostly distinct.
Three tools follow a verb_noun pattern (run_sql, reconcile_ledger, explain_exception), while vendor_balances is a noun phrase. All are snake_case and readable, but the verb-first consistency is not universal.
Four tools are well-scoped for a read-only reconciliation server: a generic query, a balance report, a reconciliation engine, and an exception inspector. Each has a distinct role and none feels redundant.
The read-only surface covers querying, balances, reconciliation, and exception drill-down. A direct structured tool for listing invoices or payments is absent, but run_sql provides a viable workaround, leaving only minor gaps.
Maintenance
Related MCP Connectors
Personal finance ledger for AI agents — query spending, track bills, forecast cash flow.
AI agents for bookkeeping, reconciliation, and financial close for SMBs.
Primary-source SEC filing intelligence and financial/disclosure reconciliation for AI agents.
Connect Claude or Cursor to books, invoices, bills, payroll, and sealed closes.
Related MCP Servers
- AlicenseNot gradedqualityBmaintenanceAn accounting-ops agent that reconciles payments against open orders, auto-books provably safe payments through a deterministic policy gate, and escalates exceptions to a human queue with audit trails.MIT
- AlicenseAqualityCmaintenanceEnables AI agents to interact with a double-entry ledger, offering tools for account management, balanced journal entries, balance queries, trial balance, and penny-perfect allocation. Built with safety by construction: no update/delete tools, idempotent posting, and an append-only journal.7MIT
- FlicenseAqualityBmaintenanceEnables an AI agent to handle accounts payable tasks against a mock ERP, including reading and writing bills and vendors, checking duplicates, matching invoices, recommending approvals, and queuing payment releases, with configurable profiles that limit available tools.11-
- FlicenseAqualityCmaintenanceProvides AI agents with deterministic, offline finance tools for commodity margin analysis, loan covenant compliance, invoice auditing, AP exception classification, and five-day close readiness.146 npm-