Skip to main content
Glama

ledger-mcp

An MCP server that lets LLM agents (Claude, ChatGPT, Cursor, any MCP client) query and reconcile an accounts-payable ledger, plus an eval suite that measures how well the server does the job.

It ships with a synthetic ledger (8 vendors, 131 invoices, 125 payments) where real-world problems are planted on purpose: underpayments, duplicate payments, payments in the wrong currency, overdue invoices, payments without a reference, and a payment that could belong to two invoices. The eval suite checks the server against that ground truth.

Why

Reconciling invoices against payments is a classic job for an agent: lots of rows, fuzzy references, and a few exceptions that matter. Letting a model write arbitrary SQL against financial data is risky, so this server gives it:

  • Safe access: the database is opened read-only, and free-form SQL is limited to a single SELECT with a row cap.

  • Deterministic tools for the hard parts: matching logic lives in tested Python, not in the model's head.

  • Measurable quality: precision and recall per exception type, a hostile-SQL safety check, and a contract check, all runnable in CI.

Related MCP server: ledgerkit-mcp

What it exposes

Type

Name

What it does

Tool

run_sql

Read-only SELECT / WITH query, max 200 rows, flags truncation

Tool

vendor_balances

Billed, paid and open balance per vendor

Tool

reconcile_ledger

Matches invoices to payments and lists exceptions by type

Tool

explain_exception

Raw invoice, payment and vendor rows behind an exception

Resource

ledger://schema

Table definitions, so the agent writes valid SQL

Prompt

month_end_review

Step-by-step month-end close workflow

Matching rules

  1. Reference match: the payment reference contains the invoice number (case-insensitive).

  2. Fuzzy match: same vendor and currency, amount within tolerance, paid within N days of issue, and exactly one candidate invoice. If more than one invoice fits, the payment is flagged as ambiguous_payment instead of guessed.

Exception types: underpaid, overpaid, duplicate_payment, currency_mismatch, overdue_unpaid, unmatched_payment, ambiguous_payment.

Quick start

pip install -e ".[dev]"
ledger-seed                      # creates data/ledger.db and evals/golden.json
ledger-mcp                       # stdio transport, for Claude Desktop / Cursor
ledger-mcp --transport streamable-http --port 8000   # HTTP transport

Inspect it interactively with the MCP Inspector:

mcp dev src/ledger_mcp/server.py

Use it from Claude Desktop

Add this to claude_desktop_config.json:

{
  "mcpServers": {
    "ledger": {
      "command": "ledger-mcp",
      "env": { "LEDGER_DB": "/absolute/path/to/data/ledger.db" }
    }
  }
}

Then ask: "Run the month-end review and tell me which vendors I need to contact."

Docker

docker build -t ledger-mcp .
docker run -p 8000:8000 ledger-mcp   # streamable HTTP on :8000/mcp

Tests and evals

pytest -q                  # unit tests + protocol tests through an in-memory MCP client
python evals/run_evals.py  # golden-set eval, exits 1 if a quality gate fails

Sample eval output:

kind                  exp  pred  precision  recall
underpaid               3     3       1.00    1.00
duplicate_payment       2     2       1.00    1.00
currency_mismatch       2     2       1.00    1.00
overdue_unpaid          4     4       1.00    1.00
unmatched_payment       3     3       1.00    1.00
ambiguous_payment       1     1       1.00    1.00
fuzzy_matched           2     2       1.00    1.00

hostile SQL rejected: 100%   data intact: True
PASS

The golden file is written by the seed script when it plants each anomaly, so the eval never grades the code against its own output. CI runs lint, tests, evals and a Docker build on every push.

Project layout

src/ledger_mcp/
  server.py      MCP tools, resource and prompt (FastMCP)
  reconcile.py   matching logic, pure functions
  db.py          read-only SQLite access and SQL guard
  seed.py        synthetic data with planted anomalies + golden file
tests/           unit and protocol tests
evals/           end-to-end eval suite over MCP

Ideas for next steps

  • Migrate to MCP Python SDK 2.x (MCPServer), currently pinned to 1.x.

  • Postgres backend behind the same tools.

  • An agent-level eval: run a model through month_end_review and score its final table against the golden file.

  • Multi-currency matching with FX rates.

License

MIT

Available Tools

4 tools
explain_exceptionC

Return the raw invoice and payment rows behind an exception, plus the vendor.

ParametersJSON Schema
NameRequiredDescriptionDefault
invoice_idNo
payment_idNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description must carry the full behavioral burden. It implies a read operation by saying 'Return,' but does not state read-only safety, authentication needs, rate limits, or side effects.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence with no wasted words. It is appropriately sized for a tool whose output schema already covers return values.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The output schema covers return values, but the description leaves out parameter usage, exception definition, and when-to-use guidance. Given two optional parameters and no annotations, it is not complete enough for reliable invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, and the description does not mention invoice_id or payment_id at all. An agent must infer their filtering behavior from the schema names alone, with no explanation of optionality or how they select an exception.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: it returns raw invoice and payment rows behind an exception plus the vendor. This distinguishes it from generic siblings like run_sql, though it does not explicitly name alternatives or define what 'exception' means in this domain.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit guidance on when to use this tool versus run_sql, vendor_balances, or reconcile_ledger. The phrase 'behind an exception' implies a context, but no alternatives or exclusions are provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

reconcile_ledgerA

Match invoices to payments and list every exception.

Exception kinds: underpaid, overpaid, duplicate_payment, currency_mismatch, overdue_unpaid, unmatched_payment, ambiguous_payment.

Args: as_of: ISO date used to decide which unpaid invoices are overdue. amount_tolerance: max difference (same currency) still treated as equal. date_window_days: how far after issue date a reference-less payment can land. include_matches: also return every matched invoice (large output).

ParametersJSON Schema
NameRequiredDescriptionDefault
as_ofNo2026-10-01
include_matchesNo
amount_toleranceNo
date_window_daysNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does well: it enumerates the seven exception kinds, explains what as_of decides, and warns that include_matches produces a 'large output'. It is silent on whether the operation is purely read-only, on required permissions, and on failure modes, which keeps it short of a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose is front-loaded, followed by the exception taxonomy and then the args, so the most selection-relevant information comes first. The 'Args:' block is a fragment list rather than prose but every line is informative and non-redundant.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values need not be re-explained, and the description covers purpose, exception vocabulary, and all four parameters. Against zero annotation coverage the only real gap is the safety/mutation profile, which is not stated.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, and it does: all four parameters get real semantics (as_of as the overdue cutoff, amount_tolerance as same-currency equality bound, date_window_days as the landing window for reference-less payments, include_matches as output expansion). This meaning could not be recovered from the bare parameter titles.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a concrete verb+resource pair: 'Match invoices to payments and list every exception', naming the resource (invoices/payments) and the output (exceptions). It is distinguishable from siblings like vendor_balances (balances) and explain_exception (explains one), but it never names or contrasts a sibling explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to invoke this versus run_sql, vendor_balances, or explain_exception, and no stated prerequisites or exclusions. Usage must be inferred entirely from the purpose sentence.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

run_sqlA

Run a read-only SELECT (or WITH ... SELECT) against the ledger.

Returns columns, rows and whether the result was truncated. Write statements, multiple statements and PRAGMAs are rejected.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
queryYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does well: it declares the operation read-only and enumerates what is rejected (write statements, multiple statements, PRAGMAs), which is exactly the safety profile an agent needs before invoking. It also notes truncation is reported. It stops short of timeout, rate-limit, or cost behavior, so it is strong but not exhaustive.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences, with the core operation front-loaded and the constraint/rejection list following. No filler, no restatement of the name.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained, and the query constraints are covered. However, for a two-parameter tool with zero schema description coverage, the omission of the 'limit' parameter leaves a real gap an agent must guess at.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It partially covers 'query' by constraining the accepted statement forms, but the second parameter 'limit' (default 50) is never mentioned, so an agent gets no guidance on row capping or how limit interacts with truncation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Run a read-only SELECT ... against the ledger') and pins down the exact accepted grammar (SELECT or WITH ... SELECT). This is clearly distinguishable from the narrower siblings vendor_balances, reconcile_ledger, and explain_exception, which cover specific analytical tasks rather than arbitrary queries.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The 'read-only SELECT' framing implies this is the ad-hoc query escape hatch, but the description never says when to prefer it over vendor_balances, reconcile_ledger, or explain_exception. Usage is inferable rather than stated, and no exclusions or alternatives are given.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

vendor_balancesA

Billed, paid and open balance per vendor, optionally filtered by currency.

Payments are counted in the vendor's own currency only, so a payment sent in the wrong currency shows up as an open balance (and as an exception in reconcile_ledger).

ParametersJSON Schema
NameRequiredDescriptionDefault
currencyNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does disclose a non-obvious behavioral rule: payments are matched only in the vendor's own currency, so a mismatched-currency payment surfaces as an open balance. This is exactly the kind of semantic caveat structured fields cannot convey, though it does not state read-only nature or any rate/latency characteristics.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences with the core output front-loaded and the caveat second; no filler, and the exception cross-reference to reconcile_ledger is placed where it is most useful.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return formatting is covered elsewhere, and the description supplies the key semantic gotcha an agent needs. Only the read-only nature and the precise meaning of the currency filter remain unstated.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% and the single currency parameter's enum is nested inside an anyOf, so the schema adds little prose. The description says the filter is optional, implying null means all currencies, but does not clarify what the filter actually constrains (vendor rows vs payments) given the currency-matching rule above.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states the resource (vendor) and the exact measures returned (billed, paid, open balance), which is specific enough to distinguish it from a generic query tool. It does not, however, explicitly differentiate itself from run_sql, the closest sibling for ad-hoc queries.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Optionally filtered by currency' and the note about wrong-currency payments imply the reconciliation use case, but there is no explicit statement of when to choose this over run_sql or reconcile_ledger. Usage is inferable but not spelled out.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 4 tool updatesv0.1.0
    • First observedexplain_exception
    • First observedreconcile_ledger
    • First observedrun_sql
    • First observedvendor_balances

TDQS

A3.5/5.0

Scored across 4 tools

Disambiguation4/5

run_sql is a generic read-only query escape hatch that could overlap with vendor_balances, but vendor_balances provides currency-aware aggregation that SQL alone wouldn't easily replicate. reconcile_ledger and explain_exception form a clear drill-down pair, so boundaries are mostly distinct.

Naming Consistency4/5

Three tools follow a verb_noun pattern (run_sql, reconcile_ledger, explain_exception), while vendor_balances is a noun phrase. All are snake_case and readable, but the verb-first consistency is not universal.

Tool Count5/5

Four tools are well-scoped for a read-only reconciliation server: a generic query, a balance report, a reconciliation engine, and an exception inspector. Each has a distinct role and none feels redundant.

Completeness4/5

The read-only surface covers querying, balances, reconciliation, and exception drill-down. A direct structured tool for listing invoices or payments is absent, but run_sql provides a viable workaround, leaving only minor gaps.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    B
    maintenance
    An accounting-ops agent that reconciles payments against open orders, auto-books provably safe payments through a deterministic policy gate, and escalates exceptions to a human queue with audit trails.
    MIT
  • A
    license
    A
    quality
    C
    maintenance
    Enables AI agents to interact with a double-entry ledger, offering tools for account management, balanced journal entries, balance queries, trial balance, and penny-perfect allocation. Built with safety by construction: no update/delete tools, idempotent posting, and an append-only journal.
    7
    MIT
  • F
    license
    A
    quality
    B
    maintenance
    Enables an AI agent to handle accounts payable tasks against a mock ERP, including reading and writing bills and vendors, checking duplicates, matching invoices, recommending approvals, and queuing payment releases, with configurable profiles that limit available tools.
    11
    -
  • F
    license
    A
    quality
    C
    maintenance
    Provides AI agents with deterministic, offline finance tools for commodity margin analysis, loan covenant compliance, invoice auditing, AP exception classification, and five-day close readiness.
    14
    6 npm
    -