Skip to main content
Glama

reconcile-mcp: An MCP Server for ISO 20022 Cash Reconciliation

Build Status PyPI version Glama MCP server score OpenSSF Scorecard License Python Version

A Model Context Protocol server that matches expected payments (from pain.001 credit transfers) against observed booked entries (from a camt.053 statement) and returns an explainable reconciliation — exact matches, short/over payments, split settlements (one-to-many), batch credits (many-to-one), and the residual unmatched items on each side, every match carrying a score and the reasons it was made.

Latest release: v0.0.7: 10 MCP tools over stdio, streamable HTTP or SSE, pure-Python matching engine, deterministic sandbox test-mode, for Python 3.10+. Part of the ISO 20022 MCP suite: you own both sides of the match.

Why this exists

Reconciliation is the treasury team's daily pain: did the money we expected actually arrive, and which invoice does each credit belong to? It is rarely one-to-one — customers underpay, settle an invoice in instalments, or a payout aggregator sends one lump covering a dozen receivables. reconcile-mcp does this matching as an agent tool, and — critically for finance — shows its work: every pairing comes with a numeric score and a plain list of the signals (reference, amount, date, counterparty) that drove it.

Related MCP server: AP Invoice Exception Review MCPB

The ISO 20022 MCP Suite

reconcile-mcp is the reconciliation workflow of eight coordinated, vendor-neutral MCP servers that together cover the ISO 20022 bank-statement workflow and the November 2026 structured-address cutover, plus a high-level orchestration layer — readiness scoring, clearing-profile linting, and audit evidence — statement depth, whole-catalogue routing, reconciliation, multi-format ingestion, and address remediation. Dependency ranges are kept aligned across the suite, so the servers co-install cleanly in a single Python environment: start with one, add the rest as your workflow grows.

Server

Scope

Surface

Install

Use it when

camt053-mcp

ISO 20022 camt.053/camt.052 bank statements: parse, validate, filter, reverse; MT940/MT942 migration; CBPR+ readiness; journal export

24 MCP tools · 4 prompts · 3 resources

pip install camt053-mcp

You work with bank-to-customer statements end to end — the suite's flagship

iso20022-mcp

Unified gateway: search / describe / validate / generate / parse meta-tools routed across the pain · pacs · camt · acmt families

7 meta-tools

pip install "iso20022-mcp[all]"

You want one entry point to every message family

reconcile-mcp

Matches expected pain.001 payments against observed camt.053 entries — exact, partial, one-to-many, many-to-one, every match scored and explained

10 MCP tools · 1 prompt · 2 resources

pip install reconcile-mcp

You need explainable statement/payment reconciliation — this package

bankstatementparser-mcp

Multi-format statement ingestion: ISO 20022 CAMT.053 and pain.001, SWIFT MT940, OFX/QFX, CSV

5 MCP tools · 1 prompt · 1 resource

pip install bankstatementparser-mcp

Your statements arrive in mixed or legacy formats

structured-address-fix-mcp

ISO 20022 postal-address classification, assessment & remediation for the November 2026 structured-address cutover (pacs.008 / pain.001 debtor & creditor addresses)

9 MCP tools

pip install structured-address-fix-mcp

You need debtor/creditor addresses cliff-ready ahead of 14 Nov 2026

iso20022-readiness-suite-mcp

Orchestration gateway: detect → structurally validate → clearing-profile lint → readiness score, plus automated remediation and pacs.002 bank-response simulation — a meta-client over the foundational servers

4 MCP tools

pip install iso20022-readiness-suite-mcp

You want one high-level readiness / orchestration entry point over the suite

iso20022-bank-profile-mcp

Manages, validates and serves bank-specific clearing profiles / rule packs (CBPR+, SEPA_Instant, FedNow, Generic); premium rule-pack entitlement gating

4 MCP tools

pip install iso20022-bank-profile-mcp

You lint payments against your own institution's market practice

iso20022-evidence-pack-mcp

Compiles readiness findings, remediation diffs and simulated responses into a sealed, Ed25519-signable audit evidence pack

6 MCP tools

pip install iso20022-evidence-pack-mcp

You need tamper-evident audit / certification artifacts

In one line each: camt053-mcp is the bank-statement flagship (deepest camt.05x surface, stdio + authenticated streamable HTTP); iso20022-mcp is the generic message toolkit (a handful of verbs over the whole catalogue); reconcile-mcp is the reconciliation workflow (did the money we expected actually arrive?); bankstatementparser-mcp is the ingestion layer (many formats in, one transaction shape out); and structured-address-fix-mcp is the postal-address specialist (debtor/creditor addresses cliff-ready for the Nov 2026 cutover).

The suite also includes per-family servers — pain001-mcp (credit transfer initiation), pacs008-mcp (FI-to-FI credit transfers), and acmt001-mcp (account management) — whose parsed output feeds straight into this server's normalize_* adapters.

Install

pip install reconcile-mcp
# or run without installing:
uvx reconcile-mcp

MCP client config (e.g. Claude Desktop claude_desktop_config.json):

{
  "mcpServers": {
    "reconcile": {
      "command": "reconcile-mcp"
    }
  }
}

Quick start (zero real data)

The server ships a sandbox test-mode: deterministic scenarios so you can run the whole flow with no setup and no real cash data. One call gets you a full, explainable result:

run_sandbox_scenario(name="month_end")

returns a realistic mixed close — one clean match, one short payment, one split settlement, and an unexpected credit correctly left unmatched:

{
  "summary": {
    "expected_count": 3, "observed_count": 5,
    "matched_expected": 3, "unmatched_observed": 1,
    "matches_by_type": {"exact": 1, "amount_mismatch": 1, "one_to_many": 1},
    "fully_reconciled": false
  },
  "matches": [
    {"type": "amount_mismatch", "expected": ["INV-6002"], "observed": ["ENT-52"],
     "amount_delta": "-99.99", "confidence": "high",
     "reasons": ["reference exact", "amount close (delta -99.99)", "date +/-0d", "counterparty exact"]},
    {"type": "exact", "expected": ["INV-6001"], "observed": ["ENT-51"], "amount_delta": "0.00"},
    {"type": "one_to_many", "expected": ["INV-6003"], "observed": ["ENT-53", "ENT-54"],
     "reasons": ["amount sum of 2 entries"]}
  ],
  "unmatched_observed": ["ENT-55"]
}

List every scenario with list_sandbox_scenarios; load one to inspect or edit its inputs with load_sandbox_scenario.

Transports

One command line, three transports:

Command

Transport

Endpoint

Protocol revisions

reconcile-mcp

stdio

the client spawns the process

2026-07-28, 2025-11-25

reconcile-mcp --transport streamable-http

Streamable HTTP

http://127.0.0.1:8000/mcp

2026-07-28 (stateless, server/discover) and 2025-11-25 (initialize, Mcp-Session-Id) on the same endpoint; responses stream as server-sent events, GET opens the server-to-client stream

reconcile-mcp --transport sse

HTTP+SSE (2024-11-05)

http://127.0.0.1:8000/sse and /messages/

for clients that still expect the older transport

--host and --port change the bind address (defaults 127.0.0.1 and 8000). The HTTP transports carry no authentication of their own: bind loopback, or put the server behind a gateway you trust before binding a routable address. Every release is verified over streamable HTTP with passmcp in both protocol eras and over SSE with the MCP SDK client; see ADR 0001.

{
  "mcpServers": {
    "reconcile": { "url": "http://127.0.0.1:8000/mcp" }
  }
}

Bring your own data

Records are small canonical objects — id and amount required, everything else optional and used to sharpen matching:

{
  "id": "INV-1001",            // your reference / end-to-end id
  "amount": 1200.00,
  "currency": "EUR",           // ISO 4217
  "date": "2026-03-02",        // ISO-8601
  "counterparty": "Acme Ltd",
  "reference": "INV-1001"      // remittance / structured reference
}

Already using the rest of the suite? Feed parsed output straight in — the adapters map it for you:

  • normalize_pain001(document) → the expected side, from pain001-mcp.

  • normalize_camt053(document) → the observed side, from camt053-mcp.

Then call reconcile(expected, observed).

Tools

  • reconcile — Match expected payments against observed entries; full explainable report.

  • explain_match — Score a single expected/observed pair with a per-signal breakdown (tuning aid).

  • normalize_pain001 — Adapt parsed pain.001 output into canonical expected records.

  • normalize_camt053 — Adapt parsed camt.053 output into canonical observed records.

  • list_sandbox_scenarios — List the built-in test-mode scenarios and magic references.

  • load_sandbox_scenario — Return one scenario's expected/observed inputs to inspect or edit.

  • run_sandbox_scenario — Load a scenario and reconcile it in one call — the fastest first run.

  • match_names_probabilistic — Score two counterparty names by Jaro-Winkler similarity and flag a match at a threshold.

  • match_amounts_with_fx_drift — Compare two amounts in different currencies through a supplied FX rate, within a percentage tolerance.

  • reconcile_many_to_many — Match each deposit to a disjoint subset of invoices whose amounts sum to it (subset-sum ILP; needs the ilp extra).

Plus one prompt and two resources:

  • Prompt reconcile_workflow — Step-by-step guidance from raw pain.001/camt.053 to an explained match report.

  • Resource reconcile://sandbox-scenarios — The built-in scenario catalogue.

  • Resource reconcile://sandbox/{scenario_id} — One scenario's expected/observed inputs.

How matching works

Each candidate pair is scored on four weighted signals, then classified:

  • Reference (0.45) — exact / partial equality of references and end-to-end ids, normalised to bare alphanumerics.

  • Amount (0.35) — exact within tolerance, or a linearly-decaying closeness with the delta reported.

  • Date (0.10) — proximity within a configurable window; neutral if unknown.

  • Counterparty (0.10) — token-set overlap of names; neutral if unknown.

Assignment is greedy, highest-score-first and fully deterministic (a total tiebreak order), so the same inputs always produce the same result. Residuals are then tested for one-to-many (a bounded subset-sum: one expected settled by several entries) and many-to-one (one entry covering several expected).

Tune any of it via the options argument: abs_tol / rel_tol, date_window_days, high_threshold, review_threshold, currency_strict, enable_one_to_many, max_combination.

Development

git clone https://github.com/sebastienrousseau/reconcile-mcp
cd reconcile-mcp
python -m venv .venv && . .venv/bin/activate
pip install -e . && pip install pytest pytest-cov ruff black mypy
pytest                      # 100% branch coverage gate
ruff check reconcile_mcp tests && black --check reconcile_mcp tests && mypy reconcile_mcp

License

Licensed under the Apache License, Version 2.0 or the MIT License, at your option.


mcp-name: io.github.sebastienrousseau/reconcile-mcp

Available Tools

10 tools
explain_matchA
Read-onlyIdempotent

Explain the match score between one expected and one observed record.

Purpose: Scores a single transaction pair across reference, amount, date, and name signals, returning individual signal weights and an explanation even for pairs below threshold.

When to use:

  • When tuning scoring weights, tolerances, or debugging match decisions.

  • When explaining match confidence for human review or exception handling.

When NOT to use:

  • Do NOT use for entire batches; use reconcile instead.

  • Do NOT use for cross-currency conversions without a known FX rate; use match_amounts_with_fx_drift.

Behavioral transparency: Pure, deterministic, side-effect-free, read-only calculation without external dependencies.

ParametersJSON Schema
NameRequiredDescriptionDefault
optionsNoOptional tuning object: 'abs_tol'/'rel_tol' (amount tolerance), 'date_window_days', 'high_threshold', 'review_threshold', 'currency_strict', 'enable_one_to_many', 'max_combination'.
expectedYesOne expected record.
observedYesOne observed record.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false, so the safety profile is covered. The description adds genuine value by noting it returns individual signal weights and an explanation even for below-threshold pairs, plus that it is pure and dependency-free. Some phrases ('side-effect-free, read-only') restate annotations rather than extend them.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The purpose sentence is front-loaded and the labeled sections (Purpose, When to use, When NOT to use, Behavioral transparency) make it scannable. Slightly verbose with some redundancy between the purpose line and the purpose section, but nothing is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be described, and annotations carry the safety profile. The description covers purpose, routing, and determinism fully, leaving nothing an agent needs in order to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the options object already enumerates its tuning keys (abs_tol/rel_tol, date_window_days, thresholds, etc.). The description mentions tuning weights and tolerances but adds no syntax or format detail beyond what the schema provides, so baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Explain the match score between one expected and one observed record') and elaborates on what signals are scored (reference, amount, date, name). This clearly distinguishes it from siblings like reconcile and match_amounts_with_fx_drift.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides explicit 'When to use' and 'When NOT to use' sections, naming concrete alternatives (reconcile for batches, match_amounts_with_fx_drift for cross-currency without FX rate). The routing guidance is unambiguous.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_sandbox_scenariosA
Read-onlyIdempotent

List built-in sandbox reconciliation scenarios and magic references.

Purpose: Returns the catalog of deterministic test-mode scenarios and synthetic magic references to demonstrate reconciliation outcomes without production data.

When to use:

  • When discovering available test scenarios (e.g. clean_match, month_end, fx_drift).

  • When testing agent decision-making workflows with zero setup.

When NOT to use:

  • Do NOT use for production reconciliation batches.

Behavioral transparency: Static catalogue retrieval; read-only and instantaneous.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, openWorldHint=false and destructiveHint=false, so the safety profile is fully covered by structured data. The description's line 'Static catalogue retrieval; read-only and instantaneous' largely restates that, adding only the minor detail that retrieval is instantaneous and involves no production data.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Content is front-loaded with the resource definition first, then usage. The labeled-section formatting ('Purpose:', 'When to use:', 'Behavioral transparency:') is slightly heavier than necessary for a short catalogue tool, but every sentence carries information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, annotations covering the safety profile, and zero parameters, the description needs only to state scope and usage - which it does. Nothing an agent needs in order to call this tool correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so there is nothing for the description to disambiguate. Baseline 4 applies; no parameter guidance is needed or missing.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource ('List built-in sandbox reconciliation scenarios and magic references') and clarifies these are deterministic test-mode items, so the agent knows exactly what is returned. It does not explicitly distinguish itself from siblings like load_sandbox_scenario or run_sandbox_scenario, relying on the verb 'list' to imply catalog-only retrieval.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives explicit when-to-use conditions with concrete examples (clean_match, month_end, fx_drift) and a clear when-NOT-to-use exclusion ('Do NOT use for production reconciliation batches'). This routes the agent well against the run/load siblings without ambiguity.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

load_sandbox_scenarioA
Read-onlyIdempotent

Load expected and observed record inputs for a named sandbox scenario.

Purpose: Retrieves the test-mode expected and observed fixture records for a named scenario so they can be inspected, customized, or passed to reconcile.

When to use:

  • When reviewing fixture data prior to executing reconciliation runs.

  • When creating custom variations of standard reconciliation scenarios.

When NOT to use:

  • Do NOT use if you want to run and reconcile the scenario in one step; use run_sandbox_scenario.

  • Do NOT use with non-existent scenario names.

Behavioral transparency: Deterministic dictionary lookup; read-only with no side effects.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYesScenario name, e.g. 'clean_match'.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false, and openWorldHint=false, so the safety profile is fully covered structurally. The description's 'Deterministic dictionary lookup; read-only with no side effects' largely restates those annotations rather than adding new behavioral context such as error behavior for unknown names.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Header-based structure with the core action front-loaded in the first sentence. Slightly more verbose than needed for a one-parameter lookup, but every section is short and purposeful.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple deterministic lookup with an output schema present (so return values need no explanation), the description covers purpose, routing, and behavior adequately. An agent has everything needed to select and call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the single 'name' parameter already carries an example ('clean_match'). The description only echoes 'named scenario' without adding format or constraint detail beyond the schema, so baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Load expected and observed record inputs') plus the scope qualifier 'for a named sandbox scenario'. It is clearly distinguishable from siblings like run_sandbox_scenario and list_sandbox_scenarios.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides explicit when-to-use bullets (reviewing fixtures, creating custom variations) and when-NOT-to-use bullets that name the alternative tool to use instead (run_sandbox_scenario). Nothing is left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

match_amounts_with_fx_driftMatch amounts (FX drift)A
Read-onlyIdempotent

Compare two amounts in different currencies within an FX drift tolerance.

Purpose: Converts currency amounts using a provided exchange rate and verifies if the percentage difference falls within an acceptable tolerance window using exact decimal arithmetic.

When to use:

  • When reconciling cross-currency payments with conversion timing or rate variance.

  • When evaluating FX tolerance bounds on cross-border transactions.

When NOT to use:

  • Do NOT use for same-currency comparisons where standard tolerance applies; use reconcile.

  • Do NOT use with non-positive FX rates.

Behavioral transparency: Pure mathematical calculation; read-only and idempotent.

ParametersJSON Schema
NameRequiredDescriptionDefault
fx_rateYesUnits of currency_a per one unit of currency_b (e.g. an EUR/USD quote of 1.08 is USD per EUR).
amount_aYesAmount denominated in currency_a.
amount_bYesAmount denominated in currency_b.
currency_aYesISO 4217 code of amount_a.
currency_bYesISO 4217 code of amount_b.
tolerance_pctNoMax percentage difference still counted as a match.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint and destructiveHint=false, so the restatement of 'read-only and idempotent' is redundant. However the description adds genuine context beyond the annotations: 'pure mathematical calculation', exact decimal arithmetic semantics, and the non-positive-rate input constraint.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with a one-line summary followed by scannable Purpose/When-to-use/When-NOT sections; every sentence is informative. Slightly penalized for restating the annotation-declared read-only/idempotent traits, which duplicate structured data.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values needn't be described, and annotations already cover the safety profile. Combined with the usage routing, edge-case rejection (non-positive rates) and arithmetic-precision note, an agent has everything needed to call this correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, including the direction of the fx_rate quote and the tolerance_pct default, so the schema carries the parameter burden. The description mentions conversion and a tolerance window conceptually but adds no syntax or format detail beyond what the schema already provides; baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: compare two amounts in different currencies within an FX drift tolerance, then explicitly says what it does (converts using a provided rate and checks percentage difference). It distinguishes itself from the sibling 'reconcile' by naming it in the when-NOT section.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides explicit 'When to use' bullets (cross-currency reconciliation, FX tolerance bounds) and 'When NOT to use' bullets that name the alternative ('use reconcile' for same-currency) and an invalid input case (non-positive FX rates). Nothing is left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

match_names_probabilisticMatch names (probabilistic)A
Read-onlyIdempotent

Score two counterparty names using Jaro-Winkler string similarity.

Purpose: Computes similarity in [0, 1] between two company or individual names, tolerating typographical variations and legal suffix abbreviations (e.g. 'Corp' vs 'Corporation Inc').

When to use:

  • When verifying whether two counterparty or entity names represent the same party.

  • When evaluating name similarity thresholds for automated matching rules.

When NOT to use:

  • Do NOT use for full multi-signal transaction matching; use reconcile or explain_match.

  • Do NOT use with non-string inputs.

Behavioral transparency: Pure, deterministic string calculation; read-only and idempotent.

ParametersJSON Schema
NameRequiredDescriptionDefault
name_aYesFirst counterparty name.
name_bYesSecond counterparty name.
thresholdNoSimilarity in [0, 1] at or above which the pair is a match.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false, and openWorldHint=false, so the safety profile is covered. The description still adds value by stating the operation is a pure, deterministic string calculation that tolerates typographical variation and legal suffix abbreviations, which explains behavior annotations cannot express.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core operation, then cleanly grouped under Purpose / When to use / When NOT to use / Behavioral transparency headers. Slightly over-structured for a three-parameter tool, and the transparency line partly restates annotations, but no sentence is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Covers purpose, positive usage, negative usage with alternatives, and behavioral character. An output schema exists, so return values need no explanation, and the remaining gaps are negligible for a deterministic two-string comparison tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and all three parameters carry descriptions, so this is the baseline case. The description adds meaningful context about what the computed similarity represents and what variations it tolerates, but adds nothing about the threshold parameter beyond what the schema already says.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Score two counterparty names using Jaro-Winkler string similarity') and pins down what the score means ([0, 1]). It explicitly names the siblings it is not (reconcile, explain_match), so an agent can distinguish it without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides explicit when-to-use cases (verifying same party, evaluating matching thresholds) and explicit when-NOT-to-use cases naming the alternative tools for multi-signal matching. Nothing about tool selection is left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

normalize_camt053A
Read-onlyIdempotent

Convert parsed camt.053 statement entries into canonical observed records.

Purpose: Adapts parsed ISO 20022 camt.053 bank statement documents or entry lists into canonical observed records ready for the reconcile engine.

When to use:

  • When ingesting parsed bank statements before calling reconcile.

  • When transforming bank booking entries into uniform id, amount, date, and counterparty fields.

When NOT to use:

  • Do NOT use for payment initiation orders (pain.001); use normalize_pain001 instead.

  • Do NOT pass raw unparsed XML text; parse or deserialize first.

Behavioral transparency: Pure, deterministic, read-only data normalization.

ParametersJSON Schema
NameRequiredDescriptionDefault
documentYesParsed camt.053 document or entry list.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint and destructiveHint=false, so the safety bar is low and 'pure, deterministic, read-only' partly restates them. However, the description adds a genuine input constraint beyond the schema/annotations (must be parsed first, not raw XML) and clarifies the deterministic transformation intent, which is useful behavioral context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded verb+resource sentence followed by clearly labeled Purpose / When / When-NOT / Behavioral sections, which is easy to scan. The 'Purpose' paragraph mildly restates the opening line, costing a little redundancy, but overall it is tight and well-organized.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained. The description covers source format, target representation, the pre-parse prerequisite, exclusions, and deterministic behavior — everything an agent needs to invoke it correctly given the annotations and output schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There is a single required parameter with 100% schema coverage, so the schema already documents it. The description restates that it accepts a parsed document or entry list but adds no format, shape, or validation detail beyond the schema. Baseline 3 is appropriate when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (convert/adapt) and resource (parsed camt.053 statement entries into canonical observed records), and names the target consumer (reconcile engine). It explicitly distinguishes itself from the sibling normalize_pain001, so an agent can route correctly without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides explicit when-to-use bullets (before reconcile, transforming booking entries) and when-NOT-to-use bullets (not for pain.001 — use normalize_pain001; not for raw XML). Both the trigger condition and the alternative sibling are named.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

normalize_pain001A
Read-onlyIdempotent

Convert parsed pain.001 payment instructions into canonical expected records.

Purpose: Adapts parsed ISO 20022 pain.001 credit transfer documents or transaction lists into canonical expected records ready for the reconcile engine.

When to use:

  • When ingesting parsed pain.001 files before calling reconcile.

  • When mapping diverse payment field structures into uniform id, amount, and reference fields.

When NOT to use:

  • Do NOT use for bank statement entries (camt.053); use normalize_camt053 instead.

  • Do NOT pass raw unparsed XML text; parse or deserialize to a dictionary or list first.

Behavioral transparency: Pure, deterministic, read-only data normalization.

ParametersJSON Schema
NameRequiredDescriptionDefault
documentYesParsed pain.001 document or transaction list.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false and openWorldHint=false, so safety is covered. The description reinforces this with 'pure, deterministic, read-only data normalization' and adds the input-hygiene constraint about not passing raw XML, but it does not say what happens on malformed/empty documents or how errors surface.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Clearly front-loaded and parsed into labeled sections (Purpose / When to use / When NOT to use / Behavioral transparency) that make it skimmable. Slight redundancy between the opening sentence and the Purpose paragraph, but no filler and nothing critical buried.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter normalizer, the description covers purpose, routing to alternatives, and input preconditions, while the output schema carries return-value details and annotations carry the safety profile. Nothing needed to invoke it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Only one parameter, and schema description coverage is 100%, so the baseline is 3. The description adds domain framing (parsed pain.001 document OR transaction list, and that fields are mapped to id/amount/reference) but no format or structure detail beyond what the schema already conveys.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource (convert parsed pain.001 instructions into canonical expected records) and names the domain (ISO 20022 credit transfers) explicitly. It also distinguishes itself from the sibling normalize_camt053 by naming it as the wrong tool for bank statements, so an agent can separate the two without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides explicit when-to-use conditions (ingesting parsed pain.001 before reconcile, mapping heterogeneous fields to uniform id/amount/reference) and when-NOT-to-use conditions with the named alternative (camt.053 → normalize_camt053). It further states a concrete prerequisite: parse/deserialize raw XML before passing it in.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

reconcileA
Read-onlyIdempotent

Reconcile expected payments against observed bank-statement entries.

Purpose: Matches expected payment instructions (e.g. from pain.001) against observed bank transactions (e.g. from camt.053) and produces an explainable reconciliation report covering exact matches, short/over adjustments, split settlements (one-to-many), batch credits (many-to-one), and unmatched residuals.

When to use:

  • When performing multi-transaction cash reconciliation between ledgers and statements.

  • When transparent per-match confidence scores and reason breakdowns are required for audit.

When NOT to use:

  • Do NOT use for raw, unparsed XML files; normalize input documents first using normalize_pain001 or normalize_camt053.

  • Do NOT use for single-pair tuning analysis; use explain_match instead.

  • Do NOT use for unconstrained combinatorial invoice subsets; use reconcile_many_to_many.

Behavioral transparency: Pure, deterministic, side-effect-free, read-only calculation without network or disk access.

ParametersJSON Schema
NameRequiredDescriptionDefault
optionsNoOptional tuning object: 'abs_tol'/'rel_tol' (amount tolerance), 'date_window_days', 'high_threshold', 'review_threshold', 'currency_strict', 'enable_one_to_many', 'max_combination'.
expectedYesList of canonical records. Each is an object with 'id' (string) and 'amount' (number) required, plus optional 'currency' (ISO 4217), 'date' (ISO-8601), 'counterparty' (name), 'reference' (remittance/end-to-end id).
observedYesList of canonical records. Each is an object with 'id' (string) and 'amount' (number) required, plus optional 'currency' (ISO 4217), 'date' (ISO-8601), 'counterparty' (name), 'reference' (remittance/end-to-end id).

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, openWorldHint=false, and destructiveHint=false; the description reiterates these with useful framing ('pure, deterministic, side-effect-free, read-only calculation without network or disk access'), which adds confidence about determinism and isolation but doesn't add much beyond structured data.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core action, then organized into Purpose / When to use / When NOT to use / Behavioral sections. Each sentence is purposeful and the structure aids scanning.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given a complex reconciliation tool with an output schema and rich annotations, the description covers purpose, usage, exclusions, and behavioral traits comprehensively. An agent has all it needs to select and call the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents the parameters thoroughly. The description references the input formats (pain.001, camt.053) which contextualizes the 'expected'/'observed' semantics, but it doesn't compensate for the high schema coverage by adding further detail.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a precise verb+resource statement ('Reconcile expected payments against observed bank-statement entries') and enumerates the specific matching scenarios handled (exact, short/over, split, batch, residuals), which clearly distinguishes it from sibling tools like reconcile_many_to_many and explain_match.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Offers explicit 'When to use' and 'When NOT to use' sections that name specific alternatives (normalize_pain001, normalize_camt053, explain_match, reconcile_many_to_many) and the conditions that select them, leaving no ambiguity about routing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

reconcile_many_to_manyReconcile many-to-manyA
Read-onlyIdempotent

Match statement deposits to disjoint subsets of invoices via subset-sum ILP.

Purpose: Solves bounded subset-sum as an integer linear program (ILP) to pair statement deposits against matching combinations of outstanding invoice records.

When to use:

  • When customer deposits consolidate multiple invoices or split remittances.

  • When one-to-one or one-to-many heuristics leave unmatched aggregates.

When NOT to use:

  • Do NOT use for straightforward 1:1 or 1:N reconciliation where reconcile is sufficient and faster.

  • Do NOT use without the optional ilp extra installed (pip install reconcile-mcp[ilp]).

Behavioral transparency: Deterministic ILP optimization; read-only and idempotent.

ParametersJSON Schema
NameRequiredDescriptionDefault
invoicesYesList of canonical records. Each is an object with 'id' (string) and 'amount' (number) required, plus optional 'currency' (ISO 4217), 'date' (ISO-8601), 'counterparty' (name), 'reference' (remittance/end-to-end id).
statementsYesList of canonical records. Each is an object with 'id' (string) and 'amount' (number) required, plus optional 'currency' (ISO 4217), 'date' (ISO-8601), 'counterparty' (name), 'reference' (remittance/end-to-end id).

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover readOnly/idempotent/non-destructive, but the description adds genuine context beyond them: deterministic ILP optimization, bounded subset-sum formulation, and the hard dependency on the optional 'ilp' extra (pip install reconcile-mcp[ilp]). It does not discuss runtime cost or result ambiguity, but the dependency disclosure is valuable.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core action, then cleanly sectioned into Purpose / When to use / When NOT to use / Behavioral transparency. Every line carries information, though the labeled headings are slightly more verbose than an inline two-sentence form would be.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Output schema exists so return values need not be explained. With usage conditions, exclusions, dependency requirement, and behavioral profile all covered for a 2-param tool, nothing an agent needs to invoke it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and both parameters are richly documented in the schema itself (id, amount, currency, date, counterparty, reference). The description adds no syntax or format detail beyond what the schema already supplies, so baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a precise verb+resource (match statement deposits to disjoint subsets of invoices) plus the mechanism (subset-sum ILP). It explicitly contrasts with the sibling 'reconcile', so an agent can distinguish the two without reading schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides explicit 'when to use' cases (consolidated deposits, split remittances, unmatched aggregates) and explicit 'when NOT to use' cases naming the alternative tool 'reconcile' and the missing-extra prerequisite. Routing is unambiguous.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

run_sandbox_scenarioA
Read-onlyIdempotent

Load a named sandbox scenario and immediately return its reconciliation report.

Purpose: One-call execution that loads built-in fixture records and executes the reconciliation pipeline, returning a complete explainable match report.

When to use:

  • When performing quick verification, smoke tests, or initial agent demonstrations.

  • When verifying tolerance tuning against standard test cases.

When NOT to use:

  • Do NOT use with live production data.

  • Do NOT use when custom expected or observed records must be supplied; use reconcile instead.

Behavioral transparency: Deterministic, in-memory execution; read-only and idempotent.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYesScenario name, e.g. 'month_end'.
optionsNoOptional tuning object: 'abs_tol'/'rel_tol' (amount tolerance), 'date_window_days', 'high_threshold', 'review_threshold', 'currency_strict', 'enable_one_to_many', 'max_combination'.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Adds useful context beyond annotations: deterministic in-memory execution and built-in fixture records. However, 'read-only and idempotent' restates readOnlyHint/idempotentHint already present in the annotations, so the incremental value is partly redundant.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded purpose with clear section headers, and each bullet earns its place. The labeled 'Behavioral transparency' section partly duplicates the annotations, adding mild redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return-value detail is unnecessary. For a deterministic fixture-runner with full annotation and schema coverage, nothing an agent needs to invoke it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so both 'name' and the tuning 'options' object are fully documented in the schema. The description adds no syntax or format detail beyond that, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (load) plus resource (named sandbox scenario) and the immediate effect (returns reconciliation report). The 'Purpose' line distinguishes it clearly from reconcile and load_sandbox_scenario.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit when-to-use bullets (smoke tests, tolerance tuning) and when-not-to-use bullets (no live data, no custom records), naming the alternative tool 'reconcile' for the custom-records case.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 3 tool updatesv0.0.5
    • Addedmatch_amounts_with_fx_drift
    • Addedmatch_names_probabilistic
    • Addedreconcile_many_to_many
  2. 7 tool updatesv0.0.2
    • First observedexplain_match
    • First observedlist_sandbox_scenarios
    • First observedload_sandbox_scenario
    • First observednormalize_camt053
    • First observednormalize_pain001
    • First observedreconcile
    • First observedrun_sandbox_scenario

TDQS

A4.4/5.0

Scored across 10 tools

Disambiguation5/5

Each tool targets a distinct operation and resource, with explicit 'When NOT to use' guidance that cleanly separates reconcile, reconcile_many_to_many, explain_match, and the specialized matchers. The sandbox tools are also clearly divided into list, load, and run.

Naming Consistency5/5

All names use snake_case and follow a consistent verb-first pattern (normalize_*, reconcile*, explain_match, match_*, list/load/run_*). Domain-specific suffixes like camt053 and pain001 are predictable and do not introduce mixed conventions.

Tool Count5/5

Ten tools is well-scoped for a reconciliation engine, covering normalization, core matching, explanation, specialized matchers, and sandbox fixtures. No tool appears to be filler, and the set is neither too thin nor too heavy.

Completeness4/5

The core reconciliation lifecycle is covered: normalize both message types, reconcile, explain, many-to-many matching, specialized matchers, and sandbox fixtures. Minor gaps exist for full batch FX-drift reconciliation and other ISO 20022 message types, but agents can work around them.

Maintenance

ActivityActive
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    B
    quality
    D
    maintenance
    AI workbench for financial contract analysis, risk analytics (VaR/CVaR, RWA Basel III), regulatory compliance (EMIR, REMIT, MiFID II, CBAM, EUDR) and counterparty due diligence (KYB/UBO, OFAC, IMO). Zero Retention. 8 MCP tools.
    8
    31 npm
    MIT
  • F
    license
    Not graded
    quality
    D
    maintenance
    Self-serve MCPB demo for accounts payable invoice exception review. It performs deterministic matching across invoice, purchase order, goods receipt, vendor master, invoice history, tax code master, and payment rules.
    -
  • A
    license
    A
    quality
    A
    maintenance
    Unified gateway for ISO 20022 message families, providing meta-tools to search, describe, validate, generate, and parse financial messages.
    7
    356 PyPI
    1
    Apache 2.0
  • F
    license
    B
    quality
    B
    maintenance
    Enables AI assistants to perform financial reconciliation with a deterministic proof engine: intake files, match transactions, verify proofs, resolve exceptions, and sign off on balanced journals under the user's authority.
    21
    -