reconcile-mcp
reconcile-mcp is an MCP server that matches expected ISO 20022 payments against observed bank-statement entries and returns explainable reconciliation results.
Reconcile expected
pain.001payments vs observedcamt.053entries, returning exact matches, short/over payments, one-to-many split settlements, many-to-one batch credits, and unmatched residuals, each with scores and reasons.Explain a single expected/observed pair with a per-signal breakdown (reference, amount, date, counterparty).
Normalize parsed
pain.001documents into canonical expected records.Normalize parsed
camt.053documents into canonical observed records.List, load, and run built-in deterministic sandbox scenarios for zero-data demos and smoke tests.
Probabilistically match counterparty names using Jaro-Winkler similarity and a configurable threshold.
Compare amounts across currencies using a supplied FX rate and percentage tolerance with exact decimal arithmetic.
Reconcile many-to-many via subset-sum ILP, matching each deposit/statement to a disjoint subset of invoices.
Tune matching via options such as amount tolerances, date window, high/review thresholds, currency strictness, one-to-many enablement, and max combination size.
Expose 10 MCP tools over stdio, streamable HTTP, or SSE, plus a reconciliation workflow prompt and sandbox scenario resources.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@reconcile-mcpReconcile my expected payments against observed bank entries."
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
reconcile-mcp: An MCP Server for ISO 20022 Cash Reconciliation
A Model Context Protocol server that matches expected payments
(from pain.001 credit transfers) against observed booked entries (from a
camt.053 statement) and returns an explainable reconciliation — exact
matches, short/over payments, split settlements (one-to-many), batch credits
(many-to-one), and the residual unmatched items on each side, every match
carrying a score and the reasons it was made.
Latest release: v0.0.7: 10 MCP tools over stdio, streamable HTTP or SSE, pure-Python matching engine, deterministic sandbox test-mode, for Python 3.10+. Part of the ISO 20022 MCP suite: you own both sides of the match.
Why this exists
Reconciliation is the treasury team's daily pain: did the money we expected
actually arrive, and which invoice does each credit belong to? It is rarely
one-to-one — customers underpay, settle an invoice in instalments, or a payout
aggregator sends one lump covering a dozen receivables. reconcile-mcp does
this matching as an agent tool, and — critically for finance — shows its
work: every pairing comes with a numeric score and a plain list of the
signals (reference, amount, date, counterparty) that drove it.
Related MCP server: AP Invoice Exception Review MCPB
The ISO 20022 MCP Suite
reconcile-mcp is the reconciliation workflow of eight coordinated,
vendor-neutral MCP servers that together cover the ISO 20022 bank-statement
workflow and the November 2026 structured-address cutover, plus a high-level orchestration layer — readiness scoring, clearing-profile linting, and audit evidence — statement depth,
whole-catalogue routing, reconciliation, multi-format ingestion, and address
remediation. Dependency ranges are kept aligned across the suite,
so the servers co-install cleanly in a single Python environment: start with
one, add the rest as your workflow grows.
Server | Scope | Surface | Install | Use it when |
ISO 20022 | 24 MCP tools · 4 prompts · 3 resources |
| You work with bank-to-customer statements end to end — the suite's flagship | |
Unified gateway: | 7 meta-tools |
| You want one entry point to every message family | |
Matches expected | 10 MCP tools · 1 prompt · 2 resources |
| You need explainable statement/payment reconciliation — this package | |
Multi-format statement ingestion: ISO 20022 CAMT.053 and pain.001, SWIFT MT940, OFX/QFX, CSV | 5 MCP tools · 1 prompt · 1 resource |
| Your statements arrive in mixed or legacy formats | |
ISO 20022 postal-address classification, assessment & remediation for the November 2026 structured-address cutover ( | 9 MCP tools |
| You need debtor/creditor addresses cliff-ready ahead of 14 Nov 2026 | |
Orchestration gateway: detect → structurally validate → clearing-profile lint → readiness score, plus automated remediation and | 4 MCP tools |
| You want one high-level readiness / orchestration entry point over the suite | |
Manages, validates and serves bank-specific clearing profiles / rule packs (CBPR+, SEPA_Instant, FedNow, Generic); premium rule-pack entitlement gating | 4 MCP tools |
| You lint payments against your own institution's market practice | |
Compiles readiness findings, remediation diffs and simulated responses into a sealed, Ed25519-signable audit evidence pack | 6 MCP tools |
| You need tamper-evident audit / certification artifacts |
In one line each: camt053-mcp is the bank-statement flagship (deepest
camt.05x surface, stdio + authenticated streamable HTTP);
iso20022-mcp is the generic message toolkit (a handful of verbs over
the whole catalogue); reconcile-mcp is the reconciliation workflow
(did the money we expected actually arrive?);
bankstatementparser-mcp is the ingestion layer (many formats in, one
transaction shape out); and structured-address-fix-mcp is the
postal-address specialist (debtor/creditor addresses cliff-ready for the
Nov 2026 cutover).
The suite also includes per-family servers — pain001-mcp
(credit transfer initiation), pacs008-mcp (FI-to-FI credit
transfers), and acmt001-mcp (account management) — whose
parsed output feeds straight into this server's normalize_* adapters.
Install
pip install reconcile-mcp
# or run without installing:
uvx reconcile-mcpMCP client config (e.g. Claude Desktop claude_desktop_config.json):
{
"mcpServers": {
"reconcile": {
"command": "reconcile-mcp"
}
}
}Quick start (zero real data)
The server ships a sandbox test-mode: deterministic scenarios so you can run the whole flow with no setup and no real cash data. One call gets you a full, explainable result:
run_sandbox_scenario(name="month_end")returns a realistic mixed close — one clean match, one short payment, one split settlement, and an unexpected credit correctly left unmatched:
{
"summary": {
"expected_count": 3, "observed_count": 5,
"matched_expected": 3, "unmatched_observed": 1,
"matches_by_type": {"exact": 1, "amount_mismatch": 1, "one_to_many": 1},
"fully_reconciled": false
},
"matches": [
{"type": "amount_mismatch", "expected": ["INV-6002"], "observed": ["ENT-52"],
"amount_delta": "-99.99", "confidence": "high",
"reasons": ["reference exact", "amount close (delta -99.99)", "date +/-0d", "counterparty exact"]},
{"type": "exact", "expected": ["INV-6001"], "observed": ["ENT-51"], "amount_delta": "0.00"},
{"type": "one_to_many", "expected": ["INV-6003"], "observed": ["ENT-53", "ENT-54"],
"reasons": ["amount sum of 2 entries"]}
],
"unmatched_observed": ["ENT-55"]
}List every scenario with list_sandbox_scenarios; load one to inspect or edit
its inputs with load_sandbox_scenario.
Transports
One command line, three transports:
Command | Transport | Endpoint | Protocol revisions |
| stdio | the client spawns the process | 2026-07-28, 2025-11-25 |
| Streamable HTTP |
| 2026-07-28 (stateless, |
| HTTP+SSE (2024-11-05) |
| for clients that still expect the older transport |
--host and --port change the bind address (defaults 127.0.0.1 and
8000). The HTTP transports carry no authentication of their own: bind
loopback, or put the server behind a gateway you trust before binding a
routable address. Every release is verified over streamable HTTP with
passmcp in both protocol
eras and over SSE with the MCP SDK client; see
ADR 0001.
{
"mcpServers": {
"reconcile": { "url": "http://127.0.0.1:8000/mcp" }
}
}Bring your own data
Records are small canonical objects — id and amount required, everything
else optional and used to sharpen matching:
{
"id": "INV-1001", // your reference / end-to-end id
"amount": 1200.00,
"currency": "EUR", // ISO 4217
"date": "2026-03-02", // ISO-8601
"counterparty": "Acme Ltd",
"reference": "INV-1001" // remittance / structured reference
}Already using the rest of the suite? Feed parsed output straight in — the adapters map it for you:
normalize_pain001(document)→ the expected side, frompain001-mcp.normalize_camt053(document)→ the observed side, fromcamt053-mcp.
Then call reconcile(expected, observed).
Tools
reconcile— Match expected payments against observed entries; full explainable report.explain_match— Score a single expected/observed pair with a per-signal breakdown (tuning aid).normalize_pain001— Adapt parsedpain.001output into canonical expected records.normalize_camt053— Adapt parsedcamt.053output into canonical observed records.list_sandbox_scenarios— List the built-in test-mode scenarios and magic references.load_sandbox_scenario— Return one scenario's expected/observed inputs to inspect or edit.run_sandbox_scenario— Load a scenario and reconcile it in one call — the fastest first run.match_names_probabilistic— Score two counterparty names by Jaro-Winkler similarity and flag a match at a threshold.match_amounts_with_fx_drift— Compare two amounts in different currencies through a supplied FX rate, within a percentage tolerance.reconcile_many_to_many— Match each deposit to a disjoint subset of invoices whose amounts sum to it (subset-sum ILP; needs theilpextra).
Plus one prompt and two resources:
Prompt
reconcile_workflow— Step-by-step guidance from rawpain.001/camt.053to an explained match report.Resource
reconcile://sandbox-scenarios— The built-in scenario catalogue.Resource
reconcile://sandbox/{scenario_id}— One scenario's expected/observed inputs.
How matching works
Each candidate pair is scored on four weighted signals, then classified:
Reference (0.45) — exact / partial equality of references and end-to-end ids, normalised to bare alphanumerics.
Amount (0.35) — exact within tolerance, or a linearly-decaying closeness with the delta reported.
Date (0.10) — proximity within a configurable window; neutral if unknown.
Counterparty (0.10) — token-set overlap of names; neutral if unknown.
Assignment is greedy, highest-score-first and fully deterministic (a total tiebreak order), so the same inputs always produce the same result. Residuals are then tested for one-to-many (a bounded subset-sum: one expected settled by several entries) and many-to-one (one entry covering several expected).
Tune any of it via the options argument: abs_tol / rel_tol,
date_window_days, high_threshold, review_threshold, currency_strict,
enable_one_to_many, max_combination.
Development
git clone https://github.com/sebastienrousseau/reconcile-mcp
cd reconcile-mcp
python -m venv .venv && . .venv/bin/activate
pip install -e . && pip install pytest pytest-cov ruff black mypy
pytest # 100% branch coverage gate
ruff check reconcile_mcp tests && black --check reconcile_mcp tests && mypy reconcile_mcpLicense
Licensed under the Apache License, Version 2.0 or the MIT License, at your option.
mcp-name: io.github.sebastienrousseau/reconcile-mcp
Available Tools
10 toolsexplain_matchARead-onlyIdempotent
Explain the match score between one expected and one observed record.
Purpose: Scores a single transaction pair across reference, amount, date, and name signals, returning individual signal weights and an explanation even for pairs below threshold.
When to use:
When tuning scoring weights, tolerances, or debugging match decisions.
When explaining match confidence for human review or exception handling.
When NOT to use:
Do NOT use for entire batches; use reconcile instead.
Do NOT use for cross-currency conversions without a known FX rate; use match_amounts_with_fx_drift.
Behavioral transparency: Pure, deterministic, side-effect-free, read-only calculation without external dependencies.
| Name | Required | Description | Default |
|---|---|---|---|
| options | No | Optional tuning object: 'abs_tol'/'rel_tol' (amount tolerance), 'date_window_days', 'high_threshold', 'review_threshold', 'currency_strict', 'enable_one_to_many', 'max_combination'. | |
| expected | Yes | One expected record. | |
| observed | Yes | One observed record. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false, so the safety profile is covered. The description adds genuine value by noting it returns individual signal weights and an explanation even for below-threshold pairs, plus that it is pure and dependency-free. Some phrases ('side-effect-free, read-only') restate annotations rather than extend them.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The purpose sentence is front-loaded and the labeled sections (Purpose, When to use, When NOT to use, Behavioral transparency) make it scannable. Slightly verbose with some redundancy between the purpose line and the purpose section, but nothing is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be described, and annotations carry the safety profile. The description covers purpose, routing, and determinism fully, leaving nothing an agent needs in order to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the options object already enumerates its tuning keys (abs_tol/rel_tol, date_window_days, thresholds, etc.). The description mentions tuning weights and tolerances but adds no syntax or format detail beyond what the schema provides, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Explain the match score between one expected and one observed record') and elaborates on what signals are scored (reference, amount, date, name). This clearly distinguishes it from siblings like reconcile and match_amounts_with_fx_drift.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides explicit 'When to use' and 'When NOT to use' sections, naming concrete alternatives (reconcile for batches, match_amounts_with_fx_drift for cross-currency without FX rate). The routing guidance is unambiguous.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_sandbox_scenariosARead-onlyIdempotent
List built-in sandbox reconciliation scenarios and magic references.
Purpose: Returns the catalog of deterministic test-mode scenarios and synthetic magic references to demonstrate reconciliation outcomes without production data.
When to use:
When discovering available test scenarios (e.g. clean_match, month_end, fx_drift).
When testing agent decision-making workflows with zero setup.
When NOT to use:
Do NOT use for production reconciliation batches.
Behavioral transparency: Static catalogue retrieval; read-only and instantaneous.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, openWorldHint=false and destructiveHint=false, so the safety profile is fully covered by structured data. The description's line 'Static catalogue retrieval; read-only and instantaneous' largely restates that, adding only the minor detail that retrieval is instantaneous and involves no production data.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Content is front-loaded with the resource definition first, then usage. The labeled-section formatting ('Purpose:', 'When to use:', 'Behavioral transparency:') is slightly heavier than necessary for a short catalogue tool, but every sentence carries information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, annotations covering the safety profile, and zero parameters, the description needs only to state scope and usage - which it does. Nothing an agent needs in order to call this tool correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so there is nothing for the description to disambiguate. Baseline 4 applies; no parameter guidance is needed or missing.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource ('List built-in sandbox reconciliation scenarios and magic references') and clarifies these are deterministic test-mode items, so the agent knows exactly what is returned. It does not explicitly distinguish itself from siblings like load_sandbox_scenario or run_sandbox_scenario, relying on the verb 'list' to imply catalog-only retrieval.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives explicit when-to-use conditions with concrete examples (clean_match, month_end, fx_drift) and a clear when-NOT-to-use exclusion ('Do NOT use for production reconciliation batches'). This routes the agent well against the run/load siblings without ambiguity.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
load_sandbox_scenarioARead-onlyIdempotent
Load expected and observed record inputs for a named sandbox scenario.
Purpose: Retrieves the test-mode expected and observed fixture records for a named scenario so they can be inspected, customized, or passed to reconcile.
When to use:
When reviewing fixture data prior to executing reconciliation runs.
When creating custom variations of standard reconciliation scenarios.
When NOT to use:
Do NOT use if you want to run and reconcile the scenario in one step; use run_sandbox_scenario.
Do NOT use with non-existent scenario names.
Behavioral transparency: Deterministic dictionary lookup; read-only with no side effects.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | Scenario name, e.g. 'clean_match'. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false, and openWorldHint=false, so the safety profile is fully covered structurally. The description's 'Deterministic dictionary lookup; read-only with no side effects' largely restates those annotations rather than adding new behavioral context such as error behavior for unknown names.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Header-based structure with the core action front-loaded in the first sentence. Slightly more verbose than needed for a one-parameter lookup, but every section is short and purposeful.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple deterministic lookup with an output schema present (so return values need no explanation), the description covers purpose, routing, and behavior adequately. An agent has everything needed to select and call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the single 'name' parameter already carries an example ('clean_match'). The description only echoes 'named scenario' without adding format or constraint detail beyond the schema, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Load expected and observed record inputs') plus the scope qualifier 'for a named sandbox scenario'. It is clearly distinguishable from siblings like run_sandbox_scenario and list_sandbox_scenarios.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides explicit when-to-use bullets (reviewing fixtures, creating custom variations) and when-NOT-to-use bullets that name the alternative tool to use instead (run_sandbox_scenario). Nothing is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
match_amounts_with_fx_driftMatch amounts (FX drift)ARead-onlyIdempotent
Compare two amounts in different currencies within an FX drift tolerance.
Purpose: Converts currency amounts using a provided exchange rate and verifies if the percentage difference falls within an acceptable tolerance window using exact decimal arithmetic.
When to use:
When reconciling cross-currency payments with conversion timing or rate variance.
When evaluating FX tolerance bounds on cross-border transactions.
When NOT to use:
Do NOT use for same-currency comparisons where standard tolerance applies; use reconcile.
Do NOT use with non-positive FX rates.
Behavioral transparency: Pure mathematical calculation; read-only and idempotent.
| Name | Required | Description | Default |
|---|---|---|---|
| fx_rate | Yes | Units of currency_a per one unit of currency_b (e.g. an EUR/USD quote of 1.08 is USD per EUR). | |
| amount_a | Yes | Amount denominated in currency_a. | |
| amount_b | Yes | Amount denominated in currency_b. | |
| currency_a | Yes | ISO 4217 code of amount_a. | |
| currency_b | Yes | ISO 4217 code of amount_b. | |
| tolerance_pct | No | Max percentage difference still counted as a match. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint and destructiveHint=false, so the restatement of 'read-only and idempotent' is redundant. However the description adds genuine context beyond the annotations: 'pure mathematical calculation', exact decimal arithmetic semantics, and the non-positive-rate input constraint.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with a one-line summary followed by scannable Purpose/When-to-use/When-NOT sections; every sentence is informative. Slightly penalized for restating the annotation-declared read-only/idempotent traits, which duplicate structured data.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists so return values needn't be described, and annotations already cover the safety profile. Combined with the usage routing, edge-case rejection (non-positive rates) and arithmetic-precision note, an agent has everything needed to call this correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, including the direction of the fx_rate quote and the tolerance_pct default, so the schema carries the parameter burden. The description mentions conversion and a tolerance window conceptually but adds no syntax or format detail beyond what the schema already provides; baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: compare two amounts in different currencies within an FX drift tolerance, then explicitly says what it does (converts using a provided rate and checks percentage difference). It distinguishes itself from the sibling 'reconcile' by naming it in the when-NOT section.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides explicit 'When to use' bullets (cross-currency reconciliation, FX tolerance bounds) and 'When NOT to use' bullets that name the alternative ('use reconcile' for same-currency) and an invalid input case (non-positive FX rates). Nothing is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
match_names_probabilisticMatch names (probabilistic)ARead-onlyIdempotent
Score two counterparty names using Jaro-Winkler string similarity.
Purpose: Computes similarity in [0, 1] between two company or individual names, tolerating typographical variations and legal suffix abbreviations (e.g. 'Corp' vs 'Corporation Inc').
When to use:
When verifying whether two counterparty or entity names represent the same party.
When evaluating name similarity thresholds for automated matching rules.
When NOT to use:
Do NOT use for full multi-signal transaction matching; use reconcile or explain_match.
Do NOT use with non-string inputs.
Behavioral transparency: Pure, deterministic string calculation; read-only and idempotent.
| Name | Required | Description | Default |
|---|---|---|---|
| name_a | Yes | First counterparty name. | |
| name_b | Yes | Second counterparty name. | |
| threshold | No | Similarity in [0, 1] at or above which the pair is a match. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false, and openWorldHint=false, so the safety profile is covered. The description still adds value by stating the operation is a pure, deterministic string calculation that tolerates typographical variation and legal suffix abbreviations, which explains behavior annotations cannot express.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core operation, then cleanly grouped under Purpose / When to use / When NOT to use / Behavioral transparency headers. Slightly over-structured for a three-parameter tool, and the transparency line partly restates annotations, but no sentence is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Covers purpose, positive usage, negative usage with alternatives, and behavioral character. An output schema exists, so return values need no explanation, and the remaining gaps are negligible for a deterministic two-string comparison tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and all three parameters carry descriptions, so this is the baseline case. The description adds meaningful context about what the computed similarity represents and what variations it tolerates, but adds nothing about the threshold parameter beyond what the schema already says.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Score two counterparty names using Jaro-Winkler string similarity') and pins down what the score means ([0, 1]). It explicitly names the siblings it is not (reconcile, explain_match), so an agent can distinguish it without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides explicit when-to-use cases (verifying same party, evaluating matching thresholds) and explicit when-NOT-to-use cases naming the alternative tools for multi-signal matching. Nothing about tool selection is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
normalize_camt053ARead-onlyIdempotent
Convert parsed camt.053 statement entries into canonical observed records.
Purpose: Adapts parsed ISO 20022 camt.053 bank statement documents or entry lists into canonical observed records ready for the reconcile engine.
When to use:
When ingesting parsed bank statements before calling reconcile.
When transforming bank booking entries into uniform id, amount, date, and counterparty fields.
When NOT to use:
Do NOT use for payment initiation orders (pain.001); use normalize_pain001 instead.
Do NOT pass raw unparsed XML text; parse or deserialize first.
Behavioral transparency: Pure, deterministic, read-only data normalization.
| Name | Required | Description | Default |
|---|---|---|---|
| document | Yes | Parsed camt.053 document or entry list. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint and destructiveHint=false, so the safety bar is low and 'pure, deterministic, read-only' partly restates them. However, the description adds a genuine input constraint beyond the schema/annotations (must be parsed first, not raw XML) and clarifies the deterministic transformation intent, which is useful behavioral context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded verb+resource sentence followed by clearly labeled Purpose / When / When-NOT / Behavioral sections, which is easy to scan. The 'Purpose' paragraph mildly restates the opening line, costing a little redundancy, but overall it is tight and well-organized.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained. The description covers source format, target representation, the pre-parse prerequisite, exclusions, and deterministic behavior — everything an agent needs to invoke it correctly given the annotations and output schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There is a single required parameter with 100% schema coverage, so the schema already documents it. The description restates that it accepts a parsed document or entry list but adds no format, shape, or validation detail beyond the schema. Baseline 3 is appropriate when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (convert/adapt) and resource (parsed camt.053 statement entries into canonical observed records), and names the target consumer (reconcile engine). It explicitly distinguishes itself from the sibling normalize_pain001, so an agent can route correctly without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides explicit when-to-use bullets (before reconcile, transforming booking entries) and when-NOT-to-use bullets (not for pain.001 — use normalize_pain001; not for raw XML). Both the trigger condition and the alternative sibling are named.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
normalize_pain001ARead-onlyIdempotent
Convert parsed pain.001 payment instructions into canonical expected records.
Purpose: Adapts parsed ISO 20022 pain.001 credit transfer documents or transaction lists into canonical expected records ready for the reconcile engine.
When to use:
When ingesting parsed pain.001 files before calling reconcile.
When mapping diverse payment field structures into uniform id, amount, and reference fields.
When NOT to use:
Do NOT use for bank statement entries (camt.053); use normalize_camt053 instead.
Do NOT pass raw unparsed XML text; parse or deserialize to a dictionary or list first.
Behavioral transparency: Pure, deterministic, read-only data normalization.
| Name | Required | Description | Default |
|---|---|---|---|
| document | Yes | Parsed pain.001 document or transaction list. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false and openWorldHint=false, so safety is covered. The description reinforces this with 'pure, deterministic, read-only data normalization' and adds the input-hygiene constraint about not passing raw XML, but it does not say what happens on malformed/empty documents or how errors surface.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Clearly front-loaded and parsed into labeled sections (Purpose / When to use / When NOT to use / Behavioral transparency) that make it skimmable. Slight redundancy between the opening sentence and the Purpose paragraph, but no filler and nothing critical buried.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter normalizer, the description covers purpose, routing to alternatives, and input preconditions, while the output schema carries return-value details and annotations carry the safety profile. Nothing needed to invoke it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Only one parameter, and schema description coverage is 100%, so the baseline is 3. The description adds domain framing (parsed pain.001 document OR transaction list, and that fields are mapped to id/amount/reference) but no format or structure detail beyond what the schema already conveys.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource (convert parsed pain.001 instructions into canonical expected records) and names the domain (ISO 20022 credit transfers) explicitly. It also distinguishes itself from the sibling normalize_camt053 by naming it as the wrong tool for bank statements, so an agent can separate the two without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides explicit when-to-use conditions (ingesting parsed pain.001 before reconcile, mapping heterogeneous fields to uniform id/amount/reference) and when-NOT-to-use conditions with the named alternative (camt.053 → normalize_camt053). It further states a concrete prerequisite: parse/deserialize raw XML before passing it in.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
reconcileARead-onlyIdempotent
Reconcile expected payments against observed bank-statement entries.
Purpose: Matches expected payment instructions (e.g. from pain.001) against observed bank transactions (e.g. from camt.053) and produces an explainable reconciliation report covering exact matches, short/over adjustments, split settlements (one-to-many), batch credits (many-to-one), and unmatched residuals.
When to use:
When performing multi-transaction cash reconciliation between ledgers and statements.
When transparent per-match confidence scores and reason breakdowns are required for audit.
When NOT to use:
Do NOT use for raw, unparsed XML files; normalize input documents first using normalize_pain001 or normalize_camt053.
Do NOT use for single-pair tuning analysis; use explain_match instead.
Do NOT use for unconstrained combinatorial invoice subsets; use reconcile_many_to_many.
Behavioral transparency: Pure, deterministic, side-effect-free, read-only calculation without network or disk access.
| Name | Required | Description | Default |
|---|---|---|---|
| options | No | Optional tuning object: 'abs_tol'/'rel_tol' (amount tolerance), 'date_window_days', 'high_threshold', 'review_threshold', 'currency_strict', 'enable_one_to_many', 'max_combination'. | |
| expected | Yes | List of canonical records. Each is an object with 'id' (string) and 'amount' (number) required, plus optional 'currency' (ISO 4217), 'date' (ISO-8601), 'counterparty' (name), 'reference' (remittance/end-to-end id). | |
| observed | Yes | List of canonical records. Each is an object with 'id' (string) and 'amount' (number) required, plus optional 'currency' (ISO 4217), 'date' (ISO-8601), 'counterparty' (name), 'reference' (remittance/end-to-end id). |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, openWorldHint=false, and destructiveHint=false; the description reiterates these with useful framing ('pure, deterministic, side-effect-free, read-only calculation without network or disk access'), which adds confidence about determinism and isolation but doesn't add much beyond structured data.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action, then organized into Purpose / When to use / When NOT to use / Behavioral sections. Each sentence is purposeful and the structure aids scanning.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given a complex reconciliation tool with an output schema and rich annotations, the description covers purpose, usage, exclusions, and behavioral traits comprehensively. An agent has all it needs to select and call the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents the parameters thoroughly. The description references the input formats (pain.001, camt.053) which contextualizes the 'expected'/'observed' semantics, but it doesn't compensate for the high schema coverage by adding further detail.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a precise verb+resource statement ('Reconcile expected payments against observed bank-statement entries') and enumerates the specific matching scenarios handled (exact, short/over, split, batch, residuals), which clearly distinguishes it from sibling tools like reconcile_many_to_many and explain_match.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Offers explicit 'When to use' and 'When NOT to use' sections that name specific alternatives (normalize_pain001, normalize_camt053, explain_match, reconcile_many_to_many) and the conditions that select them, leaving no ambiguity about routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
reconcile_many_to_manyReconcile many-to-manyARead-onlyIdempotent
Match statement deposits to disjoint subsets of invoices via subset-sum ILP.
Purpose: Solves bounded subset-sum as an integer linear program (ILP) to pair statement deposits against matching combinations of outstanding invoice records.
When to use:
When customer deposits consolidate multiple invoices or split remittances.
When one-to-one or one-to-many heuristics leave unmatched aggregates.
When NOT to use:
Do NOT use for straightforward 1:1 or 1:N reconciliation where reconcile is sufficient and faster.
Do NOT use without the optional ilp extra installed (pip install reconcile-mcp[ilp]).
Behavioral transparency: Deterministic ILP optimization; read-only and idempotent.
| Name | Required | Description | Default |
|---|---|---|---|
| invoices | Yes | List of canonical records. Each is an object with 'id' (string) and 'amount' (number) required, plus optional 'currency' (ISO 4217), 'date' (ISO-8601), 'counterparty' (name), 'reference' (remittance/end-to-end id). | |
| statements | Yes | List of canonical records. Each is an object with 'id' (string) and 'amount' (number) required, plus optional 'currency' (ISO 4217), 'date' (ISO-8601), 'counterparty' (name), 'reference' (remittance/end-to-end id). |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover readOnly/idempotent/non-destructive, but the description adds genuine context beyond them: deterministic ILP optimization, bounded subset-sum formulation, and the hard dependency on the optional 'ilp' extra (pip install reconcile-mcp[ilp]). It does not discuss runtime cost or result ambiguity, but the dependency disclosure is valuable.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action, then cleanly sectioned into Purpose / When to use / When NOT to use / Behavioral transparency. Every line carries information, though the labeled headings are slightly more verbose than an inline two-sentence form would be.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Output schema exists so return values need not be explained. With usage conditions, exclusions, dependency requirement, and behavioral profile all covered for a 2-param tool, nothing an agent needs to invoke it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and both parameters are richly documented in the schema itself (id, amount, currency, date, counterparty, reference). The description adds no syntax or format detail beyond what the schema already supplies, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a precise verb+resource (match statement deposits to disjoint subsets of invoices) plus the mechanism (subset-sum ILP). It explicitly contrasts with the sibling 'reconcile', so an agent can distinguish the two without reading schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides explicit 'when to use' cases (consolidated deposits, split remittances, unmatched aggregates) and explicit 'when NOT to use' cases naming the alternative tool 'reconcile' and the missing-extra prerequisite. Routing is unambiguous.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
run_sandbox_scenarioARead-onlyIdempotent
Load a named sandbox scenario and immediately return its reconciliation report.
Purpose: One-call execution that loads built-in fixture records and executes the reconciliation pipeline, returning a complete explainable match report.
When to use:
When performing quick verification, smoke tests, or initial agent demonstrations.
When verifying tolerance tuning against standard test cases.
When NOT to use:
Do NOT use with live production data.
Do NOT use when custom expected or observed records must be supplied; use reconcile instead.
Behavioral transparency: Deterministic, in-memory execution; read-only and idempotent.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | Scenario name, e.g. 'month_end'. | |
| options | No | Optional tuning object: 'abs_tol'/'rel_tol' (amount tolerance), 'date_window_days', 'high_threshold', 'review_threshold', 'currency_strict', 'enable_one_to_many', 'max_combination'. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Adds useful context beyond annotations: deterministic in-memory execution and built-in fixture records. However, 'read-only and idempotent' restates readOnlyHint/idempotentHint already present in the annotations, so the incremental value is partly redundant.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded purpose with clear section headers, and each bullet earns its place. The labeled 'Behavioral transparency' section partly duplicates the annotations, adding mild redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return-value detail is unnecessary. For a deterministic fixture-runner with full annotation and schema coverage, nothing an agent needs to invoke it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so both 'name' and the tuning 'options' object are fully documented in the schema. The description adds no syntax or format detail beyond that, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (load) plus resource (named sandbox scenario) and the immediate effect (returns reconciliation report). The 'Purpose' line distinguishes it clearly from reconcile and load_sandbox_scenario.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit when-to-use bullets (smoke tests, tolerance tuning) and when-not-to-use bullets (no live data, no custom records), naming the alternative tool 'reconcile' for the custom-records case.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
3 tool updates
v0.0.5- Added
match_amounts_with_fx_drift - Added
match_names_probabilistic - Added
reconcile_many_to_many
7 tool updates
v0.0.2- First observed
explain_match - First observed
list_sandbox_scenarios - First observed
load_sandbox_scenario - First observed
normalize_camt053 - First observed
normalize_pain001 - First observed
reconcile - First observed
run_sandbox_scenario
TDQS
Scored across 10 tools
Each tool targets a distinct operation and resource, with explicit 'When NOT to use' guidance that cleanly separates reconcile, reconcile_many_to_many, explain_match, and the specialized matchers. The sandbox tools are also clearly divided into list, load, and run.
All names use snake_case and follow a consistent verb-first pattern (normalize_*, reconcile*, explain_match, match_*, list/load/run_*). Domain-specific suffixes like camt053 and pain001 are predictable and do not introduce mixed conventions.
Ten tools is well-scoped for a reconciliation engine, covering normalization, core matching, explanation, specialized matchers, and sandbox fixtures. No tool appears to be filler, and the set is neither too thin nor too heavy.
The core reconciliation lifecycle is covered: normalize both message types, reconcile, explain, many-to-many matching, specialized matchers, and sandbox fixtures. Minor gaps exist for full batch FX-drift reconciliation and other ISO 20022 message types, but agents can work around them.
Maintenance
Related MCP Connectors
- FinStatOAuthai.finstat
Documents in, reconciled double-entry books out. Statements, invoices, receipts, matched and posted.
Cross-border payment & banking intelligence for AI agents: SWIFT/BIC, IBAN, sanctions, FX, tracking.
Deterministic banking, LEI, VAT, SWIFT & compliance checks via MCP with signed XDR-1 receipts.
Turn bank statement PDFs into categorized, balance-checked transactions and reports.
Related MCP Servers
- AlicenseBqualityDmaintenanceAI workbench for financial contract analysis, risk analytics (VaR/CVaR, RWA Basel III), regulatory compliance (EMIR, REMIT, MiFID II, CBAM, EUDR) and counterparty due diligence (KYB/UBO, OFAC, IMO). Zero Retention. 8 MCP tools.831 npmMIT
- FlicenseNot gradedqualityDmaintenanceSelf-serve MCPB demo for accounts payable invoice exception review. It performs deterministic matching across invoice, purchase order, goods receipt, vendor master, invoice history, tax code master, and payment rules.-
- AlicenseAqualityAmaintenanceUnified gateway for ISO 20022 message families, providing meta-tools to search, describe, validate, generate, and parse financial messages.7356 PyPI1Apache 2.0
- FlicenseBqualityBmaintenanceEnables AI assistants to perform financial reconciliation with a deterministic proof engine: intake files, match transactions, verify proofs, resolve exceptions, and sign off on balanced journals under the user's authority.21-