mcp-validation-server
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@mcp-validation-serverrun the column contract check on these rows and show me the quarantine list"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
mcp-validation-server
A Model Context Protocol server that gives a language model a set of data-validation tools it cannot argue with.
MCP in one sentence: the Model Context Protocol is an open standard that lets an AI application call tools running in a separate process, so the tool does the work and returns a result the model has to accept.
Why this exists
A model asked "is this data clean?" will answer. It will answer confidently whether or not it checked anything. The fix is not a better prompt; it is to take the question away from the model and give it to code that computes an answer and hands it back.
This server exposes four deterministic checks — contract validation, arithmetic reconciliation, distinct entity counting, and duplicate detection. Each one returns a structured result with reasons and row/column locators. None of them need a model, a confidence score, or reference data. All of them can be wrong in ways you can see, which is the only kind of check worth trusting.
Related MCP server: dq-mcp
The design rule
A validation failure is a result, not an exception.
Every tool returns a JSON object with the same envelope:
Field | Meaning |
|
|
|
|
| One entry per problem: a stable |
| Plain English: what to do next |
ok: false is a successful call reporting a failing dataset. The model reads findings, sees
exactly which row and column failed, and can say so instead of guessing.
Only a genuinely malformed tool argument raises — a contract that does not parse, a regex that does not compile, a key that is not a column name. There is nothing to validate until the call itself is valid, so the call is rejected with a message naming the problem.
The line this server draws:
The contract is the call's definition → a broken contract is an argument error.
The rows are the data → broken rows are a result, with a reason each.
The four tools
validate_rows
Applies a column contract — required, type, min, max, range, pattern, allowed,
not_blank, and per-column on_fail — and returns the passing rows and an explicit quarantine
list. Every quarantined row carries a locator (index, one-based row number, and the column that
failed), a list of human-readable reasons, and machine-readable findings.
accepted + quarantined always equals the number of rows you passed in. Nothing is ever silently
dropped. A row that is not even an object is quarantined with that as its reason.
reconcile
Given an opening balance, a list of movements, and a closing balance, reports whether
closing = opening + movements holds within a tolerance, showing both sides and the variance. This is
the check that does not depend on how plausible any single figure looked: clean extraction with quietly
wrong numbers is the failure that survives field-level validation, and arithmetic is what catches it.
When the variance exactly equals the size of one movement, the result says so — that is arithmetic, not
a guess, and it is usually the answer. When a movement cannot be read, closes becomes null rather
than false: a total computed over an incomplete set of movements must never be reported as a pass.
count_distinct
Counts distinct entities on a caller-supplied stable key and reports the difference between that and the row count, broken into its two causes: keys repeated across rows, and rows carrying no usable key. One customer can appear on fifty rows; "50" is a row count, not a customer count, and the difference is the error that survives review because the number looks plausible.
Rows with a blank key are excluded from the count and listed individually — never counted as one entity, never dropped. Nothing beyond whitespace and letter case is normalised: stripping punctuation or applying fuzzy matching changes which records count as the same entity, and that is a decision for a person.
duplicate_report
Finds suspected duplicate groups and returns them for human review. It never merges, never deletes,
and never picks a winner. Every group carries auto_merged: false, decision_required: true, the full
rows with locators, and a review_question phrased for a person to answer.
Three kinds of suspicion, each labelled:
same_key— the same normalised key on more than one row. Genuine duplicates, or legitimate repeated transactions? The tool cannot tell, so it reports rather than decides.same_signature— different keys whose values acrossmatch_fieldsare identical. The re-registration case.near_match— different keys with similar values. Only produced when you supplysimilarity_threshold, because "close enough" is a policy choice, not a fact.
Install
git clone https://github.com/devops-gm88/mcp-validation-server.git
cd mcp-validation-server
python3 -m venv .venv
.venv/bin/pip install -e .Python 3.10 or newer. mcp is the only runtime dependency; everything else is the standard library.
Register it with an MCP client
The server speaks MCP over stdio, so an MCP host launches it as a subprocess and talks to it on stdin/stdout. This is the standard configuration block used by MCP hosts that launch local servers (Claude Desktop, Claude Code, and others):
{
"mcpServers": {
"data-validation": {
"command": "/absolute/path/to/mcp-validation-server/.venv/bin/python",
"args": ["-m", "mcp_validation_server"]
}
}
}Use the absolute path to the interpreter — an MCP host does not inherit your shell's PATH or
virtualenv. If you installed with pipx or globally, the console script is equivalent:
{
"mcpServers": {
"data-validation": {
"command": "mcp-validation-server",
"args": []
}
}
}Verify the server starts and lists its tools without involving a client at all:
.venv/bin/python examples/stdio_handshake.pyOr with the official MCP Inspector CLI, which speaks to it as a real MCP host does. Point command at
the same absolute interpreter path as the block above:
cat > /tmp/mcp.json <<JSON
{
"mcpServers": {
"data-validation": {
"command": "$PWD/.venv/bin/python",
"args": ["-m", "mcp_validation_server"]
}
}
}
JSON
npx @modelcontextprotocol/inspector --cli --config /tmp/mcp.json --server data-validation --method tools/listTo call a tool through the Inspector rather than just list it:
npx @modelcontextprotocol/inspector --cli --config /tmp/mcp.json --server data-validation \
--method tools/call --tool-name count_distinct \
--tool-arg 'rows=[{"id":"C-1"},{"id":"C-1"},{"id":""}]' 'key=id'On the tool schemas
The Inspector reports 0 errors and 7 warnings for these tools, and the warnings are deliberate. Four
are on the items of a rows or columns array; three are on reconcile's opening_balance,
movements and closing_balance, which accept a number or a numeric string and report anything else as
a finding rather than refusing the call:
Schema carries no validation keyword at all, so it accepts any value.
That looseness is the design rule, applied at the protocol boundary. If rows were declared as
array of objects, a client sending a row that is a bare string would be stopped by the input schema and
the model would get an unrecoverable protocol error — when the useful answer is "row 3 of your data is a
string, not an object; here it is, held back with that reason". So row content is intentionally
unconstrained and every unusable row comes back as a quarantine finding with a locator. What is
constrained is the shape of the call: a missing key, a similarity_threshold outside 0–1, and a
column spec that will not parse are all rejected before anything runs.
The trade is stated plainly rather than papered over: it costs a schema warning, and it buys a model that can recover from bad data instead of just failing on it.
Worked example
Real output. The script that produced it is examples/worked_example.py, and the untrimmed result is
saved beside it as examples/worked_example_output.json.
A model calls validate_rows on four rows of deliberately broken data:
{
"rows": [
{"invoice": "INV-1001", "amount": "100.50", "date": "2026-04-01"},
{"invoice": "", "amount": "20", "date": "2026-04-02"},
{"invoice": "INV-1003", "amount": "-5", "date": "2026-04-03"},
{"invoice": "INV-1004", "amount": "abc", "date": "04/05/2026"}
],
"columns": [
{"name": "invoice", "type": "string", "required": true, "pattern": "INV-\\d+"},
{"name": "amount", "type": "number", "min": 0},
{"name": "date", "type": "date", "required": true}
],
"row_key": "invoice"
}and gets back (is_error: false — the call succeeded; the data failed):
{
"tool": "validate_rows",
"ok": false,
"verdict": "quarantined",
"finding_count": 4,
"findings": "<<4 findings, one per quarantine reason; see worked_example_output.json>>",
"guidance": "3 of 4 rows were quarantined and are listed in `quarantine` with a locator and reasons. Failure codes: below_minimum (1), not_a_date (1), not_a_number (1), required_blank (1). Do not use the accepted rows as if they were the whole dataset: report both counts, and resolve or explicitly exclude the quarantined rows before aggregating.",
"contract": "<<the parsed contract echoed back; see worked_example_output.json>>",
"summary": {
"rows_in": 4,
"accepted": 1,
"quarantined": 3,
"flagged": 0,
"pass_rate": 0.25,
"accounted_for": 4
},
"accepted": [
{
"invoice": "INV-1001",
"amount": 100.5,
"date": "2026-04-01"
}
],
"quarantine": [
{
"locator": {
"row_index": 1,
"row_number": 2
},
"row": {
"invoice": "",
"amount": 20.0,
"date": "2026-04-02"
},
"reasons": [
"invoice: is required but blank"
],
"findings": [
{
"code": "required_blank",
"severity": "error",
"message": "is required but blank",
"locator": {
"row_index": 1,
"row_number": 2,
"column": "invoice"
},
"expected": "a non-blank value"
}
]
},
{
"locator": {
"row_index": 2,
"row_number": 3,
"row_key": "INV-1003"
},
"row": {
"invoice": "INV-1003",
"amount": -5.0,
"date": "2026-04-03"
},
"reasons": [
"amount: -5 is below the minimum of 0"
],
"findings": [
{
"code": "below_minimum",
"severity": "error",
"message": "-5 is below the minimum of 0",
"locator": {
"row_index": 2,
"row_number": 3,
"row_key": "INV-1003",
"column": "amount"
},
"value": -5.0
}
]
},
{
"locator": {
"row_index": 3,
"row_number": 4,
"row_key": "INV-1004"
},
"row": {
"invoice": "INV-1004",
"amount": "abc",
"date": "04/05/2026"
},
"reasons": [
"amount: 'abc' is not a number",
"date: '04/05/2026' is not an ISO-8601 date (expected YYYY-MM-DD)"
],
"findings": [
{
"code": "not_a_number",
"severity": "error",
"message": "'abc' is not a number",
"locator": {
"row_index": 3,
"row_number": 4,
"row_key": "INV-1004",
"column": "amount"
},
"value": "abc",
"expected": "a value of type number"
},
{
"code": "not_a_date",
"severity": "error",
"message": "'04/05/2026' is not an ISO-8601 date (expected YYYY-MM-DD)",
"locator": {
"row_index": 3,
"row_number": 4,
"row_key": "INV-1004",
"column": "date"
},
"value": "04/05/2026",
"expected": "a value of type date"
}
]
}
],
"flagged": [],
"errors_by_column": {
"amount": 2,
"date": 1,
"invoice": 1
},
"findings_by_code": {
"below_minimum": 1,
"not_a_date": 1,
"not_a_number": 1,
"required_blank": 1
},
"columns_seen": [],
"columns_required_but_absent": []
}What a model can now do that it could not before: say "3 of your 4 rows are unusable, here is which
one and why", and be right. The one clean row came back with its string amount coerced to the number
100.5; the three broken rows came back with their original values intact and a reason each. Row 4
failed twice, on two different columns, and both reasons are listed.
Note the locator gives both row_index (0-based, for the machine) and row_number (1-based, for the
person reading the report), and quotes row_key when the row has one.
What this does not do
It does not repair data. It reports which records are untrustworthy and why. It never fills, imputes, or corrects a value.
It does not merge duplicates, ever.
duplicate_reportreturns suspicions with a question attached. Merging is a decision with consequences, and it belongs to a person.It does not score data quality. There is no 0–100 number and no dashboard. There are findings with codes and locators.
It does not prove the data is correct. Passing a contract means the rows satisfy the rules you supplied and nothing else. A contract that does not check a column says nothing about that column.
reconcileconfirms internal consistency, not completeness: movements that are missing entirely will still balance if the closing balance was derived from the same wrong set.It does not know your domain. Whether a repeated key is a defect depends on whether the table is a transaction log or an entity register. The tools report the number; you decide what it means.
It does not normalise beyond whitespace and case. No punctuation stripping, no accent folding, no nickname tables, no fuzzy matching except where you explicitly ask for a similarity threshold.
It has no persistence, network access, or file access. Every tool takes its data as an argument and returns a result. Nothing is written anywhere.
It does not cap the size of the data you send it. A large
rowsarray comes back in full, in the result.duplicate_reporthas amax_groupscap that reports what it omitted;validate_rowsandcount_distinctdo not truncate, so pass a summary-sized batch rather than a whole table.It does not check the caller. There is no authentication. It is a local stdio server for one user's machine.
Relationship to sheets-data-validation-kit
This project is a sibling, not a wrapper. It is self-contained: pip install -e . here installs
everything it needs, and nothing here imports the other package.
The ideas and the naming come from
sheets-data-validation-kit, a
standard-library-only Python library of the same primitives — Field, RowContract, the rule objects,
Min/Max/Range/Pattern, accept-or-quarantine, exact-match entity resolution, and never
auto-merging a duplicate. The kit is the library you import; this is the MCP server that puts those
checks behind a protocol so a model can call them.
Two differences are worth knowing:
|
| |
Interface | Python objects you construct in code | JSON arguments over MCP |
Contract shape |
| Column spec objects; parsed, and unknown keys rejected |
Locators | Reasons as strings | Findings with codes and row/column locators |
Failure model | Returns a result object | Returns a result, plus a |
The server's contract.py is a re-implementation for JSON-supplied contracts, not a copy: it adds type
coercion, locators, per-finding codes, and the parsing layer that turns a model's JSON into rules and
rejects a malformed contract before any data is touched.
Tests
pip install -e .
python3 -m unittest discover -s tests -v139 tests, all passing on Python 3.14.7 — the count and result are from an actual run, not an estimate. Every tool has non-MCP unit tests that call the underlying functions directly; no test needs a client, a transport, or a subprocess.
Without installing first:
PYTHONPATH=src python3 -m unittest discover -s tests -vThe tests exercise the design rule directly: that accepted + quarantined equals the input count, that
a failing dataset never raises, that an unreadable movement withholds the reconciliation verdict instead
of producing a false pass, that nothing is ever merged, and that the four tools publish one shared
envelope.
Layout
mcp-validation-server/
├── examples/
│ ├── stdio_handshake.py # start the server and list its tools, no client needed
│ ├── worked_example.py # produces the README's worked example
│ └── worked_example_output.json # the untrimmed result, as produced
├── src/mcp_validation_server/
│ ├── __init__.py # public API
│ ├── __main__.py # python3 -m mcp_validation_server
│ ├── app.py # the MCPServer instance and its instructions
│ ├── tools.py # the four @mcp.tool() functions and their docstrings
│ ├── findings.py # Finding, Locator, the shared response envelope
│ ├── contract.py # column specs, rules, RowContract, quarantine
│ ├── reconcile.py # opening + movements = closing
│ ├── entities.py # distinct counting, duplicate grouping
│ └── server.py # entry point and transport selection
└── tests/ # stdlib unittest, no MCP client requiredNotes on the SDK
Built against the official MCP Python SDK, mcp 2.2.0. Two things differ from most published examples,
which were written for mcp 1.x:
FastMCPwas renamed toMCPServer. Inmcp2.x the import isfrom mcp.server.mcpserver import MCPServer;mcp.server.fastmcpno longer exists and raises aModuleNotFoundErrorpointing at the migration guide.pyproject.tomltherefore requiresmcp>=2.0,<3.Only
ToolErrorcarries a message to the caller. Any other exception raised inside a tool is wrapped by the SDK asUnexpectedToolError, whose text is the genericError executing tool <name>, with the original logged server-side only. Since a contract error the model cannot read is a contract error the model cannot fix, malformed arguments are raised asmcp.server.mcpserver.exceptions.ToolError.Return annotations matter for structured output. Annotating a tool with
typing.Dict[str, Any]makes the SDK publish an output schema of{"result": {...}}and neststructuredContentone level down; the PEP 585dict[str, Any]annotation keeps the payload at the top level. These tools usedict[str, Any].
Licence
MIT. See LICENSE.
Available Tools
4 toolscount_distinctA
Count distinct entities on a stable key, and report the gap to a row count.
Use this instead of counting rows whenever the unit you are reporting on is an entity — customers, patients, households, sites, invoices — rather than a line in a table. One customer can appear on fifty rows; "50" is a row count, not a customer count, and the difference is the error that survives review because the number looks plausible.
This tool returns both numbers and the difference between them, broken into the two causes: keys repeated across rows, and rows that carry no usable key. Rows with a blank key are excluded from the count and listed individually — never counted as one entity, never dropped.
Args:
rows: The data, as an array of objects. Example:
[{"customer_id": "C-1", "site": "A"}, {"customer_id": "C-1", "site": "B"}].
key: The column holding the stable identifier. Choose a real identifier
(an account number, a registration key), not a name — names collide,
and two different people sharing a name are not one entity.
case_sensitive: If false (the default), keys are compared after trimming
surrounding whitespace and folding case, so "C-1" and " c-1 "
count as one entity. Set true to treat any difference as a different
entity. Nothing else is normalised: no punctuation is stripped and no
fuzzy matching is applied, because those change which records are
considered the same entity, and that is a decision for a person.
Returns:
An object with:
ok (true when the row count and the distinct count agree),
verdict ("clean" | "review_required"),
row_count, distinct_count, difference,
difference_breakdown ({from_repeated_keys, from_rows_with_blank_key}),
rows_with_a_key, rows_with_blank_key, blank_key_rows,
repeated_keys (each with its key and the row indexes it appears on),
findings, and guidance.
Report `distinct_count` as the entity count. Report `row_count` as the
row count. Never present one as the other.Raises:
ToolError: if key is not a non-empty string. Rows that are not
objects, or whose key is missing or blank, never raise — they are
held back, counted in rows_with_blank_key, and listed in
blank_key_rows.
| Name | Required | Description | Default |
|---|---|---|---|
| key | Yes | ||
| rows | Yes | ||
| case_sensitive | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does so well: it discloses that blank-key rows are excluded, counted, and listed (never dropped or merged), that normalization is limited to trim+case-fold with no fuzzy matching, and exactly which inputs raise ToolError versus which are silently held back. This is exactly the behavioral context an agent needs for a mutation of judgement calls.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the purpose and cleanly organized into Args/Returns/Raises. Some motivational prose ('the error that survives review because the number looks plausible') is rhetorical padding, but the structure is tight and every section is scannable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return-value explanation is not strictly required, yet the description still enumerates the key return fields and even the reporting obligation ('never present one as the other'). For a tool with real ambiguity around entity identity and error handling, nothing an agent needs is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, and it does: rows is illustrated with a concrete example, key is given selection guidance (real identifier, not a name, with a reason), and case_sensitive's default and exact normalization semantics are spelled out. All three parameters are fully covered.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Count distinct entities on a stable key') and immediately clarifies the crucial distinction from a row count. It sharply frames what the tool is versus generic row counting, but never names or differentiates itself from its actual siblings (duplicate_report, reconcile, validate_rows).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a clear when-to-use rule: use it instead of counting rows when the reported unit is an entity rather than a line in a table. The inverse condition is implied rather than stated, and no sibling alternative is named, so routing to duplicate_report vs. reconcile is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
duplicate_reportA
Find suspected duplicate groups and return them for human review.
Use this when the same real entity may be recorded more than once — a re-registration, a rename, a typo in a reference field — and you need to know how much of the data that affects before trusting any count.
This tool NEVER merges, deletes, or picks a winner. Every group it returns is
a suspicion with auto_merged: false and decision_required: true, plus a
review_question phrased for a person. A silent merge corrupts counts
invisibly and is very hard to unpick later, so the output is a review queue,
not a result.
Three kinds of suspicion are reported, each labelled in kind:
same_key: the same normalised key appears on more than one row. Either genuine duplicates or legitimate repeated transactions — the tool cannot tell which, so it reports rather than decides.same_signature: different keys whose values acrossmatch_fieldsare identical. This is the re-registration case.near_match: different keys whose values acrossmatch_fieldsare similar rather than identical. Only produced when you supplysimilarity_threshold, because "close enough" is a policy choice.
Args:
rows: The data, as an array of objects.
key: The column holding the identifier. Needed to tell "one entity,
several rows" from "several entities".
match_fields: Columns whose combination suggests two different keys are
the same entity, for example ["name", "postcode"] or
["given_name", "family_name", "date_of_birth"]. Comparing on
several fields is far safer than one: a single name column collides
constantly. Omit it to check only exact key repeats.
similarity_threshold: A number above 0 and at most 1. When supplied,
match_fields values are also compared with string similarity
(difflib.SequenceMatcher) and groups scoring at or above this value
are reported as near_match. 0.9 is a reasonable starting point;
lower values find more and mean less. Omit it and only exact matches
are reported.
case_sensitive: If false (the default), values are compared after trimming
whitespace and folding case.
max_groups: Cap on the number of groups returned, default 50. When the
cap bites, groups_omitted says how many were left out — the count is
never silently reduced.
Returns:
An object with:
ok (true only when no suspicion was found),
verdict ("clean" | "review_required"),
row_count, distinct_keys, duplicate_group_count,
rows_in_duplicate_groups, auto_merged (always false),
nothing_was_merged_or_deleted (always true),
blank_key_rows, groups, groups_omitted, findings, and guidance.
Each group has `group_id`, `kind`, `reason`, `size`, `keys`,
`recommended_action` ("human_review"), `decision_required`, a
`review_question`, and `members` (each with a locator and the full row).Raises:
ToolError: if key is not a non-empty string, if match_fields is not
an array of column names, or if similarity_threshold is not above
0 and at most 1. All three are malformed arguments rather than data
problems. Row content never raises.
| Name | Required | Description | Default |
|---|---|---|---|
| key | Yes | ||
| rows | Yes | ||
| max_groups | No | ||
| match_fields | No | ||
| case_sensitive | No | ||
| similarity_threshold | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the entire behavioral burden and does so richly: it never merges/deletes/picks a winner, groups are always `auto_merged: false` / `decision_required: true`, `near_match` only appears when a threshold is supplied, and `max_groups` truncation is surfaced via `groups_omitted` rather than silently applied. It also enumerates the exact argument-validation errors that raise ToolError versus data problems that never raise.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Well front-loaded: the first sentence states the action and the second gives the trigger, before any detail. The Args block is justified given 0% schema coverage, but the Returns block is somewhat redundant with the existing output schema and repeats facts already stated in the prose (e.g., `auto_merged: false`), adding length without new decision-relevant information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Although an output schema exists, the description goes further by explaining the meaning of the returned verdict and review queue rather than just its shape, and it covers the failure modes, defaults, and policy-driven behaviors an agent needs. Nothing required for correct invocation is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate fully — and it does, covering all six parameters with meaning beyond types: `key` distinguishes one-entity-many-rows from many-entities, `match_fields` is explained with concrete examples and a warning about single-column collisions, `similarity_threshold` gets its domain ('above 0 and at most 1'), algorithm, and a suggested starting value, and `case_sensitive` documents the trimming/case-folding default behavior.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('find suspected duplicate groups and return them for human review') and immediately scopes the tool's role against the obvious alternative behavior of merging. The three `kind` categories further pin down exactly what it detects, so an agent can tell it apart from siblings like reconcile or count_distinct.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit when-to-use condition ('when the same real entity may be recorded more than once — a re-registration, a rename, a typo') and the motivating context ('before trusting any count'). It does not name the sibling tools (validate_rows, reconcile, count_distinct) as alternatives or state when NOT to use it, so it stops short of full routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
reconcileA
Check that opening + movements = closing, and show both sides and the variance.
Use this whenever a period of numbers is supposed to add up: a bank reconciliation, a roll-forward, a statement extract, a ledger movement summary. It is the check that does not depend on how plausible any single figure looked. If the arithmetic does not close, the numbers are wrong even if every field looked fine on its own — that is the failure that survives field-level validation and reaches a report.
This tool is deterministic arithmetic. It has no opinion about whether the movements are the right movements; it only states whether the given opening, movements and closing are consistent with each other.
Args:
opening_balance: The balance at the start of the period, as a number or a
numeric string. Example: 1000.00.
movements: The changes during the period, as an array. Each element is
either a signed number (250.0 increases, -40.5 decreases) or an
object {"amount": 250.0, "direction": "in", "label": "INV-1001", "date": "2026-04-03"}. Use "direction": "out" for a decrease, or
pass the amount as a negative number — never both. Passing an outflow
as a positive amount is the most common cause of a variance, so state
the direction explicitly when you can.
closing_balance: The balance at the end of the period, as a number or a
numeric string. Example: 1210.00.
tolerance: The largest variance to accept as closing, in the same units
as the balances. Default 0.01, which suits two-decimal currency.
Set it to 0 for an exact check.
Returns:
An object with:
ok (true only if the arithmetic closes),
verdict ("clean" | "does_not_close" | "inputs_rejected"),
opening_balance, movements_net, expected_closing_balance,
reported_closing_balance, variance, absolute_variance, tolerance,
closes (true/false, or null when the inputs were incomplete),
arithmetic_is_reliable, movement_counts, movement_totals,
movements (each parsed movement with its signed amount),
findings, and guidance.
`findings` always shows both sides of the equation in words, so the
variance can be read without recomputing it. When the variance exactly
equals the size of one movement, an extra finding says so — that is
arithmetic, not a guess, and it is usually the answer.Raises:
Nothing for bad data. A movement that cannot be read is reported in
findings and excluded, and closes becomes null with
arithmetic_is_reliable: false, because a total over an incomplete set of
movements must never be reported as a pass. Genuinely malformed arguments
(movements that are not an array at all, or a non-numeric opening or
closing balance) return the same envelope with verdict set to
"inputs_rejected".
| Name | Required | Description | Default |
|---|---|---|---|
| movements | Yes | ||
| tolerance | No | ||
| closing_balance | Yes | ||
| opening_balance | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does so: deterministic arithmetic with no side effects, unreadable movements are excluded and reported in findings, closes becomes null with arithmetic_is_reliable false, and malformed arguments yield verdict 'inputs_rejected'. Unusually thorough error-behavior disclosure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The core equation is front-loaded and the Args/Returns/Raises structure is scannable, but several sentences are motivational rhetoric ('the failure that survives field-level validation and reaches a report') that an agent does not strictly need. Long yet mostly earning its space.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Output schema exists, yet the description still usefully flags the notable return fields and the findings behavior, including the special case where variance equals one movement. Edge cases and failure modes are covered, so nothing needed to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate and it does — opening/closing accept numbers or numeric strings, movements accept signed numbers or objects with amount/direction/label/date, the never-both direction rule is stated, and tolerance is explained with its default and 'set 0 for exact' semantics.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a precise verb and resource — verify opening + movements = closing — and immediately frames the scope ('a period of numbers'). It also carves out its distinction from sibling validate_rows by contrasting field-level validation with whole-period arithmetic.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit triggering conditions ('whenever a period of numbers is supposed to add up') with concrete examples (bank reconciliation, roll-forward, statement extract, ledger movement summary), and states the exclusion plainly: it has no opinion on whether the movements are the right ones.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
validate_rowsA
Validate rows against a column contract and quarantine what fails.
Use this when you have tabular data (spreadsheet rows, CSV records, query results) and you need to know which records are trustworthy before using them. It applies each column's rules — required, type, range, pattern, allowed values — and returns two lists: the rows that passed, and an explicit quarantine list giving a reason for every row that did not.
Nothing is ever silently dropped. accepted plus quarantine always equals
the number of rows you passed in, and each quarantined entry carries a
locator (row index, row number, and the column that failed) plus a
reasons list and machine-readable findings.
Do NOT use this to repair data. It reports; it does not fix, coerce silently,
or delete. Coercion only happens where it is unambiguous ("100.50" becomes
the number 100.5), and the cleaned value is returned so you can see it.
Args:
rows: The data, as an array of objects. Each object is one row, mapping
column name to value. Example:
[{"invoice": "INV-1", "amount": "100.50", "date": "2026-04-01"}].
columns: The contract, as an array of column specs. Example:
[{"name": "invoice", "type": "string", "required": true, "pattern": "INV-\\d+"}, {"name": "amount", "type": "number", "min": 0, "max": 1000000}, {"name": "date", "type": "date", "required": true}].
A required column applies to every row. Unknown spec keys are
rejected rather than ignored, so a typo cannot become a rule that
quietly never runs. Every column in columns must be listed once.
row_key: Optional name of a column to quote in each locator, so that
quarantined rows can be matched back to a record by a human. It does
not affect validation.
Returns:
An object with:
ok (true only if no row was quarantined),
verdict ("clean" | "flagged" | "quarantined"),
summary (rows_in, accepted, quarantined, flagged, pass_rate,
accounted_for),
accepted (the passing rows, with unambiguous coercions applied),
quarantine (each entry: locator, row, reasons, findings),
flagged (rows kept despite a problem on a flag-mode column),
errors_by_column, findings_by_code, columns_seen,
columns_required_but_absent, contract (the parsed contract, so you can
confirm which rules ran), findings, and guidance.
`columns_required_but_absent` is worth checking before anything else: a
required column missing from every row usually means the wrong sheet or
the wrong header row was read, which is a different problem from bad data.Raises:
ToolError: if columns is malformed — an unknown type, a bad regular
expression, a duplicate column name, an unparseable spec. There is
nothing to validate until the contract parses, so the call is
rejected with a message naming the offending column and listing the
keys and types that are accepted. Row content never raises: a row
that is not an object is quarantined with the reason.
The contract is the call's definition, so a broken contract is an
argument error. The rows are the data, so broken rows are a result.
That is the line this server draws.
| Name | Required | Description | Default |
|---|---|---|---|
| rows | Yes | ||
| columns | Yes | ||
| row_key | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does so thoroughly: it promises nothing is silently dropped, states the invariant accepted + quarantine = rows_in, documents that each quarantine entry carries locator/reasons/findings, bounds coercion to unambiguous cases with the cleaned value surfaced, and draws a precise line between contract errors (ToolError) and bad rows (quarantined). This is exactly the behavioral context an agent needs before calling a mutation-free validator.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
It is long, but the length is justified by zero schema description coverage and the Args/Returns/Raises structure is front-loaded with the core purpose first. A few lines (the 'contract is the call's definition' paragraph) restate an idea already conveyed, but nothing is wasted enough to hurt usability.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 3-param validator with a malformed-input failure mode, the description covers purpose, usage boundaries, error semantics, and the full Args contract. Return-value detail partly overlaps the existing output schema but adds handling advice ('check columns_required_but_absent before anything else') that the schema alone cannot convey.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, and it does: each of the three params gets a concrete example (row object, column spec with type/min/max/pattern/required), plus non-obvious rules such as unknown spec keys being rejected and every column being listed exactly once. row_key's no-op-on-validation semantics are also stated.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Validate rows against a column contract and quarantine what fails') and names the exact artifact returned (two lists: accepted and quarantine). An agent can distinguish this from reconcile, count_distinct, and duplicate_report purely from the first sentence.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a clear when-to-use trigger (tabular data / need to know which records are trustworthy) and an explicit when-not ('Do NOT use this to repair data. It reports; it does not fix, coerce silently, or delete.'). It does not name a specific sibling tool to use instead, so it falls just short of the top band.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
4 tool updates
v0.1.0- First observed
count_distinct - First observed
duplicate_report - First observed
reconcile - First observed
validate_rows
TDQS
Scored across 4 tools
Each tool targets a distinct data-quality operation: contract-based row validation, arithmetic reconciliation, entity counting, and duplicate detection. The only overlap is that count_distinct and duplicate_report both consume rows+key, but their stated purposes (counting vs. reviewing suspects) are clearly separated by the descriptions.
The set mixes conventions: validate_rows and count_distinct follow a verb_noun pattern, while reconcile is a bare verb and duplicate_report is noun_noun. All are readable and snake_case, but there is no single predictable rule.
Four tools is well-scoped for a tabular data-validation server; each covers a distinct, classic quality check (row validity, closure, distinctness, duplication) and none feels redundant or padded.
The surface covers the core data-quality checks and is deliberately read-only, with clear reporting rather than repair. A gap remains for cross-column/cross-table consistency (e.g. referential integrity between datasets) and schema inference, but core workflows are fully addressed.
Maintenance
Related MCP Connectors
Deterministic validation for AI-generated artifacts: JSON Schema, OpenAPI response, SQL syntax.
Messy spreadsheets in, clean checkable tables out. Every result carries its arithmetic proof.
Deterministic signed verification of numeric & financial claims for AI agents & spreadsheets.
Auto-discover validation rules from data — scan, profile, health-score. No rules to write.
Related MCP Servers
- AlicenseNot gradedqualityAmaintenanceDeterministic verification for AI-generated analysis. Reconciliation, consistency and Excel-integrity checks that stop the line when the numbers don't add up.45 PyPI1MIT
- AlicenseNot gradedqualityCmaintenanceEnables language models to run data-quality checks and profiling on local files, using dbt-style assertions like not_null, unique, relationships, and accepted_values.MIT
- FlicenseBqualityBmaintenanceEnables AI assistants to perform financial reconciliation with a deterministic proof engine: intake files, match transactions, verify proofs, resolve exceptions, and sign off on balanced journals under the user's authority.21-
- FlicenseAqualityBmaintenanceEnables LLMs to audit data drift and model degradation in tabular ML pipelines through deterministic statistical tests such as Kolmogorov-Smirnov and Population Stability Index, plus reusable prompts and standards resources.4-