| validate_rowsA | Validate rows against a column contract and quarantine what fails. Use this when you have tabular data (spreadsheet rows, CSV records, query
results) and you need to know which records are trustworthy before using
them. It applies each column's rules — required, type, range, pattern,
allowed values — and returns two lists: the rows that passed, and an explicit
quarantine list giving a reason for every row that did not. Nothing is ever silently dropped. accepted plus quarantine always equals
the number of rows you passed in, and each quarantined entry carries a
locator (row index, row number, and the column that failed) plus a
reasons list and machine-readable findings. Do NOT use this to repair data. It reports; it does not fix, coerce silently,
or delete. Coercion only happens where it is unambiguous ("100.50" becomes
the number 100.5), and the cleaned value is returned so you can see it. Args:
rows: The data, as an array of objects. Each object is one row, mapping
column name to value. Example:
[{"invoice": "INV-1", "amount": "100.50", "date": "2026-04-01"}].
columns: The contract, as an array of column specs. Example:
[{"name": "invoice", "type": "string", "required": true, "pattern": "INV-\\d+"}, {"name": "amount", "type": "number", "min": 0, "max": 1000000}, {"name": "date", "type": "date", "required": true}].
A required column applies to every row. Unknown spec keys are
rejected rather than ignored, so a typo cannot become a rule that
quietly never runs. Every column in columns must be listed once.
row_key: Optional name of a column to quote in each locator, so that
quarantined rows can be matched back to a record by a human. It does
not affect validation. Returns:
An object with:
ok (true only if no row was quarantined),
verdict ("clean" | "flagged" | "quarantined"),
summary (rows_in, accepted, quarantined, flagged, pass_rate,
accounted_for),
accepted (the passing rows, with unambiguous coercions applied),
quarantine (each entry: locator, row, reasons, findings),
flagged (rows kept despite a problem on a flag-mode column),
errors_by_column, findings_by_code, columns_seen,
columns_required_but_absent, contract (the parsed contract, so you can
confirm which rules ran), findings, and guidance. `columns_required_but_absent` is worth checking before anything else: a
required column missing from every row usually means the wrong sheet or
the wrong header row was read, which is a different problem from bad data.
Raises:
ToolError: if columns is malformed — an unknown type, a bad regular
expression, a duplicate column name, an unparseable spec. There is
nothing to validate until the contract parses, so the call is
rejected with a message naming the offending column and listing the
keys and types that are accepted. Row content never raises: a row
that is not an object is quarantined with the reason. The contract is the call's definition, so a broken contract is an
argument error. The rows are the data, so broken rows are a result.
That is the line this server draws.
|
| reconcileA | Check that opening + movements = closing, and show both sides and the variance. Use this whenever a period of numbers is supposed to add up: a bank
reconciliation, a roll-forward, a statement extract, a ledger movement
summary. It is the check that does not depend on how plausible any single
figure looked. If the arithmetic does not close, the numbers are wrong even
if every field looked fine on its own — that is the failure that survives
field-level validation and reaches a report. This tool is deterministic arithmetic. It has no opinion about whether the
movements are the right movements; it only states whether the given
opening, movements and closing are consistent with each other. Args:
opening_balance: The balance at the start of the period, as a number or a
numeric string. Example: 1000.00.
movements: The changes during the period, as an array. Each element is
either a signed number (250.0 increases, -40.5 decreases) or an
object {"amount": 250.0, "direction": "in", "label": "INV-1001", "date": "2026-04-03"}. Use "direction": "out" for a decrease, or
pass the amount as a negative number — never both. Passing an outflow
as a positive amount is the most common cause of a variance, so state
the direction explicitly when you can.
closing_balance: The balance at the end of the period, as a number or a
numeric string. Example: 1210.00.
tolerance: The largest variance to accept as closing, in the same units
as the balances. Default 0.01, which suits two-decimal currency.
Set it to 0 for an exact check. Returns:
An object with:
ok (true only if the arithmetic closes),
verdict ("clean" | "does_not_close" | "inputs_rejected"),
opening_balance, movements_net, expected_closing_balance,
reported_closing_balance, variance, absolute_variance, tolerance,
closes (true/false, or null when the inputs were incomplete),
arithmetic_is_reliable, movement_counts, movement_totals,
movements (each parsed movement with its signed amount),
findings, and guidance. `findings` always shows both sides of the equation in words, so the
variance can be read without recomputing it. When the variance exactly
equals the size of one movement, an extra finding says so — that is
arithmetic, not a guess, and it is usually the answer.
Raises:
Nothing for bad data. A movement that cannot be read is reported in
findings and excluded, and closes becomes null with
arithmetic_is_reliable: false, because a total over an incomplete set of
movements must never be reported as a pass. Genuinely malformed arguments
(movements that are not an array at all, or a non-numeric opening or
closing balance) return the same envelope with verdict set to
"inputs_rejected". |
| count_distinctA | Count distinct entities on a stable key, and report the gap to a row count. Use this instead of counting rows whenever the unit you are reporting on is
an entity — customers, patients, households, sites, invoices — rather than a
line in a table. One customer can appear on fifty rows; "50" is a row count,
not a customer count, and the difference is the error that survives review
because the number looks plausible. This tool returns both numbers and the difference between them, broken into
the two causes: keys repeated across rows, and rows that carry no usable key.
Rows with a blank key are excluded from the count and listed individually —
never counted as one entity, never dropped. Args:
rows: The data, as an array of objects. Example:
[{"customer_id": "C-1", "site": "A"}, {"customer_id": "C-1", "site": "B"}].
key: The column holding the stable identifier. Choose a real identifier
(an account number, a registration key), not a name — names collide,
and two different people sharing a name are not one entity.
case_sensitive: If false (the default), keys are compared after trimming
surrounding whitespace and folding case, so "C-1" and " c-1 "
count as one entity. Set true to treat any difference as a different
entity. Nothing else is normalised: no punctuation is stripped and no
fuzzy matching is applied, because those change which records are
considered the same entity, and that is a decision for a person. Returns:
An object with:
ok (true when the row count and the distinct count agree),
verdict ("clean" | "review_required"),
row_count, distinct_count, difference,
difference_breakdown ({from_repeated_keys, from_rows_with_blank_key}),
rows_with_a_key, rows_with_blank_key, blank_key_rows,
repeated_keys (each with its key and the row indexes it appears on),
findings, and guidance. Report `distinct_count` as the entity count. Report `row_count` as the
row count. Never present one as the other.
Raises:
ToolError: if key is not a non-empty string. Rows that are not
objects, or whose key is missing or blank, never raise — they are
held back, counted in rows_with_blank_key, and listed in
blank_key_rows. |
| duplicate_reportA | Find suspected duplicate groups and return them for human review. Use this when the same real entity may be recorded more than once — a
re-registration, a rename, a typo in a reference field — and you need to know
how much of the data that affects before trusting any count. This tool NEVER merges, deletes, or picks a winner. Every group it returns is
a suspicion with auto_merged: false and decision_required: true, plus a
review_question phrased for a person. A silent merge corrupts counts
invisibly and is very hard to unpick later, so the output is a review queue,
not a result. Three kinds of suspicion are reported, each labelled in kind: same_key: the same normalised key appears on more than one row. Either
genuine duplicates or legitimate repeated transactions — the tool cannot
tell which, so it reports rather than decides.
same_signature: different keys whose values across match_fields are
identical. This is the re-registration case.
near_match: different keys whose values across match_fields are similar
rather than identical. Only produced when you supply
similarity_threshold, because "close enough" is a policy choice.
Args:
rows: The data, as an array of objects.
key: The column holding the identifier. Needed to tell "one entity,
several rows" from "several entities".
match_fields: Columns whose combination suggests two different keys are
the same entity, for example ["name", "postcode"] or
["given_name", "family_name", "date_of_birth"]. Comparing on
several fields is far safer than one: a single name column collides
constantly. Omit it to check only exact key repeats.
similarity_threshold: A number above 0 and at most 1. When supplied,
match_fields values are also compared with string similarity
(difflib.SequenceMatcher) and groups scoring at or above this value
are reported as near_match. 0.9 is a reasonable starting point;
lower values find more and mean less. Omit it and only exact matches
are reported.
case_sensitive: If false (the default), values are compared after trimming
whitespace and folding case.
max_groups: Cap on the number of groups returned, default 50. When the
cap bites, groups_omitted says how many were left out — the count is
never silently reduced. Returns:
An object with:
ok (true only when no suspicion was found),
verdict ("clean" | "review_required"),
row_count, distinct_keys, duplicate_group_count,
rows_in_duplicate_groups, auto_merged (always false),
nothing_was_merged_or_deleted (always true),
blank_key_rows, groups, groups_omitted, findings, and guidance. Each group has `group_id`, `kind`, `reason`, `size`, `keys`,
`recommended_action` ("human_review"), `decision_required`, a
`review_question`, and `members` (each with a locator and the full row).
Raises:
ToolError: if key is not a non-empty string, if match_fields is not
an array of column names, or if similarity_threshold is not above
0 and at most 1. All three are malformed arguments rather than data
problems. Row content never raises. |