validate_rows
Validate tabular rows against a column contract and return accepted rows plus a quarantine list with per-row failure reasons.
Instructions
Validate rows against a column contract and quarantine what fails.
Use this when you have tabular data (spreadsheet rows, CSV records, query results) and you need to know which records are trustworthy before using them. It applies each column's rules — required, type, range, pattern, allowed values — and returns two lists: the rows that passed, and an explicit quarantine list giving a reason for every row that did not.
Nothing is ever silently dropped. accepted plus quarantine always equals
the number of rows you passed in, and each quarantined entry carries a
locator (row index, row number, and the column that failed) plus a
reasons list and machine-readable findings.
Do NOT use this to repair data. It reports; it does not fix, coerce silently,
or delete. Coercion only happens where it is unambiguous ("100.50" becomes
the number 100.5), and the cleaned value is returned so you can see it.
Args:
rows: The data, as an array of objects. Each object is one row, mapping
column name to value. Example:
[{"invoice": "INV-1", "amount": "100.50", "date": "2026-04-01"}].
columns: The contract, as an array of column specs. Example:
[{"name": "invoice", "type": "string", "required": true, "pattern": "INV-\\d+"}, {"name": "amount", "type": "number", "min": 0, "max": 1000000}, {"name": "date", "type": "date", "required": true}].
A required column applies to every row. Unknown spec keys are
rejected rather than ignored, so a typo cannot become a rule that
quietly never runs. Every column in columns must be listed once.
row_key: Optional name of a column to quote in each locator, so that
quarantined rows can be matched back to a record by a human. It does
not affect validation.
Returns:
An object with:
ok (true only if no row was quarantined),
verdict ("clean" | "flagged" | "quarantined"),
summary (rows_in, accepted, quarantined, flagged, pass_rate,
accounted_for),
accepted (the passing rows, with unambiguous coercions applied),
quarantine (each entry: locator, row, reasons, findings),
flagged (rows kept despite a problem on a flag-mode column),
errors_by_column, findings_by_code, columns_seen,
columns_required_but_absent, contract (the parsed contract, so you can
confirm which rules ran), findings, and guidance.
`columns_required_but_absent` is worth checking before anything else: a
required column missing from every row usually means the wrong sheet or
the wrong header row was read, which is a different problem from bad data.Raises:
ToolError: if columns is malformed — an unknown type, a bad regular
expression, a duplicate column name, an unparseable spec. There is
nothing to validate until the contract parses, so the call is
rejected with a message naming the offending column and listing the
keys and types that are accepted. Row content never raises: a row
that is not an object is quarantined with the reason.
The contract is the call's definition, so a broken contract is an
argument error. The rows are the data, so broken rows are a result.
That is the line this server draws.
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| rows | Yes | ||
| columns | Yes | ||
| row_key | No |
Output Schema
| Name | Required | Description | Default |
|---|---|---|---|
No arguments | |||