duplicate_report
Find suspected duplicate groups in row data and return them as a review queue for human decision — never merging, deleting, or picking a winner.
Instructions
Find suspected duplicate groups and return them for human review.
Use this when the same real entity may be recorded more than once — a re-registration, a rename, a typo in a reference field — and you need to know how much of the data that affects before trusting any count.
This tool NEVER merges, deletes, or picks a winner. Every group it returns is
a suspicion with auto_merged: false and decision_required: true, plus a
review_question phrased for a person. A silent merge corrupts counts
invisibly and is very hard to unpick later, so the output is a review queue,
not a result.
Three kinds of suspicion are reported, each labelled in kind:
same_key: the same normalised key appears on more than one row. Either genuine duplicates or legitimate repeated transactions — the tool cannot tell which, so it reports rather than decides.same_signature: different keys whose values acrossmatch_fieldsare identical. This is the re-registration case.near_match: different keys whose values acrossmatch_fieldsare similar rather than identical. Only produced when you supplysimilarity_threshold, because "close enough" is a policy choice.
Args:
rows: The data, as an array of objects.
key: The column holding the identifier. Needed to tell "one entity,
several rows" from "several entities".
match_fields: Columns whose combination suggests two different keys are
the same entity, for example ["name", "postcode"] or
["given_name", "family_name", "date_of_birth"]. Comparing on
several fields is far safer than one: a single name column collides
constantly. Omit it to check only exact key repeats.
similarity_threshold: A number above 0 and at most 1. When supplied,
match_fields values are also compared with string similarity
(difflib.SequenceMatcher) and groups scoring at or above this value
are reported as near_match. 0.9 is a reasonable starting point;
lower values find more and mean less. Omit it and only exact matches
are reported.
case_sensitive: If false (the default), values are compared after trimming
whitespace and folding case.
max_groups: Cap on the number of groups returned, default 50. When the
cap bites, groups_omitted says how many were left out — the count is
never silently reduced.
Returns:
An object with:
ok (true only when no suspicion was found),
verdict ("clean" | "review_required"),
row_count, distinct_keys, duplicate_group_count,
rows_in_duplicate_groups, auto_merged (always false),
nothing_was_merged_or_deleted (always true),
blank_key_rows, groups, groups_omitted, findings, and guidance.
Each group has `group_id`, `kind`, `reason`, `size`, `keys`,
`recommended_action` ("human_review"), `decision_required`, a
`review_question`, and `members` (each with a locator and the full row).Raises:
ToolError: if key is not a non-empty string, if match_fields is not
an array of column names, or if similarity_threshold is not above
0 and at most 1. All three are malformed arguments rather than data
problems. Row content never raises.
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| key | Yes | ||
| rows | Yes | ||
| max_groups | No | ||
| match_fields | No | ||
| case_sensitive | No | ||
| similarity_threshold | No |
Output Schema
| Name | Required | Description | Default |
|---|---|---|---|
No arguments | |||