Skip to main content
Glama
devops-gm88

mcp-validation-server

by devops-gm88

duplicate_report

Find suspected duplicate groups in row data and return them as a review queue for human decision — never merging, deleting, or picking a winner.

Instructions

Find suspected duplicate groups and return them for human review.

Use this when the same real entity may be recorded more than once — a re-registration, a rename, a typo in a reference field — and you need to know how much of the data that affects before trusting any count.

This tool NEVER merges, deletes, or picks a winner. Every group it returns is a suspicion with auto_merged: false and decision_required: true, plus a review_question phrased for a person. A silent merge corrupts counts invisibly and is very hard to unpick later, so the output is a review queue, not a result.

Three kinds of suspicion are reported, each labelled in kind:

  • same_key: the same normalised key appears on more than one row. Either genuine duplicates or legitimate repeated transactions — the tool cannot tell which, so it reports rather than decides.

  • same_signature: different keys whose values across match_fields are identical. This is the re-registration case.

  • near_match: different keys whose values across match_fields are similar rather than identical. Only produced when you supply similarity_threshold, because "close enough" is a policy choice.

Args: rows: The data, as an array of objects. key: The column holding the identifier. Needed to tell "one entity, several rows" from "several entities". match_fields: Columns whose combination suggests two different keys are the same entity, for example ["name", "postcode"] or ["given_name", "family_name", "date_of_birth"]. Comparing on several fields is far safer than one: a single name column collides constantly. Omit it to check only exact key repeats. similarity_threshold: A number above 0 and at most 1. When supplied, match_fields values are also compared with string similarity (difflib.SequenceMatcher) and groups scoring at or above this value are reported as near_match. 0.9 is a reasonable starting point; lower values find more and mean less. Omit it and only exact matches are reported. case_sensitive: If false (the default), values are compared after trimming whitespace and folding case. max_groups: Cap on the number of groups returned, default 50. When the cap bites, groups_omitted says how many were left out — the count is never silently reduced.

Returns: An object with: ok (true only when no suspicion was found), verdict ("clean" | "review_required"), row_count, distinct_keys, duplicate_group_count, rows_in_duplicate_groups, auto_merged (always false), nothing_was_merged_or_deleted (always true), blank_key_rows, groups, groups_omitted, findings, and guidance.

Each group has `group_id`, `kind`, `reason`, `size`, `keys`,
`recommended_action` ("human_review"), `decision_required`, a
`review_question`, and `members` (each with a locator and the full row).

Raises: ToolError: if key is not a non-empty string, if match_fields is not an array of column names, or if similarity_threshold is not above 0 and at most 1. All three are malformed arguments rather than data problems. Row content never raises.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
keyYes
rowsYes
max_groupsNo
match_fieldsNo
case_sensitiveNo
similarity_thresholdNo

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault

No arguments

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.1.0

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the entire behavioral burden and does so richly: it never merges/deletes/picks a winner, groups are always `auto_merged: false` / `decision_required: true`, `near_match` only appears when a threshold is supplied, and `max_groups` truncation is surfaced via `groups_omitted` rather than silently applied. It also enumerates the exact argument-validation errors that raise ToolError versus data problems that never raise.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Well front-loaded: the first sentence states the action and the second gives the trigger, before any detail. The Args block is justified given 0% schema coverage, but the Returns block is somewhat redundant with the existing output schema and repeats facts already stated in the prose (e.g., `auto_merged: false`), adding length without new decision-relevant information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Although an output schema exists, the description goes further by explaining the meaning of the returned verdict and review queue rather than just its shape, and it covers the failure modes, defaults, and policy-driven behaviors an agent needs. Nothing required for correct invocation is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate fully — and it does, covering all six parameters with meaning beyond types: `key` distinguishes one-entity-many-rows from many-entities, `match_fields` is explained with concrete examples and a warning about single-column collisions, `similarity_threshold` gets its domain ('above 0 and at most 1'), algorithm, and a suggested starting value, and `case_sensitive` documents the trimming/case-folding default behavior.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('find suspected duplicate groups and return them for human review') and immediately scopes the tool's role against the obvious alternative behavior of merging. The three `kind` categories further pin down exactly what it detects, so an agent can tell it apart from siblings like reconcile or count_distinct.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit when-to-use condition ('when the same real entity may be recorded more than once — a re-registration, a rename, a typo') and the motivating context ('before trusting any count'). It does not name the sibling tools (validate_rows, reconcile, count_distinct) as alternatives or state when NOT to use it, so it stops short of full routing guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.