Skip to main content
Glama
ianderso
by ianderso

find_duplicates

Read-only

Reports likely duplicate media, sources, citations, vital events, or person names in a Gramps tree as candidates only; it never merges, letting you verify and decide.

Instructions

Find likely-duplicate objects. Reports only -- it never merges anything.

Duplicates are not merely untidy: a duplicate SOURCE makes a single-sourced fact look corroborated, which is a false evidentiary claim. But the reverse error is just as real -- an index entry and the register page it indexes are TWO documents and must stay separate. This tool finds candidates; deciding which are truly the same document is yours. Merge with merge_objects.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
kindYesmedia_checksum (same file uploaded twice), source_title (same document entered twice), citation_page (same source+page cited more than once, flagging any graded differently), vital_events (a person with two Births), person_name (same name, possible same person).
limitNoMax groups to return.
include_privateNoShow living people and private records in full. Only when the user asks for them; they are withheld by default.

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.1.0

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, so safety is covered; the description adds operational detail beyond that by clarifying that read-only means "never merges" and that output is candidate groups, not decisions. It also explains the evidentiary rationale for why the tool errs toward caution. No mention of permissions, cost, or performance characteristics.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the action and its scope in the first sentence. The middle sentences are editorial but functional — they justify why duplicate detection is non-trivial and warn against over-merging. It runs slightly long for the functional payload, but every sentence carries usable guidance.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only search tool with full schema coverage and no output schema, the description covers the detection/decision boundary and points at merge_objects for follow-through. It does not describe the shape of the returned report (the grouping is only implied by the limit parameter's "Max groups to return"), which is a minor gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the kind enum values, limit, and include_private are all documented in the schema itself, including the privacy default. The description text adds no parameter-level detail beyond the schema, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb + resource ("Find likely-duplicate objects") and immediately scopes the behavior ("Reports only -- it never merges anything"). It explicitly names the sibling it is not (merge_objects), so an agent can separate detection from mutation without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear conditions: this tool produces candidates, the human decides which are truly the same, and merge_objects performs the merge. It supplies the false-positive caution (index entry vs. register page must stay separate), which is real usage guidance. It stops short of explicit when-not triggers against other read tools like verify_tree or list_unsourced_facts.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.