Skip to main content
Glama

laserfiche_document_find_duplicates

Find byte-identical documents in a folder tree and group them, revealing duplicate files and estimating recoverable space from deduplication.

Instructions

Find byte-identical documents in a folder tree and group them.

Use this to answer "are there duplicate files in here?" or "how much space would deduping this folder recover?" — it downloads nothing for documents whose size doesn't collide with another's, and only hashes the ones that do, so it's usually far cheaper than it sounds. This is a read-only scan: it finds duplicates, it does not delete or merge them — follow up with delete_entry/delete_edoc yourself on whichever copies you decide to remove.

Two-pass approach: first probes every document's size (headers only, no bytes transferred); only documents that share a size with another are then downloaded and hashed (SHA-256). A repository of mostly-distinct documents therefore touches a small fraction of the tree on pass two.

This is a single blocking call with no interim progress — for a large max_entries this can take a while (network round-trips per document plus the downloads pass two triggers). Start with the default and raise max_entries only once you've seen how large the tree is, e.g. via list_folder.

Sibling tools: get_entry_by_path to resolve a path to the folder_id this tool needs; get_document_edoc to download or read a specific document once you've identified which copy to keep; compare_entries to check whether two SIMILAR-but-not-identical documents differ only in metadata.

Returns {"mode": "duplicate_report", "folder_id", "recursive", "walk_truncated": bool, "documents_examined", "documents_hashed", "bytes_downloaded", "total_wasted_bytes", "max_bytes", "groups": [{"sha256", "byte_size", "wasted_bytes", "entries": [{"entry_id", "name"}, ...]}, ...], "skipped": [{"entry_id", "name", "reason"}, ...], "folders_unreadable": [<folder_id>, ...]}. groups is sorted by wasted_bytes descending (biggest recoverable space first). folders_unreadable lists subfolders the walk couldn't list (usually permissions) — entries under them are NOT included, so a non-empty list means the scan was partial even if walk_truncated is false. On failure returns {"mode": "error", "error": <slug>, "entry_id": <folder_id>} — not_found/auth_failed (bad folder_id) or not_a_folder (folder_id points at a document).

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
folder_idYesInteger entry ID of the folder to scan. Resolve a path first with get_entry_by_path if you only have a location.
max_bytesNoSkip (don't download or hash) any document larger than this many bytes; it's listed in `skipped` with a reason rather than silently treated as unique. Defaults to LF_EDOC_MAX_BYTES (25 MB).
recursiveNoScan subfolders too. False = only this folder's immediate children.
max_entriesNoStop walking the tree after this many entries (folders + documents). A stop, not a filter — hitting it sets walk_truncated=true rather than silently reporting a partial tree as complete. Defaults to 2000; raise for an exhaustive audit of a larger tree, but expect the call to take proportionally longer since it runs to completion in one request.

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault

No arguments

Schema Changelog

Changes observed during successful MCP inspections.

  1. Addedv2.3.0

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden and delivers thoroughly. It discloses that this is a read-only scan, does not delete or merge, uses a two-pass approach (size probe then SHA-256 hashing only on collisions), is a single blocking call with no interim progress, and can be partial if folders are unreadable. It also details performance implications and error modes (not_found, auth_failed, not_a_folder). No contradictions with annotations since none exist.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long but every sentence earns its place. It is front-loaded with purpose and use case, then behavior, then sibling references, then a detailed return format. Structure uses paragraphs and a bullet-like listing of the output, making it scannable. No filler or redundancy; the length is proportionate to the tool's complexity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 4-parameter tool with an output schema, the description covers everything an agent needs: return structure (mode, groups, skipped, folders_unreadable), sorting (by wasted_bytes descending), partial-scan indicators (walk_truncated and folders_unreadable), and error responses. It even warns about performance for large max_entries. Nothing essential is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds meaningful operational context: max_bytes defaults to LF_EDOC_MAX_BYTES (25 MB) and skipped entries are reported with a reason; max_entries is 'a stop, not a filter' that sets walk_truncated=true; recursive defaults to true. These clarifications help an agent make correct choices beyond what the schema states, justifying a 4.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a precise verb and resource: 'Find byte-identical documents in a folder tree and group them.' It immediately gives the two canonical use cases ('are there duplicate files in here?' and 'how much space would deduping this folder recover?') and differentiates itself from siblings like compare_entries (similar-but-not-identical) and search tools. An agent can tell exactly what this tool does and what it does not do.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly states when to use it ('Use this to answer...') and what not to do (read-only scan, must follow up with delete_entry/delete_edoc yourself). It names the sibling tools for alternatives: get_entry_by_path for resolving folder_id, get_document_edoc for downloading a specific copy, and compare_entries for similar-but-not-identical documents. This leaves no ambiguity about selection.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.