laserfiche_document_find_duplicates
Find byte-identical documents in a folder tree and group them, revealing duplicate files and estimating recoverable space from deduplication.
Instructions
Find byte-identical documents in a folder tree and group them.
Use this to answer "are there duplicate files in here?" or "how much
space would deduping this folder recover?" — it downloads nothing for
documents whose size doesn't collide with another's, and only hashes the
ones that do, so it's usually far cheaper than it sounds. This is a
read-only scan: it finds duplicates, it does not delete or merge them —
follow up with delete_entry/delete_edoc yourself on whichever
copies you decide to remove.
Two-pass approach: first probes every document's size (headers only, no bytes transferred); only documents that share a size with another are then downloaded and hashed (SHA-256). A repository of mostly-distinct documents therefore touches a small fraction of the tree on pass two.
This is a single blocking call with no interim progress — for a large
max_entries this can take a while (network round-trips per document
plus the downloads pass two triggers). Start with the default and raise
max_entries only once you've seen how large the tree is, e.g. via
list_folder.
Sibling tools: get_entry_by_path to resolve a path to the
folder_id this tool needs; get_document_edoc to download or read
a specific document once you've identified which copy to keep;
compare_entries to check whether two SIMILAR-but-not-identical
documents differ only in metadata.
Returns {"mode": "duplicate_report", "folder_id", "recursive", "walk_truncated": bool, "documents_examined", "documents_hashed", "bytes_downloaded", "total_wasted_bytes", "max_bytes", "groups": [{"sha256", "byte_size", "wasted_bytes", "entries": [{"entry_id", "name"}, ...]}, ...], "skipped": [{"entry_id", "name", "reason"}, ...], "folders_unreadable": [<folder_id>, ...]}. groups is sorted by
wasted_bytes descending (biggest recoverable space first).
folders_unreadable lists subfolders the walk couldn't list (usually
permissions) — entries under them are NOT included, so a non-empty list
means the scan was partial even if walk_truncated is false. On
failure returns {"mode": "error", "error": <slug>, "entry_id": <folder_id>} — not_found/auth_failed (bad folder_id) or
not_a_folder (folder_id points at a document).
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| folder_id | Yes | Integer entry ID of the folder to scan. Resolve a path first with get_entry_by_path if you only have a location. | |
| max_bytes | No | Skip (don't download or hash) any document larger than this many bytes; it's listed in `skipped` with a reason rather than silently treated as unique. Defaults to LF_EDOC_MAX_BYTES (25 MB). | |
| recursive | No | Scan subfolders too. False = only this folder's immediate children. | |
| max_entries | No | Stop walking the tree after this many entries (folders + documents). A stop, not a filter — hitting it sets walk_truncated=true rather than silently reporting a partial tree as complete. Defaults to 2000; raise for an exhaustive audit of a larger tree, but expect the call to take proportionally longer since it runs to completion in one request. |
Output Schema
| Name | Required | Description | Default |
|---|---|---|---|
No arguments | |||