Skip to main content
Glama
deerajd

MCP Duplicate File Manager Server

by deerajd
README.md
# MCP Duplicate File Manager Server

A Model Context Protocol (MCP) server that enables LLM-based AI agents to manage duplicate files efficiently. Designed for seamless integration with LLMs and AI agents using the MCP protocol with intelligent duplicate detection and safe deletion capabilities.

---

## Features

- Exposes file management tools for LLM agents via MCP:
  - `list_files`: List files in directories with optional filtering
  - `find_duplicates`: Find duplicate files based on content hash (MD5)
  - `delete_duplicates`: Safely delete duplicate files with intelligent copy detection
- Smart copy detection: Prioritizes deleting files with copy indicators ("copy", "(1)", "(2)", etc.)
- Content-based duplicate detection using MD5 hashing for accuracy
- Recursive directory scanning with optional pattern matching
- Dry-run mode for safe preview before actual deletion
- Simple integration with Claude, MCP clients, or other LLM-based agents

---

## Quickstart

### 1. Install & Run

Run the MCP Duplicate File Manager Server using [`uvx`](https://github.com/uvx-dev/uvx):

```bash
uvx --from git+https://github.com/deerajd/DuplicateFileRemover-MCP.git mcp-server
```

---

### 2. Integrate with Your LLM Agent

Add the following to your tool configuration (for Claude, MCP clients, etc.):

```json
{
  "mcpServers": {
    "DuplicateListor": {
      "command": "uvx",
      "args": [
        "--from",
        "git+https://github.com/deerajd/DuplicateFileRemover-MCP.git",
        "mcp-server"
      ]
    }
  }
}
```

This configuration launches the MCP Duplicate File Manager Server as a subprocess, exposing file management tools to your LLM agent.

---

### 3. Usage

From your LLM agent, you can call:

#### List Files in Directory
```
list_files(directory="./Documents", pattern="*.pdf")        # List PDF files
list_files(directory="/home/user", include_hidden=True)     # Include hidden files
list_files()                                                # List all files in current directory
```

#### Find Duplicate Files
```
find_duplicates(directory="./Downloads")                    # Find duplicates in Downloads
find_duplicates(directory="/Users/john/Desktop", recursive=False)  # Non-recursive search
find_duplicates()                                           # Find duplicates in current directory
```

#### Delete Duplicate Files
```
delete_duplicates(directory="./Documents", dry_run=True)    # Preview what would be deleted
delete_duplicates(directory="./Documents", dry_run=False)   # Actually delete duplicates
delete_duplicates(directory="./Pictures", recursive=True)   # Recursive deletion with preview
```

The server provides detailed responses including:
- List of files with metadata (name, size, type, path)
- Duplicate file groups organized by content hash
- Deletion results with preserved originals and deleted copies
- Error handling for inaccessible files or directories

---

## API Reference

### `list_files(directory=".", pattern="*", include_hidden=False)`
- **Parameters:**
  - `directory` (str): Directory path to list files from (default: current directory)
  - `pattern` (str): File pattern to match (e.g., "*.py", "*.txt") (default: "*")
  - `include_hidden` (bool): Whether to include hidden files (default: False)
- **Returns:** List of dictionaries containing file information (name, size, type, path, modified)

### `find_duplicates(directory=".", recursive=True)`
- **Parameters:**
  - `directory` (str): Directory path to search for duplicates (default: current directory)
  - `recursive` (bool): Whether to search recursively in subdirectories (default: True)
- **Returns:** Dictionary with hash as key and list of duplicate file paths as values

### `delete_duplicates(directory=".", recursive=True, dry_run=True)`
- **Parameters:**
  - `directory` (str): Directory path to search for duplicates (default: current directory)
  - `recursive` (bool): Whether to search recursively in subdirectories (default: True)
  - `dry_run` (bool): If True, only show what would be deleted without actually deleting (default: True)
- **Returns:** Dictionary with information about deleted files and preserved originals

#### Smart Copy Detection
The server intelligently identifies and prioritizes deletion of files with copy indicators:
- `copy` - Files containing "copy" in the filename
- `(1)`, `(2)`, etc. - Files with numbered suffixes
- `_copy`, `-copy` - Files with copy prefixes/suffixes
- `duplicate`, `dup` - Files marked as duplicates
- Multiple language variants (copie, kopia, copia)

---

## Development

### Requirements
- Python 3.8+
- [`hashlib`](https://docs.python.org/3/library/hashlib.html) for MD5 hashing
- [`pathlib`](https://docs.python.org/3/library/pathlib.html) for file path operations
- [`FastMCP`](https://github.com/modelcontext/fastmcp) for MCP server implementation

### Project Structure
```
DuplicateListor/
├── duplicatelist.py             # Main server implementation
├── main.py                      # Entry point
├── pyproject.toml               # Project configuration
└── README.md                    # This file
```

---

## Safety Features

- **Dry-run Mode**: Preview changes before actual deletion
- **Smart Original Detection**: Preserves files without copy indicators
- **Error Handling**: Graceful handling of inaccessible files
- **Content-based Matching**: Uses MD5 hashing for accurate duplicate detection
- **Detailed Reporting**: Comprehensive feedback on all operations

---

## License

MIT License

---

TDQS

A4.5/5.0

Scored across 3 tools

Disambiguation5/5

Each tool has a distinct purpose: list_files for directory exploration, find_duplicates for identifying duplicates, and delete_duplicates for removing them. There is no overlap or ambiguity between the tool functions.

Naming Consistency5/5

All tools follow the verb_noun pattern with consistent snake_case and plural nouns (list_files, find_duplicates, delete_duplicates). The naming is uniform and intuitive.

Tool Count5/5

With 3 tools, the server is tightly scoped to the core duplicate file management workflow. Each tool is essential and none are redundant, making the count appropriate.

Completeness5/5

The tool set covers the full lifecycle: listing files to explore, finding duplicates, and then deleting them with a dry-run option for safety. No critical operations are missing for the stated purpose.

Maintenance

ActivityInactive
ResponsivenessNo issues