Skip to main content
Glama
faiaz000

fuzzy-match-mcp

README.md
# Fuzzy Match MCP

A Python-based Model Context Protocol (MCP) server for deterministic fuzzy text matching.

It can normalize text, compare strings, rank possible matches, group duplicate values, and explain why two values match or differ. The server uses [RapidFuzz](https://github.com/rapidfuzz/RapidFuzz) for similarity scoring and can be used from MCP clients such as Cursor.

## Features

- Normalize text before comparison
- Compare two strings using multiple similarity algorithms
- Rank the best matches from a list
- Detect and group probable duplicates
- Explain matching results using token overlap
- Use specialized matching profiles for companies, products, and addresses
- Select different scoring strategies
- Run locally through MCP over `stdio`
- Testable with `pytest`

## Available MCP tools

### `normalize_text`

Normalizes a text value before fuzzy matching.

Example input:

```json
{
  "value": "  Müller & Söhne GmbH!  ",
  "profile": "general"
}
```

Example result:

```json
{
  "original": "  Müller & Söhne GmbH!  ",
  "profile": "general",
  "normalized": "muller and sohne gmbh"
}
```

### `compare_strings`

Compares two strings and returns detailed similarity scores.

Example input:

```json
{
  "first": "Deutsche Bank AG",
  "second": "Deutsche Bank Aktiengesellschaft",
  "threshold": 90,
  "profile": "company",
  "strategy": "strict"
}
```

The response includes:

- Normalized values
- Individual similarity scores
- Selected strategy
- Final selected score
- Match decision

### `find_best_matches`

Ranks candidate strings according to their similarity with a query.

Example input:

```json
{
  "query": "Samsung Galaxy S24",
  "choices": [
    "Apple iPhone 15",
    "Samsung Galaxy S24 128GB",
    "Galaxy S24 Smartphone",
    "Google Pixel 9"
  ],
  "limit": 3,
  "threshold": 40,
  "profile": "product",
  "strategy": "weighted"
}
```

### `find_duplicate_groups`

Groups values that probably represent the same entity.

Example input:

```json
{
  "values": [
    "Deutsche Bank AG",
    "Deutsche-Bank Aktiengesellschaft",
    "Deutsche Bank",
    "Commerzbank AG",
    "Commerz Bank",
    "Amazon Germany GmbH"
  ],
  "threshold": 80,
  "profile": "company",
  "strategy": "strict"
}
```

This is useful for:

- Customer-data cleanup
- Product-catalog deduplication
- Company-name matching
- Imported CSV cleanup
- Contact-list deduplication

### `explain_match`

Compares two values and explains the result using normalized tokens and score information.

Example input:

```json
{
  "first": "Samsung Galaxy S24 128GB Black",
  "second": "Galaxy S24 Black 128 GB",
  "threshold": 75,
  "profile": "product",
  "strategy": "strict"
}
```

The response includes:

- Common tokens
- Tokens found only in the first value
- Tokens found only in the second value
- Token-overlap percentage
- A deterministic explanation
- Final match decision

## Matching profiles

The server supports the following normalization profiles.

| Profile | Description |
|---|---|
| `general` | Lowercases text, removes accents and punctuation, and collapses whitespace |
| `company` | Applies general normalization and removes common legal company suffixes |
| `product` | Standardizes product units and joins model names such as `S 24` → `s24` |
| `address` | Standardizes common address terms such as `street` → `st` |

### Company-profile example

```text
Deutsche Bank AG
Deutsche Bank Aktiengesellschaft
```

Both normalize approximately to:

```text
deutsche bank
```

### Product-profile example

```text
Samsung Galaxy S 24 128 GB
```

Normalizes to:

```text
samsung galaxy s24 128gb
```

## Matching strategies

| Strategy | Description |
|---|---|
| `ratio` | Standard character similarity |
| `partial` | Finds the best matching substring |
| `token_sort` | Sorts tokens before comparison |
| `token_set` | Compares unique token sets |
| `weighted` | Uses RapidFuzz's weighted ratio |
| `strict` | Uses a composite score designed to reduce permissive partial matches |

The `strict` strategy excludes `partial_ratio` from the final composite score because partial matching can be misleading when a short string appears inside a much longer string.

## Thresholds

Similarity scores range from `0` to `100`.

- `90–100`: very strict
- `80–89`: useful default for names and product titles
- `70–79`: more permissive
- Below `70`: may create more false-positive matches

A result is considered a match when:

```python
selected_score >= threshold
```

The ideal threshold depends on the dataset and the acceptable false-positive rate.

## Project structure

```text
fuzzy-match-mcp/
├── .cursor/
│   └── mcp.json
├── fuzzy_match_mcp/
│   ├── __init__.py
│   ├── grouping.py
│   ├── matching.py
│   ├── server.py
│   ├── tools.py
│   └── validators.py
├── tests/
│   ├── __init__.py
│   └── test_matching.py
├── main.py
├── pyproject.toml
├── README.md
└── uv.lock
```

## Requirements

- Python 3.10 or newer
- `uv`
- MCP Python SDK
- RapidFuzz
- pytest for development

## Installation

Clone the repository:

```bash
git clone <your-repository-url>
cd fuzzy-match-mcp
```

Install the dependencies:

```bash
uv sync
```

If you are creating the project from scratch:

```bash
uv add "mcp[cli]" rapidfuzz
uv add --dev pytest
```

## Running the tests

```bash
uv run pytest
```

## Running with MCP Inspector

Use MCP Inspector during development:

```bash
uv run mcp dev main.py
```

This opens an MCP development interface where the registered tools can be inspected and called.

## Running as a local stdio server

```bash
uv run python main.py
```

The process waits for MCP JSON-RPC messages through standard input.

Do not type into the terminal while the server is running over `stdio`. A blank terminal input is not a valid MCP JSON-RPC message.

Do not use ordinary `print()` statements in the server because standard output is reserved for MCP communication. Use logging through standard error instead.

## Cursor configuration

Create:

```text
.cursor/mcp.json
```

Use the following configuration on Windows:

```json
{
  "mcpServers": {
    "fuzzy-match": {
      "type": "stdio",
      "command": "${workspaceFolder}/.venv/Scripts/python.exe",
      "args": [
        "${workspaceFolder}/main.py"
      ]
    }
  }
}
```

Then:

1. Open the repository root in Cursor.
2. Open **Customize → MCPs**.
3. Enable `fuzzy-match`.
4. Click **Reload** after changing the registered tools.
5. Use Cursor Agent to call the MCP tools.

Cursor starts the Python process automatically. Do not manually run `main.py` at the same time.

## Example Cursor Agent prompts

### Normalize text

```text
Use the fuzzy-match normalize_text tool to normalize:

"  Müller & Söhne GmbH!  "

Use the general profile.
```

### Compare company names

```text
Use the fuzzy-match compare_strings tool.

Compare:
- Deutsche Bank AG
- Deutsche Bank Aktiengesellschaft

Use:
- profile: company
- strategy: strict
- threshold: 90
```

### Rank product matches

```text
Use the fuzzy-match find_best_matches tool.

Query:
Samsung Galaxy S24

Choices:
- Apple iPhone 15
- Samsung Galaxy S24 128GB
- Galaxy S24 Smartphone
- Google Pixel 9

Use the product profile and return the best three matches.
```

### Group duplicates

```text
Use the fuzzy-match find_duplicate_groups tool with:

- Deutsche Bank AG
- Deutsche-Bank Aktiengesellschaft
- Deutsche Bank
- Commerzbank AG
- Commerz Bank
- Amazon Germany GmbH

Use:
- profile: company
- strategy: strict
- threshold: 80
```

### Explain a match

```text
Use the fuzzy-match explain_match tool to compare:

- Samsung Galaxy S24 128GB Black
- Galaxy S24 Black 128 GB

Use:
- profile: product
- strategy: strict
- threshold: 75
```

## How it works

The request flow is:

```text
MCP client
    ↓
MCP tool in tools.py
    ↓
Input validation in validators.py
    ↓
Matching or grouping logic
    ↓
RapidFuzz
    ↓
Structured JSON result
```

Normalization happens before similarity scoring. It can include:

- Unicode case folding
- Accent removal
- Punctuation replacement
- Whitespace cleanup
- Profile-specific transformations

## Performance notes

`find_best_matches` compares one query against each candidate.

`find_duplicate_groups` performs pairwise comparisons. Its approximate comparison count is:

```text
n × (n - 1) / 2
```

For this reason, duplicate grouping is intentionally limited to 500 values in the current version.

The grouping implementation uses union-find. If A matches B and B matches C, all three values can be placed in the same group even when A and C do not directly exceed the threshold.

## Current limitations

- Duplicate grouping is pairwise and is not intended for very large datasets.
- Company suffixes and address aliases are rule-based and may not cover every country.
- Product normalization supports only a small set of common units.
- Fuzzy matching does not prove that two real-world entities are identical.
- Thresholds should be evaluated against domain-specific examples before automatic merging.

## Planned improvements

- Structured record matching
- Batch matching
- Threshold evaluation
- Custom normalization options
- Additional international company suffixes
- CSV import and export
- Better canonical-value selection
- More matching profiles
- MCP resources and reusable prompts

## Contributing

Contributions are welcome.

Suggested workflow:

```bash
git checkout -b feature/my-change
uv sync
uv run pytest
```

Before submitting a change:

- Add tests for new behavior
- Keep MCP tools focused
- Avoid writing to standard output
- Preserve structured JSON responses
- Document new profiles and strategies

## License

Add your chosen license before publishing the project.

A common choice for an open-source MCP server is the MIT License.

TDQS

A3.6/5.0

Scored across 5 tools

Disambiguation5/5

Each tool has a clearly distinct purpose: direct comparison, normalization, candidate ranking, duplicate grouping, and detailed explanation. No overlap in functionality.

Naming Consistency5/5

All tool names follow a consistent verb_noun pattern (compare_strings, normalize_text, etc.) with clear and predictable naming.

Tool Count5/5

5 tools is well-scoped for a fuzzy matching server, covering essential operations without being excessive or sparse.

Completeness4/5

The surface covers normalization, comparison, matching, deduplication, and explanation. A minor gap is the lack of a tool to list available profiles/strategies, but descriptions provide that information.

Maintenance

ActivityStale
ResponsivenessNo issues