Skip to main content
Glama
README.md
# mcp-tool-router

Send an LLM only the tools that matter.

Agents with many MCP servers pay for every tool definition on every request:
dozens of schemas fill the context window, cost tokens, and make the model
more likely to pick the wrong tool. `mcp-tool-router` ranks your catalog
against each question, by keywords (BM25) and optionally by meaning
(embeddings), and returns the few tools that are relevant.

- **`ToolRouter`**: per-question tool selection for your own agent loop
  (OpenAI, Anthropic, Gemini or MCP tool formats). Zero dependencies.
- **Semantic matching**: "what is revenue in the last 5 days" finds
  `get_sales_data` even though its description never says "revenue".
- **`RouterSearchTransform`**: a drop-in replacement for FastMCP's
  `BM25SearchTransform` that also matches word stems, typos and meaning.
- **`mcp-tool-router eval`**: measures whether routing keeps the right tool,
  and how much schema payload it saves. Use `--min-recall` as a CI gate.
- **`mcp-tool-router serve`**: puts any `mcpServers` config behind a
  `search_tools` + `call_tool` proxy.

## Install

Not on PyPI yet. Install from GitHub:

```bash
uv add "mcp-tool-router @ git+https://github.com/dhruv1n30/mcp-tool-router"            # ToolRouter + eval, no dependencies
uv add "mcp-tool-router[semantic] @ git+https://github.com/dhruv1n30/mcp-tool-router"  # + a local embedding model (fastembed, CPU, no API key)
uv add "mcp-tool-router[fastmcp] @ git+https://github.com/dhruv1n30/mcp-tool-router"   # + RouterSearchTransform, serve, eval --config
```

With pip, use `pip install "mcp-tool-router[fastmcp] @ git+https://github.com/dhruv1n30/mcp-tool-router"`.
From a local checkout, use `uv add --editable "path/to/mcp-tool-router[fastmcp]"`.

## Route tools in your own agent loop

```python
from mcp_tool_router import ToolRouter

router = ToolRouter(tools, max_tools=8, always_include=["get_user_profile"])

relevant = router.select("how much stock is in warehouse B?")
response = client.chat.completions.create(model=..., messages=..., tools=relevant)
```

`tools` can be MCP `tools/list` results (dicts or `mcp.types.Tool`), OpenAI
function tools, Anthropic tools, or Gemini function declarations. `select`
returns your own objects, so they go straight back to the same SDK.

- **Pinned tools** (`always_include`) are sent on every question, outside the cap.
- **Follow-ups**: `router.select("and last month?", history=["sales by category"])`
  keeps the earlier topic's tools in play.
- **No match** fails open to the full catalog by default (`on_no_match="all"`);
  use `on_no_match="pinned"` to send only pinned tools instead.
- **Prompt caching**: results come back in catalog order, so the same tool set
  always serializes identically.

## Match by meaning, not just keywords

Keywords can't connect "revenue", "money we made" or "earnings" to a tool
described as "Sales totals and order counts". Pass an embedder and the router
also ranks tools by meaning:

```python
from mcp_tool_router import ToolRouter
from mcp_tool_router.embedders import fastembed_embedder

router = ToolRouter(tools, max_tools=5, embed=fastembed_embedder())
router.select("what is revenue in last 5 days")  # -> [..., get_sales_data, ...]
```

`fastembed_embedder()` runs `snowflake/snowflake-arctic-embed-s` locally on
the CPU (about 130 MB, downloaded on first use, roughly 20 ms per question).
Any function with the signature `embed(texts, *, query) -> vectors` works, so
you can use your provider's embeddings instead:

```python
def embed(texts, *, query):
    response = openai_client.embeddings.create(model="text-embedding-3-small", input=list(texts))
    return [item.embedding for item in response.data]


router = ToolRouter(tools, embed=embed)
```

How the two signals combine: the keyword ranking and the meaning ranking are
merged with Reciprocal Rank Fusion. A tool found both ways ranks highest, and
a tool found only one way still makes the list. Tool descriptions are embedded
once, on the first question; each question is embedded once and cached.

There is always a "nearest" tool by meaning, so with an embedder the router
fills its `max_tools` budget on most questions. Set `min_similarity` (tune it
with `eval`) if you want unrelated questions to fall through to `on_no_match`.

## Use it inside a FastMCP server

```python
from fastmcp import FastMCP
from mcp_tool_router import RouterSearchTransform

mcp = FastMCP("My server")
mcp.add_transform(RouterSearchTransform(max_results=8, always_visible=["ping"]))
# list_tools now returns ping, search_tools and call_tool
```

Same interface as FastMCP's built-in search transforms, with better matching.
For the query `inventroy worth`, FastMCP's `BM25SearchTransform` returns no
tools and `RouterSearchTransform` returns `get_inventory_value`. This is
checked in `tests/test_transform.py`. Pass `embed=` for semantic search too.

## Proxy existing MCP servers

```bash
mcp-tool-router serve --config claude_desktop_config.json --max-results 8 --semantic
```

Point your client at the proxy instead of the individual servers. It sees two
tools, `search_tools` and `call_tool`, in place of the whole catalog.

## Measure before you trust it

Write labelled questions as CSV (`question,expected`, with `|` between several
expected tools) or JSONL, then:

```bash
mcp-tool-router eval --tools examples/tools.json --cases examples/cases.csv --max-tools 3
mcp-tool-router eval --tools examples/tools.json --cases examples/cases.csv --max-tools 3 --semantic
```

On the bundled toy example (8 tools, 17 questions, 5 of them phrased with
words the descriptions don't contain):

| | recall | no match (sent everything) | wrong tool | tools per question |
|---|---|---|---|---|
| keywords only | 94.1% | 4 | 1 | 2.8 of 8 |
| keywords + meaning (`--semantic`) | 100% | 0 | 0 | 3.0 of 8 |

The example is a toy; run `eval` on your own tools and real user questions.
Use `--config mcp.json` to fetch the tools from live servers instead of a file,
and `--min-recall 0.95` to fail a CI job when routing starts dropping tools.

Watch the **no match** line. A question that matches nothing fails open to the
whole catalog, which counts as a hit, so recall alone can look perfect while
the token savings disappear.

## How it works

1. **Tokenize**: names, descriptions, parameter names and descriptions, and enum
   values are split into words (snake_case and camelCase aware, Unicode
   normalized, stopwords removed). Indexing enum values makes tools like
   `get_metric(metric="tax")` findable by "tax".
2. **Match**: each question word matches catalog words exactly (1.0), by
   prefix ("sale"/"sales", 0.9), or by typo similarity ("inventroy", 0.7).
   `match="exact"` or `"prefix"` turns the looser kinds off.
3. **Rank**: Okapi BM25. Words found in few tools count for more than words
   found in many, and shorter descriptions win ties.
4. **Meaning** (optional): cosine similarity between the question's embedding
   and each tool's, fused with the BM25 ranking by Reciprocal Rank Fusion.

## Limits

- Without an embedder, matching is lexical: "revenue" won't find a sales tool.
  Pass `embed=` (or `--semantic`), or put your users' words in the descriptions.
- The default embedding model is English. Pass a multilingual model name to
  `fastembed_embedder()` for other languages. The stopword list is English too.
- Small embedding models separate close topics by thin margins. Route to a
  handful of tools (`max_tools` 3 to 8), not one, and check with `eval`.
- Payload sizes are characters / 4, a rough token estimate. Exact counts
  need your model's tokenizer.
- `RouterSearchTransform` subclasses FastMCP's search-transform hooks, so the
  extra is pinned to `fastmcp>=4,<5`.

## Development

```bash
uv sync --all-extras
uv run pytest --cov      # unit, property-based (Hypothesis) and end-to-end tests
uv run pytest -m model   # also run the real embedding model (downloads ~130 MB)
uv run ruff check . && uv run ruff format --check .
```

The end-to-end tests start a real FastMCP server in a subprocess and route to
it through the proxy.

## License

MIT