Skip to main content
Glama

mcp-tool-router

Send an LLM only the tools that matter.

Agents with many MCP servers pay for every tool definition on every request: dozens of schemas fill the context window, cost tokens, and make the model more likely to pick the wrong tool. mcp-tool-router ranks your catalog against each question, by keywords (BM25) and optionally by meaning (embeddings), and returns the few tools that are relevant.

  • ToolRouter: per-question tool selection for your own agent loop (OpenAI, Anthropic, Gemini or MCP tool formats). Zero dependencies.

  • Semantic matching: "what is revenue in the last 5 days" finds get_sales_data even though its description never says "revenue".

  • RouterSearchTransform: a drop-in replacement for FastMCP's BM25SearchTransform that also matches word stems, typos and meaning.

  • mcp-tool-router eval: measures whether routing keeps the right tool, and how much schema payload it saves. Use --min-recall as a CI gate.

  • mcp-tool-router serve: puts any mcpServers config behind a search_tools + call_tool proxy.

Install

Not on PyPI yet. Install from GitHub:

uv add "mcp-tool-router @ git+https://github.com/dhruv1n30/mcp-tool-router"            # ToolRouter + eval, no dependencies
uv add "mcp-tool-router[semantic] @ git+https://github.com/dhruv1n30/mcp-tool-router"  # + a local embedding model (fastembed, CPU, no API key)
uv add "mcp-tool-router[fastmcp] @ git+https://github.com/dhruv1n30/mcp-tool-router"   # + RouterSearchTransform, serve, eval --config

With pip, use pip install "mcp-tool-router[fastmcp] @ git+https://github.com/dhruv1n30/mcp-tool-router". From a local checkout, use uv add --editable "path/to/mcp-tool-router[fastmcp]".

Related MCP server: MCP-Lens

Route tools in your own agent loop

from mcp_tool_router import ToolRouter

router = ToolRouter(tools, max_tools=8, always_include=["get_user_profile"])

relevant = router.select("how much stock is in warehouse B?")
response = client.chat.completions.create(model=..., messages=..., tools=relevant)

tools can be MCP tools/list results (dicts or mcp.types.Tool), OpenAI function tools, Anthropic tools, or Gemini function declarations. select returns your own objects, so they go straight back to the same SDK.

  • Pinned tools (always_include) are sent on every question, outside the cap.

  • Follow-ups: router.select("and last month?", history=["sales by category"]) keeps the earlier topic's tools in play.

  • No match fails open to the full catalog by default (on_no_match="all"); use on_no_match="pinned" to send only pinned tools instead.

  • Prompt caching: results come back in catalog order, so the same tool set always serializes identically.

Match by meaning, not just keywords

Keywords can't connect "revenue", "money we made" or "earnings" to a tool described as "Sales totals and order counts". Pass an embedder and the router also ranks tools by meaning:

from mcp_tool_router import ToolRouter
from mcp_tool_router.embedders import fastembed_embedder

router = ToolRouter(tools, max_tools=5, embed=fastembed_embedder())
router.select("what is revenue in last 5 days")  # -> [..., get_sales_data, ...]

fastembed_embedder() runs snowflake/snowflake-arctic-embed-s locally on the CPU (about 130 MB, downloaded on first use, roughly 20 ms per question). Any function with the signature embed(texts, *, query) -> vectors works, so you can use your provider's embeddings instead:

def embed(texts, *, query):
    response = openai_client.embeddings.create(model="text-embedding-3-small", input=list(texts))
    return [item.embedding for item in response.data]


router = ToolRouter(tools, embed=embed)

How the two signals combine: the keyword ranking and the meaning ranking are merged with Reciprocal Rank Fusion. A tool found both ways ranks highest, and a tool found only one way still makes the list. Tool descriptions are embedded once, on the first question; each question is embedded once and cached.

There is always a "nearest" tool by meaning, so with an embedder the router fills its max_tools budget on most questions. Set min_similarity (tune it with eval) if you want unrelated questions to fall through to on_no_match.

Use it inside a FastMCP server

from fastmcp import FastMCP
from mcp_tool_router import RouterSearchTransform

mcp = FastMCP("My server")
mcp.add_transform(RouterSearchTransform(max_results=8, always_visible=["ping"]))
# list_tools now returns ping, search_tools and call_tool

Same interface as FastMCP's built-in search transforms, with better matching. For the query inventroy worth, FastMCP's BM25SearchTransform returns no tools and RouterSearchTransform returns get_inventory_value. This is checked in tests/test_transform.py. Pass embed= for semantic search too.

Proxy existing MCP servers

mcp-tool-router serve --config claude_desktop_config.json --max-results 8 --semantic

Point your client at the proxy instead of the individual servers. It sees two tools, search_tools and call_tool, in place of the whole catalog.

Measure before you trust it

Write labelled questions as CSV (question,expected, with | between several expected tools) or JSONL, then:

mcp-tool-router eval --tools examples/tools.json --cases examples/cases.csv --max-tools 3
mcp-tool-router eval --tools examples/tools.json --cases examples/cases.csv --max-tools 3 --semantic

On the bundled toy example (8 tools, 17 questions, 5 of them phrased with words the descriptions don't contain):

recall

no match (sent everything)

wrong tool

tools per question

keywords only

94.1%

4

1

2.8 of 8

keywords + meaning (--semantic)

100%

0

0

3.0 of 8

The example is a toy; run eval on your own tools and real user questions. Use --config mcp.json to fetch the tools from live servers instead of a file, and --min-recall 0.95 to fail a CI job when routing starts dropping tools.

Watch the no match line. A question that matches nothing fails open to the whole catalog, which counts as a hit, so recall alone can look perfect while the token savings disappear.

How it works

  1. Tokenize: names, descriptions, parameter names and descriptions, and enum values are split into words (snake_case and camelCase aware, Unicode normalized, stopwords removed). Indexing enum values makes tools like get_metric(metric="tax") findable by "tax".

  2. Match: each question word matches catalog words exactly (1.0), by prefix ("sale"/"sales", 0.9), or by typo similarity ("inventroy", 0.7). match="exact" or "prefix" turns the looser kinds off.

  3. Rank: Okapi BM25. Words found in few tools count for more than words found in many, and shorter descriptions win ties.

  4. Meaning (optional): cosine similarity between the question's embedding and each tool's, fused with the BM25 ranking by Reciprocal Rank Fusion.

Limits

  • Without an embedder, matching is lexical: "revenue" won't find a sales tool. Pass embed= (or --semantic), or put your users' words in the descriptions.

  • The default embedding model is English. Pass a multilingual model name to fastembed_embedder() for other languages. The stopword list is English too.

  • Small embedding models separate close topics by thin margins. Route to a handful of tools (max_tools 3 to 8), not one, and check with eval.

  • Payload sizes are characters / 4, a rough token estimate. Exact counts need your model's tokenizer.

  • RouterSearchTransform subclasses FastMCP's search-transform hooks, so the extra is pinned to fastmcp>=4,<5.

Development

uv sync --all-extras
uv run pytest --cov      # unit, property-based (Hypothesis) and end-to-end tests
uv run pytest -m model   # also run the real embedding model (downloads ~130 MB)
uv run ruff check . && uv run ruff format --check .

The end-to-end tests start a real FastMCP server in a subprocess and route to it through the proxy.

License

MIT

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    D
    maintenance
    Acts as a proxy/router for multiple downstream MCP servers, exposing only meta-tools to the host to reduce token usage, enabling efficient search and invocation of tools from a fleet of servers.
    8 npm
    MIT
  • A
    license
    Not graded
    quality
    C
    maintenance
    Intelligent MCP proxy server that reduces context bloat by serving only the tools your AI actually needs through semantic search and a fixed two-tool surface.
    13
    MIT
  • A
    license
    Not graded
    quality
    B
    maintenance
    Fronts every MCP server you have with two tools: search shortlists relevant tools, and execute runs sandboxed Python that calls upstream tools as mcp..(). It enables agents to make bulk, repeated, or conditional tool calls using calibrated decision-model choices instead of per-item LLM generation.
    3
    MIT