toolrank
Ships a toolrank Docker image for amd64 and arm64 with compose files that run toolrank alongside a vLLM embedding server, and a Dockerfile combining both in one container.
Integrates with LangGraph via langgraph-bigtool, plus LlamaIndex agents and the LiteLLM proxy, so agent frameworks can rank and load tools on demand instead of stuffing every definition into the prompt.
Provides OpenAI-style client-side tool search (tool_search), letting OpenAI-based agents retrieve relevant tools before calling them.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@toolrankfind tools for creating a GitHub issue"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
toolrank
Tool retrieval for LLM agents with hundreds of tools. Instead of putting every tool definition into the prompt, toolrank picks the few a request needs, with an embedding model (Qwen3-Embedding-8B, trained further on tool retrieval), and serves them to your agent over MCP or REST. Every retriever it ships is measured on the same public benchmarks.
Status: alpha (0.2); interfaces may still change. Documentation: https://yaman.dev/toolrank/
Quick start
Index MCP servers or OpenAPI specs, then serve them to any MCP client as two tools, search_tools
and call_tool:
pip install "toolrank[mcp]"
toolrank ingest mcp --server time="uvx mcp-server-time" --out tools/
toolrank serve --data tools/ # MCP at http://127.0.0.1:8765/mcp, REST at /v1toolrank ranks with its backbone, Qwen3-Embedding-8B trained further on tool retrieval
(yasinyaman/toolrank-emb-8b), served by vLLM (--emb-url, by default
http://127.0.0.1:8091/v1). On a GPU host, Docker runs both:
cd deploy/docker && cp .env.example .env # set TOOLRANK_API_KEY
docker compose run --rm toolrank ingest mcp --config /config/toolrank.json --out /data
docker compose up -dThe quick start connects Claude Code, Claude Desktop and REST clients.
Related MCP server: fastmcp-gateway
Results
Retrieval quality on three benchmarks, every number from toolrank eval under ToolRet's protocol
(the reports are in docs/results/;
scripts/readme_table.py checks each one before it prints it).
Retriever | ToolRet NDCG@10 | ToolRet NDCG@10 cat-macro | LiveMCPBench Recall@5 | MCP-Zero top-1 |
BM25, without instruction | 29.01 | 22.24 | 31.68 | 80.44 |
BM25, with instruction | 39.27 | 36.41 | 22.92 | 45.63 |
Qwen3-Embedding-8B | 51.11 | 46.54 | 50.82 | 78.19 |
Qwen3-Embedding-8B + toolrank heads v0.1 | 54.03 | 47.13 | 53.03 | 79.87 |
Qwen3-Embedding-8B in FP8 + toolrank heads v0.1 | 53.94 | 47.27 | 53.48 | 79.51 |
toolrank backbone v0.2 (Qwen3-Embedding-8B + LoRA) | 58.90 | 54.36 | 52.06 | 88.57 |
toolrank backbone v0.2 in FP8 (the default) | 59.02 | 54.53 | 52.06 | 87.71 |
NV-Embed-v1 (ToolRet paper) | — | 42.71 | — | — |
gte-Qwen2-1.5B-instruct (ToolRet paper) | — | 45.96 | — | — |
StackOne v2, a fine-tuned 109M BGE-base (StackOne) | — | 54.40 | — | — |
ToolRet: 7,961 queries over 44,453 tools, top 100 over the whole corpus. NDCG@10 is the micro-average of the paper's released code; cat-macro is the paper's own aggregation (the mean of the web, code and customized categories) and the only column with published numbers. Our BM25 reproduces the paper's BM25s within 0.1 (22.24 / 36.41 against 22.32 / 36.46).
LiveMCPBench (94 queries, 525 tools) and MCP-Zero (2,792 tools): the tool text includes the MCP server's name (
toolrank data server-names). One LiveMCPBench query is about one point. MCP-Zero ships no queries: ours were written by Qwen3-8B, one per tool (toolrank data pull mcp-zero), so its column does not compare with the MCP-Zero paper. Top-1 is Precision@1.With instruction, each query carries its task's instruction (ToolRet) or a generic one (the MCP sets), as the embedding model is served; the generic instruction costs BM25 on the MCP sets. BM25 is bm25s without stemming, the paper's setting.
The heads (29.9M parameters,
docs/heads/MODEL_CARD.md) were trained on ToolRet's training pairs, so ToolRet is in-domain for them and the MCP sets are not. On MCP-Zero, BM25 without instruction still wins at top-1: each generated query opens with aserver:line that usually names the server, and exact matching rewards that.The toolrank backbone v0.2 (
docs/backbone/MODEL_CARD.md) is Qwen3-Embedding-8B with a LoRA trained on 20,000 of ToolRet's training pairs, served without heads (the v0.1 heads cost it 1–2 points). Its checkpoint was picked on MCP-Zero, so that column is its selection set; ToolRet is in-domain, LiveMCPBench is held out. Beyond these sets, on generated requests over a GitHub + Stripe catalogue it gains 12–20 NDCG@10 points on tasks that need two or three tools and ties with the base model on requests for one tool (the model card has the numbers).FP8: the bf16 weights quantized as vLLM loads them (
--quantization fp8). Every column is within a point of bf16, at about half the weight memory and batch-1 latency.Reproduce:
bash scripts/readme_results.shwhere the backbone is served (reports indocs/results/;EMB_URL,EMB_MODEL,TAGandROWSselect another endpoint and rows), thenuv run python scripts/readme_table.py --write.
What's inside
Ingestion of MCP servers (stdio and streamable HTTP) and OpenAPI 3.x specs; a re-run syncs only what changed. Guide
Search and serve: adaptive K, a persistent vector index (numpy, FAISS HNSW or pgvector), an MCP proxy with two tools, a REST API, API keys and a usage log. Guide
Agent platforms: toolrank as Claude's (
tool_reference) and OpenAI's (client-sidetool_search) tool search. GuideFrameworks: LangGraph (langgraph-bigtool), LlamaIndex agents and the LiteLLM proxy. Guide
Fine-tuning: heads trained on your own request-to-tool pairs, the epoch picked on a dev set. Guide
Benchmarks: ToolRet, LiveMCPBench and MCP-Zero with BM25, dense and head scorers (and CLM, for comparison). Benchmarks
Docker: the
toolrankimage for amd64 and arm64, compose files with vLLM, and a Dockerfile that puts vLLM and toolrank in one container. Guide
Why
Tool definitions are expensive context. Anthropic measured 58 tools at about 55K tokens per request and reports tool-selection accuracy falling past 30 to 50 tools. RAG-MCP lifted selection accuracy from 13.6% to 43.1% on a large MCP set by retrieving tools first.
Hosted tool searches are tied to one model provider or cloud, and most are lexical. toolrank is model-agnostic and runs on your own hardware.
Small heads on a frozen embedding model (29.9M parameters for the request and the tool side together). You embed your tools once and rank with one dot product. The heads start as the identity, so heads fine-tuned on your data start from the base model's quality, not below it.
Feedback
Tried it? Tell us how it went: what you set up, what worked and what did not. Questions and ideas go to Discussions.
Contributing
Set-up, tests and conventions are in
CONTRIBUTING.md. The weekly
reports behind every number, in Turkish, are in
docs/reports/.
License
Apache-2.0 (LICENSE). NOTICE credits what toolrank takes from CLM (the head
architecture), ToolRet (its task metadata) and MCP-Zero (its query prompts);
THIRD_PARTY_NOTICES.md lists the dependencies, models and container
images. Contributions are welcome: see CONTRIBUTING.md and the
code of conduct; report vulnerabilities as SECURITY.md says.
This server cannot be deployed
Maintenance
Related MCP Connectors
The OpenRouter for tools. One MCP connection gives any AI agent 254 hosted tools, pay per call.
Search, vet & assemble MCP servers from your agent: verified tools, risk labels, and trust scores.
The Remote MCP server acts as a standardized bridge between LLM applications (like Claude, ChatGPT, and Cursor) and external services, enabling AI agents to access external tools and resources. Its primary capability is providing a centralized search tool to discover other MCP servers and their respective tools. Unlike local implementations, it runs remotely with OAuth authentication and permission controls for security.
Pay-per-use tool marketplace for AI agents. Search, price-check, and call APIs via MCP.
Related MCP Servers
- AlicenseNot gradedqualityNot gradedmaintenanceA drop-in MCP proxy that aggregates multiple backend servers into two meta-tools for efficient tool discovery and execution. It enables AI clients to access hundreds of tools while minimizing context window usage through searchable indexing.1 npm-
- AlicenseNot gradedqualityAmaintenanceAggregates tools from multiple upstream MCP servers and exposes them through 4 meta-tools, enabling LLMs to discover and use hundreds of tools without loading all schemas upfront.2Apache 2.0
- AlicenseNot gradedqualityDmaintenanceActs as a proxy/router for multiple downstream MCP servers, exposing only meta-tools to the host to reduce token usage, enabling efficient search and invocation of tools from a fleet of servers.8 npmMIT
- AlicenseNot gradedqualityAmaintenanceEnables agents to use many MCP servers without context bloat by exposing meta-tools (search, load, call, run_code) that reduce token usage via progressive disclosure and result trimming.12 npmMIT