ToolBox
Provides read-only integration with a Zerodha Kite brokerage account, allowing read tools such as live price and account queries while blocking all trading/order tools (e.g., place_order, modify_order).
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@ToolBoxCompare HDFC Bank's P/E with its peers and show the live price"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
ToolBox
One MCP endpoint in front of all your MCP servers. For each prompt, Jev (TypeSafe's decision model) picks the right server and tool, and a risk gate decides whether the call runs, asks you first, or is blocked. Calls run over warm connections, and each decision is logged.
Why: connect an agent to many MCP servers and every tool's schema lands in its context. That's tens of thousands of tokens, and near-identical tools (the same Grafana tools for AWS and for GCP) make it pick the wrong one. ToolBox shows the agent three tools instead:
Tool | What it does |
| One Jev call routes a request to servers and tools. Returns confidence, effect (read/write), gate and input schema. |
| Runs a tool through the risk gate. Reads run, writes ask you first, blocked tools never run. |
| Browse the catalog when routing misses something. |
Install
Needs Python 3.11+ and uv.
git clone https://github.com/gaharivatsa/toolbox && cd toolbox
uv sync
uv run toolbox discover # connect to the servers in the config and cache their tools
uv run toolbox doctor # config, key, Jev latency, and one live routing callThe shipped toolbox.example.toml connects three servers:
Server | What it is | Needs |
Web search | Nothing | |
Brokerage account | A Kite login. It is read-only here: ToolBox never trades. | |
Indian stock fundamentals | A clone at |
To change anything, copy the example to toolbox.toml, which git ignores. ToolBox uses that
file when it exists.
A Jev key
Routing needs a key for Jev. Without one, ToolBox still works, but it falls back to keyword search. There are two ways to get one:
OpenRouter (the default in the example config): create a regular API key at openrouter.ai/settings/keys. A management key won't work; it returns 401 "User not found".
TypeSafe directly: set
provider = "typesafe"under[jev]and use a key from console.typesafe.ai.
Put the key in ~/.toolbox/secrets.env (and make that file private with chmod 600) or export it:
OPENROUTER_API_KEY=sk-or-v1-...Every ToolBox process reads that file, including the MCP server that Claude Code starts. An exported variable overrides the file. A routing call costs about $0.0001.
Related MCP server: smart-router
Use it
uv run toolbox route "Compare HDFC Bank's P/E with its peers and show the live price"
uv run toolbox call screener/get_peers '{"symbol": "HDFCBANK"}'
uv run toolbox ask "Is HDFC Bank cheaper than its peers on P/E?"
uv run toolbox log # every routing decision and call, with timingstoolbox ask is the "send a prompt, get an answer" path. It routes first, then runs Claude Code
headless (claude -p, with your existing login), with ToolBox as its only MCP server.
From Claude Code
This repo comes set up. .mcp.json registers the server, and .claude/settings.json adds a hook.
The hook routes each prompt before Claude's first turn and adds the picked tools to the context, so
Claude calls them directly instead of spending a turn on find_tools. Run claude inside the repo
and approve the toolbox server.
To use it everywhere, register it at user scope:
claude mcp add toolbox -s user -- uv run --project /path/to/toolbox toolbox serveThen add the hook to ~/.claude/settings.json, using the absolute path /path/to/toolbox/.venv/bin/toolbox-hook.
Without a Jev key the hook stays silent, because keyword guesses aren't worth adding to your prompt.
From other MCP clients
Cursor, Codex, or anything else that speaks MCP can use either of two options:
stdio: point the client at
uv run --project /path/to/toolbox toolbox serve.HTTP daemon: run
uv run toolbox serve --httpand connect tohttp://127.0.0.1:8765/mcp. Send the headerAuthorization: Bearer <token>, where the token is in~/.toolbox/daemon.token(mode 0600). While the daemon runs, the hook andaskuse it automatically. That keeps one warm Jev connection and warm server sessions, so Kite stays logged in between Claude sessions.
How routing works
Each prompt costs one Jev request, sent as a speculative fan-out. Jev answers every question in parallel:
needs_tools(yes/no): does the request need any tool at all?server_i(yes/no, one per server): the answers are independent, so one prompt can select several servers. Each question asks whether this service is needed "for a step none of the other services can do as well". The service list sits in the shared state, which Jev reads once per request. Without that wording, web search said yes to almost everything: 7 extra servers on 10 test prompts, versus 0 with it.tool_i_j(choice, one per server): picks among that server's tools, plusnone_of_these. Servers with more than 254 tools are split into chunks.risk(score, 0–3): how consequential the request is.
Code keeps the tool answers only for servers where P(needed) ≥ 0.5. A top tool counts as confident at ≥ 0.5. Runner-ups at ≥ 0.15 are offered too, up to 3 per server and 8 in total. If a server's picks are uncertain, it returns a shortlist. If tools are needed but no server matches, it returns "unclear" (ask the user).
If there is no key, or Jev is down or takes longer than 3 s, routing falls back to BM25 keyword search, so ToolBox is never much slower than having no router.
Measured with live Jev (via OpenRouter, September 2026)
Check | Result |
One routing call | 350–520 ms on a warm connection (one outlier at 1.1 s), ~2,700 tokens, ~$0.0001 |
15 test prompts | Exact server set 14/15, no needed server missed. 5/5 on the 5 prompts that played no part in tuning. |
Keyword fallback, same prompts | Top pick usually wrong: "latest news" → |
"Buy 10 shares of Infosys" |
|
"Thanks" or "capital of France" |
|
The prompts were hand-written. Grow the set from your own traffic in toolbox log before trusting
the thresholds.
The risk gate
Evidence is checked strongest first:
The server's
block,confirmandallowlists.The server's policy:
read_only,ask_writes(the default) ortrusted.The tool's verb, refined by MCP annotations when they carry real information.
Jev, only for tools nothing else could classify. It sees the actual arguments and must be at least 0.85 sure the call only reads.
Hard rules always beat Jev.
Kite is
read_only. All 16 read tools run, and all 5 order tools (place_order,modify_order, …) are blocked.loginis on the allow list. Kite marks every tool,get_ltpincluded,readOnlyHint=false, destructiveHint=true. Those are the MCP spec defaults, so ToolBox ignores them.Confirmations use MCP elicitation. Both protocol generations are supported: the back-channel prompt used before 2026, and the 2026-07-28 input-required round trip. The approval is bound with an HMAC to the exact tool and arguments.
A client that can't show a prompt never gets writes through. No agent approves its own actions.
Speed
Measured on one machine; your network will change these numbers.
What | Result |
Jev network round trip | ~0.25 s on a reused connection, ~0.85 s on a new one. The server and daemon keep one open. |
Downstream sessions stay warm | A stdio server's first call took 951 ms, the next one 254 ms |
Blocked call | ~1 ms, and the server is never contacted |
Hook, live Jev, via the daemon | 0.5–0.6 s per prompt |
Hook, live Jev, no daemon | ~1 s per prompt (Python start-up, SDK import and a new TLS connection) |
Hook with nothing to do | 0.1–0.2 s |
| Route 0.47 s, agent 33 s (3 turns, $0.10). The agent's time dominates. |
Files
Path | What |
| Servers, policies, thresholds. No secrets: env values use |
| API keys, loaded by every ToolBox process (keep it |
| Cached tool catalog. It is refreshed each time the server starts. |
| Every route and call, with timings. Argument values are off by default. |
| stderr of each stdio server |
To add a server, add a [[servers]] block (stdio: command/args/env, or HTTP: url/headers)
and run uv run toolbox discover <name>.
Privacy
Routing sends the prompt and the tools' names and descriptions to Jev (through OpenRouter or directly to TypeSafe). For tools the gate can't classify, it also sends the call's arguments. Think twice before putting servers for private or work data behind ToolBox.
Tests
uv run pytest runs 35 tests. They cover the router, Jev client and risk gate, a fake stdio MCP
server behind the hub, and ToolBox's own MCP server with approvals under both protocol generations.
Known gaps
Kite needs a login per ToolBox session: call
kite/loginand open the link. The daemon keeps the session.Prompts from downstream servers (for example, screener asking for cookies) aren't forwarded to your client yet.
The routing evaluation is small. Treat the thresholds as starting points.
License
MIT. See LICENSE.
Available Tools
3 toolsbrowse_toolsARead-only
Look through the tool catalog when find_tools didn't surface what you need.
No arguments: list the servers. server="kite": list that server's tools. tool_id="kite/get_ltp": full description and input_schema for one tool.
| Name | Required | Description | Default |
|---|---|---|---|
| server | No | ||
| tool_id | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=false, so safety is covered; the description adds the hierarchical browse behavior (servers -> server tools -> single tool detail) which is real behavioral context. It omits any note on catalog size, pagination, or truncation, which keeps it short of a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short lines, front-loaded with the when-to-use fallback condition followed by a compact mode table. No filler sentences and nothing repeated from the schema or annotations.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need no explanation, and the description covers when to use it and what every parameter variant returns. The only gap is operational scale — nothing about how large the server/tool listings may be or whether results are paged.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and both parameters are undocumented strings, yet the description fully compensates: no args = server list, server="kite" = that server's tools, tool_id="kite/get_ltp" = full detail. It even conveys the expected tool_id format (server/tool) by example.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('look through the tool catalog') and immediately distinguishes itself from the sibling find_tools by positioning itself as the fallback. The three drill-down modes make the tool's scope unambiguous without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly names the triggering condition ('when find_tools didn't surface what you need') and the alternative it follows, which is exactly the routing an agent needs. It also enumerates the three invocation modes, so there is no ambiguity about how to use it once selected.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
call_toolA
Run a tool on a server connected to ToolBox.
tool_id comes from find_tools or browse_tools, e.g. "screener/get_peers". arguments must match that tool's input_schema. Actions that change things may ask the user to approve them first; blocked actions never run.
| Name | Required | Description | Default |
|---|---|---|---|
| tool_id | Yes | ||
| arguments | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does add real behavioral context: mutating actions 'may ask the user to approve them first' and 'blocked actions never run.' It omits error/failure semantics and result format, but the approval and blocking behaviors are meaningful disclosures beyond the schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four short sentences, front-loaded with the core action, then provenance, then argument constraints, then behavior. No padding; every sentence adds information an agent needs.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a generic dispatcher with no output schema and no annotations, the definition covers purpose, argument sourcing, and approval/blocking behavior well. It does not describe what a call returns or how errors surface, which is the main remaining gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, and it largely does: it explains the provenance and example format of tool_id ('screener/get_peers') and states that arguments must match the target tool's input_schema. This meaningfully clarifies both parameters despite the bare schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Run a tool on a server connected to ToolBox') and implicitly distinguishes itself from the discovery siblings by noting that tool_id comes from find_tools or browse_tools, making clear this is the execution step rather than discovery.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear context for when to use it: you must first obtain tool_id via find_tools or browse_tools, and arguments must match the target tool's input_schema. It does not state explicit exclusions or when-not-to-use conditions, but the prerequisite routing is clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
find_toolsARead-only
Find the best tools for a task across every server connected to ToolBox.
Pass the user's request, or the current sub-task, in plain language. Returns tool_ids with confidence, effect (read/write), gate (allow/confirm/block) and input_schema. Then run one with call_tool.
| Name | Required | Description | Default |
|---|---|---|---|
| context | No | ||
| request | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint and openWorldHint, but the description adds behavior they don't cover: results carry confidence, effect (read/write), and a gate value (allow/confirm/block), which tells the agent results may need confirmation before execution. It doesn't say what happens on zero matches or how many results are returned.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short front-loaded sentences: what it does, what to pass, what comes back and what to do next. Every sentence earns its place with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists so return values need not be spelled out, yet the description usefully summarizes the result shape and the call_tool handoff. The one real gap is the undocumented 'context' parameter and edge-case behavior when nothing matches.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must carry the burden and only partially does. It clarifies the required 'request' as plain-language text (user request or sub-task), but the optional 'context' parameter is never addressed or distinguished from 'request'.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Find the best tools for a task') with scope ('across every server connected to ToolBox'), and explicitly names the follow-up sibling call_tool. It does not differentiate itself from browse_tools, the other discovery sibling, so sibling differentiation is only partial.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Pass the user's request, or the current sub-task, in plain language' gives a clear condition and input expectation, and 'Then run one with call_tool' makes the intended workflow explicit. There is no exclusion telling the agent when browse_tools would be the better choice.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
3 tool updates
v0.1.0- First observed
browse_tools - First observed
call_tool - First observed
find_tools
TDQS
Scored across 3 tools
find_tools and browse_tools both aid discovery, but descriptions clearly distinguish them: find_tools uses natural-language search, while browse_tools is a fallback catalog browser. call_tool is unambiguously for execution, so only minor overlap exists.
All names use snake_case and follow a verb_noun pattern (find_tools, call_tool, browse_tools). The only inconsistency is call_tool using singular 'tool' while the others use plural 'tools', which is a minor deviation.
Three tools form a minimal, well-scoped interface for a meta-server: search for tools, execute a tool, and browse the catalog when search fails. Each tool earns its place without redundancy.
The surface covers the full lifecycle of tool discovery and invocation: find_tools returns tool IDs with schemas, call_tool executes with arguments, and browse_tools provides fallback listing and detailed descriptions. No obvious operational gaps exist for this domain.
Maintenance
Related MCP Connectors
Zero-secret MCP gateway for AI agents: risk-scored, audited calls with human-in-the-loop approval.
Find, vet, and run MCP tools through a secure audited gateway with prompt-injection risk scoring
Governed MCP gateway: one endpoint for your tools, with credential custody and audit log.
- gatewayOAuthai.sealgate
MCP gateway with runtime security policy, tool-call-level control, and audit of agent actions.
Related MCP Servers
- AlicenseNot gradedqualityCmaintenanceActs as a proxy/router for multiple downstream MCP servers, exposing only meta-tools to the host to reduce token usage, enabling efficient search and invocation of tools from a fleet of servers.10 npmMIT
- AlicenseAqualityBmaintenanceA single MCP server that fronts many downstream MCP servers and Skills, exposing only four tools (search, call_tool, use_skill, admin) so that an agent's context window only ever sees search results on demand.510MIT

AISIX AI Gatewayofficial
AlicenseAqualityAmaintenanceSelf-hosted MCP gateway that registers upstream MCP servers and fronts them behind one governed Streamable HTTP endpoint: per-tool access control by caller API key, guardrails over tool arguments and results, rate limits, and usage logs. The same Rust gateway also proxies LLM and A2A agent traffic.41167Apache 2.0- AlicenseAqualityAmaintenanceEnables AI harnesses to connect to a single MCP endpoint that routes to multiple downstream MCP servers, discovering and executing capabilities on demand while keeping tool schemas out of context.466 npmApache 2.0