Skip to main content
Glama
gaharivatsa

ToolBox

by gaharivatsa

ToolBox

One MCP endpoint in front of all your MCP servers. For each prompt, Jev (TypeSafe's decision model) picks the right server and tool, and a risk gate decides whether the call runs, asks you first, or is blocked. Calls run over warm connections, and each decision is logged.

Why: connect an agent to many MCP servers and every tool's schema lands in its context. That's tens of thousands of tokens, and near-identical tools (the same Grafana tools for AWS and for GCP) make it pick the wrong one. ToolBox shows the agent three tools instead:

Tool

What it does

find_tools

One Jev call routes a request to servers and tools. Returns confidence, effect (read/write), gate and input schema.

call_tool

Runs a tool through the risk gate. Reads run, writes ask you first, blocked tools never run.

browse_tools

Browse the catalog when routing misses something.

Install

Needs Python 3.11+ and uv.

git clone https://github.com/gaharivatsa/toolbox && cd toolbox
uv sync
uv run toolbox discover      # connect to the servers in the config and cache their tools
uv run toolbox doctor        # config, key, Jev latency, and one live routing call

The shipped toolbox.example.toml connects three servers:

Server

What it is

Needs

Exa

Web search

Nothing

Zerodha Kite

Brokerage account

A Kite login. It is read-only here: ToolBox never trades.

screener-mcp

Indian stock fundamentals

A clone at ~/screener-mcp

To change anything, copy the example to toolbox.toml, which git ignores. ToolBox uses that file when it exists.

A Jev key

Routing needs a key for Jev. Without one, ToolBox still works, but it falls back to keyword search. There are two ways to get one:

  • OpenRouter (the default in the example config): create a regular API key at openrouter.ai/settings/keys. A management key won't work; it returns 401 "User not found".

  • TypeSafe directly: set provider = "typesafe" under [jev] and use a key from console.typesafe.ai.

Put the key in ~/.toolbox/secrets.env (and make that file private with chmod 600) or export it:

OPENROUTER_API_KEY=sk-or-v1-...

Every ToolBox process reads that file, including the MCP server that Claude Code starts. An exported variable overrides the file. A routing call costs about $0.0001.

Related MCP server: smart-router

Use it

uv run toolbox route "Compare HDFC Bank's P/E with its peers and show the live price"
uv run toolbox call screener/get_peers '{"symbol": "HDFCBANK"}'
uv run toolbox ask "Is HDFC Bank cheaper than its peers on P/E?"
uv run toolbox log           # every routing decision and call, with timings

toolbox ask is the "send a prompt, get an answer" path. It routes first, then runs Claude Code headless (claude -p, with your existing login), with ToolBox as its only MCP server.

From Claude Code

This repo comes set up. .mcp.json registers the server, and .claude/settings.json adds a hook. The hook routes each prompt before Claude's first turn and adds the picked tools to the context, so Claude calls them directly instead of spending a turn on find_tools. Run claude inside the repo and approve the toolbox server.

To use it everywhere, register it at user scope:

claude mcp add toolbox -s user -- uv run --project /path/to/toolbox toolbox serve

Then add the hook to ~/.claude/settings.json, using the absolute path /path/to/toolbox/.venv/bin/toolbox-hook. Without a Jev key the hook stays silent, because keyword guesses aren't worth adding to your prompt.

From other MCP clients

Cursor, Codex, or anything else that speaks MCP can use either of two options:

  • stdio: point the client at uv run --project /path/to/toolbox toolbox serve.

  • HTTP daemon: run uv run toolbox serve --http and connect to http://127.0.0.1:8765/mcp. Send the header Authorization: Bearer <token>, where the token is in ~/.toolbox/daemon.token (mode 0600). While the daemon runs, the hook and ask use it automatically. That keeps one warm Jev connection and warm server sessions, so Kite stays logged in between Claude sessions.

How routing works

Each prompt costs one Jev request, sent as a speculative fan-out. Jev answers every question in parallel:

  • needs_tools (yes/no): does the request need any tool at all?

  • server_i (yes/no, one per server): the answers are independent, so one prompt can select several servers. Each question asks whether this service is needed "for a step none of the other services can do as well". The service list sits in the shared state, which Jev reads once per request. Without that wording, web search said yes to almost everything: 7 extra servers on 10 test prompts, versus 0 with it.

  • tool_i_j (choice, one per server): picks among that server's tools, plus none_of_these. Servers with more than 254 tools are split into chunks.

  • risk (score, 0–3): how consequential the request is.

Code keeps the tool answers only for servers where P(needed) ≥ 0.5. A top tool counts as confident at ≥ 0.5. Runner-ups at ≥ 0.15 are offered too, up to 3 per server and 8 in total. If a server's picks are uncertain, it returns a shortlist. If tools are needed but no server matches, it returns "unclear" (ask the user).

If there is no key, or Jev is down or takes longer than 3 s, routing falls back to BM25 keyword search, so ToolBox is never much slower than having no router.

Measured with live Jev (via OpenRouter, September 2026)

Check

Result

One routing call

350–520 ms on a warm connection (one outlier at 1.1 s), ~2,700 tokens, ~$0.0001

15 test prompts

Exact server set 14/15, no needed server missed. 5/5 on the 5 prompts that played no part in tuning.

Keyword fallback, same prompts

Top pick usually wrong: "latest news" → kite/get_ltp, "capital of France" → screener/run_screen

"Buy 10 shares of Infosys"

kite/place_order at 0.98 with request risk 3 of 3, then blocked by the gate

"Thanks" or "capital of France"

no_tools, so nothing is injected

The prompts were hand-written. Grow the set from your own traffic in toolbox log before trusting the thresholds.

The risk gate

Evidence is checked strongest first:

  1. The server's block, confirm and allow lists.

  2. The server's policy: read_only, ask_writes (the default) or trusted.

  3. The tool's verb, refined by MCP annotations when they carry real information.

  4. Jev, only for tools nothing else could classify. It sees the actual arguments and must be at least 0.85 sure the call only reads.

Hard rules always beat Jev.

  • Kite is read_only. All 16 read tools run, and all 5 order tools (place_order, modify_order, …) are blocked. login is on the allow list. Kite marks every tool, get_ltp included, readOnlyHint=false, destructiveHint=true. Those are the MCP spec defaults, so ToolBox ignores them.

  • Confirmations use MCP elicitation. Both protocol generations are supported: the back-channel prompt used before 2026, and the 2026-07-28 input-required round trip. The approval is bound with an HMAC to the exact tool and arguments.

  • A client that can't show a prompt never gets writes through. No agent approves its own actions.

Speed

Measured on one machine; your network will change these numbers.

What

Result

Jev network round trip

~0.25 s on a reused connection, ~0.85 s on a new one. The server and daemon keep one open.

Downstream sessions stay warm

A stdio server's first call took 951 ms, the next one 254 ms

Blocked call

~1 ms, and the server is never contacted

Hook, live Jev, via the daemon

0.5–0.6 s per prompt

Hook, live Jev, no daemon

~1 s per prompt (Python start-up, SDK import and a new TLS connection)

Hook with nothing to do

0.1–0.2 s

toolbox ask, two-server question

Route 0.47 s, agent 33 s (3 turns, $0.10). The agent's time dominates.

Files

Path

What

toolbox.example.toml / toolbox.toml

Servers, policies, thresholds. No secrets: env values use ${VAR}, paths may start with ~.

~/.toolbox/secrets.env

API keys, loaded by every ToolBox process (keep it chmod 600)

~/.toolbox/catalog.json

Cached tool catalog. It is refreshed each time the server starts.

~/.toolbox/audit.jsonl

Every route and call, with timings. Argument values are off by default.

~/.toolbox/logs/<server>.log

stderr of each stdio server

To add a server, add a [[servers]] block (stdio: command/args/env, or HTTP: url/headers) and run uv run toolbox discover <name>.

Privacy

Routing sends the prompt and the tools' names and descriptions to Jev (through OpenRouter or directly to TypeSafe). For tools the gate can't classify, it also sends the call's arguments. Think twice before putting servers for private or work data behind ToolBox.

Tests

uv run pytest runs 35 tests. They cover the router, Jev client and risk gate, a fake stdio MCP server behind the hub, and ToolBox's own MCP server with approvals under both protocol generations.

Known gaps

  • Kite needs a login per ToolBox session: call kite/login and open the link. The daemon keeps the session.

  • Prompts from downstream servers (for example, screener asking for cookies) aren't forwarded to your client yet.

  • The routing evaluation is small. Treat the thresholds as starting points.

License

MIT. See LICENSE.

Available Tools

3 tools
browse_toolsA
Read-only

Look through the tool catalog when find_tools didn't surface what you need.

No arguments: list the servers. server="kite": list that server's tools. tool_id="kite/get_ltp": full description and input_schema for one tool.

ParametersJSON Schema
NameRequiredDescriptionDefault
serverNo
tool_idNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=false, so safety is covered; the description adds the hierarchical browse behavior (servers -> server tools -> single tool detail) which is real behavioral context. It omits any note on catalog size, pagination, or truncation, which keeps it short of a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short lines, front-loaded with the when-to-use fallback condition followed by a compact mode table. No filler sentences and nothing repeated from the schema or annotations.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need no explanation, and the description covers when to use it and what every parameter variant returns. The only gap is operational scale — nothing about how large the server/tool listings may be or whether results are paged.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and both parameters are undocumented strings, yet the description fully compensates: no args = server list, server="kite" = that server's tools, tool_id="kite/get_ltp" = full detail. It even conveys the expected tool_id format (server/tool) by example.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('look through the tool catalog') and immediately distinguishes itself from the sibling find_tools by positioning itself as the fallback. The three drill-down modes make the tool's scope unambiguous without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly names the triggering condition ('when find_tools didn't surface what you need') and the alternative it follows, which is exactly the routing an agent needs. It also enumerates the three invocation modes, so there is no ambiguity about how to use it once selected.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

call_toolA

Run a tool on a server connected to ToolBox.

tool_id comes from find_tools or browse_tools, e.g. "screener/get_peers". arguments must match that tool's input_schema. Actions that change things may ask the user to approve them first; blocked actions never run.

ParametersJSON Schema
NameRequiredDescriptionDefault
tool_idYes
argumentsNo

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does add real behavioral context: mutating actions 'may ask the user to approve them first' and 'blocked actions never run.' It omits error/failure semantics and result format, but the approval and blocking behaviors are meaningful disclosures beyond the schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four short sentences, front-loaded with the core action, then provenance, then argument constraints, then behavior. No padding; every sentence adds information an agent needs.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a generic dispatcher with no output schema and no annotations, the definition covers purpose, argument sourcing, and approval/blocking behavior well. It does not describe what a call returns or how errors surface, which is the main remaining gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, and it largely does: it explains the provenance and example format of tool_id ('screener/get_peers') and states that arguments must match the target tool's input_schema. This meaningfully clarifies both parameters despite the bare schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Run a tool on a server connected to ToolBox') and implicitly distinguishes itself from the discovery siblings by noting that tool_id comes from find_tools or browse_tools, making clear this is the execution step rather than discovery.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear context for when to use it: you must first obtain tool_id via find_tools or browse_tools, and arguments must match the target tool's input_schema. It does not state explicit exclusions or when-not-to-use conditions, but the prerequisite routing is clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

find_toolsA
Read-only

Find the best tools for a task across every server connected to ToolBox.

Pass the user's request, or the current sub-task, in plain language. Returns tool_ids with confidence, effect (read/write), gate (allow/confirm/block) and input_schema. Then run one with call_tool.

ParametersJSON Schema
NameRequiredDescriptionDefault
contextNo
requestYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint and openWorldHint, but the description adds behavior they don't cover: results carry confidence, effect (read/write), and a gate value (allow/confirm/block), which tells the agent results may need confirmation before execution. It doesn't say what happens on zero matches or how many results are returned.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short front-loaded sentences: what it does, what to pass, what comes back and what to do next. Every sentence earns its place with no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values need not be spelled out, yet the description usefully summarizes the result shape and the call_tool handoff. The one real gap is the undocumented 'context' parameter and edge-case behavior when nothing matches.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must carry the burden and only partially does. It clarifies the required 'request' as plain-language text (user request or sub-task), but the optional 'context' parameter is never addressed or distinguished from 'request'.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Find the best tools for a task') with scope ('across every server connected to ToolBox'), and explicitly names the follow-up sibling call_tool. It does not differentiate itself from browse_tools, the other discovery sibling, so sibling differentiation is only partial.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Pass the user's request, or the current sub-task, in plain language' gives a clear condition and input expectation, and 'Then run one with call_tool' makes the intended workflow explicit. There is no exclusion telling the agent when browse_tools would be the better choice.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 3 tool updatesv0.1.0
    • First observedbrowse_tools
    • First observedcall_tool
    • First observedfind_tools

TDQS

A4.3/5.0

Scored across 3 tools

Disambiguation4/5

find_tools and browse_tools both aid discovery, but descriptions clearly distinguish them: find_tools uses natural-language search, while browse_tools is a fallback catalog browser. call_tool is unambiguously for execution, so only minor overlap exists.

Naming Consistency4/5

All names use snake_case and follow a verb_noun pattern (find_tools, call_tool, browse_tools). The only inconsistency is call_tool using singular 'tool' while the others use plural 'tools', which is a minor deviation.

Tool Count5/5

Three tools form a minimal, well-scoped interface for a meta-server: search for tools, execute a tool, and browse the catalog when search fails. Each tool earns its place without redundancy.

Completeness5/5

The surface covers the full lifecycle of tool discovery and invocation: find_tools returns tool IDs with schemas, call_tool executes with arguments, and browse_tools provides fallback listing and detailed descriptions. No obvious operational gaps exist for this domain.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    C
    maintenance
    Acts as a proxy/router for multiple downstream MCP servers, exposing only meta-tools to the host to reduce token usage, enabling efficient search and invocation of tools from a fleet of servers.
    10 npm
    MIT
  • A
    license
    A
    quality
    B
    maintenance
    A single MCP server that fronts many downstream MCP servers and Skills, exposing only four tools (search, call_tool, use_skill, admin) so that an agent's context window only ever sees search results on demand.
    5
    10
    MIT
  • A
    license
    A
    quality
    A
    maintenance
    Self-hosted MCP gateway that registers upstream MCP servers and fronts them behind one governed Streamable HTTP endpoint: per-tool access control by caller API key, guardrails over tool arguments and results, rate limits, and usage logs. The same Rust gateway also proxies LLM and A2A agent traffic.
    4
    1
    167
    Apache 2.0
  • A
    license
    A
    quality
    A
    maintenance
    Enables AI harnesses to connect to a single MCP endpoint that routes to multiple downstream MCP servers, discovering and executing capabilities on demand while keeping tool schemas out of context.
    4
    66 npm
    Apache 2.0