Skip to main content
Glama
gaharivatsa

ToolBox

by gaharivatsa
README.md
# ToolBox

One MCP endpoint in front of all your MCP servers. For each prompt, [Jev](https://docs.typesafe.ai)
(TypeSafe's decision model) picks the right server and tool, and a risk gate decides whether the
call runs, asks you first, or is blocked. Calls run over warm connections, and each decision is logged.

**Why:** connect an agent to many MCP servers and every tool's schema lands in its context. That's
tens of thousands of tokens, and near-identical tools (the same Grafana tools for AWS and for GCP)
make it pick the wrong one. ToolBox shows the agent three tools instead:

| Tool | What it does |
|---|---|
| `find_tools` | One Jev call routes a request to servers and tools. Returns confidence, effect (read/write), gate and input schema. |
| `call_tool` | Runs a tool through the risk gate. Reads run, writes ask you first, blocked tools never run. |
| `browse_tools` | Browse the catalog when routing misses something. |

## Install

Needs Python 3.11+ and [uv](https://docs.astral.sh/uv/).

```sh
git clone https://github.com/gaharivatsa/toolbox && cd toolbox
uv sync
uv run toolbox discover      # connect to the servers in the config and cache their tools
uv run toolbox doctor        # config, key, Jev latency, and one live routing call
```

The shipped `toolbox.example.toml` connects three servers:

| Server | What it is | Needs |
|---|---|---|
| [Exa](https://exa.ai) | Web search | Nothing |
| [Zerodha Kite](https://kite.trade) | Brokerage account | A Kite login. It is read-only here: ToolBox never trades. |
| [screener-mcp](https://github.com/gaharivatsa/screener-mcp) | Indian stock fundamentals | A clone at `~/screener-mcp` |

To change anything, copy the example to `toolbox.toml`, which git ignores. ToolBox uses that
file when it exists.

### A Jev key

Routing needs a key for Jev. Without one, ToolBox still works, but it falls back to keyword search.
There are two ways to get one:

- **OpenRouter** (the default in the example config): create a regular API key at
  [openrouter.ai/settings/keys](https://openrouter.ai/settings/keys). A management key won't work;
  it returns 401 "User not found".
- **TypeSafe directly:** set `provider = "typesafe"` under `[jev]` and use a key from
  [console.typesafe.ai](https://console.typesafe.ai).

Put the key in `~/.toolbox/secrets.env` (and make that file private with `chmod 600`) or export it:

```sh
OPENROUTER_API_KEY=sk-or-v1-...
```

Every ToolBox process reads that file, including the MCP server that Claude Code starts. An exported
variable overrides the file. A routing call costs about $0.0001.

## Use it

```sh
uv run toolbox route "Compare HDFC Bank's P/E with its peers and show the live price"
uv run toolbox call screener/get_peers '{"symbol": "HDFCBANK"}'
uv run toolbox ask "Is HDFC Bank cheaper than its peers on P/E?"
uv run toolbox log           # every routing decision and call, with timings
```

`toolbox ask` is the "send a prompt, get an answer" path. It routes first, then runs Claude Code
headless (`claude -p`, with your existing login), with ToolBox as its only MCP server.

### From Claude Code

This repo comes set up. `.mcp.json` registers the server, and `.claude/settings.json` adds a hook.
The hook routes each prompt before Claude's first turn and adds the picked tools to the context, so
Claude calls them directly instead of spending a turn on `find_tools`. Run `claude` inside the repo
and approve the `toolbox` server.

To use it everywhere, register it at user scope:

```sh
claude mcp add toolbox -s user -- uv run --project /path/to/toolbox toolbox serve
```

Then add the hook to `~/.claude/settings.json`, using the absolute path `/path/to/toolbox/.venv/bin/toolbox-hook`.
Without a Jev key the hook stays silent, because keyword guesses aren't worth adding to your prompt.

### From other MCP clients

Cursor, Codex, or anything else that speaks MCP can use either of two options:

- **stdio:** point the client at `uv run --project /path/to/toolbox toolbox serve`.
- **HTTP daemon:** run `uv run toolbox serve --http` and connect to `http://127.0.0.1:8765/mcp`. Send the
  header `Authorization: Bearer <token>`, where the token is in `~/.toolbox/daemon.token` (mode 0600).
  While the daemon runs, the hook and `ask` use it automatically. That keeps one warm Jev connection and
  warm server sessions, so Kite stays logged in between Claude sessions.

## How routing works

Each prompt costs one Jev request, sent as a speculative fan-out. Jev answers every question in parallel:

- `needs_tools` (yes/no): does the request need any tool at all?
- `server_i` (yes/no, one per server): the answers are independent, so one prompt can select several
  servers. Each question asks whether this service is needed "for a step none of the other services
  can do as well". The service list sits in the shared state, which Jev reads once per request.
  Without that wording, web search said yes to almost everything: 7 extra servers on 10 test prompts,
  versus 0 with it.
- `tool_i_j` (choice, one per server): picks among that server's tools, plus `none_of_these`. Servers
  with more than 254 tools are split into chunks.
- `risk` (score, 0–3): how consequential the request is.

Code keeps the tool answers only for servers where P(needed) ≥ 0.5. A top tool counts as confident
at ≥ 0.5. Runner-ups at ≥ 0.15 are offered too, up to 3 per server and 8 in total. If a server's
picks are uncertain, it returns a shortlist. If tools are needed but no server matches, it returns
"unclear" (ask the user).

If there is no key, or Jev is down or takes longer than 3 s, routing falls back to BM25 keyword
search, so ToolBox is never much slower than having no router.

### Measured with live Jev (via OpenRouter, September 2026)

| Check | Result |
|---|---|
| One routing call | 350–520 ms on a warm connection (one outlier at 1.1 s), ~2,700 tokens, ~$0.0001 |
| 15 test prompts | Exact server set 14/15, no needed server missed. 5/5 on the 5 prompts that played no part in tuning. |
| Keyword fallback, same prompts | Top pick usually wrong: "latest news" → `kite/get_ltp`, "capital of France" → `screener/run_screen` |
| "Buy 10 shares of Infosys" | `kite/place_order` at 0.98 with request risk 3 of 3, then blocked by the gate |
| "Thanks" or "capital of France" | `no_tools`, so nothing is injected |

The prompts were hand-written. Grow the set from your own traffic in `toolbox log` before trusting
the thresholds.

## The risk gate

Evidence is checked strongest first:

1. The server's `block`, `confirm` and `allow` lists.
2. The server's policy: `read_only`, `ask_writes` (the default) or `trusted`.
3. The tool's verb, refined by MCP annotations when they carry real information.
4. Jev, only for tools nothing else could classify. It sees the actual arguments and must be at
   least 0.85 sure the call only reads.

Hard rules always beat Jev.

- **Kite is `read_only`.** All 16 read tools run, and all 5 order tools (`place_order`,
  `modify_order`, …) are blocked. `login` is on the allow list. Kite marks every tool,
  `get_ltp` included, `readOnlyHint=false, destructiveHint=true`. Those are the MCP spec defaults,
  so ToolBox ignores them.
- **Confirmations use MCP elicitation.** Both protocol generations are supported: the back-channel
  prompt used before 2026, and the 2026-07-28 input-required round trip. The approval is bound with
  an HMAC to the exact tool and arguments.
- **A client that can't show a prompt never gets writes through.** No agent approves its own actions.

## Speed

Measured on one machine; your network will change these numbers.

| What | Result |
|---|---|
| Jev network round trip | ~0.25 s on a reused connection, ~0.85 s on a new one. The server and daemon keep one open. |
| Downstream sessions stay warm | A stdio server's first call took 951 ms, the next one 254 ms |
| Blocked call | ~1 ms, and the server is never contacted |
| Hook, live Jev, via the daemon | 0.5–0.6 s per prompt |
| Hook, live Jev, no daemon | ~1 s per prompt (Python start-up, SDK import and a new TLS connection) |
| Hook with nothing to do | 0.1–0.2 s |
| `toolbox ask`, two-server question | Route 0.47 s, agent 33 s (3 turns, $0.10). The agent's time dominates. |

## Files

| Path | What |
|---|---|
| `toolbox.example.toml` / `toolbox.toml` | Servers, policies, thresholds. No secrets: env values use `${VAR}`, paths may start with `~`. |
| `~/.toolbox/secrets.env` | API keys, loaded by every ToolBox process (keep it `chmod 600`) |
| `~/.toolbox/catalog.json` | Cached tool catalog. It is refreshed each time the server starts. |
| `~/.toolbox/audit.jsonl` | Every route and call, with timings. Argument values are off by default. |
| `~/.toolbox/logs/<server>.log` | stderr of each stdio server |

To add a server, add a `[[servers]]` block (stdio: `command`/`args`/`env`, or HTTP: `url`/`headers`)
and run `uv run toolbox discover <name>`.

## Privacy

Routing sends the prompt and the tools' names and descriptions to Jev (through OpenRouter or
directly to TypeSafe). For tools the gate can't classify, it also sends the call's arguments. Think
twice before putting servers for private or work data behind ToolBox.

## Tests

`uv run pytest` runs 35 tests. They cover the router, Jev client and risk gate, a fake stdio MCP
server behind the hub, and ToolBox's own MCP server with approvals under both protocol generations.

## Known gaps

- Kite needs a login per ToolBox session: call `kite/login` and open the link. The daemon keeps the session.
- Prompts from downstream servers (for example, screener asking for cookies) aren't forwarded to your client yet.
- The routing evaluation is small. Treat the thresholds as starting points.

## License

MIT. See [LICENSE](LICENSE).

TDQS

A4.3/5.0

Scored across 3 tools

Disambiguation4/5

find_tools and browse_tools both aid discovery, but descriptions clearly distinguish them: find_tools uses natural-language search, while browse_tools is a fallback catalog browser. call_tool is unambiguously for execution, so only minor overlap exists.

Naming Consistency4/5

All names use snake_case and follow a verb_noun pattern (find_tools, call_tool, browse_tools). The only inconsistency is call_tool using singular 'tool' while the others use plural 'tools', which is a minor deviation.

Tool Count5/5

Three tools form a minimal, well-scoped interface for a meta-server: search for tools, execute a tool, and browse the catalog when search fails. Each tool earns its place without redundancy.

Completeness5/5

The surface covers the full lifecycle of tool discovery and invocation: find_tools returns tool IDs with schemas, call_tool executes with arguments, and browse_tools provides fallback listing and detailed descriptions. No obvious operational gaps exist for this domain.

Maintenance

ActivityMaintained
ResponsivenessNo issues