jev-router
OfficialFronts the Atlassian MCP server through jev-router, routing requests to its Jira and Confluence tools so agents can work with issues and docs without loading every Atlassian tool schema.
Routes requests to the Confluence MCP server's tools for pages, spaces, and content search, adding its schemas to the LLM context only when Confluence content is relevant.
Fronts the Datadog MCP server, routing requests to its monitoring, logs, metrics, and incident tools when Datadog is judged relevant to the request.
Fronts the git MCP server, making its repository operations (status, diff, commit, log, branch) available through jev-router when a request requires local version-control actions.
Fronts the GitHub MCP server, routing requests to its repositories, issues, pull requests, and Actions tools; the example config shows it added as a remote (HTTP) MCP server behind the gateway.
Routes requests to the Google Maps MCP server's tools for places, geocoding, and directions when location-based capabilities are needed.
Fronts the Grafana MCP server, routing requests to its dashboards, alerts, and on-call/user tools (e.g. get_current_oncall_users) only when Grafana is selected.
Routes requests to the Jira MCP server's tools (issues, projects, boards) through the gateway, exposing only Jira's schemas to the LLM when a request needs them.
Fronts the Kubernetes MCP server, routing requests to kubectl-style tools such as pod log retrieval, resource inspection, and cluster operations during incidents.
Routes requests to the MongoDB MCP server's tools for database, collection, and document operations when a request involves MongoDB data.
Fronts the Notion MCP server, routing requests to its pages, databases, and content tools when Notion content or updates are required.
Fronts the PagerDuty MCP server, routing requests to its incident, service, and on-call tools (e.g. list_oncalls) when incident response is needed.
Fronts the Prometheus MCP server (shipped as a mock in the repo and available as a real server), routing requests to PromQL queries, metric names/metadata, and scrape target tools.
Fronts the Sentry MCP server, routing requests to its issue search and error-tracking tools, e.g. searching issues when payments throw 500s.
Routes requests to the Slack MCP server's tools for channels and messages, e.g. posting a summary to an incidents channel.
Fronts the Supabase MCP server, routing requests to its project, database, and table tools when Supabase resources are involved.
Routes requests to the Terraform MCP server's tools for infrastructure-as-code providers, modules, and registry lookups when IaC assistance is needed.
Fronts the VictoriaMetrics MCP server, routing requests to its metrics query and metadata tools when VictoriaMetrics time-series data is relevant.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@jev-routerPayments is throwing 500s, check Sentry and the pod logs"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
jev-mcp-router
Let a decision model pick your agent's MCP tools, not the LLM.
An MCP gateway that sits between your AI agent and your MCP servers. For each request, Jev (TypeSafe AI's "System One" decision model) answers one yes/no question per server and per tool, and the LLM only ever sees the handful of tools Jev said yes to.

Why
With MCP, every tool from every connected server is put into the LLM's prompt on every turn. Connect a few real servers (GitHub, Grafana, PagerDuty, Jira, Kubernetes...) and that is hundreds of tools and ~200k tokens before the user's request is even read. The LLM then spends expensive reasoning deciding which 2 or 3 of them to call.
Choosing tools is a decision, not a generation task. Jev doesn't generate text: it takes a state plus typed questions and returns an answer per question, for about $0.04 per million input tokens.
flowchart LR
subgraph without["Without Jev"]
direction LR
U1([Request]) --> L1["LLM<br/>reads all 649 tool schemas<br/>~201k tokens / turn"]
L1 --> T1[[Tools]]
end
subgraph with["With Jev"]
direction LR
U2([Request]) --> J["Jev<br/>yes / no per server and tool<br/>~$0.0003, ~1 s"]
J -->|"~3 YES tools"| L2["LLM<br/>reads ~1k tokens"]
L2 --> T2[[Tools]]
endRelated MCP server: MCP Gateway
Results
649 real tools from 25 MCP servers, 111 hand-labeled requests (details):
All needed tools picked | Cost to choose tools | Time to choose tools | |
Jev (this repo) | 95% | $0.0003 | ~1 s idle, 2.0 s under load |
Claude Sonnet 5, full catalog in prompt (cached) | 99% | $0.0101 | 3.2 s |
Claude Haiku 4.5, same | 93% | $0.0031 | 1.8 s |
GPT-5 nano, same | 86% | $0.0003 | 3.2 s |
Embedding search, top 20 | 85% | ~$0 | <0.01 s |
Sonnet is more accurate. Jev gets close at 1/32 of the cost, and the LLM that does the actual work sees ~1k tokens of tool schemas instead of ~201k.
How it works
sequenceDiagram
autonumber
participant A as Agent / LLM
participant G as jev-router (MCP gateway)
participant J as Jev
participant S as Your MCP servers
Note over G,S: on startup: connect to every server, list its tools
A->>G: jev_route("Payments is throwing 500s, check Sentry and the pod logs")
G->>J: Step 1: one yes/no per server (25 questions)
J-->>G: YES: sentry, kubernetes
G->>J: Step 2: one yes/no per tool of the YES servers
J-->>G: YES: sentry search_issues, kubernetes kubectl_logs, ...
G-->>A: schemas of the YES tools only (+ matching playbooks)
A->>G: jev_call("sentry__search_issues", {...})
G->>S: tools/call
S-->>G: result
G-->>A: resultSmall catalogs (up to 40 tools and skills) are routed in a single Jev request.
Large catalogs are routed in two steps: servers first (kept at p >= 0.3, tuned so a needed server is rarely dropped), then the tools of the kept servers. If a server is clearly needed but none of its tools clears 0.5, its top-ranked tool is taken anyway.
Skills / playbooks (
skills/<name>/SKILL.md) are routed the same way. The chosen playbook's text goes straight into the response, so the LLM doesn't spend a turn loading it.Jev returns a probability per question; the gateway turns it into a plain YES/NO.
Your agent connects to one MCP server, jev-router, which exposes two tools:
Tool | What it does |
| Jev's decision: the input schemas of the YES tools and the text of the YES playbooks |
| Calls the tool on the underlying MCP server |
Quickstart
Requirements: Python 3.11+, uv, and an OpenRouter API key (Jev is served through OpenRouter; a direct TypeSafe key also works).
git clone https://github.com/ini8labs/jev-mcp-router.git
cd jev-mcp-router
uv sync
cp .env.example .env # then put your OPENROUTER_API_KEY in .envThe repo ships with two mock MCP servers (Prometheus and VictoriaLogs, real tool names, canned data for one incident), so everything below works out of the box:
uv run python -m jevmcp catalog # what Jev decides over
uv run python -m jevmcp route "why is checkout slow?" # Jev's YES/NO only, no LLM
uv run python -m jevmcp run "why is checkout slow?" --compare
# end to end, vs an LLM that sees every tool
uv run python -m jevmcp eval # routing accuracy on evals/queries.jsonlNo key yet? JEV_FAKE=1 LLM_FAKE=1 runs offline with keyword stand-ins. Their
decisions are not Jev's; it's only for checking the wiring.
Use it from Claude Code or Claude Desktop
Claude Code, from inside the cloned repo: .mcp.json already registers the
gateway, so just start claude there. Or add it from anywhere:
claude mcp add jev-router -- uv run --directory /path/to/jev-mcp-router python -m jevmcp serveClaude Desktop, in claude_desktop_config.json:
{
"mcpServers": {
"jev-router": {
"command": "uv",
"args": ["run", "--directory", "/path/to/jev-mcp-router", "python", "-m", "jevmcp", "serve"]
}
}
}Then remove the servers that jev-router now fronts from the client's own config,
so their schemas stop being loaded directly.
Connect your own MCP servers
Edit mcp_servers.json (same shape as Claude Desktop's mcpServers, stdio or HTTP),
or point JEV_MCP_SERVERS at another file:
{
"mcpServers": {
"prometheus": {
"description": "Prometheus: PromQL metric queries, metric names and metadata, scrape targets",
"command": "docker",
"args": ["run", "-i", "--rm", "-e", "PROMETHEUS_URL", "ghcr.io/pab1it0/prometheus-mcp-server:latest"],
"env": { "PROMETHEUS_URL": "http://host.docker.internal:9090" }
},
"github": { "description": "GitHub: repos, issues, PRs, Actions", "url": "https://example.com/mcp" }
}
}descriptionis what Jev reads when choosing servers. One plain line about what the server is for works best. Without it, the server's tool names are used.tool_hints.json(optional) replaces a tool's description for routing only, keyed byserver__tool. Use it for tools whose descriptions are boilerplate, and to mark catch-all tools ("run any kubectl command") as a last resort. The LLM still receives the original schema.mcp_servers.real.jsonhas working entries for the real Prometheus and VictoriaLogs MCP servers.
Benchmark
bench/ reproduces the numbers above.
Catalog:
bench/catalog/holds the real tool lists of 25 MCP servers, captured by starting each one and callingtools/list(bench/collect.py): GitHub, Grafana, PagerDuty, Atlassian (Jira + Confluence), Kubernetes, Datadog, Sentry, AWS CloudWatch, Prometheus, VictoriaLogs, VictoriaMetrics, Supabase, MongoDB, Notion, Slack, Playwright, Terraform, Postgres, filesystem, git, fetch, time, memory, Context7, Google Maps.Requests:
bench/requests.jsonl, 111 requests, each labeled with the tools it needs. A tool group like["pagerduty__list_oncalls", "grafana__get_current_oncall_users"]means any one of them is acceptable. 10 requests span several servers; 3 need no tool.Scoring: a request passes when every needed tool was selected. Extra tools are counted separately, since they cost schema tokens but don't break the answer.
LLM routers get the same catalog (descriptions truncated to 300 characters, as Jev gets) as a cached system prompt and reply with a JSON list of tool ids.
uv run python bench/run.py report # re-score saved results, no API spend
uv run python bench/run.py router # the shipped router end to end (~$0.04)
uv run python bench/run.py llm anthropic/claude-sonnet-5 # an LLM router (~$1.10)
uv run --group bench python bench/run.py embed # local embedding baseline
uv run python bench/latency.py # Jev latency on an idle connectionTo start all 25 servers live behind the gateway (dummy credentials: routing works, tool calls fail), build the Go-based ones first:
bench/install_servers.sh # needs Go and git; npx/uvx servers fetch on demand
JEV_MCP_SERVERS=bench/mcp_servers.bench.json uv run python -m jevmcp route \
"Payments is throwing 500s. Check Sentry, look at the pod logs in Kubernetes, post a summary to #incidents"Caveats. One run; Jev's answers vary by a few points between runs. The requests
and labels were written by the same person who tuned the router, so expect somewhat
lower numbers on your own traffic. tool_hints.json was written after seeing which
tools failed; its +2 points are within run-to-run noise.
Limitations
Jev picks tools, not arguments. It can't write a PromQL query or an email body, so an LLM still fills in arguments. It just does so with ~3 tool schemas instead of 649.
No fallback yet. If Jev misses a needed tool (about 1 request in 20 in the benchmark), the LLM has no way to ask for more. A
request_more_toolsescape hatch is the obvious next step.Where Jev misses: tools with almost no description, catch-all tools winning over the specific one, and the second action in two-part requests ("create a branch and switch to it").
Small setups don't need this. Under ~30 tools with prompt caching, the saving is fractions of a cent per request.
Configuration
Variable | Default | |
| Used for Jev and for the LLM in | |
| Alternative to OpenRouter for Jev (set | |
|
| |
|
| |
|
| Any OpenRouter model id |
|
| Server config for the CLI and the gateway |
|
|
|
Layout
jevmcp/
jev.py Jev client: POST /v1/systemone, yes/no questions -> YES/NO, offline fake
router.py one-step / two-step routing, fallback
catalog.py MCP connections (Hub), tool calls, skills and hints loading
gateway.py the jev-router MCP server: jev_route, jev_call
agent.py end-to-end runs: Jev path vs LLM-only baseline, per-step cost
llm.py OpenRouter chat for arguments and answers, live pricing
cli.py catalog / route / run / eval / serve
mock_servers/ Prometheus + VictoriaLogs mocks for the quickstart
skills/ example SRE playbooks
evals/ 17 labeled requests for the mock setup
bench/ 649-tool catalog, 111 labeled requests, benchmark runner and results
docs/ the benchmark image and its HTML source
tool_hints.json routing-only descriptions for badly described toolsLicense
MIT. Jev is a model by TypeSafe AI; this project is not affiliated with TypeSafe AI and you need your own API access to use it.
Available Tools
2 toolsjev_callB
Call a tool chosen by jev_route. tool is the id from jev_route (e.g. prometheus__execute_range_query).
| Name | Required | Description | Default |
|---|---|---|---|
| tool | Yes | ||
| arguments | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden, and it discloses almost nothing: it does not say that side effects depend entirely on the target tool, that arguments must match the target tool's schema, how errors are surfaced, or whether there is any validation. It adds only the example id format, which is a structural hint rather than behavioral transparency.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, no filler, with the core action front-loaded and the key parameter hint immediately after. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be described. What is missing for a dynamic proxy tool is the `arguments` contract and error/validation behavior when the chosen tool fails; the description is adequate as wiring but incomplete for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% for both parameters, so the description must compensate. It does explain the `tool` parameter well ('the id from jev_route', with a concrete example like prometheus__execute_range_query), but says nothing about `arguments`, which is a free-form object whose contract (keys must mirror the target tool's schema) is the crux of using this tool correctly.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb ('Call') plus the resource being called, and explicitly ties itself to the sibling by stating the tool is 'chosen by jev_route'. That differentiates it from jev_route (which routes/selects) without opening either schema, though it doesn't spell out what 'calling' entails for the agent.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'the id from jev_route' implies jev_call must be invoked after jev_route, which is useful sequencing guidance. However, there is no explicit when-to-use/when-not, no statement that jev_route should not be used directly for execution, and no mention of what to do if the id is unknown.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
jev_routeB
Decide which tools (across all connected MCP servers) and playbooks a request needs. Returns only the YES tools with their input schemas, and the instructions of each YES playbook.
| Name | Required | Description | Default |
|---|---|---|---|
| request | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It discloses return behavior ('Returns only the YES tools with their input schemas, and the instructions of each YES playbook'), which is useful. However, it omits other behavioral traits like whether the operation is read-only, side effects, or permission requirements.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the core purpose, and no wasted words. Structure is appropriate for a tool description.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so the description needn't explain return values in full, yet it does so helpfully. However, it lacks critical context for a routing tool: when to invoke it relative to sibling jev_call, and how to formulate the 'request' parameter. This makes it only moderately complete for an agent's needs.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There is one parameter ('request') with 0% schema description coverage. The description says 'a request needs' but adds no meaning beyond the schema—no format, examples, or expected content. It fails to compensate for the lack of schema documentation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Decide') and resource ('which tools and playbooks a request needs'), making the purpose clear. However, it does not distinguish itself from the sibling tool jev_call, so an agent cannot infer the division of labor from the description alone.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this tool versus alternatives. The sibling jev_call is never mentioned, and there is no explicit context about when routing should occur (e.g., before calling tools). Usage is only implied by the description of what the tool does.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
2 tool updates
v0.1.0- First observed
jev_call - First observed
jev_route
TDQS
Scored across 2 tools
jev_route handles selection/decision while jev_call handles execution. The two tools have clearly distinct roles with no overlap, so agents cannot misselect.
Both tools use the jev_ prefix and a concise verb (route, call) in snake_case. The pattern is consistent and predictable.
Two tools are exactly what a router needs: one to decide and one to invoke. No bloat or missing operation, so the count is well scoped.
The router covers the full flow of selecting relevant tools/playbooks and then calling a chosen tool. No obvious lifecycle gaps for its orchestration purpose.
Maintenance
Related MCP Connectors
The OpenRouter for tools. One MCP connection gives any AI agent 254 hosted tools, pay per call.
Zero-setup MCP gateway securely connecting AI to your tools with authentication and workflows
Connect MCP clients to 2,000+ AI models without managing provider API keys.
The Remote MCP server acts as a standardized bridge between LLM applications (like Claude, ChatGPT, and Cursor) and external services, enabling AI agents to access external tools and resources. Its primary capability is providing a centralized search tool to discover other MCP servers and their respective tools. Unlike local implementations, it runs remotely with OAuth authentication and permission controls for security.
Related MCP Servers
- AlicenseNot gradedqualityFmaintenanceA universal gateway that aggregates multiple MCP servers into a single interface while providing advanced token optimization, result filtering, and automated summarization. It enables efficient management of large tool catalogs and reduces context usage by up to 95% for major AI clients.25 npm16MIT
- AlicenseNot gradedqualityCmaintenanceAggregates multiple Model Context Protocol servers into a single gateway to provide unified search, description, and execution of tools. It reduces context limit issues by dynamically fetching specific tool schemas only when needed rather than loading all available tools at once.14 npm22MIT
- AlicenseNot gradedqualityCmaintenanceA gateway that aggregates multiple MCP servers into a single endpoint, namespacing their tools and forwarding calls, so an agent connects to one MCP to access the entire stack.MIT
- FlicenseNot gradedqualityBmaintenanceSingle gateway that aggregates dozens of upstream MCP servers, enabling AI clients to connect once and access all tools.-