Skip to main content
Glama

Agent Router MCP

A local MCP decision router for Claude Code and Codex Herdr workers. It exposes two tools, route_next_stage and advise_next_action, through a stable stdio server. The decision backend is selectable between the upstream Laya library and the locally served Kev model. Existing Claude/Codex MCP registrations named laya continue to work; the server implementation and GitHub project are named agent-router-mcp.

Why this repository exists

The source began as a machine-local custom Laya MCP wrapper. Keeping that implementation available makes it possible to reproduce, inspect, and configure routing instead of depending on a private local-only copy. MIT covers this repository's wrapper code. Laya and Kev remain separate upstream dependencies with their own licenses and model terms.

Herdr is the orchestration layer: it owns panes and starts workers. The router only recommends a client, model, effort, and bounded next action. Its routing catalog and shared behavior live in danielcaze/agent-workflow; this repository contains the runnable MCP implementation.

Related MCP server: agent-fastpath

Requirements and setup

  • Node.js 20 or newer.

  • Laya mode: the first Laya call downloads and caches its local ONNX model (about 1.7 GB).

  • Kev mode: a separate local Kev server at 127.0.0.1:8009; see Local Kev server.

Install the MCP dependencies and run the server:

npm ci
$env:ROUTER_DECISION_BACKEND = 'laya' # or 'kev'
node .\index.js

For Claude Code and Codex, register the existing MCP name laya with node /path/to/agent-router-mcp/index.js. Keeping that client-side alias avoids changing local authentication, endpoint, and MCP settings. Restart the client after changing ROUTER_DECISION_BACKEND or the Kev endpoint/model variables.

Local Kev server

Kev is a decision model, not an agent CLI. It implements the typed SystemOne request used by this wrapper. This project calls the upstream-compatible /v1/systemone endpoint on loopback. It does not upload task state to a hosted router.

Clone jaredpalmer/kev, install its documented serving extra with Python 3.12, then start a local checkpoint:

uv sync --extra serve
$env:KEV_PREFIX_CACHE = '1'
$env:KEV_PREFIX_MAX_TOKENS = '1024'
$env:KEV_TRUNCATE_STATES = '0'
uv run --extra serve python -m kev.serve --run jaredpalmer/kev-0.8b --host 127.0.0.1 --port 8009

The included scripts/start-kev.ps1 applies these conservative cache limits and launches the server hidden, logging under the user's local application data directory. It requires KEV_REPO_PATH or a sibling kev-upstream checkout. The model weights download from Hugging Face on first launch. Stop or restart the local Python process to apply server changes. The server binds only to loopback and has no API key by default; do not expose it on a public interface.

Select it for the MCP process with $env:ROUTER_DECISION_BACKEND = 'kev'. Optional variables are KEV_BASE_URL, KEV_MODEL (default kev-latest), and KEV_TIMEOUT_MS (1,000–120,000; default 30,000). ROUTER_DECISION_BACKEND defaults to laya. Kev-0.8B's upstream model card reports limited decision-routing accuracy; validate it on your own workload before changing that default.

Tools

route_next_stage

By default, asks the decision backend to compare Claude and Codex and choose the model/effort tier (fast, balanced, or deep) for a substantial Herdr stage. It returns the selected client, model, effort, confidence, probability distributions, uncertainty marker, and a worker launch instruction. The catalog comes from the shared agent-workflow repository. Haiku and Sol-family candidates are excluded there.

For Codex, the instruction passes --dangerously-bypass-approvals-and-sandbox and --no-daemon after Herdr's -- separator because Herdr invokes the executable directly and bypasses PowerShell profile functions. It retains -c model_context_window=100000 and recommends handing off to a fresh worker around 80,000 active-context tokens. YOLO mode is intentional and must be used only in an authorized Herdr worker.

advise_next_action

Ranks two to four explicit bounded actions. Low confidence recommends gathering evidence; an action marked as approval-required remains subject to user approval. This MCP never grants permission or executes the selected action.

Validation and limitations

Both backends use the same tool schemas and output shape. Backend errors are returned to the caller; the router does not silently switch providers. The recommendation is advisory, and it cannot establish that a CLI account supports the returned model. Kev's model card and benchmark results describe model performance, not verified Herdr workflow quality.

Upstream references

Copyright (c) 2026 Daniel Cazé. The wrapper source is available under MIT. Upstream dependencies and model weights retain their respective licenses and terms.

Available Tools

2 tools
advise_next_actionA

Advisory choice among 2-4 explicit, bounded development actions when the next step is uncertain. Supply evidence and constraints. Low confidence calls for more evidence; approval remains with the user.

ParametersJSON Schema
NameRequiredDescriptionDefault
actionsYes
evidenceYes
situationYes
constraintsYes

TDQS

A3.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It usefully discloses that the output is advisory and non-binding ('approval remains with the user') and that low confidence should trigger more evidence, which are meaningful behavioral traits. It says nothing about whether state is mutated, permission requirements, or what is actually returned.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences with the core purpose front-loaded, followed by input guidance and the approval caveat. No sentence is filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Hmm, this is not 5. For a 4-required-param tool with 0% schema description coverage, no annotations and no output schema, the description should say more about return values and the relationship to route_next_stage. It covers purpose, trigger and the approval boundary but leaves key call-time details to the schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It does add meaning for three of four parameters: actions are '2-4 explicit, bounded development actions', and 'evidence' and 'constraints' are named as required inputs. The 'situation' parameter and the format/bounds of evidence and constraints are left entirely to the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific function (advisory choice among 2-4 bounded development actions) and the triggering condition (next step uncertain), so an agent can tell what it does. It never names or contrasts the sibling route_next_stage, so differentiation is left to inference.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives clear context for use ('when the next step is uncertain') and a heuristic for insufficient input ('low confidence calls for more evidence'). However, with a sibling tool route_next_stage present, there is no when-not guidance or explicit routing between the two, leaving the agent to guess which tool fits.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

route_next_stageA

Let the configured decision backend pick the best-fitting worker model for one substantial Herdr stage from the whole Claude and Codex catalog. Leave compare_providers=true and omit client; the selected choice is returned even when uncertain. Pass client only with compare_providers=false. The returned client determines the worker kind. Supply a compact task, progress and acceptance summary.

ParametersJSON Schema
NameRequiredDescriptionDefault
clientNo
stage_summaryYes
compare_providersNo

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does well: it discloses that the choice is returned 'even when uncertain' (no confidence gating), that 'the returned client determines the worker kind', and the intended scope of one substantial stage. It omits auth, cost, or latency behavior, keeping it short of a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the action, then the operating instructions, then the input requirement. Every sentence carries information, though the dense domain jargon ('Herdr stage', 'worker kind') adds some parsing cost.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, annotations, or param descriptions, the description supplies the missing return semantics ('the selected choice is returned', 'returned client determines the worker kind') and parameter interplay. The main gap is the absence of any sibling-differentiation or when-to-invoke context.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, and it does: it explains the compare_providers default and its interaction with client, and it specifies what stage_summary should contain ('compact task, progress and acceptance summary'). Only the 1400-char bound is left to the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific action ('pick the best-fitting worker model') on a specific resource ('one substantial Herdr stage') and names the catalog scope ('Claude and Codex'). It is clear what the tool does, though it never distinguishes itself from the sibling 'advise_next_action'.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives explicit parameter-mode guidance ('Leave compare_providers=true and omit client', 'Pass client only with compare_providers=false'), which is strong. However, it offers no guidance on when to choose this tool over the sibling 'advise_next_action', which is the core of this dimension.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 2 tool updatesv1.1.0
    • First observedadvise_next_action
    • First observedroute_next_stage

TDQS

A3.6/5.0

Scored across 2 tools

Disambiguation3/5

Both tools are decision/routing helpers that pick among options for the 'next step,' which creates real overlap in intent. The descriptions do distinguish them (route_next_stage selects a worker model/backend for a Herdr stage, while advise_next_action recommends among bounded development actions), but an agent could plausibly reach for the wrong one when unsure what to do next.

Naming Consistency5/5

Both names follow the same verb_noun pattern (route_next_stage, advise_next_action) with snake_case throughout. The convention is predictable and readable.

Tool Count3/5

Two tools is thin for a routing/decision server, sitting at the borderline where a set feels underdeveloped. It is not extreme, but there is little surface to work with.

Completeness2/5

The domain (decision routing over model catalogs and next actions) implies supporting operations like inspecting the catalog, configuring the backend, or querying available workers, none of which exist. The surface is a narrow slice with notable gaps.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    A
    quality
    B
    maintenance
    Enables an AI agent to turn an uncertain routing decision into a confidence-scored choice among explicit routes, with a thresholded recommendation to proceed, review deeper, or ask for human input, while remaining advisory-only.
    1
    MIT
  • A
    license
    Not graded
    quality
    B
    maintenance
    Enables coding agents to offload quick judgment calls like shipping readiness, file triage, and claim verification to a fast local MCP server with calibrated confidence and safe fallbacks.
    3
    MIT
  • A
    license
    Not graded
    quality
    C
    maintenance
    Enables local, bounded routing and evidence judgments for Codex tasks, with tools for reading file context, checking outputs, and selecting registered workflows while keeping execution and review with the client.
    MIT