Skip to main content
Glama
README.md
# Agent Router MCP

A local MCP decision router for Claude Code and Codex Herdr workers. It exposes two tools, `route_next_stage` and `advise_next_action`, through a stable stdio server. The decision backend is selectable between the upstream Laya library and the locally served Kev model. Existing Claude/Codex MCP registrations named `laya` continue to work; the server implementation and GitHub project are named `agent-router-mcp`.

## Why this repository exists

The source began as a machine-local custom Laya MCP wrapper. Keeping that implementation available makes it possible to reproduce, inspect, and configure routing instead of depending on a private local-only copy. MIT covers this repository's wrapper code. Laya and Kev remain separate upstream dependencies with their own licenses and model terms.

Herdr is the orchestration layer: it owns panes and starts workers. The router only recommends a client, model, effort, and bounded next action. Its routing catalog and shared behavior live in [`danielcaze/agent-workflow`](https://github.com/danielcaze/agent-workflow); this repository contains the runnable MCP implementation.

## Requirements and setup

- Node.js 20 or newer.
- Laya mode: the first Laya call downloads and caches its local ONNX model (about 1.7 GB).
- Kev mode: a separate local Kev server at `127.0.0.1:8009`; see [Local Kev server](#local-kev-server).

Install the MCP dependencies and run the server:

```powershell
npm ci
$env:ROUTER_DECISION_BACKEND = 'laya' # or 'kev'
node .\index.js
```

For Claude Code and Codex, register the existing MCP name `laya` with `node /path/to/agent-router-mcp/index.js`. Keeping that client-side alias avoids changing local authentication, endpoint, and MCP settings. Restart the client after changing `ROUTER_DECISION_BACKEND` or the Kev endpoint/model variables.

## Local Kev server

Kev is a decision model, not an agent CLI. It implements the typed SystemOne request used by this wrapper. This project calls the upstream-compatible `/v1/systemone` endpoint on loopback. It does not upload task state to a hosted router.

Clone [jaredpalmer/kev](https://github.com/jaredpalmer/kev), install its documented serving extra with Python 3.12, then start a local checkpoint:

```powershell
uv sync --extra serve
$env:KEV_PREFIX_CACHE = '1'
$env:KEV_PREFIX_MAX_TOKENS = '1024'
$env:KEV_TRUNCATE_STATES = '0'
uv run --extra serve python -m kev.serve --run jaredpalmer/kev-0.8b --host 127.0.0.1 --port 8009
```

The included `scripts/start-kev.ps1` applies these conservative cache limits and launches the server hidden, logging under the user's local application data directory. It requires `KEV_REPO_PATH` or a sibling `kev-upstream` checkout. The model weights download from Hugging Face on first launch. Stop or restart the local Python process to apply server changes. The server binds only to loopback and has no API key by default; do not expose it on a public interface.

Select it for the MCP process with `$env:ROUTER_DECISION_BACKEND = 'kev'`. Optional variables are `KEV_BASE_URL`, `KEV_MODEL` (default `kev-latest`), and `KEV_TIMEOUT_MS` (1,000–120,000; default 30,000). `ROUTER_DECISION_BACKEND` defaults to `laya`. Kev-0.8B's upstream model card reports limited decision-routing accuracy; validate it on your own workload before changing that default.

## Tools

### `route_next_stage`

By default, asks the decision backend to compare Claude and Codex and choose the model/effort tier (`fast`, `balanced`, or `deep`) for a substantial Herdr stage. It returns the selected client, model, effort, confidence, probability distributions, uncertainty marker, and a worker launch instruction. The catalog comes from the shared `agent-workflow` repository. Haiku and Sol-family candidates are excluded there.

For Codex, the instruction passes `--dangerously-bypass-approvals-and-sandbox` and `--no-daemon` after Herdr's `--` separator because Herdr invokes the executable directly and bypasses PowerShell profile functions. It retains `-c model_context_window=100000` and recommends handing off to a fresh worker around 80,000 active-context tokens. YOLO mode is intentional and must be used only in an authorized Herdr worker.

### `advise_next_action`

Ranks two to four explicit bounded actions. Low confidence recommends gathering evidence; an action marked as approval-required remains subject to user approval. This MCP never grants permission or executes the selected action.

## Validation and limitations

Both backends use the same tool schemas and output shape. Backend errors are returned to the caller; the router does not silently switch providers. The recommendation is advisory, and it cannot establish that a CLI account supports the returned model. Kev's model card and benchmark results describe model performance, not verified Herdr workflow quality.

## Upstream references

- [Laya TypeScript package](https://github.com/receptron/laya) and [npm package](https://www.npmjs.com/package/@receptron/laya).
- [Kev source and serving documentation](https://github.com/jaredpalmer/kev), [Kev-0.8B model card](https://huggingface.co/jaredpalmer/kev-0.8b), and [Kev license](https://github.com/jaredpalmer/kev/blob/main/LICENSE).
- [Shared agent-workflow rules and distribution](https://github.com/danielcaze/agent-workflow).

Copyright (c) 2026 Daniel Cazé. The wrapper source is available under [MIT](LICENSE). Upstream dependencies and model weights retain their respective licenses and terms.

TDQS

A3.6/5.0

Scored across 2 tools

Disambiguation3/5

Both tools are decision/routing helpers that pick among options for the 'next step,' which creates real overlap in intent. The descriptions do distinguish them (route_next_stage selects a worker model/backend for a Herdr stage, while advise_next_action recommends among bounded development actions), but an agent could plausibly reach for the wrong one when unsure what to do next.

Naming Consistency5/5

Both names follow the same verb_noun pattern (route_next_stage, advise_next_action) with snake_case throughout. The convention is predictable and readable.

Tool Count3/5

Two tools is thin for a routing/decision server, sitting at the borderline where a set feels underdeveloped. It is not extreme, but there is little surface to work with.

Completeness2/5

The domain (decision routing over model catalogs and next actions) implies supporting operations like inspecting the catalog, configuring the backend, or querying available workers, none of which exist. The surface is a narrow slice with notable gaps.

Maintenance

ActivityMaintained
ResponsivenessNo issues