SkillRouter
by skarborous
README.md
# skill-router-mcp
A ZCode plugin that routes natural-language tasks to the right skill in a
local pool, using the SkillRouter 0.6B two-stage retrieve-rerank pipeline:
a bi-encoder embeds the query and the skill documents, cosine retrieval
narrows the pool to the top candidates, and a pointwise reranker scores
each candidate to produce the final ranking. The pool is any folder tree
of `SKILL.md` files; everything runs locally on your machine. Current
release: 0.2.0.
How it works:
- A **scanner** walks the pool (`<pool>/<repo>/.../SKILL.md`) and parses
each skill's YAML frontmatter (`name`, `description`) and markdown body.
- An **index** persists on disk as a SQLite WAL manifest plus an fp16
vector store, so a restart serves from the existing index instead of
re-embedding the pool.
- **Jaccard duplicate clusters** collapse near-identical skills behind
the best copy, so repeated skills surface once (with their siblings in
`duplicates[]`).
- A **filesystem watcher** re-scans in the background and applies
incremental, content-hash-based updates — edits and additions are
re-embedded, unchanged files are not touched.
- The **embedder and reranker** run on CUDA by default (fp16), with a
CPU escape hatch for debugging (`ALLOW_CPU`).
- A **FastMCP daemon** serves the tools over streamable HTTP on a
loopback port with a per-launch bearer token; a stdio shim bridges
ZCode to it, and a SessionStart hook pre-warms the daemon while your
session begins.
```
ZCode ──(stdio JSON-RPC)──▶ plugin MCP shim (plugin/mcp/server.js, Node)
│ loopback HTTP + per-launch bearer token
▼ (var/daemon.json)
.venv python -m skill_router.daemon (detached, GPU; idle-stop 30 min)
watcher (background re-scan) → indexer (SQLite manifest)
→ embedder (SkillRouter-Embedding-0.6B) + reranker (SkillRouter-Reranker-0.6B)
var/ (index, daemon state, daemon.log) · models/ (checkpoints) · SKILLS_DIR pool
```
## Tools
Four MCP tools (exact response shapes in [`SPEC.md`](SPEC.md) §6):
| Tool | What it does |
|---|---|
| `search_skills(query, top_k=5)` | Embed the query → cosine top-20 → pointwise rerank → collapse duplicate clusters → top-K results with preview + `duplicates[]` |
| `get_skill(skill_id)` | One skill's full record, re-read from the live pool: body, path, repo, complete parsed frontmatter |
| `reindex(force=false)` | Incremental index refresh (content-hash diff); `force=true` re-embeds everything |
| `index_stats()` | Index overview: per-repo counts, model IDs, index age, duplicate-cluster count |
## Prerequisites
- **OS / GPU:** developed and tested on Windows 11 with an NVIDIA RTX
2080 Ti. Linux and macOS should work but are untested (see
[Limitations](#limitations)).
- **Python 3.11** and a virtualenv at `<repo>/.venv`.
- **Project dependencies:** install the project into the venv
(`pip install -e .` — see [Development](#development)).
- **PyTorch:** not a project dependency — install it into the venv
yourself, following the official instructions at
<https://pytorch.org> for your CUDA (or CPU) setup.
- **Model checkpoints** (~1.2 GB each): `pipizhao/SkillRouter-Embedding-0.6B`
and `pipizhao/SkillRouter-Reranker-0.6B`. Pre-fetch both with
`scripts/download_models.py` — set `HF_HOME` to a cache directory of
your choice (the script refuses to run without it) and run it with the
venv python; it downloads via `huggingface_hub`, resumes partial
downloads, and retries each repository up to 3 times:
```bash
HF_HOME=<repo>/.hf-cache .venv/bin/python scripts/download_models.py
# Windows: .venv\Scripts\python.exe scripts/download_models.py
```
The plugin's derived defaults expect the two checkpoints unpacked at
`<repo>/models/SkillRouter-Embedding-0.6B` and
`<repo>/models/SkillRouter-Reranker-0.6B` — place the downloaded
snapshot directories there, or point `EMB_MODEL_OR_PATH` /
`RERANK_MODEL_OR_PATH` at wherever you keep them. (Running the daemon
directly with both variables unset resolves to the Hugging Face
repository ids, which download automatically on first load.)
- **A skills pool:** any folder tree of `<pool>/<repo>/.../SKILL.md`
files. Dot-directories inside repos are indexed too.
- **ZCode** with local plugin support.
## Environment variables
The plugin resolves the skill-router-mcp checkout at runtime, in this
order (identical in `plugin/mcp/server.js` and
`plugin/hooks/warm_daemon.py`):
1. `SKILL_ROUTER_REPO` — explicit override, validated against the repo
layout (`src/` + `plugin/` + `pyproject.toml`);
2. `ZCODE_PLUGIN_ROOT` / `ZCODE_PROJECT_DIR` — the directory itself and
its parent (covers running the plugin straight from the checkout);
3. the plugin files' own location (`<repo>/plugin/...`).
An **installed** plugin runs from the ZCode cache, which never passes the
layout check — set `SKILL_ROUTER_REPO` for that install shape.
| Variable | Default when unset | Set it when |
|---|---|---|
| `SKILL_ROUTER_REPO` | *(derived — resolution order above)* | the plugin is installed from the ZCode cache (then it is required) |
| `SKILLS_DIR` | `<ZCODE_PROJECT_DIR>/skills` if that directory exists; otherwise left unset and the daemon refuses to start with an actionable message | your skill pool lives elsewhere |
| `DATA_DIR` | `<repo>/var` (index, `daemon.json` state file, `daemon.log`) | you want the index elsewhere |
| `EMB_MODEL_OR_PATH` | `<repo>/models/SkillRouter-Embedding-0.6B` | your embedding checkpoint lives elsewhere |
| `RERANK_MODEL_OR_PATH` | `<repo>/models/SkillRouter-Reranker-0.6B` | your reranker checkpoint lives elsewhere |
Values already present in the environment always win over the derived
defaults. Set them as user-level environment variables so ZCode and its
children inherit them.
**Hook prerequisite:** the SessionStart hook runs `python` from `PATH` —
any Python works (the warmer script is stdlib-only and launches the
daemon with the repo venv's own interpreter). If `python` is not
resolvable, pre-warming is skipped; the shim still respawns the daemon
on the first tool call — expect roughly 30 s to the first served search.
Additional daemon tuning (all optional; values are parsed fail-fast at
startup and the effective configuration is logged as a table):
| Variable | Default | Meaning |
|---|---|---|
| `RERANK_TOP` | `20` | Candidates fed to the reranker |
| `TOP_K_DEFAULT` | `5` | Default result count |
| `SCAN_INTERVAL` | `60` | Background re-scan seconds (0 = disabled) |
| `ROUTER_TOKEN` | *(unset)* | Optional bearer auth on the daemon |
| `ALLOW_CPU` | `false` | Debug escape hatch; otherwise fail fast if no CUDA device |
| `IDLE_TIMEOUT_S` | `1800` | Daemon self-stop after this many idle seconds (0 = never) |
| `BIND_HOST` | `127.0.0.1` | Address the daemon binds |
## Install
Clone the repository:
```bash
git clone https://github.com/skarborous/skill-router-mcp
```
Then make it known to ZCode one of two ways:
1. **Via a local marketplace directory.** In ZCode: Plugin Marketplace →
Add → **Add Plugin Marketplace**, and paste the path of a directory
containing `marketplace.json` — either this repository's root, or a
staged copy of `marketplace.json` + `plugin/` (staging keeps the
snapshot small; the repo root also carries `models/` and `.venv/`).
Then install the **SkillRouter** plugin from that marketplace and
start a new session.
2. **Manually, from the checkout.** Point ZCode at the repository's
`plugin/` directory as a local plugin. Running straight from the
checkout needs no `SKILL_ROUTER_REPO` (the shim resolves the repo from
its own location).
Note: `plugin/.zcode-plugin/plugin.json` starts the MCP shim with the
default ZCode install path (`C:\Program Files\ZCode\ZCode.exe` with
`ELECTRON_RUN_AS_NODE=1`, so the Electron binary runs as Node). If your
ZCode binary lives elsewhere, adapt the `mcpServers.command` value
accordingly.
After install, verify in a **new** session: Settings → Plugins shows
skill-router enabled, Settings → MCP shows the plugin server connected,
and the four tools are listed. User surface: `/router
[status|reindex [force]|restart]`, `/find-skill <query>`, and the bundled
`route` skill (suggests routing a task at its start).
## Performance (measured)
Measured on the development machine (see
[Limitations](#limitations)); 1605-skill pool, RTX 2080 Ti, fp16:
- Warm search: **p50 1197.1 ms / p99 2581.3 ms**.
- Cold daemon respawn to first served search: **~31 s** client-measured
(model load dominates).
- Watcher incremental cycle: **60-90 s** live (worst case ~90 s for
delete visibility).
Retrieval eval on the 50-query labeled set — Hit@1 / MRR@10 /
Recall@20:
| Backend | Hit@1 | MRR@10 | Recall@20 |
|---|---|---|---|
| bm25 (metadata-only baseline) | 0.320 | 0.445 | 0.645 |
| encoder | 0.520 | 0.598 | 0.732 |
| pipeline (retrieve + rerank) | 0.520 | 0.624 | 0.732 |
## Limitations
- Developed and tested on Windows 11 + an NVIDIA RTX 2080 Ti (Turing
sm_75, fp16 inference path). Other GPUs and OSes are untested; a CUDA
device is expected by default (`ALLOW_CPU` is a debug escape hatch,
not a supported serving mode).
- The reranker's margin over the encoder-only baseline is MRR-only on
the current eval set (see [Performance](#performance-measured)).
## Development
```bash
python3.11 -m venv .venv # Windows: py -3.11 -m venv .venv
.venv/bin/python -m pip install -e ".[dev]" # plus PyTorch per pytorch.org
.venv/bin/python -m pytest # suite: 242 passing
```
The eval harness scores the bm25 / encoder / pipeline backends over the
labeled holdout: `scripts/eval.py` with `eval/labeled.jsonl` (schemas and
usage in [`eval/README.md`](eval/README.md)).
## Credits & Citation
- **Paper:** *SkillRouter: Skill Routing for LLM Agents at Scale* —
YanZhao Zheng, ZhenTao Zhang, Chao Ma, YuanQiang Yu, JiHuai Zhu,
Yong Wu, Tianze Xu, Baohua Dong, Hangcheng Zhu, Ruohui Huang, Gang Yu
([arXiv:2603.22455](https://arxiv.org/abs/2603.22455)). This project
implements its two-stage retrieve-rerank recipe.
- **Upstream SkillRouter repository:**
<https://github.com/zhengyanzhao1997/SkillRouter> (MIT) — source of
the architecture, model recipes, and evaluation approach.
- **Models:** `pipizhao/SkillRouter-Embedding-0.6B` and
`pipizhao/SkillRouter-Reranker-0.6B` (Apache-2.0 family).
- **Benchmark data:** the skill pools used for evaluation were assembled
from [benchflow-ai/skillsbench](https://github.com/benchflow-ai/skillsbench)
and [majiayu000/claude-skill-registry](https://github.com/majiayu000/claude-skill-registry).
Benchmark data and model weights follow their upstream licenses
(models: Apache-2.0 family; benchmark sources: their own).
- **Citation:** see [`CITATION.cff`](CITATION.cff) (the paper is the
preferred citation).
- **License:** MIT — see [`LICENSE`](LICENSE).
This server cannot be deployed
Maintenance
ActivityMaintained
ResponsivenessNo issues