Skip to main content
Glama
README.md
# skill-router-mcp

A ZCode plugin that routes natural-language tasks to the right skill in a
local pool, using the SkillRouter 0.6B two-stage retrieve-rerank pipeline:
a bi-encoder embeds the query and the skill documents, cosine retrieval
narrows the pool to the top candidates, and a pointwise reranker scores
each candidate to produce the final ranking. The pool is any folder tree
of `SKILL.md` files; everything runs locally on your machine. Current
release: 0.2.0.

How it works:

- A **scanner** walks the pool (`<pool>/<repo>/.../SKILL.md`) and parses
  each skill's YAML frontmatter (`name`, `description`) and markdown body.
- An **index** persists on disk as a SQLite WAL manifest plus an fp16
  vector store, so a restart serves from the existing index instead of
  re-embedding the pool.
- **Jaccard duplicate clusters** collapse near-identical skills behind
  the best copy, so repeated skills surface once (with their siblings in
  `duplicates[]`).
- A **filesystem watcher** re-scans in the background and applies
  incremental, content-hash-based updates — edits and additions are
  re-embedded, unchanged files are not touched.
- The **embedder and reranker** run on CUDA by default (fp16), with a
  CPU escape hatch for debugging (`ALLOW_CPU`).
- A **FastMCP daemon** serves the tools over streamable HTTP on a
  loopback port with a per-launch bearer token; a stdio shim bridges
  ZCode to it, and a SessionStart hook pre-warms the daemon while your
  session begins.

```
ZCode ──(stdio JSON-RPC)──▶ plugin MCP shim (plugin/mcp/server.js, Node)
                               │ loopback HTTP + per-launch bearer token
                               ▼ (var/daemon.json)
 .venv python -m skill_router.daemon (detached, GPU; idle-stop 30 min)
   watcher (background re-scan) → indexer (SQLite manifest)
   → embedder (SkillRouter-Embedding-0.6B) + reranker (SkillRouter-Reranker-0.6B)
   var/ (index, daemon state, daemon.log) · models/ (checkpoints) · SKILLS_DIR pool
```

## Tools

Four MCP tools (exact response shapes in [`SPEC.md`](SPEC.md) §6):

| Tool | What it does |
|---|---|
| `search_skills(query, top_k=5)` | Embed the query → cosine top-20 → pointwise rerank → collapse duplicate clusters → top-K results with preview + `duplicates[]` |
| `get_skill(skill_id)` | One skill's full record, re-read from the live pool: body, path, repo, complete parsed frontmatter |
| `reindex(force=false)` | Incremental index refresh (content-hash diff); `force=true` re-embeds everything |
| `index_stats()` | Index overview: per-repo counts, model IDs, index age, duplicate-cluster count |

## Prerequisites

- **OS / GPU:** developed and tested on Windows 11 with an NVIDIA RTX
  2080 Ti. Linux and macOS should work but are untested (see
  [Limitations](#limitations)).
- **Python 3.11** and a virtualenv at `<repo>/.venv`.
- **Project dependencies:** install the project into the venv
  (`pip install -e .` — see [Development](#development)).
- **PyTorch:** not a project dependency — install it into the venv
  yourself, following the official instructions at
  <https://pytorch.org> for your CUDA (or CPU) setup.
- **Model checkpoints** (~1.2 GB each): `pipizhao/SkillRouter-Embedding-0.6B`
  and `pipizhao/SkillRouter-Reranker-0.6B`. Pre-fetch both with
  `scripts/download_models.py` — set `HF_HOME` to a cache directory of
  your choice (the script refuses to run without it) and run it with the
  venv python; it downloads via `huggingface_hub`, resumes partial
  downloads, and retries each repository up to 3 times:

  ```bash
  HF_HOME=<repo>/.hf-cache .venv/bin/python scripts/download_models.py
  # Windows: .venv\Scripts\python.exe scripts/download_models.py
  ```

  The plugin's derived defaults expect the two checkpoints unpacked at
  `<repo>/models/SkillRouter-Embedding-0.6B` and
  `<repo>/models/SkillRouter-Reranker-0.6B` — place the downloaded
  snapshot directories there, or point `EMB_MODEL_OR_PATH` /
  `RERANK_MODEL_OR_PATH` at wherever you keep them. (Running the daemon
  directly with both variables unset resolves to the Hugging Face
  repository ids, which download automatically on first load.)
- **A skills pool:** any folder tree of `<pool>/<repo>/.../SKILL.md`
  files. Dot-directories inside repos are indexed too.
- **ZCode** with local plugin support.

## Environment variables

The plugin resolves the skill-router-mcp checkout at runtime, in this
order (identical in `plugin/mcp/server.js` and
`plugin/hooks/warm_daemon.py`):

1. `SKILL_ROUTER_REPO` — explicit override, validated against the repo
   layout (`src/` + `plugin/` + `pyproject.toml`);
2. `ZCODE_PLUGIN_ROOT` / `ZCODE_PROJECT_DIR` — the directory itself and
   its parent (covers running the plugin straight from the checkout);
3. the plugin files' own location (`<repo>/plugin/...`).

An **installed** plugin runs from the ZCode cache, which never passes the
layout check — set `SKILL_ROUTER_REPO` for that install shape.

| Variable | Default when unset | Set it when |
|---|---|---|
| `SKILL_ROUTER_REPO` | *(derived — resolution order above)* | the plugin is installed from the ZCode cache (then it is required) |
| `SKILLS_DIR` | `<ZCODE_PROJECT_DIR>/skills` if that directory exists; otherwise left unset and the daemon refuses to start with an actionable message | your skill pool lives elsewhere |
| `DATA_DIR` | `<repo>/var` (index, `daemon.json` state file, `daemon.log`) | you want the index elsewhere |
| `EMB_MODEL_OR_PATH` | `<repo>/models/SkillRouter-Embedding-0.6B` | your embedding checkpoint lives elsewhere |
| `RERANK_MODEL_OR_PATH` | `<repo>/models/SkillRouter-Reranker-0.6B` | your reranker checkpoint lives elsewhere |

Values already present in the environment always win over the derived
defaults. Set them as user-level environment variables so ZCode and its
children inherit them.

**Hook prerequisite:** the SessionStart hook runs `python` from `PATH` —
any Python works (the warmer script is stdlib-only and launches the
daemon with the repo venv's own interpreter). If `python` is not
resolvable, pre-warming is skipped; the shim still respawns the daemon
on the first tool call — expect roughly 30 s to the first served search.

Additional daemon tuning (all optional; values are parsed fail-fast at
startup and the effective configuration is logged as a table):

| Variable | Default | Meaning |
|---|---|---|
| `RERANK_TOP` | `20` | Candidates fed to the reranker |
| `TOP_K_DEFAULT` | `5` | Default result count |
| `SCAN_INTERVAL` | `60` | Background re-scan seconds (0 = disabled) |
| `ROUTER_TOKEN` | *(unset)* | Optional bearer auth on the daemon |
| `ALLOW_CPU` | `false` | Debug escape hatch; otherwise fail fast if no CUDA device |
| `IDLE_TIMEOUT_S` | `1800` | Daemon self-stop after this many idle seconds (0 = never) |
| `BIND_HOST` | `127.0.0.1` | Address the daemon binds |

## Install

Clone the repository:

```bash
git clone https://github.com/skarborous/skill-router-mcp
```

Then make it known to ZCode one of two ways:

1. **Via a local marketplace directory.** In ZCode: Plugin Marketplace →
   Add → **Add Plugin Marketplace**, and paste the path of a directory
   containing `marketplace.json` — either this repository's root, or a
   staged copy of `marketplace.json` + `plugin/` (staging keeps the
   snapshot small; the repo root also carries `models/` and `.venv/`).
   Then install the **SkillRouter** plugin from that marketplace and
   start a new session.
2. **Manually, from the checkout.** Point ZCode at the repository's
   `plugin/` directory as a local plugin. Running straight from the
   checkout needs no `SKILL_ROUTER_REPO` (the shim resolves the repo from
   its own location).

Note: `plugin/.zcode-plugin/plugin.json` starts the MCP shim with the
default ZCode install path (`C:\Program Files\ZCode\ZCode.exe` with
`ELECTRON_RUN_AS_NODE=1`, so the Electron binary runs as Node). If your
ZCode binary lives elsewhere, adapt the `mcpServers.command` value
accordingly.

After install, verify in a **new** session: Settings → Plugins shows
skill-router enabled, Settings → MCP shows the plugin server connected,
and the four tools are listed. User surface: `/router
[status|reindex [force]|restart]`, `/find-skill <query>`, and the bundled
`route` skill (suggests routing a task at its start).

## Performance (measured)

Measured on the development machine (see
[Limitations](#limitations)); 1605-skill pool, RTX 2080 Ti, fp16:

- Warm search: **p50 1197.1 ms / p99 2581.3 ms**.
- Cold daemon respawn to first served search: **~31 s** client-measured
  (model load dominates).
- Watcher incremental cycle: **60-90 s** live (worst case ~90 s for
  delete visibility).

Retrieval eval on the 50-query labeled set — Hit@1 / MRR@10 /
Recall@20:

| Backend | Hit@1 | MRR@10 | Recall@20 |
|---|---|---|---|
| bm25 (metadata-only baseline) | 0.320 | 0.445 | 0.645 |
| encoder | 0.520 | 0.598 | 0.732 |
| pipeline (retrieve + rerank) | 0.520 | 0.624 | 0.732 |

## Limitations

- Developed and tested on Windows 11 + an NVIDIA RTX 2080 Ti (Turing
  sm_75, fp16 inference path). Other GPUs and OSes are untested; a CUDA
  device is expected by default (`ALLOW_CPU` is a debug escape hatch,
  not a supported serving mode).
- The reranker's margin over the encoder-only baseline is MRR-only on
  the current eval set (see [Performance](#performance-measured)).

## Development

```bash
python3.11 -m venv .venv                 # Windows: py -3.11 -m venv .venv
.venv/bin/python -m pip install -e ".[dev]"   # plus PyTorch per pytorch.org
.venv/bin/python -m pytest               # suite: 242 passing
```

The eval harness scores the bm25 / encoder / pipeline backends over the
labeled holdout: `scripts/eval.py` with `eval/labeled.jsonl` (schemas and
usage in [`eval/README.md`](eval/README.md)).

## Credits & Citation

- **Paper:** *SkillRouter: Skill Routing for LLM Agents at Scale* —
  YanZhao Zheng, ZhenTao Zhang, Chao Ma, YuanQiang Yu, JiHuai Zhu,
  Yong Wu, Tianze Xu, Baohua Dong, Hangcheng Zhu, Ruohui Huang, Gang Yu
  ([arXiv:2603.22455](https://arxiv.org/abs/2603.22455)). This project
  implements its two-stage retrieve-rerank recipe.
- **Upstream SkillRouter repository:**
  <https://github.com/zhengyanzhao1997/SkillRouter> (MIT) — source of
  the architecture, model recipes, and evaluation approach.
- **Models:** `pipizhao/SkillRouter-Embedding-0.6B` and
  `pipizhao/SkillRouter-Reranker-0.6B` (Apache-2.0 family).
- **Benchmark data:** the skill pools used for evaluation were assembled
  from [benchflow-ai/skillsbench](https://github.com/benchflow-ai/skillsbench)
  and [majiayu000/claude-skill-registry](https://github.com/majiayu000/claude-skill-registry).
  Benchmark data and model weights follow their upstream licenses
  (models: Apache-2.0 family; benchmark sources: their own).
- **Citation:** see [`CITATION.cff`](CITATION.cff) (the paper is the
  preferred citation).
- **License:** MIT — see [`LICENSE`](LICENSE).