Skip to main content
Glama

skill-router-mcp

A ZCode plugin that routes natural-language tasks to the right skill in a local pool, using the SkillRouter 0.6B two-stage retrieve-rerank pipeline: a bi-encoder embeds the query and the skill documents, cosine retrieval narrows the pool to the top candidates, and a pointwise reranker scores each candidate to produce the final ranking. The pool is any folder tree of SKILL.md files; everything runs locally on your machine. Current release: 0.2.0.

How it works:

  • A scanner walks the pool (<pool>/<repo>/.../SKILL.md) and parses each skill's YAML frontmatter (name, description) and markdown body.

  • An index persists on disk as a SQLite WAL manifest plus an fp16 vector store, so a restart serves from the existing index instead of re-embedding the pool.

  • Jaccard duplicate clusters collapse near-identical skills behind the best copy, so repeated skills surface once (with their siblings in duplicates[]).

  • A filesystem watcher re-scans in the background and applies incremental, content-hash-based updates — edits and additions are re-embedded, unchanged files are not touched.

  • The embedder and reranker run on CUDA by default (fp16), with a CPU escape hatch for debugging (ALLOW_CPU).

  • A FastMCP daemon serves the tools over streamable HTTP on a loopback port with a per-launch bearer token; a stdio shim bridges ZCode to it, and a SessionStart hook pre-warms the daemon while your session begins.

ZCode ──(stdio JSON-RPC)──▶ plugin MCP shim (plugin/mcp/server.js, Node)
                               │ loopback HTTP + per-launch bearer token
                               ▼ (var/daemon.json)
 .venv python -m skill_router.daemon (detached, GPU; idle-stop 30 min)
   watcher (background re-scan) → indexer (SQLite manifest)
   → embedder (SkillRouter-Embedding-0.6B) + reranker (SkillRouter-Reranker-0.6B)
   var/ (index, daemon state, daemon.log) · models/ (checkpoints) · SKILLS_DIR pool

Tools

Four MCP tools (exact response shapes in SPEC.md §6):

Tool

What it does

search_skills(query, top_k=5)

Embed the query → cosine top-20 → pointwise rerank → collapse duplicate clusters → top-K results with preview + duplicates[]

get_skill(skill_id)

One skill's full record, re-read from the live pool: body, path, repo, complete parsed frontmatter

reindex(force=false)

Incremental index refresh (content-hash diff); force=true re-embeds everything

index_stats()

Index overview: per-repo counts, model IDs, index age, duplicate-cluster count

Related MCP server: ragi

Prerequisites

  • OS / GPU: developed and tested on Windows 11 with an NVIDIA RTX 2080 Ti. Linux and macOS should work but are untested (see Limitations).

  • Python 3.11 and a virtualenv at <repo>/.venv.

  • Project dependencies: install the project into the venv (pip install -e . — see Development).

  • PyTorch: not a project dependency — install it into the venv yourself, following the official instructions at https://pytorch.org for your CUDA (or CPU) setup.

  • Model checkpoints (~1.2 GB each): pipizhao/SkillRouter-Embedding-0.6B and pipizhao/SkillRouter-Reranker-0.6B. Pre-fetch both with scripts/download_models.py — set HF_HOME to a cache directory of your choice (the script refuses to run without it) and run it with the venv python; it downloads via huggingface_hub, resumes partial downloads, and retries each repository up to 3 times:

    HF_HOME=<repo>/.hf-cache .venv/bin/python scripts/download_models.py
    # Windows: .venv\Scripts\python.exe scripts/download_models.py

    The plugin's derived defaults expect the two checkpoints unpacked at <repo>/models/SkillRouter-Embedding-0.6B and <repo>/models/SkillRouter-Reranker-0.6B — place the downloaded snapshot directories there, or point EMB_MODEL_OR_PATH / RERANK_MODEL_OR_PATH at wherever you keep them. (Running the daemon directly with both variables unset resolves to the Hugging Face repository ids, which download automatically on first load.)

  • A skills pool: any folder tree of <pool>/<repo>/.../SKILL.md files. Dot-directories inside repos are indexed too.

  • ZCode with local plugin support.

Environment variables

The plugin resolves the skill-router-mcp checkout at runtime, in this order (identical in plugin/mcp/server.js and plugin/hooks/warm_daemon.py):

  1. SKILL_ROUTER_REPO — explicit override, validated against the repo layout (src/ + plugin/ + pyproject.toml);

  2. ZCODE_PLUGIN_ROOT / ZCODE_PROJECT_DIR — the directory itself and its parent (covers running the plugin straight from the checkout);

  3. the plugin files' own location (<repo>/plugin/...).

An installed plugin runs from the ZCode cache, which never passes the layout check — set SKILL_ROUTER_REPO for that install shape.

Variable

Default when unset

Set it when

SKILL_ROUTER_REPO

(derived — resolution order above)

the plugin is installed from the ZCode cache (then it is required)

SKILLS_DIR

<ZCODE_PROJECT_DIR>/skills if that directory exists; otherwise left unset and the daemon refuses to start with an actionable message

your skill pool lives elsewhere

DATA_DIR

<repo>/var (index, daemon.json state file, daemon.log)

you want the index elsewhere

EMB_MODEL_OR_PATH

<repo>/models/SkillRouter-Embedding-0.6B

your embedding checkpoint lives elsewhere

RERANK_MODEL_OR_PATH

<repo>/models/SkillRouter-Reranker-0.6B

your reranker checkpoint lives elsewhere

Values already present in the environment always win over the derived defaults. Set them as user-level environment variables so ZCode and its children inherit them.

Hook prerequisite: the SessionStart hook runs python from PATH — any Python works (the warmer script is stdlib-only and launches the daemon with the repo venv's own interpreter). If python is not resolvable, pre-warming is skipped; the shim still respawns the daemon on the first tool call — expect roughly 30 s to the first served search.

Additional daemon tuning (all optional; values are parsed fail-fast at startup and the effective configuration is logged as a table):

Variable

Default

Meaning

RERANK_TOP

20

Candidates fed to the reranker

TOP_K_DEFAULT

5

Default result count

SCAN_INTERVAL

60

Background re-scan seconds (0 = disabled)

ROUTER_TOKEN

(unset)

Optional bearer auth on the daemon

ALLOW_CPU

false

Debug escape hatch; otherwise fail fast if no CUDA device

IDLE_TIMEOUT_S

1800

Daemon self-stop after this many idle seconds (0 = never)

BIND_HOST

127.0.0.1

Address the daemon binds

Install

Clone the repository:

git clone https://github.com/skarborous/skill-router-mcp

Then make it known to ZCode one of two ways:

  1. Via a local marketplace directory. In ZCode: Plugin Marketplace → Add → Add Plugin Marketplace, and paste the path of a directory containing marketplace.json — either this repository's root, or a staged copy of marketplace.json + plugin/ (staging keeps the snapshot small; the repo root also carries models/ and .venv/). Then install the SkillRouter plugin from that marketplace and start a new session.

  2. Manually, from the checkout. Point ZCode at the repository's plugin/ directory as a local plugin. Running straight from the checkout needs no SKILL_ROUTER_REPO (the shim resolves the repo from its own location).

Note: plugin/.zcode-plugin/plugin.json starts the MCP shim with the default ZCode install path (C:\Program Files\ZCode\ZCode.exe with ELECTRON_RUN_AS_NODE=1, so the Electron binary runs as Node). If your ZCode binary lives elsewhere, adapt the mcpServers.command value accordingly.

After install, verify in a new session: Settings → Plugins shows skill-router enabled, Settings → MCP shows the plugin server connected, and the four tools are listed. User surface: /router [status|reindex [force]|restart], /find-skill <query>, and the bundled route skill (suggests routing a task at its start).

Performance (measured)

Measured on the development machine (see Limitations); 1605-skill pool, RTX 2080 Ti, fp16:

  • Warm search: p50 1197.1 ms / p99 2581.3 ms.

  • Cold daemon respawn to first served search: ~31 s client-measured (model load dominates).

  • Watcher incremental cycle: 60-90 s live (worst case ~90 s for delete visibility).

Retrieval eval on the 50-query labeled set — Hit@1 / MRR@10 / Recall@20:

Backend

Hit@1

MRR@10

Recall@20

bm25 (metadata-only baseline)

0.320

0.445

0.645

encoder

0.520

0.598

0.732

pipeline (retrieve + rerank)

0.520

0.624

0.732

Limitations

  • Developed and tested on Windows 11 + an NVIDIA RTX 2080 Ti (Turing sm_75, fp16 inference path). Other GPUs and OSes are untested; a CUDA device is expected by default (ALLOW_CPU is a debug escape hatch, not a supported serving mode).

  • The reranker's margin over the encoder-only baseline is MRR-only on the current eval set (see Performance).

Development

python3.11 -m venv .venv                 # Windows: py -3.11 -m venv .venv
.venv/bin/python -m pip install -e ".[dev]"   # plus PyTorch per pytorch.org
.venv/bin/python -m pytest               # suite: 242 passing

The eval harness scores the bm25 / encoder / pipeline backends over the labeled holdout: scripts/eval.py with eval/labeled.jsonl (schemas and usage in eval/README.md).

Credits & Citation

  • Paper: SkillRouter: Skill Routing for LLM Agents at Scale — YanZhao Zheng, ZhenTao Zhang, Chao Ma, YuanQiang Yu, JiHuai Zhu, Yong Wu, Tianze Xu, Baohua Dong, Hangcheng Zhu, Ruohui Huang, Gang Yu (arXiv:2603.22455). This project implements its two-stage retrieve-rerank recipe.

  • Upstream SkillRouter repository: https://github.com/zhengyanzhao1997/SkillRouter (MIT) — source of the architecture, model recipes, and evaluation approach.

  • Models: pipizhao/SkillRouter-Embedding-0.6B and pipizhao/SkillRouter-Reranker-0.6B (Apache-2.0 family).

  • Benchmark data: the skill pools used for evaluation were assembled from benchflow-ai/skillsbench and majiayu000/claude-skill-registry. Benchmark data and model weights follow their upstream licenses (models: Apache-2.0 family; benchmark sources: their own).

  • Citation: see CITATION.cff (the paper is the preferred citation).

  • License: MIT — see LICENSE.

Related MCP Connectors

Related MCP Servers

  • A
    license
    A
    quality
    A
    maintenance
    A Model Context Protocol (MCP) server that provides intelligent search capabilities for discovering relevant Claude Agent Skills using vector embeddings and semantic similarity. This server implements the same progressive disclosure architecture that Anthropic describes in their Agent Skills enginee
    3
    407
    Apache 2.0
  • A
    license
    A
    quality
    D
    maintenance
    Local-first RAG indexing and semantic search MCP server. Enables document retrieval and context-aware queries using local embedding models.
    3
    13 npm
    MIT
  • A
    license
    Not graded
    quality
    C
    maintenance
    Retrieval-only MCP server that turns any knowledge source (Obsidian vault, notes, reference sets) into searchable Qdrant-backed skills, exposing list_skills, search_vault, and search_skill tools for agents to query via stdio or SSE.
    MIT