SkillRouter
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@SkillRouterwhich skill should I use for generating release notes?"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
skill-router-mcp
A ZCode plugin that routes natural-language tasks to the right skill in a
local pool, using the SkillRouter 0.6B two-stage retrieve-rerank pipeline:
a bi-encoder embeds the query and the skill documents, cosine retrieval
narrows the pool to the top candidates, and a pointwise reranker scores
each candidate to produce the final ranking. The pool is any folder tree
of SKILL.md files; everything runs locally on your machine. Current
release: 0.2.0.
How it works:
A scanner walks the pool (
<pool>/<repo>/.../SKILL.md) and parses each skill's YAML frontmatter (name,description) and markdown body.An index persists on disk as a SQLite WAL manifest plus an fp16 vector store, so a restart serves from the existing index instead of re-embedding the pool.
Jaccard duplicate clusters collapse near-identical skills behind the best copy, so repeated skills surface once (with their siblings in
duplicates[]).A filesystem watcher re-scans in the background and applies incremental, content-hash-based updates — edits and additions are re-embedded, unchanged files are not touched.
The embedder and reranker run on CUDA by default (fp16), with a CPU escape hatch for debugging (
ALLOW_CPU).A FastMCP daemon serves the tools over streamable HTTP on a loopback port with a per-launch bearer token; a stdio shim bridges ZCode to it, and a SessionStart hook pre-warms the daemon while your session begins.
ZCode ──(stdio JSON-RPC)──▶ plugin MCP shim (plugin/mcp/server.js, Node)
│ loopback HTTP + per-launch bearer token
▼ (var/daemon.json)
.venv python -m skill_router.daemon (detached, GPU; idle-stop 30 min)
watcher (background re-scan) → indexer (SQLite manifest)
→ embedder (SkillRouter-Embedding-0.6B) + reranker (SkillRouter-Reranker-0.6B)
var/ (index, daemon state, daemon.log) · models/ (checkpoints) · SKILLS_DIR poolTools
Four MCP tools (exact response shapes in SPEC.md §6):
Tool | What it does |
| Embed the query → cosine top-20 → pointwise rerank → collapse duplicate clusters → top-K results with preview + |
| One skill's full record, re-read from the live pool: body, path, repo, complete parsed frontmatter |
| Incremental index refresh (content-hash diff); |
| Index overview: per-repo counts, model IDs, index age, duplicate-cluster count |
Related MCP server: ragi
Prerequisites
OS / GPU: developed and tested on Windows 11 with an NVIDIA RTX 2080 Ti. Linux and macOS should work but are untested (see Limitations).
Python 3.11 and a virtualenv at
<repo>/.venv.Project dependencies: install the project into the venv (
pip install -e .— see Development).PyTorch: not a project dependency — install it into the venv yourself, following the official instructions at https://pytorch.org for your CUDA (or CPU) setup.
Model checkpoints (~1.2 GB each):
pipizhao/SkillRouter-Embedding-0.6Bandpipizhao/SkillRouter-Reranker-0.6B. Pre-fetch both withscripts/download_models.py— setHF_HOMEto a cache directory of your choice (the script refuses to run without it) and run it with the venv python; it downloads viahuggingface_hub, resumes partial downloads, and retries each repository up to 3 times:HF_HOME=<repo>/.hf-cache .venv/bin/python scripts/download_models.py # Windows: .venv\Scripts\python.exe scripts/download_models.pyThe plugin's derived defaults expect the two checkpoints unpacked at
<repo>/models/SkillRouter-Embedding-0.6Band<repo>/models/SkillRouter-Reranker-0.6B— place the downloaded snapshot directories there, or pointEMB_MODEL_OR_PATH/RERANK_MODEL_OR_PATHat wherever you keep them. (Running the daemon directly with both variables unset resolves to the Hugging Face repository ids, which download automatically on first load.)A skills pool: any folder tree of
<pool>/<repo>/.../SKILL.mdfiles. Dot-directories inside repos are indexed too.ZCode with local plugin support.
Environment variables
The plugin resolves the skill-router-mcp checkout at runtime, in this
order (identical in plugin/mcp/server.js and
plugin/hooks/warm_daemon.py):
SKILL_ROUTER_REPO— explicit override, validated against the repo layout (src/+plugin/+pyproject.toml);ZCODE_PLUGIN_ROOT/ZCODE_PROJECT_DIR— the directory itself and its parent (covers running the plugin straight from the checkout);the plugin files' own location (
<repo>/plugin/...).
An installed plugin runs from the ZCode cache, which never passes the
layout check — set SKILL_ROUTER_REPO for that install shape.
Variable | Default when unset | Set it when |
| (derived — resolution order above) | the plugin is installed from the ZCode cache (then it is required) |
|
| your skill pool lives elsewhere |
|
| you want the index elsewhere |
|
| your embedding checkpoint lives elsewhere |
|
| your reranker checkpoint lives elsewhere |
Values already present in the environment always win over the derived defaults. Set them as user-level environment variables so ZCode and its children inherit them.
Hook prerequisite: the SessionStart hook runs python from PATH —
any Python works (the warmer script is stdlib-only and launches the
daemon with the repo venv's own interpreter). If python is not
resolvable, pre-warming is skipped; the shim still respawns the daemon
on the first tool call — expect roughly 30 s to the first served search.
Additional daemon tuning (all optional; values are parsed fail-fast at startup and the effective configuration is logged as a table):
Variable | Default | Meaning |
|
| Candidates fed to the reranker |
|
| Default result count |
|
| Background re-scan seconds (0 = disabled) |
| (unset) | Optional bearer auth on the daemon |
|
| Debug escape hatch; otherwise fail fast if no CUDA device |
|
| Daemon self-stop after this many idle seconds (0 = never) |
|
| Address the daemon binds |
Install
Clone the repository:
git clone https://github.com/skarborous/skill-router-mcpThen make it known to ZCode one of two ways:
Via a local marketplace directory. In ZCode: Plugin Marketplace → Add → Add Plugin Marketplace, and paste the path of a directory containing
marketplace.json— either this repository's root, or a staged copy ofmarketplace.json+plugin/(staging keeps the snapshot small; the repo root also carriesmodels/and.venv/). Then install the SkillRouter plugin from that marketplace and start a new session.Manually, from the checkout. Point ZCode at the repository's
plugin/directory as a local plugin. Running straight from the checkout needs noSKILL_ROUTER_REPO(the shim resolves the repo from its own location).
Note: plugin/.zcode-plugin/plugin.json starts the MCP shim with the
default ZCode install path (C:\Program Files\ZCode\ZCode.exe with
ELECTRON_RUN_AS_NODE=1, so the Electron binary runs as Node). If your
ZCode binary lives elsewhere, adapt the mcpServers.command value
accordingly.
After install, verify in a new session: Settings → Plugins shows
skill-router enabled, Settings → MCP shows the plugin server connected,
and the four tools are listed. User surface: /router [status|reindex [force]|restart], /find-skill <query>, and the bundled
route skill (suggests routing a task at its start).
Performance (measured)
Measured on the development machine (see Limitations); 1605-skill pool, RTX 2080 Ti, fp16:
Warm search: p50 1197.1 ms / p99 2581.3 ms.
Cold daemon respawn to first served search: ~31 s client-measured (model load dominates).
Watcher incremental cycle: 60-90 s live (worst case ~90 s for delete visibility).
Retrieval eval on the 50-query labeled set — Hit@1 / MRR@10 / Recall@20:
Backend | Hit@1 | MRR@10 | Recall@20 |
bm25 (metadata-only baseline) | 0.320 | 0.445 | 0.645 |
encoder | 0.520 | 0.598 | 0.732 |
pipeline (retrieve + rerank) | 0.520 | 0.624 | 0.732 |
Limitations
Developed and tested on Windows 11 + an NVIDIA RTX 2080 Ti (Turing sm_75, fp16 inference path). Other GPUs and OSes are untested; a CUDA device is expected by default (
ALLOW_CPUis a debug escape hatch, not a supported serving mode).The reranker's margin over the encoder-only baseline is MRR-only on the current eval set (see Performance).
Development
python3.11 -m venv .venv # Windows: py -3.11 -m venv .venv
.venv/bin/python -m pip install -e ".[dev]" # plus PyTorch per pytorch.org
.venv/bin/python -m pytest # suite: 242 passingThe eval harness scores the bm25 / encoder / pipeline backends over the
labeled holdout: scripts/eval.py with eval/labeled.jsonl (schemas and
usage in eval/README.md).
Credits & Citation
Paper: SkillRouter: Skill Routing for LLM Agents at Scale — YanZhao Zheng, ZhenTao Zhang, Chao Ma, YuanQiang Yu, JiHuai Zhu, Yong Wu, Tianze Xu, Baohua Dong, Hangcheng Zhu, Ruohui Huang, Gang Yu (arXiv:2603.22455). This project implements its two-stage retrieve-rerank recipe.
Upstream SkillRouter repository: https://github.com/zhengyanzhao1997/SkillRouter (MIT) — source of the architecture, model recipes, and evaluation approach.
Models:
pipizhao/SkillRouter-Embedding-0.6Bandpipizhao/SkillRouter-Reranker-0.6B(Apache-2.0 family).Benchmark data: the skill pools used for evaluation were assembled from benchflow-ai/skillsbench and majiayu000/claude-skill-registry. Benchmark data and model weights follow their upstream licenses (models: Apache-2.0 family; benchmark sources: their own).
Citation: see
CITATION.cff(the paper is the preferred citation).License: MIT — see
LICENSE.
This server cannot be deployed
Maintenance
Related MCP Connectors
Search and discover Agent Skills from the skills.sh registry. Powered by HAPI MCP server.
Personal knowledge base MCP server with semantic search, auto-categorization, metadata extraction
Cloud or self-hosted knowledge for AI agents: hybrid search, reranking, GraphRAG, scoped MCP tools.
Capability registry for the agentic economy. Semantic search over verified MCP server listings.
Related MCP Servers
- AlicenseAqualityAmaintenanceA Model Context Protocol (MCP) server that provides intelligent search capabilities for discovering relevant Claude Agent Skills using vector embeddings and semantic similarity. This server implements the same progressive disclosure architecture that Anthropic describes in their Agent Skills enginee3407Apache 2.0
- AlicenseAqualityDmaintenanceLocal-first RAG indexing and semantic search MCP server. Enables document retrieval and context-aware queries using local embedding models.313 npmMIT
- AlicenseAqualityDmaintenanceRoutes SKILL.md libraries to any MCP client, enabling task matching and skill loading with embedding-based scoring, keyword fallback, and context-window discipline.58 npmMIT
- AlicenseNot gradedqualityCmaintenanceRetrieval-only MCP server that turns any knowledge source (Obsidian vault, notes, reference sets) into searchable Qdrant-backed skills, exposing list_skills, search_vault, and search_skill tools for agents to query via stdio or SSE.MIT