SkillBench MCP Server
by redkoai
README.md
# SkillBench MCP Server
A zero-dependency harness that runs held-out coding tasks with and without a skill file, then reports behavior deltas and variance. Packaged as a mcp server so users can adopt it quickly.
## Why This Exists
Skill repos are getting daily attention, but the ecosystem has almost no proof that skills improve agent behavior.
## One-Day MVP
YAML/JSON task fixtures, paired runs, transcript diffing, pass/fail metrics, and a Markdown report.; one local MCP tool, install docs, example agent transcript.
## What It Does Today
- Loads a small JSON task or project description.
- Checks whether the required fields and examples are present.
- Produces a JSON or Markdown readiness report.
- Ships with tests and a CI workflow so the repo is immediately verifiable.
## Install
```bash
python3 -m venv .venv
. .venv/bin/activate
python -m pip install -e .
```
## Run
```bash
python -m skillbench_mcp_server.cli --input examples/sample_input.json --format markdown
```
## Validate
```bash
python -m unittest discover -s tests
```
## Validation Query
`agent skills eval`
## Comparable Repos
- [darkrishabh/agent-skills-eval](https://github.com/darkrishabh/agent-skills-eval): 656 stars - A test runner for agentskills.io-style AI agent skills
- [caohaotiantian/agent-skills-eval](https://github.com/caohaotiantian/agent-skills-eval): 8 stars -
- [sasa-fajkovic/agents-skill-eval](https://github.com/sasa-fajkovic/agents-skill-eval): 4 stars - agents-skill-eval.com
## Launch Angle
show the same workflow from Claude Code/Codex/Cursor where possible
## Publishing Gate
- must disclose permissions and local data access
- Working demo or executable artifact.
- Clear examples and tests.
- Honest limitations.
- No fake stars, duplicate repos, or automated engagement.
## Status
Generated by RepoForge Radar. Treat this as a starting implementation, not a finished product.
This server cannot be deployed
Maintenance
ActivityStale
ResponsivenessNo issues