Skip to main content
Glama
README.md
# SkillBench MCP Server

A zero-dependency harness that runs held-out coding tasks with and without a skill file, then reports behavior deltas and variance. Packaged as a mcp server so users can adopt it quickly.

## Why This Exists

Skill repos are getting daily attention, but the ecosystem has almost no proof that skills improve agent behavior.

## One-Day MVP

YAML/JSON task fixtures, paired runs, transcript diffing, pass/fail metrics, and a Markdown report.; one local MCP tool, install docs, example agent transcript.

## What It Does Today

- Loads a small JSON task or project description.
- Checks whether the required fields and examples are present.
- Produces a JSON or Markdown readiness report.
- Ships with tests and a CI workflow so the repo is immediately verifiable.

## Install

```bash
python3 -m venv .venv
. .venv/bin/activate
python -m pip install -e .
```

## Run

```bash
python -m skillbench_mcp_server.cli --input examples/sample_input.json --format markdown
```

## Validate

```bash
python -m unittest discover -s tests
```

## Validation Query

`agent skills eval`

## Comparable Repos

- [darkrishabh/agent-skills-eval](https://github.com/darkrishabh/agent-skills-eval): 656 stars - A test runner for agentskills.io-style AI agent skills
- [caohaotiantian/agent-skills-eval](https://github.com/caohaotiantian/agent-skills-eval): 8 stars - 
- [sasa-fajkovic/agents-skill-eval](https://github.com/sasa-fajkovic/agents-skill-eval): 4 stars - agents-skill-eval.com 

## Launch Angle

show the same workflow from Claude Code/Codex/Cursor where possible

## Publishing Gate

- must disclose permissions and local data access
- Working demo or executable artifact.
- Clear examples and tests.
- Honest limitations.
- No fake stars, duplicate repos, or automated engagement.

## Status

Generated by RepoForge Radar. Treat this as a starting implementation, not a finished product.