Skip to main content
Glama
sudo-ai-git

io.github.sudo-ai-git/mcp-benchmark-hygiene

mcp-benchmark-hygiene

mcp-name: io.github.sudo-ai-git/mcp-benchmark-hygiene

Deterministic detection of pytest config-leakage that silently corrupts agent-benchmark / function grading.

No LLM. No network. One question, answered reliably:

If I run python -m pytest <tests> inside this workspace, will it inherit a host coverage/abort gate that mis-scores passing code as failed?


The bug this catches

Automated agent-evaluation harnesses often run python -m pytest <hidden_tests> inside the target's workspace. If that workspace nests under a repo root carrying pytest addopts — e.g.:

[tool.pytest.ini_options]
addopts = "--cov=harness --cov-report=term-missing:skip-covered --cov-fail-under=80"

...pytest resolves that host pyproject.toml as its rootdir, inherits the addopts, and fails on the host's own coverage gate (harness collected at 0% → below threshold → non-zero exit). The harness then records functionally PASSING code as FAILED.

This is exactly the bug documented in sudo-ai-git/vulcanbench-findings: VulcanBench's declarative grader mis-scored every functional task as 0.0 for this reason; with -o addopts= neutralizing the leak, the same workspaces passed 10/10.

Related MCP server: forge-repo-mcp

The fix it hands you

When a workspace is flagged CORRUPTED, the tool returns the corrected command:

python -m pytest -o addopts= <tests>

-o addopts= strips inherited coverage/abort gates. (Or run the grader from outside the repo root.)

Tools

tool

purpose

inspect_workspace(path)

full analysis: ini chain, effective addopts, CLEAN/CORRUPTED/UNKNOWN verdict + corrected command

check_addopts(path)

thin boolean: corrupted + reasons

summarize(analysis)

one-line actionable summary string

Deterministic core (no deps)

The analysis walks the workspace directory up to filesystem root, reading pyproject.toml / pytest.ini / tox.ini / setup.cfg in pytest's first-found order, and extracts addopts. Flags:

  • coverage gates--cov, --cov-fail-under, --cov-report, --cov-config

  • abort/strict gates--maxfail, -x, --strict, --strict-markers, --pdb, --ff

Only gates that change exit codes / abort grading are flagged. A harmless addopts is reported CLEAN with the exact string.

Install & run (MCP stdio)

One command (recommended) — installs from the repo, no PyPI token needed:

uv tool install git+https://github.com/sudo-ai-git/mcp-benchmark-hygiene
mcp-benchmark-hygiene                        # run stdio server
mcp-benchmark-hygiene --http --port 8137     # or Streamable HTTP

Or with pipx: pipx install git+https://github.com/sudo-ai-git/mcp-benchmark-hygiene

Direct from source (fallback):

{ "mcpServers": {
    "benchmark-hygiene": { "command": "python3", "args": ["/abs/path/to/mcp_server.py"] }
}}

Requires the official mcp python package (pip install mcp). The deterministic core (inspect_workspace / check_addopts / summarize) imports and runs with zero dependencies — the mcp package is only needed for the stdio server.

Streamable HTTP (remote/Smithery-publishable)

python3 mcp_server.py --http --port 8137   # serves on http://<host>:8137/mcp/

Run with --http to serve over Streamable HTTP (a remote MCP endpoint) instead of stdio. This is the transport smithery mcp publish <url> expects for URL-based publishing — so once a Smithery service token exists, the server deploys as-is.

Example

inspect_workspace(path="/home/runner/vulcanbench/workspace/task-1")
→ {
    "ok": true,
    "workspace": "/home/runner/vulcanbench/workspace/task-1",
    "ini_chain": [{"file": "/home/runner/vulcanbench/pyproject.toml",
                   "addopts": "--cov=harness ... --cov-fail-under=80"}],
    "effective_addopts": "--cov=harness ... --cov-fail-under=80",
    "will_corrupt_grading": true,
    "verdict": "CORRUPTED",
    "fixed_command": ["python3", "-m", "pytest", "-o", "addopts=", "<tests>"],
    "reasons": ["coverage gate(s) present: ['--cov', '--cov-fail-under']"]
}

Verification

  • python3 test_detector.py — 5/5 core detection checks (root gate, nested inheritance, clean, abort gate, pyproject-no-pytest)

  • python3 test_e2e.py — drives the real MCP stdio transport (initialize → tools/call) and asserts CORRUPTED / CLEAN thread through the wire

Part of a family

This is one of three deterministic, no-LLM agent-trust MCP servers by sudo-ai-git:

Sibling product: mcp-token-saver — token-cost proxy + analyzer for agent conversations (dedupes redundant tokens before they're billed; live-proven 74% cut). Discussion

Also in the family (a free CLI, not an MCP server): harness-audit — deterministic agent-eval / benchmark-grading hygiene audit that catches the same silent config-leakage mis-scoring class. Free lead-magnet; the same verification discipline, zero dependencies, auditable line-by-line.

License & provenance

MIT. Independently derived from the documented VulcanBench #79 finding; no endorsement by or affiliation with morganlinton/VulcanBench implied.

Hire a custom integration

Need this connected to your internal system (auth, logging, security-scan pass, hosted)? Open a custom-build request. MIT reference assets are free to use either way.

Related MCP Connectors

Related MCP Servers

  • F
    license
    Not graded
    quality
    D
    maintenance
    Enables AI assistants to run and analyze pytest tests for desktop applications through interactive commands. Supports test execution, filtering, result analysis, and debugging for comprehensive test automation workflows.
    2
    -
  • A
    license
    Not graded
    quality
    B
    maintenance
    Local-first MCP and coding-agent reliability harness that captures bounded, sanitized failure evidence and generates deterministic executable regression tests. Capture is opt-in; no API key or hosted service is required.
    2
    Apache 2.0
  • A
    license
    A
    quality
    B
    maintenance
    Provides structured, sandboxed test and lint feedback for coding agents, returning compact typed verdicts with failure fingerprints instead of raw runner output. It enables impact-selected test execution in isolated containers, distinguishing pre-existing failures from regressions.
    4
    Apache 2.0