code-evidence
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@code-evidencerun tests and lint, then give me the failure summary"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Code Evidence
Give coding agents compact failure reports and evidence tied to the code they checked.
Code Evidence combines Failure Lens and Proof of Change in a local Python core, CLI, and STDIO MCP server. It needs no LLM or paid API. Python context selection is available; sandbox execution remains a future extension.
What Works Today
Select Python symbols and heuristic callers/callees for a task, with relative paths, line spans, file hashes, inclusion reasons, and a serialized JSON byte budget.
Execute explicitly configured check names, with no arbitrary command argument exposed through MCP.
Store run history and redacted logs in local SQLite, retaining the latest 50 runs.
Extract and deduplicate diagnostic lines from Python/unittest/pytest, Ruff-style, and TypeScript-style output; retain a tail when no pattern matches.
Read log fragments on demand instead of returning full logs to the model.
Compare sampled diagnostics between attempts, including newly observed and absent messages.
Track included file content, configured check policy, and the server's Python/environment fingerprint.
Mark results
fresh,stale, orunknown; never reuse a cached result as a new execution.Distinguish nonzero exits, launch errors, timeouts, and capture errors.
This is trusted local execution, not a sandbox. Configure it only for projects and check commands you trust. A test or build hook can run arbitrary code with your user permissions, access files, contact the network, or alter its own logs.
Related MCP server: verdict
Install
Python 3.12+ is required. Install into the Python environment you want the checks to use:
git clone https://github.com/yarakrot/code-evidence.git
cd code-evidence
py -3.12 -m venv .venv
.\.venv\Scripts\python.exe -m pip install -e ".[dev,mcp]"The core has no runtime dependencies. The optional mcp extra uses the official MCP Python SDK; the dev extra installs Ruff.
CLI Workflow
.\.venv\Scripts\code-evidence.exe --project . inspect
.\.venv\Scripts\code-evidence.exe --project . run tests lint format --execute
.\.venv\Scripts\code-evidence.exe --project . summaryinspect and summary do not execute project checks. run requires --execute. Checks run sequentially with stdin disabled, shell=False, a filtered environment, and a configured timeout. CLI returns 1 if a selected check fails, and 2 for invalid input/policy.
Compare attempts or read the stored log:
code-evidence --project . summary --run-id NEW_RUN_ID --compare-to OLD_RUN_ID
code-evidence --project . evidence RUN_ID tests --start-line 1 --max-lines 40Output is JSON. run_id comes from the returned receipt. Do not copy the placeholder IDs literally.
Context Budget
code-evidence --project . context "get_run_summary freshness" --budget-bytes 12000Selection uses Python AST and lexical keyword matching, not embeddings or semantic understanding. Include symbol names or paths when describing a task in another language. Call-name links are heuristic and may be ambiguous. Every request rereads included source; no stale persistent index is reused. Overlapping ranges are avoided, the highest-ranked oversized symbol can return an explicitly marked partial range; other oversized symbols can be omitted, and the report exposes coverage and omitted matches.
The budget covers the compact ASCII-escaped JSON object, including metadata, rather than an estimated token count. MCP transport wrappers and client rendering can add overhead. Maximum budget is 60,000 bytes; indexing is limited to 300 Python files, 1 MB per parsed file, and 2500 symbols within the source inventory limits. Known secrets are redacted from fragments, but unrecognized private content may remain.
Configure Another Project
Create code-evidence.toml in the trusted project's root:
schema_version = 1
[checks.tests]
argv = ["{python}", "-m", "unittest", "discover", "-s", "tests", "-v"]
timeout_seconds = 60
[checks.lint]
argv = ["{python}", "-m", "ruff", "check", "."]
timeout_seconds = 60{python} resolves to the interpreter running Code Evidence. Install the project's test dependencies in that same environment or explicitly configure a trusted executable. The policy is loaded once at server startup; an agent cannot replace it with a new argv through a tool call. Review edited policy before restarting the server. Commands are not automatically generated from a README.
MCP Tools
Tool | Behavior |
| Select bounded redacted Python fragments for a task |
| Show configured checks and bounded changes since the last run |
| Run selected configured check names if enabled at startup |
| Read diagnostics, freshness, and optional comparison |
| Read a bounded fragment of a stored redacted log |
Launch with code-evidence --project /path/to/project serve. This defaults to read-only tool behavior: run_checks returns a policy error. To permit configured commands, add --allow-execution when starting the server.
Example configuration for an MCP client that supports mcpServers:
{
"mcpServers": {
"code-evidence": {
"command": "C:/path/to/project/.venv/Scripts/python.exe",
"args": [
"-m", "code_evidence.cli",
"--project", "C:/path/to/project",
"serve"
]
}
}
}Replace both example paths. This configuration is read-only; enabling execution is an explicit startup decision. Different clients use different configuration formats. The server does not modify your client settings or register itself automatically. Protocol output uses stdout; diagnostic logging uses stderr.
Evidence Semantics
fresh: the included source, startup policy, and server environment still match the run, and included source did not change during it.stale: one of those observed fingerprints changed.unknown: inventory was incomplete or the run did not finish normally.
Freshness is not proof of correctness, a reproducible build, or verification of external services, ignored secrets, separate tool environments, OS binaries, or transient file changes restored before the final snapshot. Success comes from process exit status, not a reassuring line in stdout. A process can exit 0 without adequate tests.
Source inventory excludes Git metadata, virtual environments, runtime folders, .env files, and private-key extensions. Files are bounded to 2 MB each, 30 MB total, and 5000 included entries. Logs retain at most the last 1 MB per check and report truncation. Diagnostic summaries retain up to 12 distinct messages; run comparisons operate on those sampled messages. Evidence fragments are capped at 100 lines and 16,000 characters. Changed-file lists are capped at 50 entries per category with full counts.
SQLite and logs remain in .code-evidence/, excluded from Git. Output redaction is best effort: review exports before sharing. The tool filters inherited environment variables but does not prevent trusted commands from reading your files or retrieving credentials themselves.
Development
.\.venv\Scripts\python.exe -m unittest discover -s tests -v
.\.venv\Scripts\ruff.exe check .
.\.venv\Scripts\ruff.exe format --check .CI tests Windows/Linux on Python 3.12 and 3.13, including a real STDIO MCP discovery and tool-call session. Tests use synthetic trusted commands, never private repository snapshots or real keys.
Next Steps
See architecture, security boundaries, and roadmap. Planned: richer context selection, diagnostic adapters, README Reality integration, isolated execution, and measured agent benchmarks. No token-saving percentage is claimed by this release.
License
MIT. See LICENSE.
This server cannot be deployed
Maintenance
Related MCP Connectors
Evidence-backed architecture-quality analysis for Python agent applications.
Deterministic context layer for your codebase: change impact, blast radius, answers with receipts.
Change-aware CI validation and affected-test guidance for coding agents.
Ask a codebase what calls what: search, blast radius, paths between symbols, and diffs.
Related MCP Servers
- AlicenseNot gradedqualityBmaintenanceProvides local static-analysis tools for Python that let coding agents trace call paths, branch guards, side effects, and change impact, producing citation-ready answers with explicit uncertainty signals.2MIT
- AlicenseAqualityBmaintenanceProvides structured, sandboxed test and lint feedback for coding agents, returning compact typed verdicts with failure fingerprints instead of raw runner output. It enables impact-selected test execution in isolated containers, distinguishing pre-existing failures from regressions.4Apache 2.0
- AlicenseNot gradedqualityBmaintenanceEnables deterministic evaluation of coding agents by exposing controlled repository tools and returning structured verification reports with pattern checks and repeat-run comparisons.MIT
- AlicenseNot gradedqualityAmaintenanceProvides local-first repository architecture analysis with file-level proof, enabling agents to map codebases, locate implementations, trace call paths, and assess change impact with deterministic evidence.4 npm2MIT