evidence-lint
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@evidence-lintrun a QC pass on the draft memo and flag unsourced claims"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
evidence-lint
A QC gate that catches AI fabrication and jargon before your research ships.
AI research tools produce confident, well-formatted output — and quietly invent things: source
URLs that 404, "facts" with no source, numbers that claim triangulation but trace to one press
release, stale figures dressed up as current, and dense consultant-speak that hides weak thinking.
evidence-lint checks all of that mechanically, so nothing reaches a model, deck, memo or filing
on a failing gate.
Built for MBB-style strategy work, legal/compliance research, due diligence, and any evidence-heavy deliverable where every number must trace to a real, live source and the prose must be plain.
Zero dependencies — Python 3.8+ standard library only.
CI-friendly — exit codes
0PASS /1WARN /2FAIL.Config-driven — tune the word lists, thresholds and trusted publishers per engagement.
Three inputs — a simple generic JSON, an agent-swarm journal directory, or a markdown vault.
Quickstart
python3 -m evidence_lint --input examples/sample_input.jsonOr install as a package (zero runtime dependencies; [mcp] adds the MCP server):
pip install "evidence-lint @ git+https://github.com/Dankaro-projects/evidence-lint.git"
evidence-lint --input evidence.jsonYou'll see dead-URL FAILs, a sourceless-fact FAIL, jargon/plain-language WARNs, and a stale-figure
WARN — then a GATE: line and lint_report.json.
Skip the network probe (fast, offline):
python3 -m evidence_lint --input examples/sample_input.json --no-urlsPoint it at agent-swarm journals instead:
python3 -m evidence_lint --journal ./run1_dir ./refresh_dir --out report.jsonLint a markdown knowledge vault (Obsidian-style or plain folders):
python3 -m evidence_lint --vault examples/sample_vaultYou'll see missing-metadata and broken-wiki-link WARNs, plus stub/orphan INFOs.
(Staleness and inbox-TTL checks key off file age — backdate a file with
touch -d "2025-01-01" file.md to watch them fire.)
Related MCP server: Recommend Agentic Trust Layer
Input format (generic)
One JSON file. Every section optional; include what you have.
{
"sources": [{"title","url","type","grade"}],
"claims": [{"claim","claim_class","evidence_grade","confidence","source_urls":[],"unit","period","is_hypothesis"}],
"metrics": [{"metric_name","value","unit","period","basis","method","source_urls":[]}],
"accepted": [{"metric_name","entity","period","confidence","triangulation","supersedes","sources":[{"title","url","grade"}]}],
"prose": [{"where","text"}]
}claim_class: F fact · E estimate · H hypothesis · A anecdote · I inference · R recommendation.
grade: A primary/current · B credible secondary · C single/dated · D unverified/marketing.
What it checks
Integrity — can FAIL the gate (fabrication risk)
Check | Fires when |
| a source URL 404s or won't resolve. Dead path on an unknown domain = FAIL; on a known publisher = WARN (repair the link). |
| example.com / not a real URL. |
| an |
| a triangulated metric has no source. |
| a claim flagged |
Rigor — WARN
weak_triangulation (claims independence but all sources one domain) · grade_source_mismatch
(A/B claim on press-release/aggregator only) · estimate_without_method · stale_load_bearing
(high-confidence figure older than the threshold, no refresh) · metric_no_unit / metric_no_period.
The stale threshold rolls by default: current_year - stale_year_offset (offset 2), so the
config never goes stale itself. Pin it absolutely with an integer load_bearing_stale_year or the
LOAD_BEARING_STALE_YEAR env var.
Vault hygiene — WARN / INFO (--vault)
The same QC philosophy pointed at the knowledge base instead of the deliverable: a stale
or unqueryable note passes silently into a deck until the evidence checks catch the symptom
downstream. --vault catches it at the source.
Check | Fires when |
| note has no metadata header, or lacks required keys (default: type, summary, tags, status, date). |
| note untouched past |
| file sat in an inbox folder past |
|
|
| no other note links here (INFO). |
| body under |
| note crosses two entries of |
Vault findings never FAIL the gate — FAIL stays reserved for fabrication risk.
Plain language & prose — WARN / INFO
jargon_plain_word (utilise→use, facilitate→help, prior to→before …) · nominalization
(make a decision→decide) · long_sentence · hard_to_read (Flesch reading ease) ·
adjective_as_analysis (strong/leading/mature without a number) · ai_tell (WARN tier + an
ai_tells_soft INFO tier; longest-phrase-first with overlap suppression) · participle_padding
(", underscoring/reflecting …" fake-depth tails) · bare_acronym (acronym with no gloss on first
use — expansion, appositive, or descriptive phrase all count as glossed; common_acronyms exempts
everyday ones) · not_answer_first · hedge_density.
Plain-language principle: in high-stakes writing the goal is not to sound smart — it is to use simple words for complex ideas. Short sentences. Say the thing.
Configuration
All word lists, thresholds and trusted publishers live in config.default.json. Copy it, edit, and
pass with --config myconfig.json to specialise per discipline (e.g. a legal profile with different
jargon and a stricter stale-year). Keys: banned_adjectives, hedges, ai_tells, jargon_swap,
nominalizations, known_publishers, placeholder_domains, weak_types, long_sentence_words,
hard_readability_fre, load_bearing_stale_year, url_timeout, url_workers, exit codes, and the
vault_* family (vault_required_keys, vault_stale_days, vault_inbox_dirs, vault_inbox_ttl_days,
vault_min_words, vault_statuses, vault_entities, vault_ignore_dirs, vault_extensions).
Output
lint_report.json: totals, findings by severity, findings by check, URL-status summary, the gate,
and the full findings list (severity, check, where, detail). Wire it into a build to render a
QC sheet/section inside the deliverable, or read the exit code in CI. For pipelines that must flag
but never abort, --advisory prints and writes everything as usual but always exits 0.
Fixtures — the gate's own regression gate
fixtures/ holds inputs pinned to an outcome: must-fail/ inputs must keep failing,
must-warn/ warning, must-pass/ passing. Run python3 fixtures/run_fixtures.py before any
release. If an engine change flips a must-fail fixture to PASS, the change softened the gate —
the change is wrong, not the fixture.
MCP server
The gate is also an MCP server, so any agent (Claude, Cursor, your own pipelines) can call it mid-workflow — check evidence before it becomes output, not after. All checks stay deterministic and reproducible; no LLM judges an LLM.
The MCP layer needs Python 3.9+ (asyncio.to_thread); the core engine stays 3.8+.
pip install mcp # only dependency, core engine stays stdlib-only
python3 -m evidence_lint.mcp.server # stdio transportClaude Desktop / Claude Code config:
{
"mcpServers": {
"evidence-lint": {
"command": "python3",
"args": ["-m", "evidence_lint.mcp.server"],
"cwd": "/absolute/path/to/evidence-lint"
}
}
}Installed as a package (pip install "evidence-lint[mcp] @ git+https://…/evidence-lint.git"),
skip cwd and use "command": "evidence-lint-mcp" with no args.
Tools: evidence_lint_evidence (full gate on a claims/sources/metrics payload),
evidence_lint_prose (fast plain-language pass, no network), evidence_lint_vault
(knowledge-base hygiene), evidence_probe_urls (citation spot-check),
evidence_list_profiles, evidence_explain_finding (what fired, why it matters, how to fix).
Profiles live in evidence_lint/mcp/profiles/ — a pack is a partial config override plus a
_profile description block. Shipped: generic, maritime, science (journals, trial
registries, health agencies, courts), nonprofit-news (brochure-tone adjectives banned,
syndicated PR treated as weak sourcing), luxury-retail (catalogue-cringe + heritage AI
tells). Add an industry (legal, medical, financial-DD) or role (analyst-coaching vs.
partner-gate severity) by dropping in one JSON file; no code changes.
URL probing is off by default in the MCP tools (check_urls: false) — hosted contexts often
lack egress; turn it on per call when the environment allows.
Roadmap (productization)
More discipline profiles (legal, medical, financial-DD) shipped as ready configs.
CSV/Markdown report writers and an HTML dashboard.
Adapters for common research-tool exports.
Optional LLM prose-reviewer pass for "is it compelling", layered on top of these deterministic checks.
License
MIT — see LICENSE.
This server cannot be deployed
Maintenance
Related MCP Connectors
Preflight QA for AI-agent deliverables with structured verdicts and repair guidance.
Pre-execution governance for AI agents. Deterministic PASS/FAIL/REVIEW verdicts, replayable proof.
Deterministic fact verification for AI agents — checksums & curated data, not guesses.
Real-time fact-check, citation verification, and source-freshness for AI agents.
Related MCP Servers
- AlicenseCqualityDmaintenanceEnables acceptance gates for AI coding-agent runs by recording evidence, running deterministic validation, applying a quality gate, and rendering auditable outcomes.7Apache 2.0
- AlicenseNot gradedqualityBmaintenanceEnables agents to verify claims with evidence-based truth scores and confidence levels by running a deterministic pipeline of evidence lanes and adversarial checks.26MIT
- AlicenseNot gradedqualityBmaintenanceEnables agents to verify their own output mid-task by checking every claim against provided sources, returning supported, partial, unsupported, or contradicted verdicts with exact citations.MIT
- AlicenseBqualityCmaintenanceEnables AI agents to verify technical claims against supplied evidence, identify unsupported assumptions and contradictions, and recommend the smallest next check before acting.5MIT