Skip to main content
Glama

Evidence Gate

Small local decisions, with visible evidence and an explicit review path.

Evidence Gate is an experimental, open-source workbench around the Laya decision model. Use it to assess whether a short passage supports a claim, classify a task, or define your own question and possible answers. It runs locally on Apple Silicon and exposes the same engine through a web interface, JSON API, CLI, and optional MCP tools for AI assistants.

This is a working prototype, not a new trained model. No superiority to Laya or Jev has been established. The contribution is the inspectable workflow, exact token budget checks, passage provenance, caching, review policy, and reproducible tests.

Try it

Requirements: Apple Silicon Mac, macOS 14+, Python 3.11+ (3.12 recommended), and uv. Tested on an M5 Pro, 24 GB RAM, macOS 27.0 with Python 3.12.13. Other OS versions have not been tested. The current model adapter is Apple-specific; Windows/Linux/CUDA support is a planned adapter, not a feature of this release.

From the downloaded or cloned project folder:

uv sync --python 3.12 --extra apple --extra mcp --extra dev
uv run evidence-gate serve

Open http://127.0.0.1:8765. Alternatively, on a Mac with uv installed, double-click Launch.command.

First launch downloads approximately 850 MB of model weights from Hugging Face. The base checkpoint is pinned to commit 20aed815fc6acde75733882e7ec0e3f28aeb9717. Weights are not included in this repo. After that download, inference can run offline:

HF_HUB_OFFLINE=1 uv run evidence-gate serve

Stop the server with Ctrl-C. If port 8765 is occupied, use serve --port 8766. The server binds only to loopback; a friend runs their own copy on their computer.

Related MCP server: jev-mcp

A three-minute demo

  1. Run CRISPR · What was actually measured? This uses a short attributed excerpt from a paper coauthored by Giulia Palermo. The question concerns reported methods.

  2. Select CRISPR · Catch an unsupported leap. The same passage does not establish universal absence of off-target effects. Inspect the model's judgment and probabilities.

  3. Try Research · Conflicting reports, an explicitly synthetic example. Each source is evaluated separately; there is no majority vote masquerading as truth.

  4. Switch to Decision playground. Replace the request and define your own choices.

  5. Paste your own short passage, then export the decision record as JSON.

The default review threshold is 0.80 for top probability and 0.20 for the winning margin. These are arbitrary starting policies, not calibrated accuracy guarantees. The research view always requires human review; a green policy pass is only a model/policy result. Paper URLs are attribution: the app does not fetch or authenticate them. Passage hashes identify exact supplied text, not whether it is genuine.

The model has a 512-token context limit including question, choices and input. Oversized inputs are rejected without truncation. This release accepts short text passages and JSON, not full PDFs, scanned papers or automatic literature searches.

Use in other software

uv run evidence-gate decide examples/route.json
uv run evidence-gate evidence examples/paper.json

While the server runs:

curl http://127.0.0.1:8765/api/decide \
  -H 'Content-Type: application/json' \
  --data-binary @examples/route.json

/api/evidence accepts the paper example. /api/status reports the loaded model. The decision API currently supports choice questions with 2–8 named options. It returns the entire probability distribution, proposal, policy status, margin, input digest, backend version/revision, local token count, cache status and timing. No tool action is automatically executed.

Connect an AI assistant

Install the mcp extra above, then configure your assistant to launch:

/absolute/path/to/evidence-gate/.venv/bin/evidence-gate mcp

This exposes local_decide and check_evidence through MCP stdio. For Codex:

codex mcp add evidence-gate -- /absolute/path/to/evidence-gate/.venv/bin/evidence-gate mcp

Use your actual installation path. The server loads the model on first tool use; the client may need a longer tool timeout on the first download. Download using the web workbench first to avoid that delay. See the official Codex MCP documentation.

Suggested instruction: “For short routing and classification tasks, use local_decide. Treat review results as unresolved. For scientific claims, use check_evidence on exact supplied passages and inspect the original source before drawing a conclusion.”

An assistant connection does not automatically replace its internal reasoning. MCP calls themselves add request/response context to a conversation. Savings are most plausible when application code invokes the decision engine before calling a large model, or handles many decisions in a batch. No net token or monetary savings have been measured. Zero generated output tokens means this model emits class probabilities rather than text; the JSON response still consumes tokens if included in an LLM prompt.

Tests and evidence

uv run pytest -q
uv run evidence-gate benchmark --output reports/my-run.json
uv run evidence-gate benchmark --dataset examples/evaluation.json --output reports/custom.json

Every custom benchmark case needs id, state, question, choices, and expected (a choice label). The bundled set contains 23 synthetic development examples. Its labels were authored during implementation, with two ambiguous absence-of-evidence examples clarified before the final runs. It is neither held out nor expert validated. The raw outputs, failures, backend revisions and dataset digest are in reports/. See RESULTS.md.

The test suite verifies software behavior using a test double; it does not establish model accuracy. Benchmarks use real downloaded model weights. Failures remain in the reports, including a prompt-injection example and ambiguous requests.

Privacy and limitations

  • The app makes no cloud inference calls and includes no telemetry or remote assets. Package installation and first model download require network access. Opening a source link contacts that source site.

  • No passages are saved to disk automatically. Up to 128 decisions live in an in-memory cache until the server exits. User-requested exports contain the input text and results; review them before sharing.

  • Model outputs can be wrong, including high-probability outputs. Evidence-checking is not scientific validation. The app does not propose or execute experiments.

  • Research assessments operate on individual passages. Missing context, figures, supplements, negation and scope differences can change the right answer.

  • More powerful hardware can improve throughput; it does not automatically improve the accuracy of unchanged weights. Better judgment requires better data, models or evaluation, not just more memory.

Development direction

Next steps are expert-labeled evaluation across unrelated tasks, a fixed held-out test set, domain calibration, comparisons against a simple rules baseline and Jev, a portable CPU/CUDA backend, longer-document retrieval with page provenance, and an optional independently evaluated stronger-model fallback. Fine-tuning should follow a stable evaluation, not precede it.

Original code: MIT. Dependencies, model weights and paper text retain their own licenses. See NOTICE.md. No affiliation with the cited researchers, UCLA, Convai Innovations, or TypeSafe AI is implied.

Related MCP Connectors

Related MCP Servers

  • A
    license
    A
    quality
    B
    maintenance
    Enables AI writing assistants to retrieve citable evidence from local PDFs and verify draft citations against their sources locally, providing verifiable support for claims.
    2
    1
    MIT
  • A
    license
    A
    quality
    A
    maintenance
    Enables agents to verify claims against cited evidence, screen content for prompt injection and relevance before reading it, and rank candidates by meaning, all with calibrated probability verdicts.
    11
    8,716 npm
    415
    MIT
  • A
    license
    A
    quality
    A
    maintenance
    Enables local AI models to structure and reason through debates using argument maps, with claims and supporting/attacking arguments automatically validated and organized.
    8
    AGPL 3.0
  • A
    license
    Not graded
    quality
    C
    maintenance
    Enables agents to fact-check claims, conduct verified research, make typed decisions with calibrated probabilities, and scan token risks through a hosted server with no API key required.
    MIT