Evidence Gate
README.md
# Evidence Gate
**Small local decisions, with visible evidence and an explicit review path.**
Evidence Gate is an experimental, open-source workbench around the Laya decision
model. Use it to assess whether a short passage supports a claim, classify a task,
or define your own question and possible answers. It runs locally on Apple Silicon
and exposes the same engine through a web interface, JSON API, CLI, and optional
MCP tools for AI assistants.
This is a working prototype, not a new trained model. No superiority to Laya or Jev
has been established. The contribution is the inspectable workflow, exact token
budget checks, passage provenance, caching, review policy, and reproducible tests.
## Try it
Requirements: Apple Silicon Mac, macOS 14+, Python 3.11+ (3.12 recommended), and
[uv](https://docs.astral.sh/uv/getting-started/installation/). Tested on an M5 Pro,
24 GB RAM, macOS 27.0 with Python 3.12.13. Other OS versions have not been tested.
The current model adapter is Apple-specific; Windows/Linux/CUDA support is a
planned adapter, not a feature of this release.
From the downloaded or cloned project folder:
```sh
uv sync --python 3.12 --extra apple --extra mcp --extra dev
uv run evidence-gate serve
```
Open **http://127.0.0.1:8765**. Alternatively, on a Mac with uv installed,
double-click `Launch.command`.
First launch downloads approximately 850 MB of model weights from Hugging Face.
The base checkpoint is pinned to commit
`20aed815fc6acde75733882e7ec0e3f28aeb9717`. Weights are not included in this repo.
After that download, inference can run offline:
```sh
HF_HUB_OFFLINE=1 uv run evidence-gate serve
```
Stop the server with Ctrl-C. If port 8765 is occupied, use `serve --port 8766`.
The server binds only to loopback; a friend runs their own copy on their computer.
## A three-minute demo
1. Run **CRISPR · What was actually measured?** This uses a short attributed excerpt
from a paper coauthored by Giulia Palermo. The question concerns reported methods.
2. Select **CRISPR · Catch an unsupported leap**. The same passage does not establish
universal absence of off-target effects. Inspect the model's judgment and probabilities.
3. Try **Research · Conflicting reports**, an explicitly synthetic example. Each
source is evaluated separately; there is no majority vote masquerading as truth.
4. Switch to **Decision playground**. Replace the request and define your own choices.
5. Paste your own short passage, then export the decision record as JSON.
The default review threshold is 0.80 for top probability and 0.20 for the winning
margin. These are arbitrary starting policies, not calibrated accuracy guarantees.
The research view always requires human review; a green policy pass is only a
model/policy result. Paper URLs are attribution: the app does not fetch or authenticate
them. Passage hashes identify exact supplied text, not whether it is genuine.
The model has a 512-token context limit including question, choices and input.
Oversized inputs are rejected without truncation. This release accepts short text
passages and JSON, not full PDFs, scanned papers or automatic literature searches.
## Use in other software
```sh
uv run evidence-gate decide examples/route.json
uv run evidence-gate evidence examples/paper.json
```
While the server runs:
```sh
curl http://127.0.0.1:8765/api/decide \
-H 'Content-Type: application/json' \
--data-binary @examples/route.json
```
`/api/evidence` accepts the paper example. `/api/status` reports the loaded model.
The decision API currently supports **choice** questions with 2–8 named options.
It returns the entire probability distribution, proposal, policy status, margin,
input digest, backend version/revision, local token count, cache status and timing.
No tool action is automatically executed.
## Connect an AI assistant
Install the `mcp` extra above, then configure your assistant to launch:
```sh
/absolute/path/to/evidence-gate/.venv/bin/evidence-gate mcp
```
This exposes `local_decide` and `check_evidence` through MCP stdio. For Codex:
```sh
codex mcp add evidence-gate -- /absolute/path/to/evidence-gate/.venv/bin/evidence-gate mcp
```
Use your actual installation path. The server loads the model on first tool use;
the client may need a longer tool timeout on the first download. Download using the
web workbench first to avoid that delay. See the
[official Codex MCP documentation](https://developers.openai.com/codex/mcp).
Suggested instruction: “For short routing and classification tasks, use local_decide.
Treat review results as unresolved. For scientific claims, use check_evidence on exact
supplied passages and inspect the original source before drawing a conclusion.”
An assistant connection does not automatically replace its internal reasoning.
MCP calls themselves add request/response context to a conversation. Savings are
most plausible when application code invokes the decision engine before calling
a large model, or handles many decisions in a batch. **No net token or monetary
savings have been measured.** Zero generated output tokens means this model emits
class probabilities rather than text; the JSON response still consumes tokens if
included in an LLM prompt.
## Tests and evidence
```sh
uv run pytest -q
uv run evidence-gate benchmark --output reports/my-run.json
uv run evidence-gate benchmark --dataset examples/evaluation.json --output reports/custom.json
```
Every custom benchmark case needs `id`, `state`, `question`, `choices`, and
`expected` (a choice label). The bundled set contains 23 synthetic development
examples. Its labels were authored during implementation, with two ambiguous
absence-of-evidence examples clarified before the final runs. It is neither held
out nor expert validated. The raw outputs, failures, backend revisions and dataset
digest are in `reports/`. See [RESULTS.md](RESULTS.md).
The test suite verifies software behavior using a test double; it does not establish
model accuracy. Benchmarks use real downloaded model weights. Failures remain in
the reports, including a prompt-injection example and ambiguous requests.
## Privacy and limitations
- The app makes no cloud inference calls and includes no telemetry or remote assets.
Package installation and first model download require network access. Opening a
source link contacts that source site.
- No passages are saved to disk automatically. Up to 128 decisions live in an
in-memory cache until the server exits. User-requested exports contain the input
text and results; review them before sharing.
- Model outputs can be wrong, including high-probability outputs. Evidence-checking
is not scientific validation. The app does not propose or execute experiments.
- Research assessments operate on individual passages. Missing context, figures,
supplements, negation and scope differences can change the right answer.
- More powerful hardware can improve throughput; it does not automatically improve
the accuracy of unchanged weights. Better judgment requires better data, models
or evaluation, not just more memory.
## Development direction
Next steps are expert-labeled evaluation across unrelated tasks, a fixed held-out
test set, domain calibration, comparisons against a simple rules baseline and Jev,
a portable CPU/CUDA backend, longer-document retrieval with page provenance, and
an optional independently evaluated stronger-model fallback. Fine-tuning should
follow a stable evaluation, not precede it.
Original code: MIT. Dependencies, model weights and paper text retain their own
licenses. See [NOTICE.md](NOTICE.md). No affiliation with the cited researchers,
UCLA, Convai Innovations, or TypeSafe AI is implied.
This server cannot be deployed
Maintenance
ActivityMaintained
ResponsivenessNo issues