benthic-mcp
OfficialClick on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@benthic-mcpfind signed datasets about ocean acidification"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
benthic-mcp
A read-only MCP server over signed Benthic Data Provenance datasets, plus the evaluation harness that measures whether the playbook it serves actually improves an agent's answers.
The server exposes four tools:
tool | what it does |
| finds the smallest relevant signed relation, its columns, and its join paths |
| runs a bounded single-relation query with compact filter, aggregate, and |
| runs a multi-relation plan using only signed BDP join paths |
| runs three allowlisted spatial RPCs |
Plain-language questions are interpreted by the calling chat model, which turns a question into compact tool arguments. This service validates and executes those arguments without embedding a model of its own, so what an agent is allowed to ask is decided by a signed manifest rather than by a prompt.
Why the trust model is the point
The catalog is signed. The server pins an Ed25519 authority key, verifies RFC 8785 canonical JSON, SHA-256 payload hashes, and every member manifest against the collection hash, and then serves only what the signature covers:
only relations marked
queryable: truein a signed manifestonly join paths present in the signed join graph
only an allowlist of RPC endpoints, hand-written in
src/benthic_mcp/catalog.pyno arbitrary URLs, relations, SQL, or write operations
API and RPC traffic restricted to HTTPS endpoints on
benthic.io
The playbook this server serves to the model is prose, and prose is model-authored. It can therefore
never widen any of the above. A fact the signed catalog already carries is always read from the
verified manifest at serve time, so guidance cannot contradict it. See
SECURITY.md for the trust boundary in one page.
Related MCP server: Vela MCP Server
Requirements
Python 3.12 or newer
Network access to
https://benthic.ioA local
llama-serverfor the agent that calls the tools
git clone https://github.com/benthic-io/benthic-mcp
cd benthic-mcp
uv sync --frozenRunning it
Two transports are supported, and they are not interchangeable with the two MCP clients in llama.cpp:
client | configured by | format | transport |
llama-server |
| object with | stdio only |
llama.cpp Web UI | its MCP settings dialog | array of server objects | HTTP and SSE |
Set BENTHIC_MCP_TRANSPORT=stdio to serve over stdio, which is what llama-server's own client needs.
A cold spawn is about two seconds. config/llama-mcp.README.md has both config formats, the accepted
schema, and the two ways to fail silently; config/llama-mcp-servers.example.json and
config/llama-mcp.example.json are the two files.
For the HTTP service, which the Web UI uses and which needs a URL and a token because a browser cannot spawn a subprocess:
mkdir -p ~/.config/benthic-mcp
cp config/benthic-mcp.env.example ~/.config/benthic-mcp/env
chmod 600 ~/.config/benthic-mcp/env
# Replace the token placeholder, and set MCP_HOST / LLAMA_HOST to your own hosts:
openssl rand -base64 36scripts/start-benthic-mcp.shThe service listens on port 8082 by default and exposes the Streamable HTTP MCP provider at
http://<mcp-host>:8082/mcp. /health requires the same bearer token.
examples/ holds two starting points that need editing for your machine:
start-llama-server.sh is one tuned llama-server configuration (the one the numbers below were
measured under), and benthic-mcp.service is a user-level systemd unit.
Set enable_thinking=false on the calling model. It cuts completion tokens by about 64% and makes
the result reproducible, which is a deployment instruction for whatever calls the MCP rather than a
server setting. Its effect on the pass rate is about one case and is not established.
docs/findings.md records that asking the model in a system prompt not to deliberate does not work.
The playbook
benthic_playbook returns dataset-specific instructions so the agent does not rediscover the schema
each session. Three modes, set by BENTHIC_PLAYBOOK_MODE:
off- no playbook; the tool reports that it is disabledseed(default) - the curated seed insrc/benthic_mcp/seed.pyactive- the promoted playbook, falling back to the seed if the file is missing
An agent can also report a mistake it made with benthic_report. A reported lesson is never served
until it has been measured to help. Accumulation used to promote on catalog verification alone,
which proves a statement is true and says nothing about whether it changes behaviour; twelve
accumulated lessons were later measured and none earned a place, while one made a case measurably
worse. Lessons are now attributed against the case they came from before they can be served, and a
regression is quarantined. The measurement, the dead ends, and the limits of the instrument are in
docs/findings.md.
What is actually verified, and how
The statistical A/B is no longer the gate. It could not be: the suite's minimum detectable effect is 7.81 cases against 4.5 cases of total headroom, so a gate built on it can only ever detect harm, which is why it kept reporting that nothing worked. Classifying all 1534 stored case-runs found that 2 of 33 cases have ever moved because server code changed. Three things replaced it.
Contracts. 24 properties over the catalog and join subsystems, stated in the code's own
vocabulary, with no model and no network, in tests/test_contracts_catalog.py. Each is falsifiable,
and each one was: six arrived violated and are now fixed, so the file's xfail list is empty. A
contract that arrives already green is asserting nothing.
Answer checking. The scorer reads the result of a call, not just its route, and score_case
returns asserted - which checks this capability actually ran. Every check defaults to True, so a
row_count_check: true on a discovery case read as an assertion that was made and passed when
nothing was checked.
A canary that was tested by breaking it. eval/canary/questions.json is six cases: one per
signed join path, plus the one non-join case that expects rows.
server | verdict |
correct | 18/18 green |
| 8/12, red on exactly the two district cases |
The four zero-row guards stayed green against the broken server, because silence satisfies a zero expectation. That is the predicted behaviour and the reason the canary keeps cases in both directions. It runs in 184s/rep against 778s/rep for the full suite.
uv run python eval/run_eval.py --questions eval/canary/questions.json --reps 3Report it as green or red, never as a percentage: six cases over four paths is a step function. The 33-case suite is tier 2, reported alongside it but not gating.
What they found
Five real defects, none of which the model-in-the-loop apparatus saw:
A signed join between a
stringcolumn holding'03'and anintegercolumn holding3compared them in Python, matched nothing, and returned 0 rows where 127 exist. 55 of 55 observed calls. The two affected cases scored as passes for 51 stored runs, and abenthic_reporthad diagnosed it correctly while the harness could not see it.The columns a
context_conditionsentry compares were never fetched, so no partial signed join could ever match a row, for any input - the unit test hand-built rows that already had the column.numberwas missing from the numeric type set. It is the catalog's own name for a decimal and the second most common declaration in the signed manifest, at 427 of 3419 columns.order=andhavingeach computed a comparison the caller never asked for, on a column whose declared type PostgREST did not honour, and raised aTypeErrorthe tool wrapper does not catch.Two of the suite's cases asked for a chain whose first hop is provably empty, and the scorer counted a truthful dead end as a failure. They now ask where the chain terminates, 4/4.
What has been measured
33 generated cases against a 35B MoE coder model, 6-turn budget. Read this as a count of cases answered correctly, not as a measurement of anything: the suite's minimum detectable effect is larger than the whole remaining failure set, so a difference of one or two cases between rows here is not distinguishable from noise.
configuration | result |
seed, | 32/33 both reps, holdout 8/8 and 7/8 |
completion tokens, thinking off | 19,092 / 20,226 per rep, against 53,231 / 56,917 with it |
enable_thinking=false cuts completion tokens by about 64%, and nothing in the run-to-run spread
comes close to that. Its effect on the pass rate is not claimed, because this suite cannot resolve an
effect that size. docs/findings.md records the arithmetic and the three measurements that got there.
The honest position on self-improvement is worth stating plainly: the loop does not learn. Its
8 recorded lessons were never measured - source_ref was added after all of them were written, so
the attribution gate had nothing to read and had run zero times. Provenance has been reconstructed
from the round log, and 6 lessons are measurable for the first time. The loop also has no way to
express a code change, which is the only kind that has ever worked here - five of five kept
interventions were code, four of four reverted ones were prose. The mechanism that works is
contracts plus a reviewer.
Evaluation
# Tier 1, the gate: six cases, one per signed join path, 184s/rep
uv run python eval/run_eval.py --questions eval/canary/questions.json --reps 3
# Tier 2, report only: the full 33, 778s/rep. Measures turn discipline more than server correctness
uv run python eval/run_eval.py # run it
uv run python eval/generate_cases.py # rebuild it from the signed catalog
uv run python eval/run_eval.py --case-filter multi_step --reps 5Each run writes the cases, model transcript, tool arguments and results, token usage, timings, and a
report under eval/runs/<run-id>/. Expected values come from direct signed PostgREST and RPC probes,
so the evaluator does not depend on the MCP query implementation it is scoring.
Four hand-verified cases in eval/golden/questions.json must pass with the seed playbook alone. They
are a check on the harness, the scorer, and the tool surface, so a failure there is a bug rather than a
result. tests/test_golden.py re-checks their expected values against the signed catalog, so a catalog
change cannot leave the suite quietly stale. Run it at --max-turns 6, not the 5 the generated suite
uses: one case needs six to test what it is for and was failing about one run in five at five, which is
a tripwire firing on a capability it is not watching. At six it is 12/12.
Long runs belong in tmux so they can be watched:
scripts/tmux-run.sh my-run logs/run.log uv run python eval/run_eval.py
scripts/tmux-ls.shTwo attribution tools sit above the suite, and they answer different questions:
# Does this one lesson fix the case it came from? Cheap, but biased against general advice.
uv run python eval/attrib.py --lesson-id <id> --case <case-id> --playbook eval/harness/cache/playbook.json
# Does this rule earn a place in the always-on core? Two full runs of the tuning split.
uv run python eval/attribute_suite.py --rule "Never end the turn without a final answer." --reps 1Configuration
Variable | Default | Purpose |
|
| Signed BDP document root |
|
| Signed collections to load |
| Current pinned key | Trusted Ed25519 public keys |
|
| Verified catalog cache |
|
| Refresh interval |
|
| Maximum stale-cache fallback age |
|
| Upstream request timeout |
|
| Maximum PostgREST page size |
|
| Default result limit |
|
| Maximum complete source scan |
|
| Per-response byte limit |
|
| HTTP bind address |
|
| HTTP port |
|
| Streamable HTTP path |
| loopback only | Accepted Host headers |
| loopback only | Accepted browser origins |
|
|
|
| unset | Required for the HTTP transport only |
|
|
|
|
| Promoted playbook location |
|
| Token cap for the always-on core slice |
|
| Record objective tool-call traces |
|
| Trace retention |
|
| Lesson retention |
|
| Store question text from |
| local llama-server | Model used by the consolidator |
The shipped allow-lists name loopback only, deliberately: they gate the Host and Origin headers,
so a default that named a particular machine would silently authorise that host for anyone who
installed the server without reading the configuration.
The fetched keys.json is not used as a trust root. Add rotated trusted keys through
BENTHIC_TRUSTED_KEYS only after out-of-band verification.
The API does not expose server-side grouped aggregates. Large calendar-year scans can exceed the complete-scan or upstream timeout limit; the service reports that limitation rather than returning partial totals. Narrow the date range or use a signed pre-aggregated relation when available.
If the signed catalog changes after a playbook was generated, the fingerprint no longer matches and the server serves the static core plus a staleness notice instead of stale dataset detail, until the candidate is re-consolidated and re-gated.
Development
uv run ruff check .
uv run ruff format --check .
uv run pyright
uv run pytest -m "not live"The non-live marker excludes tests that need a real BDP endpoint and a running llama-server:
BENTHIC_LIVE_TESTS=1 uv run pytest -m live
BENTHIC_LIVE_TESTS=1 BENTHIC_LIVE_REGRESSION=1 uv run pytest -m livetests/test_loop_invariants.py is worth reading before changing the improvement loop. It states the
properties that make a self-modifying system trustworthy - nothing is served without a measured
effect, a verdict is auditable, the document cannot self-perpetuate - rather than examples of them.
Long-running changes to the playbook belong under eval/harness*/ with BENTHIC_CACHE_DIR pointed
somewhere disposable, so a run never touches the live store:
BENTHIC_CACHE_DIR=eval/harness/cache uv run python eval/harness.py --rounds 6Documentation
docs/findings.md- what the harness measured, what did not work, and the limits of the measurementSECURITY.md- reporting a vulnerability, deployment notes, and the trust boundary
License
MIT. See LICENSE.
This server cannot be deployed
Maintenance
Related MCP Connectors
- memoricOAuthio.memoric
Provenance-first database for teams and agents: every value carries sources, rules and coverage.
Machine-native research commons for agent evidence, discovery, rooms, and bounded research quests.
Query your org's data in natural language — read-only MCP access to SQL, NoSQL, files & warehouses.
A public commons for agents to search and share reusable findings and open research questions.
Related MCP Servers
- AlicenseBqualityAmaintenanceA research-informed MCP server that enables natural language question answering over local dataframes (CSV, Parquet, or Pandas) with safe, read-only execution and typed analysis plans.3MIT
- AlicenseNot gradedqualityBmaintenanceEnables governed, agent-agnostic data exploration by allowing users to ask natural language questions through MCP-compatible agents, executing safe, permission-scoped queries against data sources and returning interactive charts.15 npmApache 2.0
- AlicenseNot gradedqualityAmaintenanceEnables natural-language queries over data warehouses with catalog-grounded semantics and per-query authorization, returning answers with attached reasoning.Apache 2.0
- AlicenseNot gradedqualityBmaintenanceEnables natural language queries to be converted into policy-verified SQL, vector search, and knowledge graph plans, with evidence-backed answers and an audit log.1Apache 2.0