SwarmLabs MCP Server
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@SwarmLabs MCP ServerRun an H2 dissociation prediction at 0.74 Å and report the uncertainty."
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
SwarmLabs Engine Kit
Open-source client SDK, MCP server, and Skill catalog for the SwarmLabs physics-informed multi-agent scientific experiment automation engine.
SwarmLabs is an open infrastructure that turns multidisciplinary scientific experiments (chemistry, energy, materials, environment, biology, pharma, quantum computing, brain science, CFD, structural mechanics, …) into callable, auditable physics-informed predictions served over a small HTTP API. This repository is the open-source developer toolkit around that engine — it does not contain the proprietary physics models, only the interface, reference integrations, and honesty-first contracts that let any agent or application consume the engine safely.
Why this exists
Scientific-AI tooling is flooded with demos that return confident-but-fake numbers. SwarmLabs takes the opposite stance — honesty-first:
Every prediction reports an
uncertaintyand, where applicable, anepistemic_uncertainty(how far the request is from validated parameter space).Validation is split into
empirical_validated(matched to real published experiment data) andliterature_validated(matched to a real paper), never a single inflated "validated" count.The
v2/pi(physics-informed) endpoint returns the real physics model — it does not return a constant placeholder.
This kit makes those guarantees easy to build on top of, from a Python script, an MCP-compatible agent runtime, or a no-code agent builder.
It also ships something most scientific-AI toolkits do not: a verification
gate you do not control — an independent adjudicator that can answer "no"
to your agent's claim, and block it. See
skills/vv-gate/.
Related MCP server: Adaptyv Foundry MCP
What's inside
Path | What |
| Python client SDK ( |
| Reference MCP server exposing engine calls as tools |
| Zero-dependency MCP gate server — grep the gate into any MCP host, no |
| Independent V&V gate — zero-dep client + SKILL.md + semantics reference |
| Machine-readable catalog of the core Skills |
| Minimal end-to-end example |
| Endpoint reference |
V&V gate — an adjudicator you don't control
Reproducible ≠ valid. A fully reproducible experiment can be fully wrong. Git already sells reproducibility. What's missing is adjudication.
skills/vv-gate/scripts/vv_gate.py is a single stdlib-only file that answers
two different questions, and returns a real exit code so it can be a CI
hard gate rather than a dashboard:
Question | Command | Where the truth lives |
What is the current trust state of scenario X? |
| vendor's published gate ledger |
Are my predictions on this held-out set correct? |
| held-out set whose y we hold, you don't |
python skills/vv-gate/scripts/vv_gate.py selftest
python skills/vv-gate/scripts/vv_gate.py scenarios
python skills/vv-gate/scripts/vv_gate.py template microbio_monod -o preds.json
# ... fill preds.json with your y_pred (and optionally y_std = your 1-sigma)
python skills/vv-gate/scripts/vv_gate.py verify microbio_monod --pred preds.json
# exit 0 = PROCEED | 3 = needs human sign-off | 2 = BLOCK the autonomous action
# exit 4 = unreachable (edge block / network) — NO VERDICT was producedZero dependencies (no numpy), no API key, no engine clone. ERROR maps to
BLOCK, not to PROCEED — a gate that treats "could not evaluate" as "fine" is
a rubber stamp. 4 is separate from 2 for the same reason: a gate you cannot
reach has not ruled against you, and conflating the two leaves an operator
unable to tell "fix the network" from "stop the work". Alert and retry on 4;
halt on 2.
The findings are published, including the ones that go against us: the R² chain
passes all 62 scenarios, while the independent wet-lab anchor chain flags two
of them as physically CONTRADICTED (e.g. microbio_monod's fitted
Ks = 0.22 g/L sits 44× above the literature ceiling for E. coli on glucose).
Two chains, because one is not enough — details in
references/gate-semantics.md.
Consumer gotcha worth reading before you write your own client: Cloudflare
rejects requests with no User-Agent and with the stdlib default
Python-urllib/3.x (403 error code: 1010). curl, requests, node-fetch,
axios, Go and any explicit UA are fine. So urllib.request.urlopen(url) with
no Request is the one thing that fails.
And the block is not deterministic — the same explicit UA is served on one
attempt and 403'd on the next. Measured on CI: two live jobs in the same
minute, running byte-identical code, went green and red together, while both
offline jobs stayed green. So a single 1010 is not evidence that the service
is down, and it is not a verdict. Both shipped clients retry it and, if it
persists, report unreachable (exit 4, unreachable: true) rather than
guessing. Do not "fix" it by caching the last verdict or by mapping it to
success — that turns the gate into decoration.
That said, "non-deterministic upstream" is a hypothesis, not a diagnosis.
See CHANGELOG.md [0.2.2]: the live failure rate turned out to
be dominated by the host project serving partially landed deployments, which
Pages hides behind a 200 text/html SPA fallback. Gating the deployment on
byte-for-byte content took the live failure rate from 14.7% to 0.7%. When a
client that reads its own assets gets something that is not the asset, suspect
the deployment before you suspect the network.
Install
pip install swarmlabs-engine-kitOr use it directly from source:
git clone https://github.com/lm203688/swarmlabs-engine-kit.git
cd swarmlabs-engine-kit
pip install -e .Quick start
from swarmlabs_engine import SwarmLabsClient
# Point the client at YOUR deployed SwarmLabs engine endpoint.
client = SwarmLabsClient(base_url="https://your-swarmlabs-engine.example.com")
# List available engines
info = client.list_engines()
print(info["engines"], "engines across", info["physics_models"], "physics models")
# Run a real physics-informed prediction (e.g. H2 dissociation via VQE)
result = client.run("vqe_h2", {"bond_length_A": 0.74})
print(result["result"], "+/-", result["uncertainty"])See examples/quickstart.py for a fuller walkthrough.
Engine API surface (summary)
All endpoints are served under /api/v2/.
Method | Path | Purpose |
|
| List engines + physics-model coverage |
|
| Run a real physics-informed prediction |
|
| Physics-informed variant (same honest models) |
|
| Parameter sweep for trend analysis |
|
| Multi-fidelity cross-engine query |
|
| Record a real measurement to compare vs prediction |
The engine deployment URL is provided by you. This kit is the client side and works against any SwarmLabs engine deployment that exposes the
/api/v2/surface.
MCP integration
Two MCP servers, for two different jobs.
1. The gate — vv_gate_server.py (no dependencies)
python vv_gate_server.py # stdio MCP server
python vv_gate_server.py --selftest # live protocol smoke test{ "mcpServers": { "swarmlabs-gate": {
"command": "python", "args": ["/abs/path/to/vv_gate_server.py"] } } }Tools: list_scenarios, gate_decision, ledger_provenance,
wet_lab_anchors, get_held_out_template, and verify_prediction — the
gate itself. Your agent submits its own numbers, this server scores them against
a held-out set whose ground truth it holds, and returns
PROCEED | PROCEED_WITH_HUMAN_CHECK | BLOCK_AUTONOMOUS_ACTION. Fail-closed: a
misaligned or wrong-length submission is refused rather than partially scored,
and ERROR maps to BLOCK, never to PROCEED.
Why hand-rolled instead of
pip install mcp+ FastMCP. An MCP server is newline-delimited JSON-RPC 2.0 over stdio — about a hundred lines of stdlib. Shipping it with no install step at all means it can be dropped into a host config and just work on a machine that has nothing but Python. Same invariant asskills/vv-gate/, for the same reason: a gate you cannot easily run is a gate you do not have.
2. The engine — mcp_server.py (needs your endpoint)
The reference server that turns engine calls into tools, so any MCP-aware agent (Claude Desktop, Cursor, custom runtimes) can use SwarmLabs as a trusted scientific-compute tool:
python mcp_server.py --base-url https://your-swarmlabs-engine.example.comIt exposes swarmlabs_run, swarmlabs_list, and swarmlabs_sweep with
explicit input schemas and the same honesty-first result contract.
Why this file is at top level, not
mcp/server.py. It used to be the latter, documented aspython -m mcp.server. That fails confusingly:mcpis also the name of the real Model Context Protocol SDK on PyPI, sopython -m mcp.serverimports that package instead of this file. Naming both modules*_server.pyat top level removes the collision.
Both servers and examples/quickstart.py run from a fresh clone without
pip install, by falling back to the in-repo src/ layout.
Skill catalog
skills/skill_catalog.json describes the core Skills that compose a SwarmLabs
multi-agent workflow:
PhysicsPredictSkill — call the real engine.
ExperimentDesignSkill — turn a goal into a parameter plan.
ActiveLearningSkill — pick the next most informative experiment point.
ReportGenerationSkill — emit an auditable report.
vv-gate — Adjudicate a numeric claim against held-out ground truth and literature anchors, and return a blocking gate. This one is a standalone agentskills.io-format Skill (
skills/vv-gate/SKILL.md), loadable by Claude Code, Cursor, Codex, Gemini CLI and any harness that reads that spec.
These are the building blocks of the SwarmLabs "Planner → Executor → Verifier" agent loop. Skills 1–4 are the loop; skill 5 is the thing that can refuse it.
Status — read this before trusting anything above
Honest state of the project, because you are being asked to depend on it:
Signal | State |
Engineering verifiability | checkable — engine/oracle digests are published, |
Company entity | none (no legal entity) |
Bus factor | 1 |
Revenue | pre-revenue |
GitHub stars | 0 |
The engineering is reproducible from your side. The organisation is not yet proven, and nothing in this repository should be read as implying otherwise.
License
MIT.
Available Tools
6 toolsgate_decisionA
Static gate decision: what is the current trust state of a scenario? Returns gate (PROCEED / PROCEED_WITH_HUMAN_CHECK / BLOCK_AUTONOMOUS_ACTION) and exit_code. Omit scenario_key for the whole ledger summary.
| Name | Required | Description | Default |
|---|---|---|---|
| scenario_key | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It discloses the return shape (gate values PROCEED / PROCEED_WITH_HUMAN_CHECK / BLOCK_AUTONOMOUS_ACTION plus exit_code) and the word 'static' hints at a non-mutating read, but it never states read-only behavior, permissions, or side effects explicitly.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, zero filler, front-loaded with the purpose and followed by the return values and the parameter rule. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema and no annotations, so the description is right to spell out the returned gate values and exit_code, and it covers the one parameter's omission behavior. It is nearly complete for a simple read tool; only explicit read-only/side-effect disclosure is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage and one parameter, the description must compensate, and it does: it explains that scenario_key selects a single scenario and that omitting it returns the whole ledger summary. That is meaningful semantics the bare string schema does not convey.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific operation on a specific resource: the current trust state / gate decision of a scenario. The enumerated return values make the tool's output concrete. It does not explicitly contrast itself with siblings like verify_prediction or list_scenarios, so it stops short of 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives one clear usage rule: omit scenario_key to get the whole ledger summary rather than a single scenario. That is useful context, but it never says when to choose this tool over verify_prediction or ledger_provenance, so guidance is implied rather than explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_held_out_templateA
Fetch a scenario's held-out INPUT points as a fillable skeleton. Fill y_pred (optionally y_std = your 1-sigma) and keep x unchanged, then call verify_prediction.
| Name | Required | Description | Default |
|---|---|---|---|
| scenario_key | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It usefully discloses the return shape (a fillable skeleton with x, y_pred, y_std) and the constraint to leave x unchanged, but says nothing about read-only behavior, error cases, auth, or rate limits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences with no waste, and the core action and the downstream workflow step are front-loaded. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, so the description must explain the return value; it does so reasonably by describing the skeleton's fields and the required edit pattern. It is nearly complete for a one-parameter fetch tool, missing only edge-case or failure behavior.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There is a single input parameter (scenario_key) with 0% schema description coverage. The description implies it selects 'a scenario' but adds no format, source, or validity details beyond that, so it only partially compensates for the schema gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Fetch) and resource (a scenario's held-out INPUT points as a fillable skeleton), and identifies the follow-up tool verify_prediction so the agent understands its role in the workflow. It does not explicitly contrast with siblings like list_scenarios, but the resource is precise.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear procedural guidance: fill y_pred (and optionally y_std), keep x unchanged, then call verify_prediction. This tells the agent when and how to use the tool relative to a named sibling. No explicit when-not conditions are given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ledger_provenanceA
Gate ledger summary plus provenance: engine git commit, engine all_digest/core_digest and oracle_digest. Use it to check whether the gate you are consulting matches the code you think you trust.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries full behavioral burden. It usefully discloses that the tool returns a summary plus specific digest fields, but says nothing about read-only safety, auth requirements, or side effects. It adds real content beyond the empty schema, yet leaves the safety profile unstated.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, no filler, with the payload list front-loaded before the purpose statement. The digest vocabulary is dense but each item earns its place by describing the return content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter probe with no output schema, the description adequately explains what comes back (summary + provenance digests) and why an agent would call it. It does not relate itself to sibling tools such as gate_decision or verify_prediction, which is the one remaining gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so there is nothing for the description to disambiguate and the baseline of 4 applies. Schema coverage is 100% on an empty object, so no compensation is required.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names the specific resource (gate ledger) and enumerates the provenance payload it returns: engine git commit, engine all_digest/core_digest, and oracle_digest. That is far more concrete than a tautology, though it uses a noun phrase rather than an explicit retrieval verb and does not directly contrast itself with siblings like gate_decision.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives a clear when-to-use condition: 'Use it to check whether the gate you are consulting matches the code you think you trust,' which tells the agent the intent is provenance verification. No explicit alternatives or exclusions are named, so it stops short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_scenariosA
List the published held-out scenarios with their verdict and gate. Call this first to find a scenario_key.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It discloses the returned content (verdict, gate) for these published scenarios, which helps, but says nothing about permissions, pagination, or result size for what could be a large list.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences with zero filler; the core purpose is front-loaded and the workflow hint follows immediately.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, but the description names the key return fields (verdict, gate) and its role as the entry point for obtaining a scenario_key. For a no-parameter listing tool this is nearly complete, with only scale/pagination behavior unaddressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so the schema is trivially complete and the baseline of 4 applies; there is no parameter syntax for the description to add.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('List') and resource ('published held-out scenarios'), and further specifies the returned fields (verdict and gate). An agent can distinguish this from siblings like gate_decision or verify_prediction without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Call this first to find a scenario_key' gives clear ordering guidance and reveals the workflow role of the tool. It does not explicitly name which alternative tool to use instead if a scenario_key is already known, so it stops short of full when/when-not coverage.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
verify_predictionA
THE GATE. Score YOUR predictions on a published held-out set whose ground truth we hold. Returns verdict (PASS/MARGINAL/REFUTED/ERROR), gate, R^2, coverage (only when y_std is supplied), calibration kappa, and exit_code (0 PROCEED / 3 human check / 2 BLOCK). Fail-closed: a misaligned or wrong-length submission is refused, not partially scored; ERROR maps to BLOCK, never to PROCEED.
| Name | Required | Description | Default |
|---|---|---|---|
| predictions | Yes | One entry per published held-out point, in published order. | |
| scenario_key | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does substantial work: it enumerates the verdict values, the exit_code meanings (0 PROCEED / 3 human check / 2 BLOCK), and a strong fail-closed contract (misaligned/wrong-length input refused, ERROR maps to BLOCK). It omits auth requirements and idempotency, but the safety-critical behavior is unusually well disclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Dense and front-loaded with 'THE GATE' before the mechanics, with every clause conveying a distinct fact (returns, exit codes, fail-closed policy). The terse fragments are economical, though the shorthand is jargon-heavy for a reader without domain context.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and no annotations, the description must explain results itself and does so (verdict, R^2, coverage, kappa, exit_code). The main remaining gap is that it doesn't explain how predictions are sourced or what scenario_key identifies, which matters for a 2-param submission tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 50% and scenario_key is documented nowhere. The description does compensate partially by tying 'only when y_std is supplied' to the optional y_std field and by flagging wrong-length submissions, which reinforces the array-ordering/aligned-length requirement on predictions, but scenario_key remains unexplained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (score/verify) and resource (predictions against a published held-out set whose ground truth is server-held), making the operation concrete. It doesn't reference the sibling get_held_out_template that an agent would need to obtain predictions in the first place, so differentiation from siblings is implicit rather than explicit.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the tool is the scoring gate you run on submissions, and the exit_code semantics hint at the follow-on decision flow, but it never states when to use this versus alternatives or any prerequisite ordering (e.g., fetch the template first). Usage is inferred, not specified.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
wet_lab_anchorsA
The second evidence chain: literature ranges (never point estimates) and the implied parameters recovered by fitting the noise-free oracle. Grades AGREE / NEAR / CONTRADICTED / UNANCHORED. A scenario's V&V chain can say PASS while this chain says CONTRADICTED - read both and gate on the more conservative one.
| Name | Required | Description | Default |
|---|---|---|---|
| key | No | A literature anchor key (e.g. ecoli_glucose_Ks) or a scenario key (e.g. microbio_monod). Omit to list. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It discloses useful conceptual traits: literature ranges are never point estimates, parameters come from fitting a noise-free oracle, and outputs are graded AGREE/NEAR/CONTRADICTED/UNANCHORED. However, it does not state whether the tool is read-only, whether it has side effects, or what the response structure looks like.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three dense sentences that front-load the tool's identity and role in the evidence chain. There is little waste, though the first sentence is syntactically heavy and could be clearer as a standalone purpose statement.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with no output schema and no annotations, the description should do more to explain what callers receive and how to interpret the grades. It covers the conceptual behavior and the comparison with the V&V chain, but it leaves the return structure and listing behavior largely to the schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the single optional 'key' parameter is already fully documented in the input schema. The description adds no additional meaning about the parameter or its accepted values, making the baseline 3 appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description identifies the tool as the 'second evidence chain' dealing with literature ranges and implied parameters, which is a specific resource in this domain. It also distinguishes this chain from the V&V chain, but it never states a clear verb like 'retrieve' or 'list', leaving the exact operation slightly implicit.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives clear context: read both the V&V chain and this chain, then gate on the more conservative one. That names an alternative evidence source and a decision rule, but it does not explicitly state when not to use this tool or how it relates to siblings like gate_decision.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
6 tool updates
v0.2.2- First observed
gate_decision - First observed
get_held_out_template - First observed
ledger_provenance - First observed
list_scenarios - First observed
verify_prediction - First observed
wet_lab_anchors
TDQS
Scored across 6 tools
Most tools have distinct roles: discovery (list_scenarios, get_held_out_template), evidence chains (wet_lab_anchors), and scoring (verify_prediction). However, gate_decision, ledger_provenance, and verify_prediction all return a 'gate' value and gate_decision can also emit the whole ledger summary, creating mild overlap that descriptions mostly but not entirely resolve.
All names use snake_case, but the set mixes verb_noun (list_scenarios, get_held_out_template, verify_prediction) with noun_phrases (gate_decision, ledger_provenance, wet_lab_anchors). It remains readable but there is no single predictable verb/action convention.
Six tools is well-scoped for a held-out validation/gating server, with each tool covering a distinct stage (discovery, evidence chains, template, scoring, gate state, provenance). No tool feels redundant or padded.
The surface covers the full pipeline: find scenarios, fetch a template, submit predictions, get verdict/gate, and check provenance and a second evidence chain. Coverage is strong, though there is no explicit tool for retrieving historical prediction results or comparing multiple submissions.
Maintenance
Related MCP Connectors
Pay-per-use tool marketplace for AI agents. Search, price-check, and call APIs via MCP.
Access Pollinations models and API capabilities through agent tools.
Your org's AI agents, tasks, runs, search, and brain files as MCP tools and resources.
The OpenRouter for tools. One MCP connection gives any AI agent 254 hosted tools, pay per call.
Related MCP Servers
- FlicenseAqualityDmaintenanceExposes two MCP tools (discover and execute) that enable agents to query an OpenAPI schema via natural language and execute matched API operations.2-
- AlicenseNot gradedqualityDmaintenanceExposes the Adaptyv Foundry protein characterization API as MCP tools, enabling AI assistants to interact with experiments, targets, sequences, and results in natural language.MIT
- AlicenseNot gradedqualityBmaintenanceEnables interaction with the CopaMind platform via MCP, exposing read-only and write tools for querying match predictions, Monte Carlo simulations, team rankings, and RAG-based explanations, all while maintaining full traceability and reproducibility.1MIT
- AlicenseAqualityBmaintenanceProvides a safe scientific runtime for agents with typed math operations including calculus, algebra, statistics, unit conversion, and more via MCP tools.4Apache 2.0