SwarmLabs MCP Server
Server Configuration
Describes the environment variables required to run the server.
| Name | Required | Description | Default |
|---|---|---|---|
No arguments | |||
Instructions
Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.
This server publishes no instructions, or was last inspected before Glama recorded them.
Capabilities
Features and capabilities supported by this server
Protocol revision2025-11-25
| Capability | Details |
|---|---|
| tools | {
"listChanged": false
} |
Tools
Functions exposed to the LLM to take actions
| Name | Description |
|---|---|
| list_scenariosA | List the published held-out scenarios with their verdict and gate. Call this first to find a scenario_key. |
| gate_decisionA | Static gate decision: what is the current trust state of a scenario? Returns gate (PROCEED / PROCEED_WITH_HUMAN_CHECK / BLOCK_AUTONOMOUS_ACTION) and exit_code. Omit scenario_key for the whole ledger summary. |
| ledger_provenanceA | Gate ledger summary plus provenance: engine git commit, engine all_digest/core_digest and oracle_digest. Use it to check whether the gate you are consulting matches the code you think you trust. |
| wet_lab_anchorsA | The second evidence chain: literature ranges (never point estimates) and the implied parameters recovered by fitting the noise-free oracle. Grades AGREE / NEAR / CONTRADICTED / UNANCHORED. A scenario's V&V chain can say PASS while this chain says CONTRADICTED - read both and gate on the more conservative one. |
| get_held_out_templateA | Fetch a scenario's held-out INPUT points as a fillable skeleton. Fill y_pred (optionally y_std = your 1-sigma) and keep x unchanged, then call verify_prediction. |
| verify_predictionA | THE GATE. Score YOUR predictions on a published held-out set whose ground truth we hold. Returns verdict (PASS/MARGINAL/REFUTED/ERROR), gate, R^2, coverage (only when y_std is supplied), calibration kappa, and exit_code (0 PROCEED / 3 human check / 2 BLOCK). Fail-closed: a misaligned or wrong-length submission is refused, not partially scored; ERROR maps to BLOCK, never to PROCEED. |
Prompts
Interactive templates invoked by user choice
| Name | Description |
|---|---|
No prompts | |
Resources
Contextual data attached and managed by the client
| Name | Description |
|---|---|
No resources | |
TDQS
Scored across 6 tools
Most tools have distinct roles: discovery (list_scenarios, get_held_out_template), evidence chains (wet_lab_anchors), and scoring (verify_prediction). However, gate_decision, ledger_provenance, and verify_prediction all return a 'gate' value and gate_decision can also emit the whole ledger summary, creating mild overlap that descriptions mostly but not entirely resolve.
All names use snake_case, but the set mixes verb_noun (list_scenarios, get_held_out_template, verify_prediction) with noun_phrases (gate_decision, ledger_provenance, wet_lab_anchors). It remains readable but there is no single predictable verb/action convention.
Six tools is well-scoped for a held-out validation/gating server, with each tool covering a distinct stage (discovery, evidence chains, template, scoring, gate state, provenance). No tool feels redundant or padded.
The surface covers the full pipeline: find scenarios, fetch a template, submit predictions, get verdict/gate, and check provenance and a second evidence chain. Coverage is strong, though there is no explicit tool for retrieving historical prediction results or comparing multiple submissions.