Skip to main content
Glama
lm203688

SwarmLabs MCP Server

Server Configuration

Describes the environment variables required to run the server.

NameRequiredDescriptionDefault

No arguments

Instructions

Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.

This server publishes no instructions, or was last inspected before Glama recorded them.

Capabilities

Features and capabilities supported by this server

Protocol revision2025-11-25

CapabilityDetails
tools
{
  "listChanged": false
}

Tools

Functions exposed to the LLM to take actions

NameDescription
list_scenariosA

List the published held-out scenarios with their verdict and gate. Call this first to find a scenario_key.

gate_decisionA

Static gate decision: what is the current trust state of a scenario? Returns gate (PROCEED / PROCEED_WITH_HUMAN_CHECK / BLOCK_AUTONOMOUS_ACTION) and exit_code. Omit scenario_key for the whole ledger summary.

ledger_provenanceA

Gate ledger summary plus provenance: engine git commit, engine all_digest/core_digest and oracle_digest. Use it to check whether the gate you are consulting matches the code you think you trust.

wet_lab_anchorsA

The second evidence chain: literature ranges (never point estimates) and the implied parameters recovered by fitting the noise-free oracle. Grades AGREE / NEAR / CONTRADICTED / UNANCHORED. A scenario's V&V chain can say PASS while this chain says CONTRADICTED - read both and gate on the more conservative one.

get_held_out_templateA

Fetch a scenario's held-out INPUT points as a fillable skeleton. Fill y_pred (optionally y_std = your 1-sigma) and keep x unchanged, then call verify_prediction.

verify_predictionA

THE GATE. Score YOUR predictions on a published held-out set whose ground truth we hold. Returns verdict (PASS/MARGINAL/REFUTED/ERROR), gate, R^2, coverage (only when y_std is supplied), calibration kappa, and exit_code (0 PROCEED / 3 human check / 2 BLOCK). Fail-closed: a misaligned or wrong-length submission is refused, not partially scored; ERROR maps to BLOCK, never to PROCEED.

Prompts

Interactive templates invoked by user choice

NameDescription

No prompts

Resources

Contextual data attached and managed by the client

NameDescription

No resources

TDQS

A3.8/5.0

Scored across 6 tools

Disambiguation4/5

Most tools have distinct roles: discovery (list_scenarios, get_held_out_template), evidence chains (wet_lab_anchors), and scoring (verify_prediction). However, gate_decision, ledger_provenance, and verify_prediction all return a 'gate' value and gate_decision can also emit the whole ledger summary, creating mild overlap that descriptions mostly but not entirely resolve.

Naming Consistency3/5

All names use snake_case, but the set mixes verb_noun (list_scenarios, get_held_out_template, verify_prediction) with noun_phrases (gate_decision, ledger_provenance, wet_lab_anchors). It remains readable but there is no single predictable verb/action convention.

Tool Count5/5

Six tools is well-scoped for a held-out validation/gating server, with each tool covering a distinct stage (discovery, evidence chains, template, scoring, gate state, provenance). No tool feels redundant or padded.

Completeness4/5

The surface covers the full pipeline: find scenarios, fetch a template, submit predictions, get verdict/gate, and check provenance and a second evidence chain. Coverage is strong, though there is no explicit tool for retrieving historical prediction results or comparing multiple submissions.

Maintenance

ActivityMaintained
ResponsivenessNo issues