Skip to main content
Glama

Server Configuration

Describes the environment variables required to run the server.

NameRequiredDescriptionDefault

No arguments

Instructions

Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.

This server publishes no instructions, or was last inspected before Glama recorded them.

Capabilities

Features and capabilities supported by this server

Protocol revision2025-11-25

CapabilityDetails
tools
{
  "listChanged": false
}
prompts
{
  "listChanged": false
}
resources
{
  "subscribe": false,
  "listChanged": false
}
experimental
{}

Tools

Functions exposed to the LLM to take actions

NameDescription
create_programmeA

Create a research programme with a goal, constraints, and budget.

Constraints are typed fields, not strings (commitment 4). Wires the optimizer role for this programme.

Structured params (constraints, allowed_variables, budget) may be sent as JSON-encoded strings if your client cannot emit objects.

metric_direction: 'minimize' (e.g. perplexity, loss) or 'maximize' (e.g. accuracy). The optimizer uses this to steer search and rank best trials.

candidate_version_id: optional RSI Phase 0 correlation — which registered candidate created this programme. Must exist if given.

investigation_id: optional zetesis correlation — which open investigation declared this programme as its obligation (open_investigation requires_programme → link_programme). Recorded as claimed provenance: episteme has no read path into zetesis, so the id is stored, not verified.

list_active_programmesA

List all active research programmes.

Call this BEFORE create_programme to check whether a programme already covers your goal. If one exists, add a hypothesis to it (formulate_hypothesis) instead of creating a duplicate programme.

list_programmesA

List programmes with optional attribution/status filters (read-only).

candidate_version_id filters to programmes created under that candidate (the RSI Phase-0 correlation) — this is how a candidate's descendants are enumerated. Each row includes candidate_version_id. Unlike list_active_programmes this is a paginated general listing across all statuses.

update_metric_directionA

Update the optimization direction for an existing programme.

Use 'minimize' for loss-like metrics (perplexity, error rate) or 'maximize' for quality metrics (accuracy, F1). The optimizer uses this to rank best trials and steer the search.

This is needed for programmes created before the metric_direction parameter was added to create_programme.

conclude_hypothesisA

Conclude a hypothesis by accepting/rejecting it.

Creates an immutable conclusion (commitment 8: programmes, not runs). When a claims role is wired, also mints a semantic claim with provenance edges — a byproduct, never a gate on the verdict. Enforcement: commitment 8 — rejects duplicate conclusions for the same hypothesis. Enforcement: commitment 10 — rejects if hypothesis not found or not in this programme. Enforcement: commitment 1 — rejects if no completed trials for the hypothesis. Enforcement: commitment 7 — rejects if any completed trial has no observations.

close_programmeA

Close a programme by setting its status to completed or abandoned.

For status="completed": rejects if any hypothesis is still under_test — the agent must call conclude_hypothesis first. Auto-marks proposed → abandoned (no effort was expended).

For status="abandoned": auto-marks all unconcluded hypotheses as abandoned (the whole programme is being abandoned).

Enforcement: state machine — rejects if programme is not active. Enforcement: commitment 1 — the loop is the unit (no closing with unconcluded hypotheses when completing).

register_candidateA

Register a researcher version with parent lineage (RSI Phase 0).

The thing doing the research becomes a durable object. parent_id links this candidate to its parent — omit for a genesis candidate. Lineage is never "latest wins": every candidate keeps its ancestry.

capability_profile describes the tools/roles available; it may be sent as a JSON-encoded string.

Enforcement: parent_id must name an existing candidate; code/harness digests must resolve to an ingested blob or be 'none' — a hash of nothing is a claim with no referent.

list_candidatesA

Enumerate the candidate population, newest first (read-only).

The population listing the Loop-1 roster pulls — candidates are the researchers; champion/challenger status is derived from promotion_decisions (see get_incumbent), never stored.

get_candidateB

Read one candidate version (read-only).

get_incumbentA

Derive the incumbent candidate from decisions (read-only).

Champion/challenger is derived from promotion_decisions, never stored: the latest 'promote' not superseded by a 'rollback' for the same candidate is the incumbent. Empty until the first promote — registered-but-undecided candidates are candidates, not the champion.

list_promotion_decisionsA

List a candidate's verdict trail, oldest first (read-only).

Decisions are insert-only — the trail is the record; reversal is a new 'rollback' verdict, never a mutation.

get_candidate_scorecardA

Descendant quality over attributed programmes (read-only).

For each programme created under this candidate (the Phase-0 candidate_version_id correlation), reports trial/observation counts and the best observed value per metric — 'best' read against the programme's metric_direction.

contract_id scopes the report to that contract's declared metrics (primary + secondary); without it all recorded metrics are reported per programme.

get_candidate_lineageA

Walk a candidate's parent chain to genesis (self first).

Returns {"lineage": [{id, parent_id, model_ref, code_artifact_digest, ...}, ...]} — the full ancestry that any result attributed to this candidate inherits.

create_evaluation_contractA

Create an evaluation contract for a programme (RSI Phase 0).

Contracts are insert-only and versioned per programme: each new contract for the same programme gets version = max+1 and supersedes the previous. metrics and promotion_policy may be sent as JSON-encoded strings.

The powered-policy gate enforces presence of the design keys (sesoi_d, target_power, min_evidence_rung), not adequacy — target_power: 0.5 pre-registers fine; adequacy surfaces downstream as underpowered / power_acknowledged on the campaign, not as a mint refusal.

Enforcement: programme must exist.

get_evaluation_contractA

Read one evaluation contract (read-only).

Contracts are insert-only and versioned per programme — the promotion pipeline reads them to learn which metrics a campaign is scored under.

record_promotion_decisionA

Record an attributed verdict on a candidate (RSI Phase 0).

verdict: promote | reject | hold | rollback. Decisions are insert-only — reversal is a new decision ('rollback'), never a mutation. evidence_refs may be sent as a JSON-encoded string list. decided_by is required: attribution is first-class.

Enforcement: verdict enum, non-empty rationale/decided_by, candidate (and contract, if given) must exist.

formulate_hypothesisB

Formulate a falsifiable hypothesis (commitment 3).

Rejects a hypothesis with no failure criterion. variables_involved may be sent as a JSON-encoded string list.

abandon_hypothesisA

Abandon a hypothesis — attributed, rationaled, never erased.

The exit for a hypothesis that can never conclude: falsified premises, superseded questions, trials that can only fail. Abandoning is NOT a verdict — no conclusion is recorded and no claim is minted; the hypothesis simply leaves the loop as 'abandoned', stamped with who decided and why. Trials already recorded stay on the record.

Applies to proposed and under_test hypotheses; terminal states (accepted|rejected|inconclusive|abandoned) refuse. To abandon every hypothesis in a programme at once, use close_programme(status="abandoned").

list_hypothesesA

List all hypotheses in a programme with their status.

Read-only enumeration — the tool-level counterpart of the programme://{id}/hypotheses resource. Use it to recover a hypothesis id (e.g. in a new session) rather than querying the state database directly.

design_experimentB

Design an experiment: create a trial (data item) within the programme.

config may be sent as a JSON-encoded string if your client cannot emit objects.

Enforcement: commitment 5 — budget is an epistemic resource (trial count + wall time). Enforcement: commitment 10 — the agent is a scientist (hypothesis must exist).

capture_bundleA

Seal the auxiliary bundle for a DESIGNED trial (commitment 6).

ORDERING: call AFTER design_experiment and BEFORE run_trial. The seal is pre-registration — the bundle fixes the auxiliary assumptions (code/env/seeds/splits) before any observation, so they cannot be retro-fitted to results. Trials that have left 'designed' are rejected. The post-run counterpart is executed_code.json — what actually ran, captured at finalization from the strace read-trace.

code_ref MUST be a path to a Python file that exposes: def run_training(config: dict) -> dict returning {"metrics": {...}, "variance": {...}}. The config is the same dict passed to design_experiment. Read the executor://contract resource for the full contract.

data_refs is an optional list of DataRef IDs (from prepare_data). When provided, the bundle records structured data provenance. When omitted, splits is used (backward-compatible).

extra_code_refs is an optional list of additional .py file paths that the trial depends on but cannot be discovered by AST import analysis — e.g. scripts invoked via subprocess.run(). These are captured into code_snippets and stored in code_hash_extra_json so the bundle is fully self-contained and rerunnable from archive (commitment 2 + 5 — the bundle must contain ALL code needed to reproduce, not just the executor).

Enforcement: commitment 1 — the loop is the unit (trial must exist). Enforcement: commitment 6 — the bundle must be controlled (code_ref validated). Concurrency: rejects if trial already has a bundle (no double capture).

Structured params (seeds, splits, data_refs, extra_code_refs) may be sent as JSON-encoded strings; seeds also accepts a bare int.

capture_bundle_from_code_hashA

Capture a bundle using a content address (code_hash) instead of a file path.

This is the rerun-from-archive path: the LLM reads an archived programme, gets the code_hash from the bundle, and captures a new bundle for a new trial using the same code content. No filesystem access is required — the code is loaded from code_snippets by hash, materialized as a real file next to the execution wrapper, and imported at run_trial time.

NOTE: code_hash is a "sha256:..." content address returned by a prior capture_bundle — it is NOT a bundle_id. To re-use an existing bundle's code, pass its code_hash field.

The code_hash must already exist in code_snippets (captured by a prior capture_bundle or prepare_data call). If it doesn't, the tool returns an error.

The bundle's code_ref is set to "code://{code_hash}" — a content address, not a file path. This is the carrier/content separation (Rule 5.4) made explicit: the bundle references the ICE directly.

code_hash_extra is an optional list of additional content addresses for locally-imported or subprocess-dispatched modules captured alongside the primary. Each must already exist in code_snippets. At run_trial time, these are materialized as real files under a _deps/ dir on sys.path, reconstructing each file's path suffix — so from src.mod import x resolves through the real import machinery (commitment 2 + 5).

All other parameters are identical to capture_bundle.

Structured params (seeds, splits, data_refs, code_hash_extra) may be sent as JSON-encoded strings; seeds also accepts a bare int.

Returns: {"bundle_id": "bundle-...", "status": "captured", "code_hash": ...}

run_trialA

Run a trial by calling the executor role.

Imports run_training from the bundle's code_ref and calls it with the trial config. The code_ref must be a path to a .py file exposing def run_training(config: dict) -> dict returning {"metrics": {...}, "variance": {...}}. Read the executor://contract resource for the full contract.

When the bundle carries data_refs, the config handed to run_training is extended with two injected keys: 'data_paths' ({split: resolved read-only path}) and 'data_ref_paths' ({data_ref_id: resolved read-only path}). The stored config_json keeps the designed config verbatim — the injection is runtime-only.

For long-running jobs, this returns quickly with status "running". Use get_trial_status to poll for completion. Note the return is not immediate — there is a consistent ~10 s handshake/settle before status "running" comes back; that settle window is what makes cancel-race behaviour reproducible.

Enforcement: commitment 6 — the bundle must be controlled. Rejects if the bundle is not fully captured. Concurrency: rejects if trial is already running (no double execution).

get_trial_statusA

Check the status of a running or completed trial.

Returns: {"trial_id": ..., "status": "running"|"completed"|"failed", "executor_output": ...} # raw output once finalized

executor_output is persisted on the trial row at finalize time — it survives server restarts (the executor's async cache does not).

artifact_path is a transient staging directory — its contents are captured content-addressed into SQLite at finalize and the files deleted. An empty artifact_path is by design, not lost data: recover contents via get_blob / the code:// resource.

Enforcement: commitment 1 — the loop is the unit (orphan check).

wait_trialA

Wait for a trial to reach a terminal state — one call instead of a polling loop.

Polls get_trial_status internally every poll_seconds until the trial is completed/failed/retryable/abandoned or timeout_seconds elapses (server-side cap: 60s). Returns the final status payload plus 'waited_seconds' and 'timed_out'. Caps the get_trial_status amplification loop: a 60s wait replaces ~30 round-trips.

cancel_trialA

Cancel a running trial, or abandon a designed one.

running → failed: kills the subprocess, marks failed. The record's failed here reflects the cancellation act, not an observed execution failure — cancelled: true in the payload marks the distinction, and an already finished executor reply means the executor completed first. A cancel landing before the task registers tombstones the id so no child can spawn under the terminal row. designed → abandoned: no executor to kill, no evidence lost — the honest terminal for a design that will never run (e.g. a bundle locked to code that no longer exists). The FSM has always permitted designed → abandoned; close_programme uses the same transition for programme-scoped sweeps.

Enforcement: commitment 1 — the loop is the unit (orphan check).

mark_retryableA

Mark a running trial as retryable (infrastructure failure).

This is for trials that are stuck in 'running' due to infrastructure issues (server restart, executor state lost, timeout) — NOT for scientific failures. A retryable trial is terminal and does not count as evidence.

It does NOT unlock re-capture or re-run: a terminal trial's record is finished. To retry the work, design_experiment a new trial and capture_bundle on it (file-path or code://).

Enforcement: commitment 1 — the loop is the unit (orphan check).

correct_trial_statusA

Correct a terminal trial's status — the recorded repair act.

For mislabeled records: e.g. the executor's outer wrapper exited 0 while the recorded output documents a crash ("status": "error" / nonzero inner exit code in the payload). Source must be completed|failed|retryable (all terminal — a retryable source is a correction made in error or evidence re-read after the fact); target must be completed|failed|retryable. reason is mandatory. Correcting TO completed requires the record to evidence a completed run — a cancellation receipt, reaper note, or failure output does not qualify, so that direction is refused unless the executor record documents an actual completion.

The correction is appended to the trial's executor_output_json under 'corrections' — the record shows both what was claimed and what it was corrected to. Never hand-edit the database: a direct sqlite3 UPDATE bypasses this audit trail.

Enforcement: commitment 1 — the loop is the unit (orphan check).

list_trialsA

List all trials in a programme with configs and status.

Read-only enumeration — the tool-level counterpart of the programme://{id}/trials resource. Use it to recover a trial id (e.g. in a new session) rather than querying the state database directly.

record_observationA

Record an observation (data item about a quality).

metrics and variance may be sent as JSON-encoded strings.

Only completed trials can be observed — a failed trial produced no measurement; its failure lives in its status, executor_output, and artifacts, not here. All-zero variance is rejected (n identical outcomes are one effective measurement, not reproducibility). If nothing is concludable, close the programme 'abandoned' rather than fabricating variance.

Enforcement: commitment 7 — reproducibility is the price of admission. Rejects a single point estimate with no variance.

update_beliefB

Update the belief state for a programme.

Calls the optimizer role's tell, then reads the updated posterior. Enforcement: commitment 2 — memory precedes optimization. Enforcement: commitment 9 — belief is a tracked quantity.

assess_programmeB

Assess programme health: progressive vs degenerating (commitment 8).

Reads observations from state.db and best_trials from the optimizer.

get_next_experimentB

Get the next experiment configuration from the optimizer role.

Reports remaining budget (commitment 5: budget is an epistemic resource). Enforcement: commitment 5 — rejects when budget is exhausted.

prepare_dataA

Prepare a dataset and return a DataRef ID.

Two regimes:

  • generated: runs the generator (via executor), stores output, computes hash. Requires generator_code_ref, generator_seed. Generator contract: the file must expose generate_data(config, output_path) — config carries 'seed' AND 'generator_seed' (same value — either spelling works) plus generator_params flattened at the top level; a seed/generator_seed key inside generator_params is superseded so the recorded seed always equals the seed delivered. The dataset is written to output_path. A generator may print a single-line {"error": "..."} JSON object to stdout to report a structured failure. The generator runs under the executor's default interpreter ([executor] python) — not a bundle env; only packages installed there are importable.

  • captured: registers an external URI, computes hash if accessible. Requires source_uri.

For stream capture (captured + temporal), provide capture_window_start and capture_window_end.

The DataRef is stored in the state DB and the data is stored in data storage (separate from code storage). Access is read-only.

Returns: {"data_ref_id": "data-ref-...", "split": ..., "regime": ...}

generator_params and capture_source_metadata may be sent as JSON-encoded strings.

verify_dataA

Verify data provenance by re-computing the hash.

Returns: {"verified": bool, "recorded_hash": ..., "computed_hash": ...}

If verified is false, the data was modified after preparation — a provenance violation that invalidates any trial using this data.

list_archivesA

List all archive files (open and sealed).

Returns a list of archives with their IDs, paths, programme counts, and sealed status. Sealed archives are read-only; open archives are still accepting new programmes.

get_archiveB

Get details about an archive, including its programmes.

Returns the archive metadata (path, sealed status, hash) and the list of programmes archived in it.

get_archived_programmeA

Get an archived programme's details from its archive file.

Returns the programme metadata, hypotheses, trials, beliefs, and conclusions from the archive SQLite file. The programme is read-only — it cannot be modified or restored to the live DB.

verify_archiveA

Verify an archive's integrity.

For sealed archives: recomputes SHA-256 of the .db file, compares to the recorded hash.

For open archives: verifies structural integrity (DB can be opened, row counts match the registry).

Returns: {verified: bool, recorded_hash, computed_hash (sealed only)}

archive_pending_programmesA

Archive terminal programmes that are still in the live DB.

This is a migration tool for programmes that were closed (completed or abandoned) before the automatic archiver was implemented. It scans the live DB for terminal programmes not yet archived and archives them.

dry_run=true reports the scan result ({would_archive, would_mark_archived, skipped}) without writing anything.

Under normal operation, this is a no-op — close_programme already archives automatically. This tool only catches programmes that slipped through (e.g. abandoned before Phase 3 was deployed).

Returns: {archived: [...], errors: [...], skipped: int}

capture_pending_artifactsA

Capture artifact files for trials that predate artifact capture.

Scans all trials with an artifact_path but no rows in trial_artifacts. Reads files from disk and stores them in SQLite (gzip-compressed). Files over 50MB are skipped and marked oversized. Files that no longer exist on disk are marked lost.

dry_run=true reports the pending trial ids without reading or storing anything.

cleanup: when True, staging files are deleted after successful capture (the SQLite copy becomes authoritative). Default False keeps the originals — conservative for migration of old trials.

Under normal operation this is a no-op — _finalize_trial already captures artifacts automatically. This tool only catches trials that slipped through (finalized before artifact capture was implemented).

Returns: {captured: [...], oversized: [...], lost: [...], trials_scanned: int, trials_with_existing: int}

check_invariantsA

Audit the experiment store against the loop's invariants.

Returns {status: ok|violations, checks: [{name, ok, violations, detail}]}. Detection complement to write-time enforcement: orphaned running trials, completed trials lacking observations, unsealed executions, strace divergence, undigested input data, budget overruns, stuck hypotheses, stalled trials. Report-only — nothing is repaired or mutated. The run is logged to /logs/ (retention: [integrity] log_max_files).

acknowledge_violationA

Acknowledge an open integrity violation — insert-only.

The check log is append-only; this records the disposition (remediated | accepted-with-reason) against (check_name, object_ref) so the finding stops gating writes and leaves the digest's blockers. The underlying record is never touched — remediation itself is done by the corrective tools first.

describe_blobA

Describe a content-addressed blob by digest (read-only).

The resolution check behind digest claims: returns {exists, resolved_in, size_bytes, content_type, captured_at} across the content stores — artifact_files (HTTP ingest), code_snippets, bundles.code_hash. resolved_in names every store holding the hash; a digest may live in more than one. Existence + metadata only — bytes are never returned. This is the read upstream servers use to verify a digest before accepting it on a registration call. For the bytes themselves, get_blob.

get_blobA

Retrieve a content-addressed blob's bytes by digest (read-only).

The read-back half of describe_blob: returns the stored bytes (base64 in content_b64) plus {digest, size_bytes, content_type, captured_at, resolved_in, served_from}. The returned bytes are re-hashed and verified against the requested digest — a blob whose stored bytes don't match its key is refused with the computed digest reported, never served.

Errors: malformed_digest | not_found | no_bytes | digest_mismatch | too_large. Read-only: no state is mutated.

read_resourceA

Read one of this server's resources by URI (read-only).

The MCP surface exposes resources only through resources/read — this tool gives tool-only clients the same read surface. Ownership is by construction: a URI outside this server's registered schemes resolves to nothing and errors rather than crossing the loop boundary. Never a write path.

Prompts

Interactive templates invoked by user choice

NameDescription
start_research_programmeScaffold a research programme given a goal, constraints, and budget. Produces a create_programme call with typed fields.
design_falsifiable_hypothesisGiven a claim, produce a hypothesis with a real failure criterion.
review_programme_healthGiven a programme ID, assess progressive vs degenerating.
monitor_running_trialHow to poll a running trial — adaptive, ETA-driven, never a blind fixed sleep.
status_reportWhere the loop stands and what to do next — read at session start and after any context compaction.

Resources

Contextual data attached and managed by the client

NameDescription
get_executor_contractThe executor contract — how run_training must be implemented. This resource documents the interface that the bundle's code_ref must satisfy for run_trial to execute successfully.
get_session_protocolThe session protocol — everything the agent needs to recover loop context after a client compaction or session restart. Contains: 1. The loop steps (the scientific method as a tool sequence) 2. The tool catalog (what tools are available) 3. The state machine (legal transitions) 4. Resources (what URIs the agent can read) 5. Dynamic current state (active programmes, hypotheses, trials)
get_statusCompact status digest — open work, blockers, next action.
get_constantsLive grounded-constants registry — disclosed from the running module so the audit surface is the code, not a file.
get_graphLoop-0 entity graph — programmes, hypotheses, trials as nodes + edges. Read-only projection for the agora hub.

TDQS

A3.7/5.0

Scored across 46 tools

Disambiguation4/5

Most tools target clearly distinct resources and lifecycle stages (programme/hypothesis/trial/candidate/contract/archive/blob). A few boundaries blur: list_active_programmes vs list_programmes, describe_blob vs get_blob, and the terminal-status cluster (mark_retryable, correct_trial_status, cancel_trial) require reading descriptions to separate. Descriptions do resolve the ambiguity, but an agent could misselect without care.

Naming Consistency4/5

Strong overall verb_noun snake_case convention (create_programme, conclude_hypothesis, record_observation, cancel_trial). Minor deviations: the get_ vs describe_ vs read_ prefix split (describe_blob/get_blob, read_resource) is semantically inconsistent, and capture_bundle vs capture_bundle_from_code_hash is a long variant. Still predictable and readable.

Tool Count2/5

46 tools is heavy for a single server, exceeding the 25+ threshold for concern. While the domain is genuinely broad (research loop + archives + integrity + blob store), operations like capture_pending_artifacts and archive_pending_programmes are migration one-offs that could be folded, and the trial-status family could be consolidated.

Completeness5/5

The surface covers the full research lifecycle: programme creation/closing, hypothesis formulate/conclude/abandon, experiment design/run/observe, candidate registration and lineage, evaluation contracts, promotion decisions, data prep/verify, archiving/verification, integrity auditing, and a resource/blob read path. Few obvious gaps remain.

Maintenance

ActivityMaintained
ResponsivenessNo issues