Skip to main content
Glama

Server Configuration

Describes the environment variables required to run the server.

NameRequiredDescriptionDefault
SIMLAB_HOMENoMove the location where runs, experiments, and floor plans are stored (default is ~/.simlab)

Instructions

Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.

This server publishes no instructions, or was last inspected before Glama recorded them.

Capabilities

Features and capabilities supported by this server

Protocol revision2025-11-25

CapabilityDetails
tools
{
  "listChanged": false
}

Tools

Functions exposed to the LLM to take actions

NameDescription
catalogueA

The worlds the lab can simulate: what each one is and is not, every knob with unit and range, the metrics and which are better when higher, and example questions with the variants that answer them. Read this first.

list_experimentsC

Experiments (shipped and the person's own) with their parameters and named variants.

get_experimentA

One experiment: description, every parameter with its default, the variants, the metrics it writes, and its YAML.

run_experimentA

Run an experiment to completion with optional parameter overrides or a named variant; returns the run id, status, headline metrics and the report (and the tail of its log when it failed). The call blocks until the run ends: a swarm run of 400 ticks takes a few seconds, 3000 ticks with 60 agents about a minute. Values outside the catalogue's ranges are refused.

run_campaignB

Run several named variants of one experiment one after another and return the comparison table. Blocks until all of them end.

tuneA

Search a parameter space for the best headline metric: strategy grid | random | bayes (Gaussian process + expected improvement). Each trial is a run and the call blocks until the budget is spent, so keep budget small (max 40). The direction comes from the catalogue unless minimize is given.

list_runsC

Recent runs, newest first, with status and headline metrics.

get_runB

One run in full: manifest (command, params, timings, versions, file hashes), metrics, report, the last lines of its stdout.log (the world's own output and errors) and the files in its folder.

compare_runsB

Headline metrics side by side and the best run per metric (lower is better unless the catalogue says otherwise; a tie names no best). Unknown run ids are an error.

list_plansA

Floor plans for the swarm worlds (use as layout: plan:): size, walls, doors and the ASCII text.

save_planA

Save an ASCII floor plan under the person's lab folder (~/.simlab/plans): # wall, D door, . free, at least 3x3, at least one free cell; it becomes layout plan:. A shipped plan's name is refused unless overwrite is true.

save_experimentA

Save an experiment YAML of the person's own under ~/.simlab/experiments (validated: name, command with {run_dir}, params for every placeholder, variants that only set known params). The command is any local program the person wants the lab to run and score; it runs on their machine with their rights, so only save a command the person asked for. A shipped experiment's name is refused unless overwrite is true (then their file replaces it for every later run).

Prompts

Interactive templates invoked by user choice

NameDescription

No prompts

Resources

Contextual data attached and managed by the client

NameDescription

No resources

TDQS

A3.6/5.0

Scored across 12 tools

Disambiguation5/5

Each tool maps to a distinct resource and action: catalogue for discovery, experiment CRUD/run (list/get/run/save), campaign batching, tune for search, run inspection (list/get/compare), and plan listing/saving. run_experiment, run_campaign, and tune all execute but are cleanly separated by scope (one, several variants, parameter search).

Naming Consistency4/5

Strong verb_noun snake_case pattern throughout (list_experiments, get_experiment, run_campaign, compare_runs, save_plan). Two outliers—catalogue and tune—lack a noun and break the pattern, but the convention is otherwise predictable and readable.

Tool Count5/5

12 tools is well-scoped for a simulation lab covering discovery, experiment lifecycle, campaign execution, tuning, run inspection, and plan authoring. Every tool earns its place with no redundant surface.

Completeness4/5

Core lifecycle is covered: discover, define, run, batch, tune, inspect, and compare. Minor gaps exist—no delete for experiments/plans, no single get_plan (only list_plans), and no run cancellation—but these are workable around for most workflows.

Maintenance

ActivityMaintained
ResponsivenessNo issues