Skip to main content
Glama

nimbus-mcp

MCP server that lets AI agents (Claude Code, Cursor, Claude Desktop) build, validate, run, and analyze Nimbus BCI pipelines — upload their own EEG data, persist pipelines into studio projects, run multi-configuration experiment campaigns, watch live EEG sessions, and (explicitly confirmed) start live streaming — through your local Nimbus backend or the hosted deployment, with one nimbus-mcp login.

Install

pip install nimbus-mcp   # or: uvx nimbus-mcp

(Also installable from source: pip install -e .)

Related MCP server: eeg-mcp

Authentication

Three ways to give the server a credential — tried in this order at startup:

  1. nimbus-mcp login (recommended — hosted API, no token pasting). A device-code login: the command prints a URL and an 8-character code, opens your browser, you approve in Nimbus Studio, and the minted token is stored at ~/.nimbus/credentials.json (0600) and picked up automatically on every future start.

    nimbus-mcp login     # options: --api-url URL, --ttl-days 7..90, --name NAME
    nimbus-mcp status    # doctor: credential source, plan, quota, days-to-expiry, live probe
    nimbus-mcp logout    # remove the stored credential

    After a login, MCP client configs need no secret at all:

    {
      "mcpServers": {
        "nimbus": {
          "command": "uvx",
          "args": ["nimbus-mcp"],
          "env": { "NIMBUS_API_URL": "https://nimbus-studio.fly.dev" }
        }
      }
    }
  2. Desktop app — zero config. Just have the Nimbus Studio desktop app running: its local key file is auto-discovered (macOS ~/Library/Application Support/Nimbus Studio/mcp-key.json, Linux ~/.config/Nimbus Studio/…, Windows %APPDATA%\Nimbus Studio\…) and the server talks to the local backend at http://127.0.0.1:8080. Nothing to paste or configure.

  3. Environment variables (advanced / CI). NIMBUS_TOKEN (a hosted API token nimb_… minted in Nimbus Studio → Account → API tokens) or NIMBUS_TOKEN_FILE (a 0600 JSON file {"token": "…"} — keeps the secret out of process env and MCP configs), or the local pair NIMBUS_MCP_KEY / NIMBUS_MCP_KEY_FILE (must match MCP_LOCAL_KEY on a local backend — see below). Explicit env always beats files on disk.

No credential anywhere? The server still starts — in setup mode. Every tool call returns {ok: false, setupRequired: true, message, options} with the three paths above, so your agent walks you through onboarding instead of the server crashing. A hosted token rejected mid-session (expired or revoked) returns the same shape, including how many days ago it expired and a nimbus-mcp login first option.

Checking who you are

whoami() → {userId, email, plan: {isPro, pioneerAccess},
            freeRuns: {monthlyLimit, remaining},
            token: {name, expiresAt, daysLeft} | null,   # hosted-token mode only
            source}                                      # store | env | token_file | …

Call whoami() from the agent to see the account, plan, this month's free-run quota, and (in token mode) the token's days-to-expiry; nimbus-mcp status is the terminal equivalent with a live backend probe.

What a hosted token means

  • The token IS you. Requests run under your account: executions appear in your studio history and your plan's quotas and limits apply — there is no separate agent allowance. When the free monthly quota is exhausted, run errors carry the upgrade link https://studio.nimbusbci.com/pricing?reason=mcp-quota.

  • CPU-only in v0.4. Token-authenticated runs do not hydrate cloud GPUs.

  • Rotation. Tokens live at most 90 days (30 by default). Plan changes are snapshotted at mint time — after an upgrade, re-login (or revoke and re-create the token) to pick up the new plan. An expired token surfaces as setup guidance with the day count, not a dead end.

Requirements (local mode)

  • A Nimbus backend running locally: the desktop app, or the dev server (cd nimbus-studio/backend-py && python -m nimbus_backend.server.app) with DEBUG=1.

  • The backend started with MCP_LOCAL_KEY=<some-secret> (never set this on Fly — it is refused there).

  • Desktop app users: open Settings → MCP & Agents — no manual key setup (the app creates the key, injects it into its backend, and hands you copy-ready configs).

Configure the backend

Desktop/dev env (e.g. backend-py/data/.env or the dev shell):

MCP_LOCAL_KEY=choose-a-long-random-string
MCP_LOCAL_USER_ID=user_your_clerk_user_id
DEBUG=1   # dev server only; the desktop app qualifies automatically

MCP_LOCAL_USER_ID sets the principal the MCP key authenticates as. Set it to your own Clerk user id (user_…) so everything the agent creates — projects, saved pipelines, executions — appears in your studio UI as yours. Pick one owner and stick with it: switching the id mid-life splits ownership of agent-created work across two principals, and neither identity then sees the whole history.

Watchdog default: streaming sessions started through MCP are auto-stopped after 15 minutes with no one watching (every stream_status / get_live_session poll resets the timer). Pass idle_timeout_sec=0 to start_stream to disable it for a session.

When enabling MCP_LOCAL_KEY on a machine connected to an untrusted network, also set HOST=127.0.0.1 on the backend. The 0.0.0.0 default (settings.host) applies to the bare dev server (python -m nimbus_backend.server.app), so with it the key would otherwise be accepted from the LAN; backend-py/scripts/run_server.py already defaults to 127.0.0.1, and the desktop app pins loopback itself.

Run the server

cd nimbus-studio/mcp
python -m venv .venv && source .venv/bin/activate
pip install -e ".[test]"
NIMBUS_MCP_KEY=choose-a-long-random-string python -m nimbus_mcp

Env vars: NIMBUS_API_URL (default http://127.0.0.1:8080, or the store's api_url after a login), NIMBUS_TOKEN / NIMBUS_TOKEN_FILE (hosted API token — see Authentication), NIMBUS_MCP_KEY (must match MCP_LOCAL_KEY), NIMBUS_MCP_KEY_FILE (path to a 0600 JSON file {"key": "…"} — the desktop app's one-click MCP setup writes it; consulted only when NIMBUS_MCP_KEY is unset), NIMBUS_EXPORT_DIR (default ~/nimbus-exports). With none of the token/key vars set, the login store and then the desktop key file are auto-discovered; with nothing found, the server runs in setup mode (every tool returns onboarding guidance).

Claude Code

# --env flags go BEFORE the -- separator (everything after it is the literal
# server command, so the after-form would feed --env to python/uvx):
claude mcp add nimbus --env NIMBUS_MCP_KEY=choose-a-long-random-string \
  -- <path-to-mcp-venv>/bin/python -m nimbus_mcp

Cursor / Claude Desktop (stdio)

{
  "mcpServers": {
    "nimbus": {
      "command": "<path-to-mcp-venv>/bin/python",
      "args": ["-m", "nimbus_mcp"],
      "env": { "NIMBUS_MCP_KEY": "choose-a-long-random-string" }
    }
  }
}

Tools (32)

Auth: whoami (account, plan, quota, token expiry) Discovery: list_nodes, get_node_schema, list_templates, get_template, list_datasets, get_leaderboard Data: upload_data Inspect: inspect_dataset, inspect_file (EDA: channels, class balance, band powers, PSD) Build: validate_pipeline, validate_node_config Run: run_pipeline (non-blocking), get_execution, list_executions, get_results, cancel_execution Campaigns: run_experiment (non-blocking, 1-25 paced runs), get_experiment Artifacts: list_artifacts, download_artifact, export_python Live: list_devices, test_device, start_stream (needs confirm=true), stream_status, get_live_session, stop_stream Projects: create_project, list_projects, save_pipeline, load_pipeline

Not sure which pipeline to build? get_leaderboard() ranks benchmarked pipelines per dataset (meanAccuracyPct desc, 95% CI) under the canonical within_session protocol — agents pick templates by ranking there and pull the winner with get_template(pipelineId).

Look at your data first

Before building any pipeline, agents can see the data: channels, sampling rate, trial/class balance, per-channel µV stats, canonical band powers, and a PSD overview — for a public dataset or a file on disk.

"Inspect BNCI2014_001 subject S01 before we pick a pipeline."

inspect_dataset(dataset="BNCI2014_001", subject="S01", mode="all")
# → {channels: {count: 22, names: [...], flatlined: []}, samplingRate: 250,
#    trials: {count: 288, classLabels: [...], classCounts: {...}},
#    channelStats: [...], bandPowers: {...}, psd: {...},
#    computedFrom: {nSamples: ..., sampleStrategy: "stratified_sample_seed42_4_of_20"}}

subject is required (e.g. "S01"; get the list via list_datasets) and a comma-list like "S01,S03" loads a cohort; mode is training | evaluation | all.

The same works for files: inspect_file picks the source from the path shape. An ABSOLUTE path reads the file from disk (local backend only — desktop app / MCP local mode: no upload step, the data never leaves the machine); a RELATIVE path — the one upload_data returns — describes the uploaded file on ANY backend (hosted or local):

"Look at ~/recordings/session-01.edf and tell me if the montage is sane."

inspect_file(path="/Users/you/recordings/session-01.edf")  # absolute → local file
inspect_file(path="uploads/<user>/session-01.edf")         # upload_data path → upload

On a hosted backend absolute paths are refused — inspect_file then returns guidance (upload_data the file and pass the returned path back, or point the server at a local backend) instead of a dead end.

Why inspect first: class balance drives stratification (imbalanced classes skew accuracy), and flatlined channels mean a montage/reference problem worth fixing before training. And a units caveat: the loader assumes volts — a µV-native CSV reads 1e6x too large; set unitsScale (e.g. 1e-6 with units: "uV") in the pipeline's custom_data config when needed.

Uploading data

Bring your own recordings instead of (or alongside) the public datasets.

"I have a .edf recording at ~/recordings/session-01.edf — upload it and build a pipeline around it."

The agent calls upload_data(file_path=…), which registers the file with the backend and returns the stored path; that path goes into a custom_data node's config ({"filePath": "<path>", "format": "edf", …}) for validate_pipeline / run_pipeline / run_experiment. For plain CSV/TSV/TXT without embedded metadata, pass sampling_rate (Hz) — the backend silently assumes 250 Hz otherwise; format overrides extension-based detection.

Experiment campaigns

One run_experiment call = a paced sweep of 1-25 pipelines (at most 2 training runs in flight) with aggregated metrics, instead of the agent babysitting 25 individual run_pipeline polls.

"Compare CSP-LDA vs EEGNet on BNCI2014_001 across subjects 1-3."

The agent builds six train graphs, calls run_experiment(runs=[{name: "csp-lda-s1", train_graph: …}, …]), gets an experimentId back immediately, then polls get_experiment(experiment_id) until status is completed — per-run status and, at the end, aggregates like {"kappa": {"mean": 0.61, "std": 0.08, "best": {name, value}}} (mean/std/best over completed runs only).

Working with projects

Agent builds, human inspects. Pipelines the agent saves land in real studio projects, so you can open the canvas and see exactly what ran.

"Save this pipeline as a project called 'motor-imagery-baseline' — I'll review it in the studio."

create_project(name) makes the container, save_pipeline(project_id, train_graph) writes the graph (layout auto-generated, revision conflicts retried once) and load_pipeline(project_id) reads it back for editing or re-running. With MCP_LOCAL_USER_ID set to your user id, the project shows up in your studio project list.

Watching a live session

While a streaming session runs, the agent can watch its telemetry and tell you what it sees.

"Watch my focus session and tell me when signal quality drops."

The agent polls get_live_session(session_id) — latest prediction, the recent window, signal quality (meanChannelQuality, snrDb, artifactProbability) and running stats — and warns when quality degrades. Each poll also resets the idle watchdog, so a session under active watch is never auto-stopped; an abandoned one is shut down after 15 minutes.

Safety

start_stream refuses to run without confirm=true — it connects an EEG device and starts a live session on a human. The X-MCP-Key path is machine-local only (never accepted on Fly deployments); hosted mode authenticates with a personal Authorization: Bearer token from nimbus-mcp login (see Authentication above). Sessions started via MCP are stopped automatically after 15 idle minutes (see the watchdog note above).

Available Tools

32 tools
cancel_executionCancel ExecutionCInspect

Cancel a running execution.

ParametersJSON Schema
NameRequiredDescriptionDefault
execution_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden of behavioral disclosure, and it discloses almost nothing. It does not state whether cancellation is immediate, whether it is idempotent, what happens to already-produced results or artifacts, or what the execution's resulting state will be. The only hint is the adjective 'running', which is helpful but far short of the needed context for a destructive operation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The sentence is front-loaded and free of padding, which is good, but at four words it is under-specified rather than truly concise. There is room to add scope and behavioral detail without becoming verbose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be described. Still, for a destructive mutation with no annotations and an entirely undocumented required parameter, the definition omits information an agent needs to invoke it safely and correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the single parameter execution_id has no documented format, source, or example anywhere. The description does not compensate, so an agent gets no help on where to obtain a valid execution_id (presumably from list_executions or get_execution).

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb (cancel) and resource (execution), so the operation is immediately identifiable. However, it offers no differentiation from sibling tools such as get_execution, list_executions, or stop_stream, which an agent must distinguish between when deciding what to call.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives like get_execution or list_executions, nor any stated prerequisites. The word 'running' hints that only in-flight executions are cancellable, but that is left implicit rather than stated as a condition.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_projectCreate ProjectBInspect

Create a project (container for one pipeline document). Returns projectId.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYes
descriptionNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It mentions that a project is a container for one pipeline document, which is useful structural context, but omits critical behavioral details for a mutation tool: no mention of required permissions, no disclosure of side effects (e.g., whether it triggers pipeline creation), and no description of what happens if a duplicate name exists.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single compact sentence that front-loads the core action and scope, then names the return value. No wasted words, appropriate for a simple definition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Although an output schema exists (so return value details needn't be explained), the description is inadequate for a mutation tool with no annotations and 0% parameter documentation. It should at least clarify required permissions, side effects, and parameter semantics, especially given the rich set of pipeline-related siblings.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so neither the name nor the optional description parameter is documented anywhere. The description does not compensate, providing no syntax, format, or constraints for either parameter. This is a severe gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Create a project') and adds a clear definition: a container for one pipeline document. This distinguishes it from save_pipeline, list_projects, and load_pipeline without ambiguity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description identifies the tool's role but provides no explicit when-to-use instructions or exclusions relative to siblings like save_pipeline or list_projects. A competent agent can infer the usage from context, but nothing is spelled out.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

download_artifactDownload ArtifactBInspect

Download one artifact file to NIMBUS_EXPORT_DIR/executions// and return its path.

ParametersJSON Schema
NameRequiredDescriptionDefault
execution_idYes
artifact_nameYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden, and it does disclose the write destination (NIMBUS_EXPORT_DIR/executions/<id>/) and that a path is returned. It omits overwrite behavior, auth/permission requirements, and any size or rate constraints, so it is only partially transparent for a filesystem-writing operation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence that conveys action, target, destination, and return value with zero waste.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so explaining the return value is unnecessary, and the description still covers destination and single-file scope. The remaining gap is the lack of any hint about where artifact_name comes from (list_artifacts) before calling this tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. The path template implies execution_id selects the output subdirectory and that artifact_name identifies a single file, but no format, naming convention, or valid-value guidance is given for either parameter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Download) and resource (one artifact file), and the destination path makes the effect concrete. It is distinguishable from list_artifacts, though the description never names that sibling explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance, no prerequisites, and no mention of the alternative (list_artifacts) that supplies artifact_name. The agent must infer the workflow from the parameter names alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

export_pythonExport PythonCInspect

Export the pipeline as a standalone runnable Python bundle (zip saved locally).

ParametersJSON Schema
NameRequiredDescriptionDefault
nameNo
train_graphYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden. It does disclose one side effect (a zip is saved locally), but omits where the file lands, whether it overwrites, what permissions or runtime requirements apply, and whether the exported bundle depends on the current environment.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single sentence with the action front-loaded and the output artifact in parentheses. Nothing is padded or repeated.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The presence of an output schema means return values need not be described, but for a tool with an undocumented nested required parameter and no annotations, the description leaves too much unspecified: export location, naming via `name`, and what the underlying `train_graph` object must contain.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate and does not: neither the required `train_graph` object (nested, typed as a free-form object) nor the optional `name` parameter is explained. 'The pipeline' loosely hints at train_graph but leaves its expected shape and the purpose of `name` unknown.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Export) and resource (the pipeline) plus the concrete output form: a standalone runnable Python bundle saved as a local zip. That distinguishes it from siblings like save_pipeline or download_artifact, though it never explicitly names those alternatives.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No statement of when to use this versus save_pipeline, load_pipeline, or download_artifact, and no prerequisites or exclusions. The only implied guidance is that it produces a Python bundle rather than a pipeline file.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_executionGet ExecutionBInspect

Execution status summary (status: running/completed/failed/cancelled).

ParametersJSON Schema
NameRequiredDescriptionDefault
execution_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It does add genuinely useful behavioral context by enumerating the possible status values (running/completed/failed/cancelled), which the schema does not. However, it says nothing about read-only safety, permissions, or behavior for an unknown execution_id.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single short sentence with the core purpose front-loaded and zero filler. It is efficient, though the brevity comes at the cost of the missing usage and parameter detail noted above.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need no explanation, and the status enumeration is a helpful addition. Still, with no annotations and no parameter documentation, an agent lacks the provenance of execution_id and any assurance that this is a safe read.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% and the single required parameter execution_id has no description in either place. The description does not explain the id's format or where to obtain it, so the parameter remains under-specified despite the tool being simple.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a clear verb+resource ('Execution status summary') and scopes it to the status field. An agent can distinguish it from cancel_execution, list_executions, and get_results, though the description never explicitly names those siblings. The word 'summary' leaves some ambiguity about whether the full execution record or just a status is returned.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this versus list_executions or get_results, and no hint that an execution_id must first be obtained elsewhere. The only usage signal is the implied 'look up one execution by id', which is left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_experimentGet ExperimentBInspect

Experiment snapshot: status (running/completed/failed), per-run rows ({name, executionId, status, error?, metrics?}) and, once finished, aggregates {metric: {mean, std, best: {name, value}}} over completed runs only (std = population; None below 2 values).

ParametersJSON Schema
NameRequiredDescriptionDefault
experiment_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden, and it does disclose genuinely non-obvious semantics: aggregates are computed over completed runs only, std is population standard deviation, and a metric is None below 2 values. What it omits is the read-only/side-effect profile and any failure behavior when experiment_id does not exist.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence that leads with the payload rather than preamble, and every clause conveys a distinct fact. The brace notation is dense but compact and readable for an agent parsing structure.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so the description did not need to restate return values, yet it spends its entire length doing so. For a one-parameter read tool that is largely sufficient, but it leaves usage guidance and the read-only assurance unaddressed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The single parameter experiment_id is self-describing by name, so the low 0% schema description coverage is a small risk in practice. However, the description adds nothing about the parameter: no format, no statement of whether a name or ID is accepted, and no behavior on a missing or unknown ID.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names the resource (an experiment) and specifies exactly what the snapshot contains: status, per-run rows, and aggregates. It is clear what the tool returns, but it never states the operation as a verb (e.g. 'retrieve') and does nothing to distinguish it from siblings like get_execution, get_results, or run_experiment.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to call this versus the many adjacent tools (get_execution, get_results, run_experiment, list_executions). The only implied usage is 'once finished, aggregates are available', which hints at timing but not tool selection.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_leaderboardGet LeaderboardAInspect

Public benchmark leaderboard: pipeline rankings per dataset.

Rankings are per-dataset under the canonical within_session protocol (see protocol). Within each dataset, rows are sorted desc by meanAccuracyPct (95% CI in ciLoPct/ciHiPct). Use pipelineId as the template id hint for get_template when building a pipeline. updated marks each dataset's most recent run; packFingerprint identifies the exact dataset pack the scores came from.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the burden and does fairly well: it discloses the ranking protocol (within_session), sort order (desc by meanAccuracyPct), CI fields, and that 'updated'/'packFingerprint' annotate recency and provenance. It stops short of stating access/auth or rate characteristics, but for a read-only public benchmark that is minor.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the one-line purpose, then layered detail on protocol, sorting, and field meaning. Every sentence carries signal, though the field-by-field catalog is slightly dense and could be trimmed given an output schema exists.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a no-arg read tool with an output schema present, the description supplies enough context to call it correctly: what the leaderboard ranks, how rows are ordered, and how to reuse pipelineId downstream. Missing only explicit invocation context and access expectations.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so the baseline is 4. The description correctly adds no parameter detail because none exist, and instead spends its budget on return-field semantics, which is the right trade-off.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific resource ('Public benchmark leaderboard: pipeline rankings per dataset') with clear scope and per-dataset framing. It is confidently distinguishable from siblings like get_results or get_experiment, though it never explicitly names what it is not.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is only implied — the description assumes the agent knows this is the entry point for benchmark rankings. It offers a downstream chaining hint ('use pipelineId as the template id hint for get_template'), but no explicit when-to-use or when-not-to-use guidance versus other result-retrieval tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_live_sessionGet Live SessionAInspect

Live snapshot of a streaming session: latest prediction + recent window, signal quality (meanChannelQuality, snrDb, artifactProbability), indicators, running stats. Poll this while a session runs. Live telemetry requires a DEPLOYED model session (hub deploy / playback with a classifier); modelless hardware streams have no telemetry — use stream_status for those. Expect low confidence during filter/ASR warm-up (first seconds); quality < 0.5 or high artifactProbability means the signal is poor. 404 => session not active in this backend. Each poll also feeds the idle watchdog (see start_stream's idle_timeout_sec), keeping an actively watched session alive.

ParametersJSON Schema
NameRequiredDescriptionDefault
windowNo
session_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description fully carries the burden and does so well: it discloses a side effect (each poll feeds the idle watchdog, keeping the session alive), error semantics (404 => session not active in this backend), warm-up behavior (low confidence during filter/ASR warm-up), and interpretation thresholds (quality < 0.5 or high artifactProbability = poor signal). Only auth requirements are unstated, which is minor.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Dense but well-structured: payload contents, usage, prerequisites, thresholds, errors, and side effects in order. Slightly packet-heavy, but essentially every sentence carries actionable information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be restated; the description focuses on what an agent can't get elsewhere — prerequisites, error codes, signal-quality interpretation, and the watchdog side effect. Complete for a polling telemetry tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and neither parameter is documented. The 'window' parameter (default 50) is never explained — the phrase 'recent window' is not tied to it, nor does the description say how its value changes behavior. At 0% coverage the description must compensate and it largely does not.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource (live snapshot of a streaming session) and enumerates the payload contents. It explicitly distinguishes itself from the sibling stream_status by noting modelless hardware streams have no telemetry, so an agent can route without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit when-to-use ('Poll this while a session runs'), an explicit when-not ('modelless hardware streams have no telemetry — use stream_status for those'), and precondition (requires a DEPLOYED model session). Nothing is left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_node_schemaGet Node SchemaBInspect

Full config JSON schema + input/output ports for one node type.

ParametersJSON Schema
NameRequiredDescriptionDefault
node_typeYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It usefully discloses that the return contains both a config schema and port definitions, which is real content beyond the input schema, but says nothing about whether node_type values must come from an existing pipeline, what happens for unknown types, or permission requirements.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single dense sentence that front-loads the key noun phrase ('Full config JSON schema') and adds the port detail without padding. Nothing wasted, though 'Full' is slightly vague about what is included.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is simple (one required param) and an output schema exists, so return values need not be described. However, with zero annotation coverage and zero parameter description coverage, the description leaves the agent without guidance on valid node_type values or error behavior.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the single node_type parameter has no description in the schema. The description partially compensates by framing the tool as operating on 'one node type', implying the parameter is a node-type identifier, but it gives no format, valid values, or source for that identifier.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (get) and resource (node schema) and specifies exactly what is returned: the full config JSON schema plus input/output ports for a single node type. This clearly separates it from list_nodes (enumeration) and validate_node_config (checking a config), though it never names those siblings explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is only implied: the phrase 'for one node type' suggests fetching this before configuring or validating a node, but there is no explicit when-to-use statement, no prerequisite (e.g., obtain node_type from list_nodes), and no mention of alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_resultsGet ResultsAInspect

Metrics for a completed run. Trimmed by default (accuracy, kappa, ITR, confusion matrix, per-class); full=True returns the complete result object.

ParametersJSON Schema
NameRequiredDescriptionDefault
fullNo
execution_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden. It usefully discloses that output is trimmed by default and that full=True returns the complete object, which is real behavior beyond the schema. It says nothing about permissions, behavior on an incomplete/failed run, or error conditions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, front-loaded with the purpose and immediately followed by the default-vs-full behavior. No filler or repetition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be re-explained; the description nonetheless names the trimmed fields helpfully. For a 2-param read tool the main residual gap is the unexplained `execution_id` and lack of routing against sibling retrieval tools.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It explains the `full` flag well (default trimmed vs complete object) but says nothing about `execution_id`, leaving half the parameters undocumented in both schema and description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource: retrieves metrics for a completed run, and even enumerates the trimmed metric fields. It does not, however, distinguish itself from close siblings like get_execution or get_experiment, leaving some ambiguity about which retrieval tool to pick.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'for a completed run' implies the precondition (run must have finished) and the context in which to use it, but there is no explicit when-to-use vs alternatives guidance or when-not guidance against get_execution/list_executions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_templateGet TemplateBInspect

Full template incl. the 'train' execGraph needed by run_pipeline/validate_pipeline.

ParametersJSON Schema
NameRequiredDescriptionDefault
template_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It usefully discloses that the return is the full template including the 'train' execGraph, but says nothing about error behavior (e.g., unknown template_id), permissions, or whether other execGraphs exist.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence that leads with the key scope ('Full template') and packs in the important downstream dependency. No waste, though it is very terse for the information an agent may need.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained, and the description does identify the execGraph content. Still, with no annotations and an undocumented parameter, the definition is only minimally complete for a fetch-by-id tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% for the single required parameter template_id, and the description adds no meaning about it (format, ID source, or lookup semantics). It neither compensates for the schema gap nor clarifies the identifier.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (get) and resource (template) and clarifies scope as the 'Full template incl. the train execGraph', which distinguishes it from list_templates. It does not explicitly name list_templates, but the 'full' qualifier and execGraph content imply a single complete object.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description ties the tool to run_pipeline/validate_pipeline, implying the execGraph is fetched ahead of running or validating a pipeline. However, it gives no explicit 'use this when/when not' or an alternative, so usage is only implied rather than stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

inspect_datasetInspect DatasetAInspect

Exploratory summary of a public EEG dataset (MOABB pack): channels, sampling rate, trial/class balance, per-channel µV stats, band powers and a PSD overview.

Look at the data BEFORE building pipelines: class balance drives stratification choices (imbalanced classes skew accuracy), and flatlined channels mean a montage/reference problem worth fixing first. subject is REQUIRED (the backend 400s without it) — get the subject list via list_datasets, e.g. "S01"; a comma-list like "S01,S03" loads a cohort. mode: training | evaluation | all. Units note: values are ASSUMED volts by the loader — a µV-native file reads 1e6x too large; set unitsScale in a pipeline's custom_data config when needed.

ParametersJSON Schema
NameRequiredDescriptionDefault
modeNoall
datasetYes
subjectNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden and does well: it discloses that the backend 400s without subject, that a comma-list loads a cohort, the valid mode values, and a non-obvious units caveat (loader assumes volts; µV-native files read 1e6x too large, fix via unitsScale). It omits cost/runtime characteristics and any auth expectations, keeping it short of a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the purpose before the guidance and caveats; every sentence carries actionable information (the units note is dense but materially useful). Slightly long, but no filler sentences.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need no explanation; the description instead covers the operational gaps an agent needs — the required subject argument, the mode values, cohort syntax, and the units gotcha. Nothing essential is missing for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate and largely does: subject is explained (required, format 'S01', comma-list cohort), mode is enumerated (training | evaluation | all), and the units caveat affects interpreting the data. The dataset parameter itself is left unexplained beyond its name.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Exploratory summary of a public EEG dataset (MOABB pack)') and enumerates exactly what is returned (channels, sampling rate, trial/class balance, per-channel µV stats, band powers, PSD). This clearly separates it from siblings like list_datasets (enumeration) and inspect_file (single file).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly states when to use it ('Look at the data BEFORE building pipelines'), explains the downstream consequences that motivate it (class balance drives stratification; flatlined channels indicate a montage/reference problem), and names the alternative for obtaining subject IDs (list_datasets). No inference required.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

inspect_fileInspect FileAInspect

Exploratory summary of an EEG file (.edf/.bdf/.mat/.csv/.tsv/.txt/.h5): channels, sampling rate, trial/class balance, per-channel µV stats, band powers and a PSD overview.

The path shape picks the source: ABSOLUTE path → read the file from disk (only on a LOCAL backend: desktop app / MCP local mode — no upload needed); RELATIVE path (the one upload_data returns) → describe the uploaded file, which works on ANY backend (hosted or local).

Look at the data BEFORE building pipelines: class balance drives stratification choices, and flatlined channels mean a montage/reference problem worth fixing first. Units note: values are ASSUMED volts by the loader — a µV-native CSV reads 1e6x too large; set unitsScale in a pipeline's custom_data config when needed. On a hosted backend absolute paths are refused and this returns guidance (upload the file first or switch to a local backend).

ParametersJSON Schema
NameRequiredDescriptionDefault
pathYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does so: it discloses backend-specific path rules, that hosted backends refuse absolute paths and return guidance, that disk reads work only on local backends, and that values are assumed volts with a 1e6 scaling hazard for µV-native CSVs. This is exactly the kind of operational context the schema cannot convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the one-line summary and then layers guidance in a logical order. It is somewhat long and the closing sentence about hosted backends restates the earlier local-backend constraint, adding a little redundancy without new meaning.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be re-explained, yet the description already frames what the summary contains. Combined with backend, upload, and unit-scaling caveats, an agent has everything needed to call this correctly on either backend.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema has 0% description coverage for the single `path` parameter, so the description must compensate, and it does: absolute paths are read from disk on local backends while relative paths (from upload_data) resolve the uploaded file on any backend. That is precise semantic guidance well beyond the bare string type.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Opens with a specific verb and resource ('Exploratory summary of an EEG file') and enumerates the exact contents returned (channels, sampling rate, class balance, per-channel µV stats, band powers, PSD overview) plus the supported formats. The resource is precisely delimited to a single file, which keeps it distinct from the dataset-oriented siblings despite no explicit naming.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly tells the agent when to run it ('Look at the data BEFORE building pipelines') and why each output matters, and it spells out the absolute-vs-relative path decision and which backends each branch works on. Backend refusal conditions are stated so the agent knows the failure mode in advance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_artifactsList ArtifactsCInspect

Trained artifacts (models/filters, e.g. *.pkl) saved by an execution.

ParametersJSON Schema
NameRequiredDescriptionDefault
execution_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.5/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden and it delivers almost nothing behavioral: no auth requirements, no ordering or pagination notes, no statement that this is a read-only listing. The one useful hint is that artifacts belong to an execution.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single terse sentence fragment with no waste, but it is under-specified rather than genuinely concise — it omits the verb and all operational context.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained. But for a tool with zero annotation coverage and an undocumented required parameter, the description should say more about usage and scoping than it does.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% for the single execution_id parameter, so the description must compensate. It does partially, by implying that artifacts are scoped to a specific execution, but it never explains the format or provenance of the execution_id itself.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names the resource (trained artifacts such as *.pkl models/filters) and ties it to an execution, so the domain is clear. However, the verb is only implied by the name 'list_artifacts' and never stated, and it does not explicitly distinguish itself from the sibling download_artifact.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance is given. With a sibling download_artifact present, the agent is left to infer whether this tool enumerates artifacts, downloads them, or both.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_datasetsList DatasetsCInspect

Curated public EEG datasets (MOABB packs) available to pipelines.

ParametersJSON Schema
NameRequiredDescriptionDefault
only_on_diskNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden, and it discloses almost nothing: no indication of whether this is a read-only enumeration, whether results are paginated, whether authentication or a project context is required. The presence of an output schema covers return shape, but the behavioral profile is essentially undocumented.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is a single short fragment with no wasted words, but it is under-specified rather than genuinely concise — the brevity comes at the cost of missing guidance and parameter meaning. Front-loading is fine given the length.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a list tool with one undocumented boolean parameter, no annotations, and no usage guidance, the description leaves the agent short of what it needs to invoke the tool correctly. The output schema excuses it from explaining return values, but the only_on_disk semantics and its interaction with the listing are still missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the single parameter only_on_disk (boolean, default true) is left completely unexplained in both schema and description. The description never mentions filtering or disk state, so an agent has to guess what the flag does and that it defaults to true.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description identifies the resource precisely (curated public EEG datasets / MOABB packs) and its scope (available to pipelines), which helps distinguish it from siblings like inspect_dataset. However, it is a noun phrase rather than a verb+resource statement, so it never explicitly says the tool lists or returns them — the agent must infer the action from the name.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no statement of when to use this tool, when not to, or how it relates to inspect_dataset, which is the obvious adjacent sibling. No prerequisites or context for choosing it are given.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_devicesList DevicesAInspect

EEG devices supported by this backend (OpenBCI, Muse, BrainBit, LSL, PiEEG...).

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full disclosure burden. It usefully scopes the result to 'supported by this backend' and lists example vendors, but for a tool sitting next to test_device it never clarifies that listing is passive/non-probing rather than an active hardware check.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with the resource first and the vendor examples immediately after. No filler, no redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Because an output schema exists, the description need not describe return values, and for a zero-argument listing tool the remaining burden is small. The only gap is the missing tie-breaker against test_device, which matters more than usual given how similar the two names are.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema has zero parameters, so there is no parameter semantics to convey and the baseline of 4 applies. The description adds nothing here, but nothing is missing either.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific resource (EEG devices supported by this backend) and enumerates concrete examples (OpenBCI, Muse, BrainBit, LSL, PiEEG), so an agent knows exactly what set is returned. It does not, however, distinguish itself from the sibling test_device, which is the most likely confusion point.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance and no mention of alternatives. An agent must infer from the name alone that this is the discovery step to run before test_device or start_stream, and nothing in the text confirms or excludes that.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_executionsList ExecutionsCInspect

Recent executions. Optional status filter (running/completed/failed/cancelled).

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
statusNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It hints at recency/ordering and names valid status values, which is useful, but says nothing about pagination, the default limit behavior, ordering guarantees, or result volume for what is effectively an unpaginated-friendly list tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two telegraphic fragments with zero filler, and the recency scope is front-loaded ahead of the filter detail. It is efficient, though arguably too sparse to be fully self-sufficient.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained. However, with no annotations and no usage routing, and with the limit parameter undocumented, the definition is only minimally complete for an agent to invoke it confidently.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It usefully enumerates the status values (running/completed/failed/cancelled) that the schema declares only as a bare string, but the limit parameter is entirely undocumented in both places, leaving half the parameter surface unexplained.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states the resource ('executions') with a recency scope ('Recent'), but the verb is only carried by the tool name and title, making it close to a restatement. It does not differentiate this list tool from the sibling get_execution or explain what 'recent' means in scope.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this versus list alternatives such as get_execution or cancel_execution, nor any prerequisites or context. Usage is only implied by the name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_nodesList NodesAInspect

List Nimbus pipeline node types (data, preprocessing, features, models...).

Use get_node_schema(node_type) for one node's full config schema and ports.

ParametersJSON Schema
NameRequiredDescriptionDefault
categoryNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden, and it only implies a safe read ('List'). It says nothing about pagination, ordering, result size, or whether the list is static or registry-driven. Adequate for a simple enumeration, but thin for a tool with zero annotation coverage.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences with zero filler; the primary purpose is front-loaded and the routing hint follows. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return-value explanation is not required, and the description covers purpose plus sibling routing. The only meaningful gap is the unclarified 'category' filter, which with 0% schema coverage is left partly to inference.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% and the single 'category' parameter is undocumented in the schema. The parenthetical ('data, preprocessing, features, models...') effectively illustrates valid category values, which adds real meaning, but it never states that these are the accepted filter values or what an omitted category returns.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('List Nimbus pipeline node types') and even enumerates the domain categories, so the agent immediately knows this returns the catalog of node types rather than one node's details. It is clearly distinguishable from the sibling get_node_schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly routes the agent: use this to list node types, use get_node_schema(node_type) when you need one node's full config schema and ports. That is a clear when-to-use-this-vs-that statement, though it stops short of stating exclusions or prerequisites.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_projectsList ProjectsBInspect

List projects owned by the current principal (agent work included).

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the disclosure burden. It helpfully discloses the ownership scope and the non-obvious fact that agent-created work is included, but says nothing about result ordering, pagination, or limits. It does at least imply a read-only operation via 'List'.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler; the scope qualifier is placed where it matters most. It is arguably too terse for the gaps it leaves, but there is zero wasted text.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists and no parameters are needed, so the description need not explain return values. However, with no annotations, it leaves the agent guessing about ordering, pagination, and result volume for a list operation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so there is no parameter semantics to document; the schema is trivially complete. Baseline 4 applies, and the description's scope note adds a filtering caveat an agent might expect to be able to control.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (list) and resource (projects) plus the ownership scope ('owned by the current principal'). It does not name or differentiate itself from any sibling, though no sibling is a competing project-list tool, so ambiguity risk is low.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit when-to-use guidance, no prerequisites, and no mention of alternatives (e.g., how this relates to create_project or to filtering projects some other way). Usage is only implied by the name and scope clause.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_templatesList TemplatesAInspect

List built-in starter pipelines (MI/P300/SSVEP...). get_template(id) for the graph.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the burden, but for a zero-parameter read-only listing there is little destructive or authentication behavior to disclose. The 'built-in' qualifier usefully signals that user-created templates are out of scope, and the output schema covers the return shape, so remaining gaps are minor.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences with zero waste; the listing scope is front-loaded and the sibling pointer follows. Every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, the description need not explain return values, and with zero parameters there is no schema gap to fill. The description supplies the domain examples and the sibling route; the only modest omission is any note on ordering, pagination, or whether templates are user-scoped.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes no parameters, so the baseline is 4. There are no argument semantics the description needs to compensate for, and the schema is fully self-describing.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (List) and resource (built-in starter pipelines) and grounds it with concrete examples (MI/P300/SSVEP) that an agent can map to the domain. It also names the sibling get_template, so the agent can distinguish the enumeration tool from the detail tool without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly routes the agent to get_template(id) for the graph, which clarifies the split between listing and retrieving details. It implies when to use this tool (enumerate built-ins) but doesn't state exclusions such as saved/user templates, so it stops short of full when/when-not guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

load_pipelineLoad PipelineBInspect

Load a project's saved pipeline (train graph + meta) for editing/re-running.

ParametersJSON Schema
NameRequiredDescriptionDefault
project_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It discloses what is loaded (graph + meta) and its purpose, but says nothing about failure modes (e.g., no saved pipeline), permissions, or whether loading mutates session state versus returning the artifact.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single dense sentence with the verb and scope front-loaded and useful clarifications in parentheses. Nothing is wasted, though the parenthetical is terse.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained. However, given zero annotations and a mutation-adjacent workflow role, the description omits behavioral details such as what happens when no pipeline is saved or how the loaded state is exposed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% for the single project_id parameter. The phrase 'a project's saved pipeline' loosely ties project_id to the project, but gives no format, ID scheme, or example, so it barely compensates for the coverage gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb (load) and resource (a project's saved pipeline), and clarifies the payload as 'train graph + meta' with the intent 'for editing/re-running.' It implicitly contrasts with siblings like save_pipeline and run_pipeline, though it never names them.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'for editing/re-running' hints at the intended workflow and implies this is a prerequisite to editing, but there is no explicit when-to-use vs. alternative guidance and no mention of what to do if no saved pipeline exists.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

run_experimentRun ExperimentAInspect

Run 1-25 pipelines as ONE paced experiment (NON-BLOCKING). Returns an experimentId immediately; a background thread submits at most 2 runs at a time (min(max_concurrent, 2)), retries queue-full up to 3 times per run, and polls each execution to completion. Poll get_experiment() for per-run status and, once finished, aggregated metrics.

ParametersJSON Schema
NameRequiredDescriptionDefault
runsYes
max_concurrentNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden and does so well: non-blocking return of an experimentId, a background thread capped at min(max_concurrent, 2), retry of queue-full up to 3 times per run, and polling to completion. It omits auth requirements and what happens on partial failure, but the concurrency/retry semantics are unusually well disclosed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the critical behavioral fact '(NON-BLOCKING)' and the 1-25 bound, and every clause carries operational information. It is dense to the point of being slightly run-on, but no sentence is filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a batch-execution tool with an output schema, the description correctly covers non-blocking semantics, scheduling behavior, and the handoff to get_experiment, so return-value detail is not needed. The only real gap is the shape of the `runs` array elements, which is not covered anywhere.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, and one parameter (`runs`) is an untyped array of open objects whose per-run shape is never explained. The description partially compensates by bounding `runs` at 1-25 pipelines and by clarifying that `max_concurrent` is effectively clamped to 2, but the structure of each run entry remains undocumented in both schema and description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb (run) and resource (1-25 pipelines as one paced experiment), and explicitly contrasts with the sibling polling tool get_experiment(). An agent can immediately tell this apart from run_pipeline (single run) and get_experiment (status polling) without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It names the follow-up tool and the condition that selects it ('Poll get_experiment() for per-run status and, once finished, aggregated metrics'), and '(NON-BLOCKING)' tells the agent it can continue working after the call. It stops short of saying when to prefer run_pipeline over this batch tool, so it is clear context rather than full when/when-not guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

run_pipelineRun PipelineAInspect

Start a pipeline run (NON-BLOCKING). Returns executionId — poll with get_execution() until status is completed/failed, then get_results(). layout is optional canvas positions ({nodes: {id: {x, y}}}); a grid is synthesized when omitted.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameNo
layoutNo
subjectNo
descriptionNo
train_graphYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does well: it discloses the async/non-blocking contract, that an executionId is returned, and the exact polling lifecycle including terminal states. It omits permission/auth requirements and any note that this creates a persistent execution record, which a mutation tool should ideally state.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the most decision-relevant fact (NON-BLOCKING) and the follow-up chain, then handles the one parameter worth explaining. No filler sentences.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The async workflow and return contract are complete, and the output schema covers return values. However, the single required parameter train_graph is entirely undocumented, which is a meaningful gap for a tool whose core input is a graph object.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate, but it only explains one of five parameters (layout's shape and the synthesized-grid default). The required train_graph payload, plus name, subject and description, receive no semantic guidance anywhere.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Start a pipeline run') and immediately signals the non-blocking execution model, which cleanly separates it from synchronous siblings like validate_pipeline and run_experiment.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit follow-up workflow — poll get_execution() until completed/failed, then call get_results() — naming the sibling tools and the terminal states. It lacks an explicit when-not (e.g., 'validate_pipeline before running'), so it falls just short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

save_pipelineSave PipelineBInspect

Save a pipeline graph into a project (visible on the studio canvas). Handles revision conflicts automatically (one retry).

ParametersJSON Schema
NameRequiredDescriptionDefault
nameNo
subjectNo
project_idYes
train_graphYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden, and it does add one genuine behavioral detail: revision conflicts are resolved automatically with a single retry. However, it omits whether an existing pipeline is overwritten or versioned, what permissions are required, and what happens if the retry also fails.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences, with the core action front-loaded and the retry behavior appended. Every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The existence of an output schema means return values need not be described, and the retry note is useful. Still, a mutation tool with no annotations, 0% parameter coverage, and a nested required object leaves important semantics (overwrite behavior, name/subject meaning) undocumented.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across four parameters, and the description explains none of them. In particular, the required nested 'train_graph' object, plus 'name' and 'subject', are left entirely for the agent to infer from the schema structure alone.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource: 'Save a pipeline graph into a project', which cleanly separates it from load_pipeline, validate_pipeline and run_pipeline by verb. The parenthetical adds scope ('visible on the studio canvas'), but no sibling is named explicitly for differentiation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no prerequisites, and no mention of alternatives such as load_pipeline or validate_pipeline. Usage is only implied by the verb 'save'.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

start_streamStart StreamAInspect

Connect an EEG device and START a live streaming session on the user's head. Requires confirm=True; call test_device first. Track with stream_status(). Idle watchdog: if no stream_status()/get_live_session() poll happens for idle_timeout_sec (default 900), the session is stopped and the device disconnected automatically — an abandoned stream never keeps running on the user's head. Any poll resets the timer; idle_timeout_sec=0 disables the watchdog.

ParametersJSON Schema
NameRequiredDescriptionDefault
portNo
confirmNo
ip_portNo
source_idNo
chunk_sizeNo
ip_addressNo
n_channelsNo
session_idNo
device_typeYes
mac_addressNo
stream_nameNo
serial_numberNo
connection_typeNo
idle_timeout_secNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.3/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does so well: it discloses the confirm gate, the prerequisite test_device call, and the idle watchdog (auto-stop + device disconnect after idle_timeout_sec, empty stream never left running, any poll resets, idle_timeout_sec=0 disables). This is substantial behavioral context about an operation on the user's body.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the purpose, then the confirm/prerequisite requirement, then the watchdog caveat. Dense but each sentence earns its place; the watchdog explanation is slightly long but warns of a real safety-relevant behavior.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained, and the safety/lifecycle behavior is covered. However, for a complex 14-parameter device-connection tool with 0% schema coverage, the omission of any device configuration semantics leaves a meaningful gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% for 14 parameters, so the description must compensate, yet it only explains confirm and idle_timeout_sec. Crucially, the required device_type parameter and all connection fields (port, ip_port, ip_address, mac_address, serial_number, connection_type, n_channels, chunk_size, source_id, session_id, stream_name) are left completely undefined.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource — connect an EEG device and start a live streaming session — and names related siblings (test_device, stream_status, stop_stream implicitly), so the agent can distinguish it from other device tools without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly states prerequisites and ordering: 'Requires confirm=True; call test_device first' and 'Track with stream_status()'. It routes the agent to the right siblings and conditions rather than leaving sequencing to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

stop_streamStop StreamAInspect

Stop a streaming session and disconnect the device (always safe to call). Also removes the session from the idle watchdog so it cannot fire after an explicit stop.

ParametersJSON Schema
NameRequiredDescriptionDefault
session_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does well: it declares the operation safe, names the device-disconnect effect, and discloses the non-obvious watchdog side effect that prevents a stale timer from firing after an explicit stop. Gaps remain around behavior when session_id is unknown/expired and whether the call is idempotent, but the disclosed traits are substantive.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tightly written sentences with the action and safety guarantee front-loaded and the subtle watchdog behavior second. No filler; every clause adds information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained, and the description covers the key side effects of a stop operation. It is slightly short on error/edge-case behavior and omits any guidance on sourcing the session identifier, which is the one thing the agent must know to call it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the single required parameter session_id is completely undocumented in both schema and description. The description never says what a session_id is or where the agent obtains one (e.g., from start_stream, stream_status, or get_live_session), so it fails to compensate for the coverage gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource ('Stop a streaming session') plus the concrete side effect ('disconnect the device'), so an agent knows exactly what happens. It is not, however, explicitly differentiated from siblings like cancel_execution or stream_status beyond the self-evident 'stop' semantics.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'always safe to call' and the note about removing the session from the idle watchdog imply this is the correct explicit-termination path, but no alternative (e.g., stream_status for inspection, cancel_execution for a different resource) is named or contrasted. Usage is inferable but never stated as when-to-use vs when-not.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

stream_statusStream StatusAInspect

Live snapshot of a streaming session (running, deviceConnected). Polling this also feeds the idle watchdog: each call resets the session's idle timer (see start_stream's idle_timeout_sec).

ParametersJSON Schema
NameRequiredDescriptionDefault
session_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden, and it discloses a genuinely non-obvious side effect: each call resets the session's idle timer, so a read-looking status poll actually keeps the session alive. That is high-value behavioral context, though permissions, error behavior for unknown sessions, and any rate limits remain unstated.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences, front-loaded with what the tool returns and followed by the polling side effect. No filler and no repetition of structured fields.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained, and the description covers purpose plus the idle-watchdog interaction. The remaining gap is session_id provenance and behavior on an invalid/expired session, minor for a one-parameter status tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the single session_id parameter is undocumented. The phrase 'the session's idle timer' only implicitly hints that session_id identifies a session; nothing states its format or where the agent obtains it (presumably from start_stream).

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource: a 'live snapshot of a streaming session' with example fields (running, deviceConnected). It does not explicitly distinguish itself from the sibling get_live_session or stop_stream, so an agent must infer the boundary, but the purpose itself is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied rather than stated: 'Polling this...' signals the intended call pattern and the description references start_stream's idle_timeout_sec, but it never says when to choose this over get_live_session or what happens if the session has already ended.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

test_deviceTest DeviceAInspect

Test a device connection WITHOUT starting a stream (safe, no confirm needed).

ParametersJSON Schema
NameRequiredDescriptionDefault
portNo
ip_portNo
source_idNo
ip_addressNo
device_typeYes
mac_addressNo
stream_nameNo
serial_numberNo
connection_typeNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden. It usefully discloses the key trait that no stream is started and no confirmation is required, implying a non-destructive probe. However, it says nothing about authentication requirements, error/failure reporting, or what happens to existing streams, which matters for a 9-parameter device tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with zero filler; the safety constraint is emphasized and placed where it is read first.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained, but for a 9-parameter tool with 0% schema coverage the description leaves the agent unable to select the right connection parameters for a given device_type. It is adequate only for the simplest invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across 9 parameters, so the description must compensate and it does not. It provides no guidance on device_type (the only required field), nor on which connection identifiers (port vs ip_address/ip_port, serial_number, mac_address, source_id, stream_name) apply to which device types.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Test a device connection') and immediately distinguishes itself from the streaming siblings by clarifying it does NOT start a stream. An agent can tell this apart from start_stream/stream_status without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The parenthetical establishes clear usage context: use this to validate a connection without the side effect of starting a stream, and that it requires no confirmation. It stops short of naming the alternative tool explicitly or stating when-not to use it, but the condition for choosing it is unambiguous.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

upload_dataUpload DataAInspect

Upload an EEG file (.edf/.bdf/.mat/.csv/.txt/.tsv/.h5/.hdf5, <=500MB) to the backend and get the registered path for a custom_data node.

sampling_rate (Hz, e.g. 250.0) is REQUIRED for plain CSV/TSV/TXT files without embedded metadata — the backend silently assumes 250 Hz otherwise, which mis-times epochs, filters and spectral features. format overrides extension-based detection (auto, mat, csv, tsv, txt, edf, bdf, h5, hdf5).

ParametersJSON Schema
NameRequiredDescriptionDefault
formatNo
file_pathYes
dataset_nameNo
sampling_rateNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does well: it discloses the 500MB cap, the supported format set, the silent 250 Hz fallback that mis-times epochs/filters/spectral features, and that format overrides extension detection. It omits auth/permission requirements and whether an upload overwrites or versions an existing dataset.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short paragraphs, front-loaded with the action and constraints, then the parameter warning. Every sentence adds information an agent needs; nothing is restated from the schema or title.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return-value explanation is unnecessary, and the description still tells the agent the practical result is a registered path usable as a custom_data node. Only the dataset_name parameter and any auth/overwrite behavior are left uncovered.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate and largely does: it enumerates the format values, explains sampling_rate's unit and silent-default hazard, and implies file_path is the local file to send. dataset_name is never mentioned, leaving one of four parameters undocumented in both places.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (upload) and resource (EEG file), names the accepted extensions and size cap, and states the outcome: a registered path for a custom_data node. No sibling tool uploads files, so the purpose is unambiguous against the list.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives a clear conditional rule for when sampling_rate is required (plain CSV/TSV/TXT without embedded metadata) and when format should be used (to override extension detection). It does not say when not to use this tool or name an alternative path, but the operational guidance for the risky parameter is explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

validate_node_configValidate Node ConfigBInspect

Validate one node's config object against its schema (get_node_schema).

ParametersJSON Schema
NameRequiredDescriptionDefault
configYes
node_typeYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It does convey that this is a validation/check operation (not a mutation) and names the schema source, but says nothing about required permissions, what happens on failure, or whether it is purely read-only. Return values are covered by the output schema, which lowers but does not remove this gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler. It is efficient, though the parenthetical schema reference is slightly terse and could be sharpened.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 2-param tool with an output schema, the description covers purpose and schema source but omits any routing between this tool and validate_pipeline and gives no config-shape guidance, leaving it only minimally adequate.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It references the config object and, via 'its schema', implicitly the node_type that selects that schema, but adds no format or structure detail for either parameter. Partially compensating at best.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource: validate one node's config object against its schema. This is clearly distinguishable from the sibling validate_pipeline (node-level vs pipeline-level), though the description never names that sibling explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The parenthetical reference to get_node_schema implies where the schema comes from, suggesting a usage flow, but there is no explicit when-to-use statement or guidance on choosing this over validate_pipeline. Usage must be inferred.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

validate_pipelineValidate PipelineAInspect

Validate a pipeline graph before running. ExecGraphSnapshot: {nodes: [{id, type, config}], connections: [{from, to}]}. Build it from get_template(id).train or from scratch using list_nodes().

ParametersJSON Schema
NameRequiredDescriptionDefault
train_graphYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It implies a non-mutating pre-flight check ('before running'), which is useful context, but it says nothing about what happens on failure, whether it blocks execution, or any auth/permission needs. The presence of an output schema lowers the bar since return values need not be described here.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose is front-loaded in the first clause, followed by the input structure and how to build it. Three compact sentences, each earning its place with no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, the description only needs to cover purpose, input construction, and execution context, which it does. It is nearly complete, though it leaves unstated what kinds of problems validation catches and whether it is purely a dry-run.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the single parameter is a loosely-typed object, so the description must compensate. It does so by spelling out the ExecGraphSnapshot shape (nodes with id/type/config, connections with from/to), which meaningfully exceeds the bare 'type: object' in the schema, though it doesn't detail the per-node config fields.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (validate) and resource (a pipeline graph), and adds the timing qualifier 'before running'. This naturally distinguishes it from the sibling validate_node_config, which operates on a single node, without the agent needing to open either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear context for when to reach for it ('before running') and even routes the agent to sibling tools for building the input ('get_template(id).train', 'list_nodes()'). It stops short of explicit when-not guidance or a direct comparison to alternatives like validate_node_config.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

whoamiWhoamiAInspect

Who you are authenticated as: account email, plan (isPro / pioneer), this month's free-run quota, and — with a hosted token — the token name and days until it expires. Call this first when setup guidance appears or to check which credential a session uses.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It transparently describes the returned credential data, including conditional hosted-token details and expiry timing, but does not explicitly state that the operation is read-only or describe authentication failure behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, both front-loaded and purposeful: the first declares the returned identity data, the second declares when to call it. There is no wasted wording.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter identity tool with an output schema and no risky side effects, the description covers what it returns, when to call it, and hosted-token behavior. Nothing an agent needs to invoke it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, so the baseline is 4. The description correctly does not discuss parameters, and there is no parameter surface to clarify beyond the empty input schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific identity action and enumerates the exact account details returned: email, plan, quota, and hosted-token expiration. This clearly distinguishes it from all sibling tools, which operate on pipelines, datasets, devices, or artifacts rather than credentials.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives explicit usage contexts: 'Call this first when setup guidance appears' and 'to check which credential a session uses.' This is clear when-to-use guidance, though no when-not-to-use condition or alternative tool is named, which is acceptable because no sibling overlaps with auth introspection.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 32 tool updatesv0.5.1
    • First observedcancel_execution
    • First observedcreate_project
    • First observeddownload_artifact
    • First observedexport_python
    • First observedget_execution
    • First observedget_experiment
    • First observedget_leaderboard
    • First observedget_live_session
    • First observedget_node_schema
    • First observedget_results
    • First observedget_template
    • First observedinspect_dataset
    • First observedinspect_file
    • First observedlist_artifacts
    • First observedlist_datasets
    • First observedlist_devices
    • First observedlist_executions
    • First observedlist_nodes
    • First observedlist_projects
    • First observedlist_templates
    • First observedload_pipeline
    • First observedrun_experiment
    • First observedrun_pipeline
    • First observedsave_pipeline
    • First observedstart_stream
    • First observedstop_stream
    • First observedstream_status
    • First observedtest_device
    • First observedupload_data
    • First observedvalidate_node_config
    • First observedvalidate_pipeline
    • First observedwhoami

TDQS

B3.2/5.0

Scored across 32 tools

Disambiguation4/5

Most tools target distinct actions/resources, but a few pairs overlap: stream_status vs get_live_session both poll a live session, and inspect_dataset vs inspect_file give near-identical exploratory summaries for different sources. Descriptions help disambiguate, but an agent may still hesitate between them.

Naming Consistency4/5

Names are predominantly consistent snake_case verb_noun (list_nodes, get_template, run_pipeline). Minor deviations like whoami and stream_status break the exact verb_noun pattern but remain readable and conventional.

Tool Count2/5

32 tools is heavy for a single MCP server; several subdomains (streaming, pipelines, datasets, projects, artifacts) could be consolidated. While each tool has a purpose, the set exceeds the 25-tool threshold where navigation becomes cumbersome.

Completeness4/5

The surface covers major workflows: device streaming, pipeline building/validation/execution, experiments, datasets, projects, and artifacts. Minor gaps exist (no delete/update for projects, pipelines, or artifacts), but agents can work around them.

Related MCP Connectors

Related MCP Servers

  • A
    license
    A
    quality
    C
    maintenance
    A Model Context Protocol server that provides AI agents with a unified interface for real-time EEG acquisition, replay, processing, visualization, recording, and stimulation from over 66 BrainFlow boards.
    47
    2
    BSD 3-Clause
  • A
    license
    B
    quality
    C
    maintenance
    An MCP server that gives AI assistants conversational access to MNE-Python for analyzing neurophysiology data (EEG, MEG, sEEG, ECoG, fNIRS). Enables plain-language analysis pipelines, from loading recordings to generating figures and explanations.
    41
    8
    MIT
  • A
    license
    Not graded
    quality
    B
    maintenance
    An MCP server that gives AI agents a unified interface to clinical/research EEG workflows: signal processing, source imaging, persistent dataset/EHR storage, and web visualization.
    59 PyPI
    1
    BSD 3-Clause