Skip to main content
Glama

Server Details

AI agents build, train, and analyze BCI/EEG pipelines: data, models, experiments, live sessions.

If you are the author of this connector, you can claim ownership by verifying the domain or GitHub account it belongs to. Claimed connector authors can inspect health checks, view analytics, and manage their listing.
Status
Healthy
Last Tested
Transport
Streamable HTTP · MCP 2025-11-25
URL
Repository
nimbusbci/nimbus-mcp
GitHub Stars
0
Server Listing
nimbus-mcp

TDQS

A3.5/5.0

Scored across 37 tools

Disambiguation4/5

Namespaces (account, calibration, catalog, data, device, execution, experiment, pipeline, project, stream) cleanly partition purposes, and most tools within a namespace are distinct. Minor potential overlap exists between data.inspect_dataset/data.inspect_file and between execution.results/experiment.get aggregates, but descriptions disambiguate these.

Naming Consistency5/5

Uniform namespace.verb_noun convention throughout (catalog.template, execution.run, project.save, stream.stop), with dot-namespacing consistently applied across all 37 tools. No mixing of camelCase or stray verb styles.

Tool Count2/5

At 37 tools this is heavy even for a broad BCI lifecycle (device, streaming, calibration, pipelines, executions, experiments, projects, catalog). While each namespace is individually coherent, the sheer surface risks overwhelming an agent's selection process.

Completeness4/5

Strong lifecycle coverage: catalog discovery, data upload/inspect, pipeline build/validate/export, project save/load, execution lifecycle, experiment runs, calibration flow, and streaming control. Minor gaps like project.delete and device enumeration beyond list/test keep it short of perfect.

Available Tools

37 tools
account.whoamiAccount.WhoamiA
Read-onlyIdempotent
Inspect

Who you are authenticated as: account email, plan (isPro / pioneer), this month's free-run quota, and — with a hosted token — the token name and days until it expires. Call this first when setup guidance appears or to check which credential a session uses.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly/idempotent/non-destructive, so the safety bar is low. The description adds real behavioral context beyond that: output varies by credential type ('with a hosted token' surfaces token name and expiry), which is a conditional behavior an agent cannot infer from annotations or the empty schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with the returned fields and closing with the usage trigger. The field enumeration is dense but each item (plan values, token expiry) earns its place by telling the agent what it will get.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return format is already covered, and the tool takes no params. For a read-only identity check, the description supplies everything an agent needs: what it does, what it returns, and when to reach for it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Zero parameters, so the baseline is 4; there is nothing to document. The description does not need to add parameter meaning and correctly does not invent any.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Names the specific operation (identify the authenticated account) and enumerates exactly what it returns: email, plan, free-run quota, and hosted-token name/expiry. This is fully distinguishable from every sibling (catalog.*, execution.*, stream.*, etc.), none of which touch identity or credentials.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives concrete triggers: 'Call this first when setup guidance appears' and 'to check which credential a session uses.' That is clear when-to-use guidance. It stops short of explicit when-not or named alternatives, but no sibling competes for this role, so the guidance is effectively complete.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

calibration.pauseCalibration.PauseAInspect

Pause a running calibration between trials (cues hold; resume anytime).

ParametersJSON Schema
NameRequiredDescriptionDefault
execution_idYesThe calibration run to pause.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations cover the safety profile (destructiveHint=false, idempotentHint=false, openWorldHint=true), so the bar is lower. The description still adds real behavioral value: 'cues hold' describes the state effect on the run and 'resume anytime' confirms reversibility, which is more than the annotations convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One short sentence, front-loaded with the action and resource, with the parenthetical adding effect detail rather than filler. 'Cues hold' is slightly jargon-ish and could be clearer, but nothing is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need no explanation, and annotations carry the safety/idempotency profile. The description covers the action and its reversible state effect, leaving only minor gaps such as interaction with sibling tools.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There is a single parameter with 100% schema description coverage, so the schema already documents execution_id fully. The description adds no format, source, or retrieval detail beyond that, making 3 the correct baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Specific verb (pause) plus resource (calibration) and a scope qualifier ('between trials') that distinguishes it from a hard stop. It does not explicitly name calibration.resume as the counterpart, though 'resume anytime' gestures at it.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Resume anytime' implies the resume path and that the action is reversible, which is implicit guidance on when this is appropriate. However, it never states when to prefer pause over execution.cancel or calibration.status, so usage is left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

calibration.resumeCalibration.ResumeAInspect

Resume a paused calibration session.

ParametersJSON Schema
NameRequiredDescriptionDefault
execution_idYesThe calibration run to resume.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare that this is a non-read-only, non-idempotent, non-destructive, open-world operation. The description adds the useful precondition that the session must already be paused, but it does not describe side effects, required permissions, or what happens if the session is not paused.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single front-loaded sentence with no wasted words. It communicates the essential action and precondition immediately, which is appropriate for a simple one-parameter tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the simple one-parameter schema, rich annotations, and presence of an output schema, the description supplies the key state precondition and action needed to invoke the tool correctly. It only omits explicit usage routing among sibling calibration tools, which is a minor remaining gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and the single parameter execution_id is fully documented as 'The calibration run to resume.' The description adds no parameter-level detail beyond what the schema already provides, so it meets the baseline for high schema coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Resume') and resource ('paused calibration session'), so the core action is unmistakable. It distinguishes the tool from calibration.pause and calibration.start through the word 'resume' and the precondition 'paused', but does not explicitly name or compare sibling tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'paused calibration session' implies the context for use: only call this when a calibration session is currently paused. However, it gives no explicit when-not guidance or alternative routing, such as whether to use calibration.start instead when no session exists.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

calibration.startCalibration.StartAInspect

Start a guided subject calibration session (NON-BLOCKING; confirm-gated — the device goes on a human's head). The Nimbus Studio app shows the cues on its calibration dashboard automatically; poll calibration.status. Requires a Pro plan (hosted token or Pro session): calibration nodes and custom-data training are gated by the freemium node policy; local X-MCP-Key principals get 403 by policy.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameNoDisplay name recorded on the execution.
portNoSerial/COM port for wired devices.
classesNoOverride class list [{id,label,cue}] (MI default: left/right hand).
confirmNoMUST be true — explicit user go-ahead for a session on their head.
ip_portNoPort for network devices.
paradigmNomi | p300 | sart | target_hit.mi
source_idNoLSL source id.
ip_addressNoDevice IP for network/wifi devices.
device_typeNoDevice id from device.list (omit → template default synthetic).
mac_addressNoBluetooth MAC (BT devices).
stream_nameNoLSL stream name (LSL devices).
serial_numberNoDevice serial (some BLE stacks).
connection_typeNoDevice selector when several exist (e.g. serial vs wifi).
trials_per_classNoOverride the template's trial count (e.g. 3 for smoke tests).

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare readOnly=false and idempotent=false; the description adds substantial non-obvious behavior: NON-BLOCKING execution, a confirm gate because the device goes on a human's head, automatic cue display in the app, the Pro-plan requirement, and a specific 403 policy for local X-MCP-Key principals. These are precisely the traits annotations cannot express.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, front-loaded with the purpose and the critical non-blocking/confirm facts before the plan-gating detail. Dense but every clause carries information; minor line-broken parentheticals add slight reading friction.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists and annotations cover the safety profile, so the description only needs to add the operational context it does: non-blocking semantics, the confirm requirement, the follow-up polling step, and the authorization gate. For a 14-parameter, policy-gated session tool, nothing an agent needs before invoking is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and each parameter (confirm, paradigm, classes, connection selectors, device_type, trials_per_class) is fully documented in the schema itself. The description echoes the confirm-gating but adds no new syntactic or constraint detail beyond what the schema already provides, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Start a guided subject calibration session') and clearly distinguishes itself from the sibling lifecycle tools (calibration.status/pause/resume/train). An agent knows immediately this begins a session rather than querying or adjusting one.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear operational guidance: it is non-blocking, the app displays cues automatically, and the caller should follow up by polling calibration.status. It also states a hard prerequisite (Pro plan / hosted token). It does not explicitly contrast with calibration.train or state when-not to use it, keeping it short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

calibration.statusCalibration.StatusA
Read-onlyIdempotent
Inspect

Live snapshot of a calibration session (phase, current trial, progress, paused). Once complete, carries the recorded upload — call calibration.train to turn it into the subject's own classifier.

ParametersJSON Schema
NameRequiredDescriptionDefault
execution_idYesThe calibration run to inspect (from calibration.start).

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly/idempotent/non-destructive, so the safety profile is covered. The description adds real behavioral context beyond the annotations: that the snapshot is 'live' and that after completion it 'carries the recorded upload,' which tells the agent the object's payload changes with session state.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences: the resource and its returned fields are front-loaded, followed by the completion/follow-up behavior. No wasted phrasing.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Output schema exists, so return values need not be spelled out; the description covers purpose, state contents, and the follow-up action. Nothing an agent needs to invoke it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Single parameter with 100% schema description coverage; the schema already explains execution_id including its origin (from calibration.start). The description adds no syntax or format detail beyond that, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific resource and action ('Live snapshot of a calibration session') and enumerates what it exposes (phase, current trial, progress, paused), so the agent knows exactly what this returns. It is clearly the read/status member of the calibration.* family, distinct from pause/resume/start/train.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides clear forward context: 'Once complete... call calibration.train to turn it into the subject's own classifier,' which routes the agent to the correct next tool. It stops short of stating when to use this vs other status-like siblings (e.g., stream.status, execution.get), so no explicit exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

calibration.trainCalibration.TrainAInspect

Train the subject's own classifier from a COMPLETED calibration session (NON-BLOCKING). Fetches the recorded upload, wires it into a train pipeline as a custom_data source, and starts the run. Requires a Pro plan (custom_data training is freemium-gated). The calibrate→train handoff requires a Postgres-backed backend (hosted or local dev); a desktop-local session completes and records, but its upload can't be resolved by MCP train today.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameNoDisplay name recorded on the training run.
paradigmNoParadigm of the recording (mi | p300 | sart | target_hit); only mi has a default train template.mi
template_idNoTrain template to use (e.g. from catalog.templates); required for non-mi paradigms (mi defaults to mi_headband_csp_lda).
train_graphNoExplicit train graph instead of a template; its first data node (custom_data/public_data) is rewired onto the recording.
execution_idYesThe COMPLETED calibration run (from calibration.start).

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only supply the generic mutation profile (readOnlyHint=false, openWorldHint=true, idempotentHint=false, destructiveHint=false). The description adds what annotations cannot: the call is NON-BLOCKING (it starts a run rather than returning a model), it is freemium-gated on a Pro plan, and it has an infrastructure dependency that can make it silently unusable. That is meaningful behavioral context beyond structured fields.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the action and the source requirement, and every sentence carries real information (non-blocking, plan gate, backend constraint). It is somewhat dense with parenthetical caveats, but nothing is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a five-parameter, async, gated mutation with an output schema, the description covers the source state, the pipeline behavior, the entitlement requirement, and the backend limitation. Return values are handled by the output schema, so nothing an agent needs before calling is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all five parameters are already documented, including the paradigm default, the template_id requirement for non-mi paradigms, and the train_graph rewiring rule. The description restates the custom_data rewiring idea but adds no syntax or format detail beyond the schema, so baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Train the subject's own classifier') and immediately scopes it to input from a COMPLETED calibration session, which cleanly separates it from calibration.start and execution.run. An agent can identify the tool without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit preconditions (session must be COMPLETED, Pro plan required, Postgres-backed backend required) and an explicit when-not (desktop-local sessions record but their upload can't be resolved by MCP train). This is exactly the when/when-not guidance that prevents failed calls.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

catalog.datasetsCatalog.DatasetsB
Read-onlyIdempotent
Inspect

Curated public EEG datasets (MOABB packs) available to pipelines.

ParametersJSON Schema
NameRequiredDescriptionDefault
only_on_diskNoOnly return datasets whose data packs are present on this backend (True by default; False also lists known-but-missing sets).

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, openWorldHint, and destructiveHint=false, so the safety profile is covered structurally. The description adds domain context ('curated', 'MOABB packs', 'available to pipelines'), but says nothing about result size, filtering behavior, or how the on-disk vs known-but-missing distinction surfaces — though the output schema likely covers the shape.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single tight sentence with the resource front-loaded and no filler. It is appropriately sized for a simple listing tool, though it is arguably too terse to convey the listing action explicitly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given an output schema exists and annotations declare the read-only safety profile, the description need not explain return values. However, for a catalog entry point it provides no indication of scope (how many datasets, filtering options) or relationship to sibling catalog tools, leaving a modest gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and the only parameter (only_on_disk) is well documented in the schema with its default and behavior. The description adds no parameter-level meaning beyond that, so the baseline of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description identifies the specific resource — curated public EEG datasets (MOABB packs) — and the phrase 'available to pipelines' implies a listing operation. It distinguishes from siblings like catalog.nodes or catalog.templates by naming the resource domain, though it never states an explicit verb like 'list' or 'retrieve'.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance and no mention of alternatives such as data.inspect_dataset for drilling into a specific set. An agent must infer that this is the entry point for discovering datasets, without any routing help.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

catalog.leaderboardCatalog.LeaderboardA
Read-onlyIdempotent
Inspect

Public benchmark leaderboard: pipeline rankings per dataset.

Rankings are per-dataset under the canonical within_session protocol (see protocol). Within each dataset, rows are sorted desc by meanAccuracyPct (95% CI in ciLoPct/ciHiPct). Use pipelineId as the template id hint for catalog.template when building a pipeline. updated marks each dataset's most recent run; packFingerprint identifies the exact dataset pack the scores came from.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already establish the safe read-only/idempotent profile, so the bar is lower; the description adds real context beyond that: the ``within_session`` protocol, that ``rows`` are sorted descending by ``meanAccuracyPct`` with CI bounds, and what ``updated``/``packFingerprint`` signify.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is front-loaded with the core purpose, then layers protocol/sorting/field meaning efficiently. The density of backtick-referenced fields is slightly heavy, but each sentence earns its place with actionable information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, return-value explanation is not strictly required, yet the description goes further by naming key fields and their meaning. For a zero-parameter read tool this is quite complete, though it omits framing like freshness or scope limits of the leaderboard.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so there is nothing parameter-level to document; the baseline for a no-param tool is 4. The description's field references (``pipelineId``, ``meanAccuracyPct``, etc.) are return-value semantics rather than input parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The opening clause states a specific verb+resource: a public benchmark leaderboard returning pipeline rankings per dataset. This is clearly distinguishable from siblings like catalog.datasets, catalog.templates, and execution.results, which don't describe benchmark rankings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives an explicit downstream usage link ('Use ``pipelineId`` as the template id hint for ``catalog.template`` when building a pipeline'), which tells the agent a concrete follow-up action. It does not state when NOT to use this tool or name alternatives, so it falls short of the top band.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

catalog.nodesCatalog.NodesA
Read-onlyIdempotent
Inspect

List Nimbus pipeline node types (data, preprocessing, features, models...).

Use catalog.node_schema(node_type) for one node's full config schema and ports.

ParametersJSON Schema
NameRequiredDescriptionDefault
categoryNoOptional filter — e.g. "data", "preprocessing", "features", "models" (exact category ids from the unfiltered list).

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false and openWorldHint=true, so the safety profile is fully covered structurally. The description adds no behavioral context beyond that — nothing about return shape or filtering semantics — so with annotations doing the heavy lifting a 3 is appropriate.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, front-loaded with the verb and resource, followed immediately by the disambiguation. Zero filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values need not be described, and the input schema is fully documented. The only mild gap is that nothing tells the agent what the unfiltered call returns versus a filtered one, but that is covered by the output schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the single `category` parameter is documented in the schema, including that it takes exact category ids from the unfiltered list. The description's category examples duplicate that, so baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('List Nimbus pipeline node types') and enumerates the domain ('data, preprocessing, features, models...'), so an agent immediately knows the scope. It also names the sibling it is not (catalog.node_schema) which handles single-node detail.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The second sentence routes the agent: use this tool to enumerate node types, use catalog.node_schema for one node's full config schema and ports. That is clear context selection guidance, though there are no explicit exclusions or prerequisite conditions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

catalog.node_schemaCatalog.Node SchemaA
Read-onlyIdempotent
Inspect

Full config JSON schema + input/output ports for one node type.

ParametersJSON Schema
NameRequiredDescriptionDefault
node_typeYesNode id from catalog.nodes (e.g. "csp", "nimbus_lda").

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already establish readOnlyHint=true, idempotentHint=true and destructiveHint=false, so the safety profile is covered. The description adds useful content-level detail (config schema, ports) but says nothing about error behavior for an unknown node_type or how large the payload can be, leaving behavioral gaps.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence-fragment with zero filler, sized appropriately for a one-parameter lookup tool. It is a noun phrase rather than a sentence, which slightly reduces immediate readability.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, the description need not explain return values, and it correctly signals the scope (one node type). For a read-only, single-param tool this is nearly complete, missing only failure behavior for an unrecognized node_type.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and the single parameter is documented in the schema, including examples ('csp', 'nimbus_lda') and its provenance from catalog.nodes. The description adds no meaning beyond that, which is the expected baseline when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific resource ('node type') and states precisely what is returned: the full config JSON schema plus input/output ports. It is clearly distinct from the list-oriented siblings catalog.nodes and catalog.templates, though it never names those siblings to make the contrast explicit.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is only implied: the node_type parameter's reference to catalog.nodes hints at a discover-then-inspect workflow, but the description never states when to call this versus catalog.template, pipeline.validate_node, or the list tools, and gives no preconditions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

catalog.templateCatalog.TemplateA
Read-onlyIdempotent
Inspect

Full template incl. the 'train' execGraph needed by execution.run/pipeline.validate.

ParametersJSON Schema
NameRequiredDescriptionDefault
template_idYesTemplate id from catalog.templates (e.g. "mi_csp_lda").

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, openWorldHint, and destructiveHint=false, so the safety profile is covered. The description adds the key content detail that the 'train' execGraph is included, but does not disclose authentication needs, rate limits, or other behavioral traits beyond the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single efficient sentence with no wasted words, front-loading what the template contains. It could be slightly clearer with an explicit action verb, but it is appropriately sized.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained, and annotations cover the safety profile. The description states the key content and usage context, making it complete enough for a simple read-only retrieval, though it stops short of contrasting with the similar catalog.templates sibling.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and the single template_id parameter is fully documented with an example in the schema. The description adds no further parameter meaning, so the baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description identifies the resource as the full template and specifies that it includes the 'train' execGraph. It is clear what the tool returns, but it lacks an explicit verb and does not directly distinguish itself from the sibling catalog.templates.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It states the tool is needed by execution.run and pipeline.validate, giving a concrete usage context. It does not provide exclusions or explicitly compare against alternatives such as catalog.templates.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

catalog.templatesCatalog.TemplatesA
Read-onlyIdempotent
Inspect

List built-in starter pipelines (MI/P300/SSVEP...). catalog.template(id) returns the graph.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, openWorldHint and destructiveHint=false, so the safety profile is covered. The description adds only that the pipelines are 'built-in' and static; it says nothing about pagination, caching, or auth needs, but with annotations and an output schema present the burden is lower.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, purpose front-loaded, zero filler. Every clause either names the resource or routes to the correct sibling.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a no-arg list tool with an output schema and full annotation coverage, the description supplies what is needed to select and invoke it. A note on whether the list is static versus user-extensible would have closed the remaining gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so per the rubric the baseline is 4. There is nothing for the description to disambiguate on the input side.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (List) and resource (built-in starter pipelines) with concrete examples (MI/P300/SSVEP). The second sentence routes the agent to the singular sibling catalog.template for the actual graph, which helps distinguish it from catalog.template.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied rather than stated: the parenthetical about catalog.template(id) hints that this tool is for browsing and the singular is for fetching a graph, but there is no explicit when/when-not guidance or named alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

data.inspect_datasetData.Inspect DatasetA
Read-onlyIdempotent
Inspect

Exploratory summary of a public EEG dataset (MOABB pack): channels, sampling rate, trial/class balance, per-channel µV stats, band powers and a PSD overview.

Look at the data BEFORE building pipelines: class balance drives stratification choices (imbalanced classes skew accuracy), and flatlined channels mean a montage/reference problem worth fixing first. subject is REQUIRED (the backend 400s without it) — get the subject list via catalog.datasets, e.g. "S01"; a comma-list like "S01,S03" loads a cohort. mode: training | evaluation | all. Units note: values are ASSUMED volts by the loader — a µV-native file reads 1e6x too large; set unitsScale in a pipeline's custom_data config when needed.

ParametersJSON Schema
NameRequiredDescriptionDefault
modeNoWhich split to summarize — training | evaluation | all.all
datasetYesDataset id from catalog.datasets (e.g. "BNCI2014_001").
subjectNoREQUIRED subject code ("S01") or comma-list cohort ("S01,S03").

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already establish readOnly/idempotent/non-destructive/openWorld, so safety is covered. The description goes beyond them with real operational context: subject is REQUIRED and the backend 400s without it, and values are ASSUMED volts so µV-native files read 1e6x too large, with the unitsScale fix named. It does not describe cost or latency, but the added caveats are substantial.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with what the summary contains, then rationale, then hard preconditions. Dense and slightly parenthesis-heavy (the units aside is long), but every sentence carries information an agent needs and nothing is redundant.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return structure needn't be explained, and annotations cover the safety profile. The description still supplies the precondition (subject required), the subject-sourcing path, the mode enumeration, and the units pitfall — everything needed to call it correctly the first time.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% so the baseline is 3, but the description adds meaning beyond the schema: the 400 error consequence of omitting subject, the cohort semantics of a comma-list ('S01,S03'), and the units-scale caveat tied to downstream config. That is a genuine increment over the schema's own text.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Exploratory summary of a public EEG dataset') and enumerates exactly what is produced: channels, sampling rate, trial/class balance, per-channel µV stats, band powers, PSD overview. This distinguishes it from the sibling data.inspect_file, which covers files rather than packed datasets.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly tells the agent to run it BEFORE building pipelines and gives the reason (class balance drives stratification; flatlined channels signal a montage/reference problem). It also routes to catalog.datasets for the subject list, naming both the when and the alternative source.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

data.inspect_fileData.Inspect FileA
Read-onlyIdempotent
Inspect

Exploratory summary of an EEG file (.edf/.bdf/.mat/.csv/.tsv/.txt/.h5): channels, sampling rate, trial/class balance, per-channel µV stats, band powers and a PSD overview.

The path shape picks the source: ABSOLUTE path → read the file from disk (only on a LOCAL backend: desktop app / MCP local mode — no upload needed); RELATIVE path (the one data.upload returns) → describe the uploaded file, which works on ANY backend (hosted or local).

Look at the data BEFORE building pipelines: class balance drives stratification choices, and flatlined channels mean a montage/reference problem worth fixing first. Units note: values are ASSUMED volts by the loader — a µV-native CSV reads 1e6x too large; set unitsScale in a pipeline's custom_data config when needed. On a hosted backend absolute paths are refused and this returns guidance (upload the file first or switch to a local backend).

ParametersJSON Schema
NameRequiredDescriptionDefault
pathYesABSOLUTE filesystem path (local backends only) or the RELATIVE upload path returned by data.upload (works on any backend).

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations cover the safety profile (readOnly, idempotent, non-destructive), and the description adds substantive behavior beyond them: absolute paths fail on hosted backends and return guidance instead of data, relative upload paths work everywhere, and the loader assumes volts so µV-native CSVs read 1e6x too large. These are exactly the failure modes an agent needs pre-warned.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose and contents are front-loaded in the first sentence, and the path/backend guidance follows logically. It is dense with useful content, though the units note and backend caveat could be tightened slightly without loss.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values need no explanation, yet the description still previews the summary contents. Combined with the units pitfall, backend restrictions, and preprocessing motivation, an agent has everything needed to call this correctly on either backend.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description contributes real meaning beyond the schema wording by spelling out the consequence of each path form (disk read vs. uploaded file) and the backend-dependent failure behavior. The only redundancy is repeating absolute/relative, which the schema already states.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Exploratory summary of an EEG file') and enumerates exactly what the summary contains (channels, sampling rate, trial/class balance, per-channel µV stats, band powers, PSD). The file-level scope is clearly separable from the sibling data.inspect_dataset and data.upload.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly says to run this BEFORE building pipelines and explains why ('class balance drives stratification choices, and flatlined channels mean a montage/reference problem'). It also gives the alternative path when a backend refuses absolute paths: 'upload the file first or switch to a local backend.'

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

data.uploadData.UploadAInspect

Upload an EEG file (.edf/.bdf/.mat/.csv/.txt/.tsv/.h5/.hdf5, <=500MB) to the backend and get the registered path for a custom_data node.

sampling_rate (Hz, e.g. 250.0) is REQUIRED for plain CSV/TSV/TXT files without embedded metadata — the backend silently assumes 250 Hz otherwise, which mis-times epochs, filters and spectral features. format overrides extension-based detection (auto, mat, csv, tsv, txt, edf, bdf, h5, hdf5).

ParametersJSON Schema
NameRequiredDescriptionDefault
formatNoOverride extension-based detection (auto|mat|csv|tsv|txt|edf|bdf|h5|hdf5).
file_pathYesLocal file to upload (.edf/.bdf/.mat/.csv/.txt/.tsv/.h5/.hdf5, <=500MB).
dataset_nameNoOptional label for the uploaded dataset.
sampling_rateNoHz for headerless CSV/TSV/TXT (REQUIRED there, e.g. 250.0).

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare the write/non-idempotent/open-world profile, so the bar is lower, yet the description adds real context: the backend silently assumes 250 Hz and this mis-times epochs, filters and spectral features, plus the 500MB cap and format-override behavior. It does not cover auth/permission needs or whether re-uploads overwrite, which keeps it from a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight paragraphs, front-loaded with the core action and constraints, no filler. The extension list in paragraph one lightly duplicates the schema, which is the only minor redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, return-value detail is unnecessary, and the description still covers side effects, hard limits, format resolution and the sampling_rate pitfall. Nothing material is missing for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so baseline is 3, but the description goes beyond the schema by explaining *why* sampling_rate matters (the silent 250 Hz default and its downstream effects) and that format overrides extension-based detection. This rationale is not present in the field descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (upload) on a specific resource (EEG file) with supported extensions and size limit, and names the output artifact (registered path for a custom_data node). It is clearly distinguishable from read-only siblings like data.inspect_file or execution.download_artifact.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives a strong conditional rule: sampling_rate is REQUIRED for headerless CSV/TSV/TXT, with the consequence spelled out. It stops short of explicitly naming alternative tools (e.g. data.inspect_file to pre-validate a file), so it is clear context rather than full when/when-not routing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

device.listDevice.ListA
Read-onlyIdempotent
Inspect

EEG devices supported by this backend (OpenBCI, Muse, BrainBit, LSL, PiEEG...).

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true and destructiveHint=false, so the safety profile is covered. The description does add one useful piece of context, that this is the set of devices the backend supports (a static capability list) rather than currently connected hardware, but it says nothing about pagination or refresh behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single short sentence with the resource stated first and examples in parentheses; nothing is wasted and nothing is buried.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, the description need not enumerate return values, and it correctly identifies the content of that output. For a zero-parameter read-only listing tool this is nearly complete, missing only a one-line pointer to when the list should be consulted.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so there is no parameter semantics to explain and the baseline of 4 applies. The description correctly neither invents inputs nor misleads about filtering options.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names the specific resource (EEG devices supported by this backend) and enumerates concrete examples (OpenBCI, Muse, BrainBit, LSL, PiEEG), so an agent can tell it apart from device.test at a glance. It is a noun phrase rather than a verb+resource, but the intent is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use statement, no mention of prerequisites, and no reference to alternatives such as device.test or stream.start. An agent must infer from the name alone that this is a discovery call made before configuring a stream.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

device.testDevice.TestA
Read-onlyIdempotent
Inspect

Test a device connection WITHOUT starting a stream (safe, no confirm needed).

ParametersJSON Schema
NameRequiredDescriptionDefault
portNoSerial/COM port for wired devices.
ip_portNoPort for network devices.
source_idNoLSL source id.
ip_addressNoDevice IP for network/wifi devices.
device_typeYesDevice id from device.list (e.g. "brainbit", "muse").
mac_addressNoBluetooth MAC (BT devices).
stream_nameNoLSL stream name (LSL devices).
serial_numberNoDevice serial (some BLE stacks).
connection_typeNoDevice-specific selector when several exist (e.g. serial vs wifi).

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly, idempotent, non-destructive and openWorld, so the safety profile is covered by structured data. The description adds only the 'no confirm needed' nuance, which largely restates the readOnly/non-destructive hints, and says nothing about latency, failure modes, or what a successful test implies.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single sentence with the key constraint front-loaded and no filler. The capitalized WITHOUT emphasizes the discriminating behavior efficiently.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With a full output schema and complete annotation set, the description doesn't need to explain returns or safety. It is nearly sufficient, only lacking a pointer to device.list for valid device_type values and any note on what a failed test looks like.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and every one of the 9 parameters carries its own documentation (port, ip_address, mac_address, connection_type, etc.), so the schema does the heavy lifting. The description contributes no parameter-level meaning beyond that, which is the baseline case.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (test) and resource (device connection) and immediately contrasts with stream.start by saying it does NOT start a stream. That is a clear, distinguishable purpose, though it doesn't reference the device.list prerequisite that supplies device_type.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit context for when to use it (probing a connection without a stream) and implicitly routes away from stream.start for actual data capture. No explicit when-not or named-alternative phrasing, so it falls short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

execution.artifactsExecution.ArtifactsB
Read-onlyIdempotent
Inspect

Trained artifacts (models/filters, e.g. *.pkl) saved by an execution.

ParametersJSON Schema
NameRequiredDescriptionDefault
execution_idYesRun whose artifacts to list.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, openWorldHint, and destructiveHint=false, so the safety profile is covered. The description adds useful domain context about artifact types (models/filters, *.pkl) but does not disclose pagination, auth requirements, or other behavioral traits beyond annotations and the output schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single tight sentence with no wasted words, and the core resource meaning is front-loaded. It is slightly under-specified as a fragment rather than a complete operational sentence.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple read-only listing tool with a full parameter schema, rich annotations, and an output schema, the description supplies the key domain context about what artifacts are. It is complete enough for correct invocation, though sibling routing is absent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the single execution_id parameter is fully documented as the run whose artifacts to list. The description does not add syntax, format, or constraints beyond what the schema already provides, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific resource and scope: trained artifacts (models/filters, e.g. *.pkl) saved by an execution. It is clear what the tool returns, though it does not explicitly state a verb like 'list' and does not differentiate itself from sibling execution.download_artifact.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance is provided. The description implies artifacts belong to an execution but never says when to call this tool versus execution.download_artifact, execution.results, or execution.get.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

execution.cancelExecution.CancelB
Destructive
Inspect

Cancel a running execution.

ParametersJSON Schema
NameRequiredDescriptionDefault
execution_idYesThe run to terminate (from execution.run/execution.list).

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, idempotentHint=false, and readOnlyHint=false, so the safety profile is covered. The description adds nothing beyond that — no mention of irreversibility, auth requirements, or what happens to partial results of the cancelled run.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with zero waste. It is arguably too terse for a destructive operation, but it is appropriately sized and structured.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present and annotations covering the safety profile, the description need not explain return values. However, for a destructive, non-idempotent operation it omits what cancellation actually does to the run, leaving a meaningful gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the single parameter already documents its origin (from execution.run/execution.list). The description adds no meaning beyond the schema, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Cancel) and resource (running execution), which is clearly distinct from siblings like execution.run and execution.list. It does not explicitly name or contrast with siblings, so it falls short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The word 'running' implies a precondition (only in-flight executions can be cancelled), which is a weak usage signal. There is no explicit when-to-use, when-not-to-use, or reference to an alternative such as stream.stop.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

execution.download_artifactExecution.Download ArtifactA
Read-onlyIdempotent
Inspect

Download one artifact file to NIMBUS_EXPORT_DIR/executions// and return its path.

ParametersJSON Schema
NameRequiredDescriptionDefault
execution_idYesRun that produced the artifact.
artifact_nameYesFile name from execution.artifacts (e.g. "nimbus_lda.pkl").

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare the safety profile (readOnly, idempotent, non-destructive), so the bar is lower. The description adds genuinely useful behavior beyond that: the file is written to a fixed local directory NIMBUS_EXPORT_DIR/executions/<id>/, which tells the agent this has a local filesystem side effect. It omits overwrite/collision behavior, which would have pushed it higher.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single sentence that is front-loaded with the action and destination and ends with the return value. Every clause earns its place with zero filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so the description needn't explain the returned path, and annotations cover the safety profile. The description supplies the destination template and confirms the return value, making it largely complete; only edge behavior like filename collisions or missing artifacts is unaddressed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, with execution_id and artifact_name both documented in the schema (including a filename example). The description adds no parameter-level detail beyond the schema, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Download) and resource (one artifact file) plus the exact landing path, so an agent immediately knows what it does. It does not explicitly contrast with the sibling execution.artifacts, but the singular-file framing implies the distinction clearly enough.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is only implied: you use this when you need the artifact file locally rather than just its metadata. There is no explicit when-to-use statement, no mention of execution.artifacts as the listing counterpart, and no prerequisites such as requiring the execution to be finished.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

execution.getExecution.GetB
Read-onlyIdempotent
Inspect

Execution status summary (status: running/completed/failed/cancelled).

ParametersJSON Schema
NameRequiredDescriptionDefault
execution_idYesThe run to check (from execution.run/execution.list).

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, openWorldHint and destructiveHint=false, so the safety profile is covered. The description's only added behavior is enumerating the returned status values (running/completed/failed/cancelled), which is useful but thin given an output schema also exists.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single short sentence with zero padding, front-loaded on the resource and its output. It is efficient, though arguably too terse to carry routing information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, return-value detail needn't be repeated, and the single required parameter is fully documented. The notable omission is sibling routing (when to pick this over execution.list/results), which keeps it short of a 5.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and the schema already explains execution_id and its provenance (from execution.run/execution.list). The description adds nothing about the parameter, so the baseline 3 for a fully documented schema applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Execution status summary') and gives the possible status values, so the agent knows it retrieves run state rather than logs or artifacts. It does not, however, distinguish itself from close siblings like execution.list, execution.results, or stream.status.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no statement of when to use this tool versus execution.list, execution.results, or stream.status. The only hint of context is the parenthetical status list, which implies polling for a run's state but never says so explicitly.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

execution.listExecution.ListC
Read-onlyIdempotent
Inspect

Recent executions. Optional status filter (running/completed/failed/cancelled).

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMaximum number of runs to return (default 20).
statusNoFilter by run status: running/completed/failed/cancelled.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly=true, idempotent=true, destructive=false, so the safety profile is covered. The description adds the recency scope and enumerates valid status values, which is useful, but says nothing about pagination, ordering, or whether 'recent' is time-bounded.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short fragments, front-loaded and free of filler. It is efficient, though the terseness borders on under-specification rather than tight writing.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, return values need not be explained, and annotations cover safety. However, the description never mentions the limit parameter or defines what 'recent' means, leaving gaps for a listing tool with no required params.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both limit and status are already documented in the schema; the description repeats the status enum rather than adding syntax or default behavior. Baseline 3 applies when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

'Recent executions' identifies the resource but never uses an explicit verb like 'list' or 'retrieve'. It is distinguishable from execution.get (single run) only by inference from the sibling name, not from the description itself.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no statement of when to use this versus execution.get, execution.results, or execution.artifacts. The word 'Recent' hints at a recency scope but no alternative is named and no condition is given.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

execution.resultsExecution.ResultsA
Read-onlyIdempotent
Inspect

Metrics for a completed run. Trimmed by default (accuracy, kappa, ITR, confusion matrix, per-class); full=True returns the complete result object.

ParametersJSON Schema
NameRequiredDescriptionDefault
fullNoReturn the backend's complete result object (all fields).
execution_idYesThe completed run to fetch metrics for.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false, so the safety profile is covered. The description adds a genuinely non-obvious behavioral trait beyond that: the default response is trimmed, and the parameters state exactly which fields the trimmed view contains and when to opt into the full object.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences, front-loaded with the core purpose, followed immediately by the default-vs-full behavior. Every clause earns its place with no repetition of the name or title.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return structure need not be explained, and annotations cover the safety profile. The description supplies the one thing structured fields cannot: the trimmed-by-default behavior and the field list of the default view, making it sufficient for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description goes beyond the schema by explaining what the default (full=false) actually yields — the trimmed field list — which the schema does not document. It thus characterizes the negative case of a boolean flag rather than merely repeating its description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names the specific resource (metrics) and scopes it to a completed run, and even enumerates the default payload fields (accuracy, kappa, ITR, confusion matrix, per-class). It is clear what the tool returns, though it does not explicitly contrast itself with the nearby execution.get or execution.artifacts siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implicitly guides the agent on the default vs. detailed retrieval mode ("Trimmed by default... full=True returns the complete result object"), which is useful. However, it gives no explicit when-to-use guidance relative to alternatives such as execution.get, nor any prerequisites (e.g., that the run must have finished).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

execution.runExecution.RunAInspect

Start a pipeline run (NON-BLOCKING). Returns executionId — poll with execution.get() until status is completed/failed, then execution.results().

ParametersJSON Schema
NameRequiredDescriptionDefault
nameNoDisplay name for the run (shown in the Runs list).
layoutNoOptional canvas positions {nodes: {id: {x, y}}}; a deterministic grid is synthesized when omitted.
subjectNoOptional dataset subject code (e.g. "S01") recorded with the run.
descriptionNoOptional longer description of the experiment.
train_graphYesPipeline graph {nodes: [{id, type, config}], connections: [{from, to}]} as built by catalog.template/pipeline.validate.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Adds real behavioral context beyond annotations: the call is NON-BLOCKING and returns an executionId for polling. Annotations already cover the safety profile (not read-only, not idempotent, not destructive). It doesn't mention auth/permission requirements, which would be the remaining gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences, front-loaded with the core action and the non-blocking constraint, followed by the exact follow-up sequence. Every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need no elaboration, yet the description still surfaces the critical executionId and the async workflow. Combined with 100% schema coverage and annotations, nothing an agent needs to invoke this correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all five parameters including train_graph, layout, and subject. The description adds no parameter-level meaning beyond that, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Start a pipeline run') plus a key behavioral qualifier (NON-BLOCKING). It clearly distinguishes itself from execution.get/results by describing the start-vs-poll relationship, but does not differentiate from the sibling experiment.run, which also appears to initiate execution.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit next-step guidance: poll execution.get() until terminal status, then call execution.results(). This is strong forward-routing. It does not, however, state when to prefer this over experiment.run or any preconditions (e.g., a validated graph).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

experiment.getExperiment.GetA
Read-onlyIdempotent
Inspect

Experiment snapshot: status (running/completed/failed), per-run rows ({name, executionId, status, error?, metrics?}) and, once finished, aggregates {metric: {mean, std, best: {name, value}}} over completed runs only (std = population; None below 2 values).

ParametersJSON Schema
NameRequiredDescriptionDefault
experiment_idYesThe experiment to inspect (from experiment.run).

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover safety (readOnly, idempotent, non-destructive), so the bar is lower; the description adds real behavioral detail beyond them: aggregates are computed over completed runs only, std is population standard deviation, and the value is None below 2 samples. These semantics are not derivable from the annotations or schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Compact and front-loaded with the returned content; the parenthetical field lists and formula caveats are dense but each clause carries information. It is telegraphic rather than wasteful, though the nesting of braces makes it slightly harder to scan.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so explaining return values is not strictly required, yet the description does so coherently and covers the finished-vs-running distinction. Combined with full schema coverage and annotations, an agent has enough to call it correctly; only the absence of usage routing is a gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Only one parameter exists and schema description coverage is 100%, with the schema already noting it comes from experiment.run. The description adds nothing about the experiment_id beyond what the schema states, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states the specific resource (an experiment snapshot) and enumerates what it contains: status, per-run rows, and aggregates. It is clearly a read/inspect tool, distinguishable from experiment.run and execution.get, though it never uses an explicit verb like 'retrieve' or 'inspect' and reads more like a return-shape spec.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit when-to-use guidance and no mention of alternatives such as execution.get or execution.results. The only implicit cue is 'once finished', hinting at polling, but the agent is left to infer the calling context entirely.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

experiment.runExperiment.RunAInspect

Run 1-25 pipelines as ONE paced experiment (NON-BLOCKING). Returns an experimentId immediately; a background thread submits at most 2 runs at a time (min(max_concurrent, 2)), retries queue-full up to 3 times per run, and polls each execution to completion. Poll experiment.get() for per-run status and, once finished, aggregated metrics.

ParametersJSON Schema
NameRequiredDescriptionDefault
runsYes1-25 entries, each {name: str, train_graph: {nodes, connections}} (same graph shape as execution.run's train_graph).
max_concurrentNoParallel submissions cap, clamped to 1-2 (default 2).

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only mark this as non-read-only, non-idempotent and non-destructive; the description adds substantial behavior beyond that — non-blocking return of an experimentId, a background thread, concurrency clamped via min(max_concurrent, 2), up to 3 queue-full retries per run, and polling each execution to completion. This is exactly the async/retry context an agent needs.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The non-blocking, batch-scoped nature is front-loaded in the first clause, and every subsequent sentence carries operational detail with no filler. The middle sentence is dense and slightly run-on, but nothing is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be described, and the description still explains the async model and points to experiment.get() for per-run status and aggregated metrics. Nothing needed to invoke or follow up on the tool is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so both parameters are already documented, but the description reinforces the meaningful semantics: the 1-25 bound on the runs array and the min(max_concurrent, 2) clamping rule. It adds marginal clarification over the schema rather than new syntax.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Run) and resource (1-25 pipelines as one paced experiment), plus the async contract. It distinguishes itself from execution.run by scoping to batch runs that are submitted as a single experiment, which an agent can tell apart without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies the use case (multi-pipeline batch, paced) and directs the agent to poll experiment.get() afterward, but never states when to prefer this over execution.run for a single run, nor any exclusion conditions. Usage is implied rather than spelled out.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

pipeline.exportPipeline.ExportB
Read-onlyIdempotent
Inspect

Export the pipeline as a standalone runnable Python bundle (zip saved locally).

ParametersJSON Schema
NameRequiredDescriptionDefault
nameNoOptional name recorded inside the bundle.
train_graphYesPipeline graph {nodes, connections} to export.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, destructiveHint=false, idempotentHint=true, and openWorldHint=true, covering the safety profile. The description adds that the output is a zip saved locally, which is useful behavioral context about the result format and side effect. It does not mention permissions or rate limits, so a 3 is appropriate.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, front-loaded sentence that states the action and result with zero wasted words. It is appropriately sized for a simple export tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be described. The schema fully documents parameters, and annotations cover safety. The description could be improved by adding usage context, but for a straightforward export tool it is largely complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema fully documents both parameters (name and train_graph). The description adds no additional meaning about parameter formats or constraints beyond what the schema provides. Baseline 3 applies when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb (Export), resource (the pipeline), and output format (standalone runnable Python bundle, zip saved locally). It implicitly distinguishes from siblings like pipeline.validate by the export action, but does not explicitly name or contrast with them. A 4 reflects clear purpose without sibling differentiation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool versus alternatives, nor any prerequisites or exclusions. The description only states what it does, leaving the agent to infer that it is used to obtain a bundle. This matches the 'no guidance' level.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

pipeline.validatePipeline.ValidateA
Read-onlyIdempotent
Inspect

Validate a pipeline graph before running. ExecGraphSnapshot: {nodes: [{id, type, config}], connections: [{from, to}]}. Build it from catalog.template(id).train or from scratch using catalog.nodes().

ParametersJSON Schema
NameRequiredDescriptionDefault
train_graphYesThe graph to validate ({nodes, connections}).

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly, idempotent, non-destructive behavior, so the safety profile is covered. The description adds the pre-execution intent and the graph shape, but says nothing about what a failed validation returns or how validation is scoped (e.g., config-level checks).

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core purpose, then supporting input documentation. Three compact lines with no filler, though the 'ExecGraphSnapshot:' shorthand is a bit cryptic for an agent that hasn't seen the type elsewhere.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values need not be explained. Given annotations covering safety and the description covering purpose, input structure, and construction path, an agent has enough to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% but the schema is loose (additionalProperties: true, no field definitions). The description compensates by specifying the actual structure - nodes with {id, type, config} and connections with {from, to} - which the schema does not provide.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Specific verb+resource combination: validate a pipeline graph, with the added scope of 'before running'. It clearly separates from generic validation, though it doesn't explicitly name pipeline.validate_node to disambiguate graph-level vs node-level validation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'before running' establishes clear context for when to call it, and it suggests the input source (catalog.template(id).train or catalog.nodes()). No explicit when-not or alternative routing to validate_node, so it stops short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

pipeline.validate_nodePipeline.Validate NodeB
Read-onlyIdempotent
Inspect

Validate one node's config object against its schema (catalog.node_schema).

ParametersJSON Schema
NameRequiredDescriptionDefault
configYesThe node's config object to check.
node_typeYesNode type id from catalog.nodes (e.g. "csp", "nimbus_lda").

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, openWorldHint and destructiveHint=false, so the safety profile is fully covered. The description adds one useful behavioral fact — validation is performed against the catalog.node_schema definition rather than an embedded schema — but says nothing about failure modes or how errors surface.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler; the verb and scope arrive first and the schema source is parenthesized. Slightly terse for a validation tool, but nothing is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained, and annotations cover safety. What is missing is the relationship to sibling tools — when to validate a single node versus the whole pipeline, and whether this is a prerequisite for execution.run — leaving the definition minimally viable.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% with two required parameters, both documented in-schema (config object, node_type id with examples). The description adds no format or semantic detail beyond what the schema already states, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a specific verb (validate), a precise resource (one node's config object), and the reference schema (catalog.node_schema). The word "one" distinguishes it from the sibling pipeline.validate, which presumably covers a whole pipeline, though that contrast is implied rather than stated.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit when-to-use guidance, no prerequisites, and no named alternative. An agent must infer that this is a pre-flight check before execution.run, and must infer the split of responsibility with pipeline.validate and catalog.node_schema on its own.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

project.createProject.CreateBInspect

Create a project (container for one pipeline document). Returns projectId.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYesProject display name (e.g. "MI CSP-LDA sweep").
descriptionNoOptional longer description shown in the studio.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, idempotentHint=false, destructiveHint=false, and openWorldHint=true, so the safety profile is covered structurally. The description adds useful chaining context by noting it returns projectId, but says nothing about required permissions, whether name collisions are rejected, or side effects on servers. Modest added value against an already-covered annotation set.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two terse sentences with zero waste; the core action leads and the return value follows. Every clause carries information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a two-parameter create tool with full schema coverage, annotations, and an output schema, the definition covers the essentials. The projectId mention is slightly redundant with the output schema, and the absence of any prerequisite or duplicate-name note is the only real gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% for both parameters, so the schema already documents name (with an example) and description. The description contributes only the 'container for one pipeline document' framing and adds no syntax, format, or constraint detail beyond the schema. Baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Create a project') and goes further by defining what a project is ('container for one pipeline document') plus a return value. It clearly differs from sibling verbs like project.list, project.load, and project.save, though it doesn't name them explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance, no prerequisites, and no mention of the alternative project operations (list/load/save) that an agent might confuse this with. The create action is largely self-evident, which keeps this above a 1, but nothing is offered about sequencing or preconditions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

project.listProject.ListA
Read-onlyIdempotent
Inspect

List projects owned by the current principal (agent work included).

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly, idempotent, openWorld and non-destructive, so the safety profile is covered. The description adds the useful scoping constraint that results are limited to projects 'owned by the current principal (agent work included)', but says nothing about pagination, limits, or ordering.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single tight sentence with the list action and its scope front-loaded; every clause carries information and nothing is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists and annotations cover safety, so return values and risk profile are already documented. The description supplies the ownership scope, but leaves pagination/result-size behavior unaddressed, keeping it just short of fully complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so the baseline is 4. There is nothing in the schema for the description to clarify further.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (list) and resource (projects) plus an ownership scope, and the verb alone separates it from project.create/load/save. It never names or contrasts those siblings explicitly, so it stops short of the top tier.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no mention of alternatives like project.load or project.create, and no prerequisites. Usage is only implied by the verb 'list'.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

project.loadProject.LoadA
Read-onlyIdempotent
Inspect

Load a project's saved pipeline (train graph + meta) for editing/re-running.

ParametersJSON Schema
NameRequiredDescriptionDefault
project_idYesProject whose pipeline document to load.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true and destructiveHint=false, so the safety profile is covered structurally. The description adds only that the loaded document is the saved pipeline for editing/re-running; it says nothing about failure modes (unsaved project, missing pipeline), permissions, or whether the load is cached.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One sentence with zero filler; the resource and its payload are front-loaded before the purpose clause.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be described, and annotations cover the safety profile. For a single-parameter read tool this is nearly sufficient; only failure conditions and the relationship to project.save/pipeline.validate are unaddressed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% for the single project_id parameter, so the schema already carries the semantics. The description adds no format, id-shape, or lookup guidance beyond the schema, so baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Load) and resource (a project's saved pipeline) and clarifies the payload as train graph + meta with the intent (editing/re-running). It is clear on its own, but it does not distinguish itself from adjacent siblings such as project.save or pipeline.export.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'for editing/re-running' implies the context in which to call it, but there is no explicit when-to-use versus alternatives (e.g., project.list to discover ids, pipeline.validate to check a loaded pipeline) and no stated prerequisites such as whether the project must already have been saved.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

project.saveProject.SaveAInspect

Save a pipeline graph into a project (visible on the studio canvas). Handles revision conflicts automatically (one retry).

ParametersJSON Schema
NameRequiredDescriptionDefault
nameNoOptional pipeline name stored on the document.
subjectNoOptional dataset subject code (e.g. "S01") for this pipeline.
project_idYesTarget project (from project.create/project.list).
train_graphYesPipeline graph {nodes, connections} to persist.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare this is a non-read-only, non-idempotent, open-world write. Beyond that the description adds a genuinely useful behavioral fact the annotations do not convey: revision conflicts are handled automatically with one retry, so the agent knows a failed save may self-heal once. It still doesn't say whether an existing graph is overwritten or a new revision created, which matters for a non-idempotent mutation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences with no filler; the core action is front-loaded and the retry behavior follows immediately as supporting detail.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Annotations cover the safety profile and an output schema exists, so return values need not be explained. The description supplies the retry semantics and the canvas-visibility outcome, leaving only overwrite-vs-new-revision behavior unaddressed for a mutation that is explicitly non-idempotent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all four parameters (name, subject, project_id, train_graph) are already documented in the schema; the description adds no syntax, format, or constraint detail beyond it. Baseline 3 is appropriate when the schema carries the full parameter burden.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Save a pipeline graph into a project') plus the observable effect ('visible on the studio canvas'), which separates it from read-side siblings like project.load. It does not explicitly distinguish itself from project.create or pipeline.export, so it stops short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Implicit context is present – the parenthetical about canvas visibility hints this is the persistence step after building a graph – and the schema points at project.create/project.list for obtaining a project_id. However, there is no explicit statement of when to prefer this over project.load or pipeline.export, so usage remains inferred rather than stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

stream.startStream.StartAInspect

Connect an EEG device and START a live streaming session on the user's head. Requires confirm=True; call device.test first. Track with stream.status(). Idle watchdog: if no stream.status()/stream.telemetry() poll happens for idle_timeout_sec (default 900), the session is stopped and the device disconnected automatically — an abandoned stream never keeps running on the user's head. Any poll resets the timer; idle_timeout_sec=0 disables the watchdog.

ParametersJSON Schema
NameRequiredDescriptionDefault
portNoSerial/COM port for wired devices.
confirmNoMUST be true to start — the explicit user go-ahead for a live session on their head; anything else is refused with zero requests.
ip_portNoPort for network devices.
source_idNoLSL source id.
chunk_sizeNoSamples per streamed chunk (default 125).
ip_addressNoDevice IP for network/wifi devices.
n_channelsNoChannel count to open (default 8).
session_idNoOptional existing session to resume/reuse.
device_typeYesDevice id from device.list (e.g. "brainbit").
mac_addressNoBluetooth MAC (BT devices).
stream_nameNoLSL stream name (LSL devices).
serial_numberNoDevice serial (some BLE stacks).
connection_typeNoDevice-specific selector (e.g. serial vs wifi).
idle_timeout_secNoWatchdog: stop+disconnect after this many seconds without a status poll (default 900; 0 disables).

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the annotations (readOnlyHint=false, destructiveHint=false, openWorldHint=true), the description discloses a rich behavioral trait: the idle watchdog that auto-stops the stream and disconnects the device after idle_timeout_sec of no polling, that any poll resets the timer, and that 0 disables it. It also explains the safety rationale (a live stream on the user's head) and the confirm gate, which the annotations alone do not convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The action and confirm gate are front-loaded in the first two sentences, and the watchdog explanation follows logically. It is somewhat longer than strictly necessary given the schema already documents idle_timeout_sec, but every sentence carries usable information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values need not be explained, and the annotations cover the safety profile. Combined with prerequisites, the tracking workflow, and the watchdog behavior, an agent has everything needed to invoke this 14-parameter tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all 14 parameters, including confirm and idle_timeout_sec. The description reinforces those two with rationale but adds no new syntax or format information, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource ('Connect an EEG device and START a live streaming session') and is clearly distinguishable from stream.stop and stream.status by name. An agent knows exactly what this tool initiates without reading the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives explicit prerequisites ('Requires confirm=True; call device.test first') and names the tracking alternative ('Track with stream.status()'), which is strong routing guidance. It stops short of an explicit when-not (e.g. what to do if a session is already running, or when to prefer stream.stop), so it is not a full 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

stream.statusStream.StatusB
Read-onlyIdempotent
Inspect

Live snapshot of a streaming session (running, deviceConnected). Polling this also feeds the idle watchdog: each call resets the session's idle timer (see stream.start's idle_timeout_sec).

ParametersJSON Schema
NameRequiredDescriptionDefault
session_idYesThe streaming session to inspect.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.2/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=true and idempotentHint=true, but the description says each call resets the session's idle timer. That is a state mutation and a cumulative effect across repeated calls, directly contradicting both annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with the core snapshot behavior and followed by the important side effect. Every sentence carries useful information with no waste.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, the description needn't explain return values. It covers the main behavior and a key side effect, though it omits sibling differentiation and the annotation contradiction means the overall definition is not fully coherent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the single session_id parameter is fully documented in the schema. The description adds no parameter meaning beyond that, making 3 the baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific resource (streaming session) and what is returned (live snapshot with running, deviceConnected). It does not explicitly differentiate from sibling stream.telemetry, so it is clear but lacks sibling routing.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Implies polling as the usage pattern and notes the idle-watchdog side effect, but gives no explicit when-to-use versus stream.telemetry or stream.start, and no when-not-to-use guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

stream.stopStream.StopA
Destructive
Inspect

Stop a streaming session and disconnect the device (always safe to call). Also removes the session from the idle watchdog so it cannot fire after an explicit stop.

ParametersJSON Schema
NameRequiredDescriptionDefault
session_idYesThe streaming session to stop.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, idempotentHint=false, and openWorldHint=true. The description adds non-obvious behavior beyond them: that the session is removed from the idle watchdog so it cannot fire after an explicit stop. 'Always safe to call' is a mild tension with idempotentHint=false but reads as 'will not error', not a contradiction.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, front-loaded with the core action and immediately followed by the safety note and side effect. No filler, though the watchdog sentence could arguably be tightened.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values need not be explained, and annotations cover the destructive/read-only profile. For a single-parameter stop tool the description supplies enough, with only the absence of alternative-routing guidance as a gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% with a single well-documented session_id, so the schema carries parameter meaning. The description adds nothing about the parameter beyond what the schema states, which is the expected baseline here.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (stop) and resource (streaming session) plus a secondary effect (disconnect the device). It is clearly distinguishable from siblings stream.start, stream.status, and stream.telemetry without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The parenthetical 'always safe to call' implies this can be invoked freely, but there is no explicit statement of when to use it versus stream.start or stream.status, nor any preconditions or when-not guidance. Usage is only implied by the obvious start/stop pairing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

stream.telemetryStream.TelemetryA
Read-onlyIdempotent
Inspect

Live snapshot of a streaming session: latest prediction + recent window, signal quality (meanChannelQuality, snrDb, artifactProbability), indicators, running stats. Poll this while a session runs. Live telemetry requires a DEPLOYED model session (hub deploy / playback with a classifier); modelless hardware streams have no telemetry — use stream.status for those. Expect low confidence during filter/ASR warm-up (first seconds); quality < 0.5 or high artifactProbability means the signal is poor. 404 => session not active in this backend. Each poll also feeds the idle watchdog (see stream.start's idle_timeout_sec), keeping an actively watched session alive.

ParametersJSON Schema
NameRequiredDescriptionDefault
windowNoHow many recent predictions/chunks to include (default 50).
session_idYesThe streaming session to read telemetry for.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare the safe-read profile; the description adds substantial context beyond them: 404 means the session is not active, low confidence is expected during filter/ASR warm-up, quality<0.5 or high artifactProbability signals a poor signal, and each poll feeds the idle watchdog keeping a watched session alive. These are non-obvious operational behaviors an agent could not infer from structured fields.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose is front-loaded in the first clause, and every subsequent sentence carries distinct operational value (routing, thresholds, error semantics, watchdog). It is a dense single block rather than being visibly chunked, which slightly hurts scannability but there is no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Output schema exists, so return values need not be explained, and annotations cover safety. The description still supplies the operationally critical context (prerequisites, error meaning, warm-up caveats, watchdog side effect), leaving nothing an agent needs in order to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and both parameters are already documented in the schema, so the baseline is 3. The description clarifies the output semantics ('recent window', 'latest prediction') but adds no syntax or format detail about the window parameter beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Live snapshot of a streaming session') and enumerates what the snapshot contains: latest prediction, recent window, signal quality metrics, indicators, running stats. It explicitly distinguishes itself from the sibling stream.status by naming the condition that routes to each.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit when-to-use ('Poll this while a session runs'), an explicit when-not ('modelless hardware streams have no telemetry — use stream.status for those'), and a prerequisite ('requires a DEPLOYED model session'). The alternative sibling is named with the selecting condition.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 5 tool updates
    • Addedcalibration.pause
    • Addedcalibration.resume
    • Addedcalibration.start
    • Addedcalibration.status
    • Addedcalibration.train
  2. 64 tool updates
    • Addedaccount.whoami
    • Removedcancel_execution
    • Addedcatalog.datasets
    • Addedcatalog.leaderboard
    • Addedcatalog.node_schema
    • Addedcatalog.nodes
    • Addedcatalog.template
    • Addedcatalog.templates
    • Removedcreate_project
    • Addeddata.inspect_dataset
    • Addeddata.inspect_file
    • Addeddata.upload
    • Addeddevice.list
    • Addeddevice.test
    • Removeddownload_artifact
    • Addedexecution.artifacts
    • Addedexecution.cancel
    • Addedexecution.download_artifact
    • Addedexecution.get
    • Addedexecution.list
    • Addedexecution.results
    • Addedexecution.run
    • Addedexperiment.get
    • Addedexperiment.run
    • Removedexport_python
    • Removedget_execution
    • Removedget_experiment
    • Removedget_leaderboard
    • Removedget_live_session
    • Removedget_node_schema
    • Removedget_results
    • Removedget_template
    • Removedinspect_dataset
    • Removedinspect_file
    • Removedlist_artifacts
    • Removedlist_datasets
    • Removedlist_devices
    • Removedlist_executions
    • Removedlist_nodes
    • Removedlist_projects
    • Removedlist_templates
    • Removedload_pipeline
    • Addedpipeline.export
    • Addedpipeline.validate
    • Addedpipeline.validate_node
    • Addedproject.create
    • Addedproject.list
    • Addedproject.load
    • Addedproject.save
    • Removedrun_experiment
    • Removedrun_pipeline
    • Removedsave_pipeline
    • Removedstart_stream
    • Removedstop_stream
    • Removedstream_status
    • Addedstream.start
    • Addedstream.status
    • Addedstream.stop
    • Addedstream.telemetry
    • Removedtest_device
    • Removedupload_data
    • Removedvalidate_node_config
    • Removedvalidate_pipeline
    • Removedwhoami
  3. 32 tool updates
    • First observedcancel_execution
    • First observedcreate_project
    • First observeddownload_artifact
    • First observedexport_python
    • First observedget_execution
    • First observedget_experiment
    • First observedget_leaderboard
    • First observedget_live_session
    • First observedget_node_schema
    • First observedget_results
    • First observedget_template
    • First observedinspect_dataset
    • First observedinspect_file
    • First observedlist_artifacts
    • First observedlist_datasets
    • First observedlist_devices
    • First observedlist_executions
    • First observedlist_nodes
    • First observedlist_projects
    • First observedlist_templates
    • First observedload_pipeline
    • First observedrun_experiment
    • First observedrun_pipeline
    • First observedsave_pipeline
    • First observedstart_stream
    • First observedstop_stream
    • First observedstream_status
    • First observedtest_device
    • First observedupload_data
    • First observedvalidate_node_config
    • First observedvalidate_pipeline
    • First observedwhoami

Related MCP Connectors

Related MCP Servers

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.