Skip to main content
Glama

evaluate_cross_session

Evaluate decoder stability across sessions: fit on a training session, then score unchanged on test sessions to verify it survives to later days.

Instructions

Fit a decoder on one session and score it, unchanged, on other sessions: the question benchmarks such as FALCON ask (does a decoder survive to a later day?).

All sessions must be open and share the same units/channels. `mask` names a boolean
behaviour signal restricting which bins are scored (FALCON files carry `eval_mask`; pass
null when the sessions have none). Also reports same-session held-out R² for comparison.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
kindNoridge
maskNoeval_mask
bin_sNo
targetNo
test_session_idsYes
train_session_idYes

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.5.1

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It discloses that the decoder is 'unchanged' across sessions, that mask restricts which bins are scored, that FALCON files carry eval_mask, and that same-session held-out R² is also reported. This is meaningful behavioral context beyond what the schema shows. It doesn't mention side effects (e.g., whether it writes anything), but the description implies a read/analysis operation, and the disclosed behavior is substantial.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and front-loaded: the core operation is stated in the first sentence, followed by the benchmark context, prerequisites, and parameter clarifications. Every sentence adds value, and there is no repetition of schema fields or filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (6 params, 0% schema coverage, no output schema, no annotations), the description covers the essential usage context: what it does, prerequisites, mask semantics, and the extra same-session R² output. It doesn't explain the target parameter or the return format, but the description is strong enough that an agent could likely call it correctly for the primary FALCON use case. A 4 is appropriate; a 5 would require explicit mention of target and output structure.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It explains the mask parameter's semantics ('boolean behaviour signal restricting which bins are scored', 'pass null when sessions have none') and clarifies the relationship between train_session_id and test_session_ids. It also mentions the kind parameter implicitly via 'decoder' and the bin_s parameter via 'bins'. However, it doesn't explain target or bin_s explicitly, so it doesn't fully compensate for the 0% coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('fit a decoder on one session and score it, unchanged, on other sessions') and names the exact resource and operation. It also explicitly references the FALCON benchmark question, which distinguishes it from sibling tools like fit_decoder and decode_window. This is a clear, specific purpose statement.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear context for when to use this tool: when benchmarking cross-session decoder survival, as FALCON does. It also specifies prerequisites ('All sessions must be open and share the same units/channels') and explains the mask parameter's role. However, it doesn't explicitly name alternative tools or state when NOT to use it, so it falls just short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.