Skip to main content
Glama

Preview what the anomaly gate would flag (beta)

preview_anomalies
Read-onlyIdempotent

Simulate anomaly detection on recorded robot logs to preview which signals and thresholds would flag events, enabling informed configuration before deploying an anomaly pipeline.

Instructions

Dry-run the on-robot anomaly screen over a recorded log without calling any decision model or writing artifacts. Learns a rolling baseline the way the src.pipeline.gates.anomaly gate does, then reports every window the screen would flag (mean shift, extreme sample, topic dropout) with the signal, value and z-score, flag counts per signal, and plain-language advice (signals that drift by design, warm-up longer than the log). Inspect topics first and pass signals as rates and errors (accelerations, angular rates, currents), never positions or orientations. Use it to choose signals and thresholds before saving an anomaly pipeline. It models screen mode with the decision model confirming every flag, every: N seconds cadences, and a gate that sees every fire (list the anomaly gate first).

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
argsNo
pathYes
topicsNo
signalsNo
max_signalsNo
z_thresholdNo
cadence_topicNo
warmup_minutesNo
window_secondsYes
cadence_secondsNo
dropout_secondsNo
baseline_window_minutesNo

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault

No arguments

Schema Changelog

Changes observed during successful MCP inspections.

  1. Addedv2.3.1

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already provide readOnlyHint=true, idempotentHint=true, destructiveHint=false, so the description correctly aligns with these (no contradiction). It adds valuable behavioral context beyond annotations: it mentions learning a rolling baseline, modeling screen mode, every: N cadences, and gate ordering (list anomaly gate first). It also notes what it does NOT do (calling decision model, writing artifacts), which helps set expectations. Minor deduction for not explicitly stating output schema details, but the output schema is provided separately.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but efficiently packed, with multiple clauses per sentence. It front-loads the purpose and main output, then adds usage guidance. However, it is a long paragraph that might be better split into sections (e.g., purpose, usage, behavior) for readability. It's not overly verbose relative to the complexity, so it earns a 4.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (12 params, output schema rich, many siblings), the description covers the essential decision points: when to use (before saving an anomaly pipeline), what it returns (flag counts, z-scores, advice), what inputs to avoid (positions/orientations), and how it models the anomaly gate. The output schema likely lists the return structure, so return values aren't needed in the description. It is complete for an agent to decide and call correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema provides default values and types but zero descriptions for parameters (0% coverage). The description compensates by explaining the meaning of key parameters: 'signals' must be rates/errors, not positions/orientations; 'window_seconds' is part of the required params; mentions cadence and warm-up concepts, which map to cadence_seconds and warmup_minutes. It doesn't explicitly define every parameter (like max_signals, z_threshold), but the context is enough for an agent to infer their roles from the defaults and overall description. Given the high parameter count (12) and low schema coverage, this is a strong effort.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool does a dry-run of the anomaly gate over a recorded log, specifying the verb (preview/dry-run), resource (anomaly screen/gate), and scope (on a log, no decision model or artifacts). It distinguishes from siblings like preview_pipeline by emphasizing on-robot real-time screen behavior and the anomaly gate specifically, and it provides a concrete list of output types (flags, z-scores, advice).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly says 'Inspect topics first and pass signals as rates and errors... never positions or orientations' and 'Use it to choose signals and thresholds before saving an anomaly pipeline'. It also hints at when not to use it implicitly by requiring a recorded log and the specific screening mode. This gives clear prerequisites and a purpose that distinguishes it from other preview/run tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.