Skip to main content
Glama

Bagel by Extelligence lets you ask questions about robotics, drone, and IoT data in plain English. Every calculation over your message data is DuckDB SQL, not model guesswork, and Bagel shows you the query so you can audit it.

Is my IMU sensor overheating?

Bagel also has an intelligent edge data reduction pipeline: describe an event and Bagel runs the detection on the robot, keeping the windows that matter and dropping the rest. An MCP server puts all of it in your LLM's hands: Claude Code, Gemini, Cursor, or a fully local model.

Bagel was the first MCP server to ship a real analysis toolkit for robotics data, and it keeps the LLM where it belongs: in front of your logs, never in your robot's control loop.

🥯 Key Features

  • Ask in plain language: No deep domain expertise needed.

  • Transparent calculations: Deterministic SQL queries. No black-box LLM math.

  • Natural-language pipelines: "Keep 10s around every hard brake, drop the rest": one sentence becomes an auditable pipeline: previewed before a byte is written, then run once, across a fleet, or standing at the edge.

  • Broad LLM support: Claude Code, Gemini, Cursor, Codex, and more.

  • Dockerized environments: No local dependencies required.

  • Extensible capabilities: Bagel can learn new tricks.

  • Wide format coverage: Missing your data format? Open a ticket.

🥯 Try it in 60 seconds

No MCP client, no LLM, no config: run the same deterministic checks against a bundled sample log and get a robot-health report card straight to your terminal.

docker run -it --rm ghcr.io/extelligence-ai/bagel/px4:latest demo
sample.ulg - 41.5s, 2018 messages, 77 topics

Power      ⚠️  min 21.07V, largest drop 2.37V at ~t=+4.8s, end 23.45V (battery_status_0)
IMU        ✅  accel_z stddev 1.6x the log baseline at ~t=+36.8s (sensor_combined_0)
GPS        —  skipped: no GPS topic
Data gaps  ✅  no gap > 1.05x median interval (checked battery_status_0, sensor_combined_0)
...

The ROS2 images (ros2-kilted, ros2-jazzy, ros2-iron, ros2-humble) run demo the same way, against a lighter bundled sample (px4 is the one that ships with a flight log rich enough to show every check). Point it at your own log with demo /path/to/log (mount it with -v first), or keep reading for the full MCP setup below.

Related MCP server: Robotics MCP Server

⚡️ Quickstart

TIP

Already have Claude Code? Just paste the link to this repo and tell Claude what environment you want:

Set up https://github.com/Extelligence-ai/bagel for ROS2 Kilted.

Claude will clone the repo, start Docker, and wire up the MCP connection for you.

📋 Prerequisites

Install Docker Desktop and Claude Code (or another MCP-enabled LLM).

NOTE

arm64 hosts (Apple Silicon, Raspberry Pi, Jetson, Graviton): ros2-kilted — the default service, and the one server.json pins — ships as a multi-arch image, so Docker pulls a native arm64 build. No extra setup.

The other services are published for amd64 only. Docker Desktop emulates them automatically, so they run on Apple Silicon (slower, but they work). On arm64 Linux with plain Docker Engine there is no emulation by default and they fail immediately with exec format error — install QEMU/binfmt first:

docker run --privileged --rm tonistiigi/binfmt --install amd64

1. Clone and start Bagel

git clone https://github.com/Extelligence-ai/bagel.git && cd bagel
docker compose run --service-ports ros2-kilted
TIP

Port 8000 already in use? SetMCP_SERVER_PORT to something else, for example MCP_SERVER_PORT=8100 docker compose run --service-ports ros2-kilted, and use that port in step 2.

Pick the service that matches your environment:

Service

Use case

ros2-kilted

ROS2 Kilted (latest)

ros2-jazzy

ROS2 Jazzy

ros2-jazzy-jev

ROS2 Jazzy + on-robot decision model (GPU, beta)

ros2-iron

ROS2 Iron

ros2-humble

ROS2 Humble

ros1-noetic

ROS1 Noetic

ros1-noetic-cv

ROS1 Noetic + CV

px4

PX4 flight logs

ardupilot

ArduPilot flight logs

betaflight

Betaflight flight logs

iot

IoT / MQTT (live)

The -jev image (beta) adds PyTorch for running a decision model on the robot (backend: local in the anomaly gate). Build any other service the same way with --build-arg JEV_MODE=true. CPU-only robots don't need it: the hosted Jev backend works in every image.

TIP

To give Bagel access to your local files, editcompose.yaml before starting Docker: uncomment and update the volumes section under your chosen service.

Wait for this output:

INFO:     Uvicorn running on http://0.0.0.0:8000 (Press CTRL+C to quit)

2. Connect Claude Code

In a new terminal:

claude mcp add --transport sse bagel http://localhost:8000/sse
NOTE

The MCP endpoint is bound tolocalhost only (not exposed to the LAN) for security. To share it with other machines, drop the 127.0.0.1 prefix in compose.yaml and put an authenticated proxy in front: see SECURITY.md.

3. Prompt

claude

Summarize the metadata of the ROS2 bag "./data/sample/ros2/mcap".

That’s it: you’re chatting with your data.

🔒 Prefer fully offline?

Swap step 2 for a local model: your data and your LLM stay on the machine:

brew install ollama && ollama serve &                                  # or ollama.com
ollama pull qwen3:8b
uvx ollmcp --mcp-server-url http://localhost:8000/sse --model qwen3:8b

Model picks, expectations, and troubleshooting: Local LLMs guide.

Bagel works with any MCP-enabled LLM. Setup runbooks for tested alternatives:

Can’t find your LLM? Open a ticket.

🔌 Agent plugins (Claude Code and Codex)

Bagel ships an agent plugin: four skills that teach the agent when and how to drive the server (log triage, pipeline authoring, live sinks, visualization export) plus the MCP connection, wired automatically. The same plugin/ directory serves both Claude Code and OpenAI Codex.

/plugin marketplace add Extelligence-ai/bagel
/plugin install bagel@bagel

Codex and ChatGPT users: install bagel from the OpenAI Plugins Directory (one click), or clone the repo and add it as a plugin marketplace (the repo carries .agents/plugins/marketplace.json). Directory installs bundle the skills only, so also connect the server once in ~/.codex/config.toml:

[mcp_servers.bagel]
url = "http://localhost:8000/mcp"

Repo-marketplace and Claude Code installs wire this connection automatically.

Then start the container for your data format (see Quickstart): the plugin connects to http://localhost:8000/mcp by default. Any other MCP client can discover the same workflows server-side via the list_agent_capabilities tool.

Keep what matters, drop the rest

A robot records more data than you can afford to move. Bagel turns a question into a detector, runs it where the data is recorded, and ships only the windows around real events.

Here it is in one conversation:

Don't know the event in advance? The anomaly gate (beta) learns what normal looks like on the robot, asks Jev to name whatever isn't, and keeps only those slices, each with a JSON label, for any bucket: S3, GCS, Azure, MinIO or R2.

The session above: a 20-minute (1,200 s) recording and the prompt "keep 10 seconds before and after every deceleration harder than −10 m/s²". The preview detects 7 events, merges them into 4 windows, and keeps 92 s of the 1,200 (7.6%); the run writes a 2.1 GB bag down to 161 MB. These figures are illustrative demo output, not a measured benchmark: the ratio is event-window duration over total duration, so it depends entirely on your workload.

✅ Supported Data Formats

Industry

Formats

Robotics

ROS1, ROS2, MCAP (any profile), Copper (via MCAP export), ROS text logs (~/.ros/log)

Robot learning

Gantry Bench evidence bundles — a dataset verdict's working (per-clip signal checks, robot-test ladder, findings) as queryable tables

Drones

PX4, ArduPilot, Betaflight

Automotive

ASAM MDF4 (.mf4), CAN captures (.blf/.asc + DBC) · beta

IoT

MQTT (live, Sparkplug B), PostgreSQL / TimescaleDB, InfluxDB 3

Hardware state

WaffleForm snapshots (.waffleform.yaml), auto-detected via waffle-iron · beta

🆚 Bagel vs. the Tools You Already Use

You already have ros2 *, PlotJuggler, and grep. Bagel doesn't replace them: it answers the questions they make you work for, then hands off to them:

You do this today

Ask Bagel instead

ros2 bag info for metadata

"Summarize this bag": same prompt works on PX4, ArduPilot, MCAP, MQTT, Postgres

ros2 topic echo /imu and eyeball raw values

"What's the peak z-deceleration in /imu? Running average over 5 s?" · real SQL underneath: peaks, running averages, percentiles, cross-topic correlations

Scrub PlotJuggler timelines hunting for the event

"Find every deceleration under −10 m/s² and cut ±30 s snippets": then open the result in PlotJuggler with a pre-framed layout

rqt_console, or grep ~/.ros/log

"Read the ERRORs from ~/.ros/log and tell me what went wrong": tracebacks included, no bag needed

Echo two topics in two terminals, correlate in a spreadsheet

"What's the correlation between current and voltage?": topics live in one SQL relation, so joins and corr() are one question

ros2 bag record -a and babysit the disk

A standing edge pipeline: record continuously, keep only event windows, drop the rest

A bash loop over 200 bags

"Run this pipeline on every bag in the folder": one pipeline, whole fleet, with a combined report

scp/aws s3 sync scripts to ship data off the robot

Upload to S3, GCS, or Azure as a pipeline step, checksum-skipping files already there

A different viewer per format: FlightPlot for PX4, MAVExplorer for ArduPilot, Blackbox Explorer for Betaflight

The same conversation for all of them, and ROS, MCAP, MQTT, Postgres, InfluxDB

Write a one-off pandas script per question

Ask the question; Bagel writes and runs the query

One sentence of plain language, one answer, instead of a pipeline of commands and a script you'll delete tomorrow.

💬 What Can I Prompt?

You can ask Bagel almost anything. For example:

What’s the correlation between current and voltage in the /spot/status/battery_states topic?

I think the robot hit a pothole. Can you check for sudden deceleration on the z-axis to confirm?

Every time the drone decelerates harder than -10 m/s², keep 10 seconds before and after. Drop everything else.

Did anything change on this robot since last week?

Time to put Bagel to the test: can it catch a drone doing barrel rolls? Spoiler: 🎉 It totally can.

💡 How Bagel Works

When you ask a question, Bagel analyzes your data source’s metadata and topics to build a high-level understanding.

Based on your prompt, if further inspection is needed, Bagel identifies the most relevant topics and interprets their meaning and structure. Bagel then writes the relevant topic messages to an Apache Arrow file and uses DuckDB to generate and execute queries against it.

This process is repeated as needed, running new queries until Bagel finds the best answer to your question.

LLMs excel at language but struggle with math. Bagel overcomes this by generating deterministic DuckDB SQL queries. These queries are displayed for you to audit, and you can guide Bagel to correct any errors.

🐶 Teach Bagel a New Trick

Bagel learns new capabilities through POML files: a structured set of instructions that describe a “trick,” such as computing latency statistics.

✍️ Create a .poml file

For example, let’s define ./src/agent/examples/woof.poml.

<poml>
    <task>
        Count the topics in the data source.
        If the count is odd, say "woof", else say "meow".
    </task>

    <output-format>
        Return the sound, the topic count, and a few cute emojis. Nothing else.
    </output-format>
</poml>

🗣️ Use the capability

Prompt Bagel:

Run the POML capability "./src/agent/examples/woof.poml" on the ROS2 bag "./data/sample/ros2/mcap".

Result:

meow 🐱 4 topics 🐱💤🎯

Teach it your own tricks (no rebuild)

Bagel discovers your own capabilities from ~/.bagel/capabilities/:

  • In conversation: do a workflow once, then say "save that as a capability called battery-triage" — Claude calls save_agent_capability and it's reusable in any future session.

  • As a file: drop a markdown file with your steps (or a POML file, if you want parameterized templates — see src/agent/compose/pipeline.poml for the house style) into ~/.bagel/capabilities/.

Either way it shows up in list_agent_capabilities as user/<name> and runs with run_poml_capability — from Claude Code, Claude Desktop, or any MCP client. Teams: keep the directory in your own git repo and sync it to every robot; it's just files. On Linux, run mkdir -p ~/.bagel/capabilities once before starting the container so the mount is owned by you, not root.

📚 Guides

📦 Integrations

  • Rerun · "show me that event in Rerun": any time window as a ready-to-open recording

  • Lichtblick / Foxglove · event windows as MCAP + pre-framed layouts for either viewer

  • PlotJuggler · open Bagel's MCAP outputs directly; one-sentence pre-framed sessions, flattened CSV/Parquet exports

  • Cloudini · decode cloudini-compressed pointclouds, or compress a bag's PointCloud2 topics into CompressedPointCloud2

  • Slack · pipelines post to your ops channel when they fire: "🚨 hard brake on {asset}"

  • LeRobot (beta) · detected events become training episodes: a LeRobotDataset v3.0

🚧 Limitations

Rough edges we know about, so you don't find them the hard way:

  • Two formats are beta. The automotive MDF4/CAN readers are verified against files we generate with the same libraries that read them (asammdf, python-can); real CANape/INCA/Vector-produced captures haven't crossed our test bench yet. LeRobot exports load-test clean with the real lerobot package, but no policy has been trained from a Bagel export yet.

  • The Jev integration is beta, and stays beta while we learn from real deployments. That covers the anomaly and decide gates (recorded logs and live subscriptions), preview_anomalies, the on-robot local backend and the ros2-jazzy-jev image. Its Jev backend has been run against live Jev through Vercel AI Gateway on a real drive, a synthetic fault log and a live MQTT stream; a direct TypeSafe key is not yet exercised, and detection quality has not been measured on logs with known incidents. On a live subscription the flagged window is kept as Parquet (write_topics_to_file); MCAP/rosbag snippets need a recorded log and are refused there. The baseline is learned per run and restarts with the process, so in screen mode the first minutes of each run are never flagged. Labels, settings and defaults may change between releases.

  • Reduction ratios are workload-dependent, and unbenchmarked. The ratio is event-window duration over total duration: quiet recordings reduce dramatically, eventful ones much less. The figures in this README are illustrative demo output, not a measured benchmark.

  • No authentication on the MCP endpoint. By design it binds to localhost only; treat it like a database socket and see SECURITY.md before sharing it beyond your machine.

  • Small local models struggle with multi-step pipelines. A 4-8B model handles tool selection and simple SQL; event-windowed reduction and multi-topic joins want a bigger model. See the Local LLMs guide.

  • Live-database end-to-end tests run outside CI. The InfluxDB and Postgres suites' pure tests run in CI; their live end-to-end cases only execute against an instance you point them at. Everything else, including the ROS bag write paths, runs in CI.

🫶 Contributing

We’d love your help! The easiest way to support the project is by giving it a ⭐ on GitHub.

Other great ways to contribute:

  • Request new features

  • Report bugs

  • Improve documentation

  • Add new capabilities

Before contributing, please review the guidelines.

Join the conversation in our Discord server. We hang out there regularly.

📄 License

Bagel is open source under the Apache License 2.0.

Agent discovery and reproducible workflows

For maintainers: discovery audits and evaluation, listing maintenance, and reproducible user reports.

Available Tools

25 tools
delete_capabilityDelete a user capabilityA
DestructiveIdempotent

Delete a capability previously saved with save_agent_capability, by the exact, full name list_agent_capabilities reports (user/-prefixed) -- a bare slug is rejected, since a user capability's name can shadow a builtin of the same stem. Only user-saved capabilities can be deleted -- builtins shipped with Bagel refuse with a clear message. An unknown name raises rather than silently no-op-ing, listing the user capabilities that do exist.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already mark destructiveHint and idempotentHint. The description adds valuable non-obvious behaviors: bare slugs are rejected due to shadowing, builtins refuse deletion, and unknown names raise instead of silently no-op-ing while listing existing user capabilities. No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three dense sentences each add necessary information: the operation and source, the exact-name requirement and rationale, and the error behavior. No filler or repetition of the schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With a single required parameter degraded by 0% schema coverage, the description fully compensates by explaining valid input, scope restrictions, and failure modes. Since an output schema exists, return-value documentation is not required here.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must carry the semantic burden for the name parameter. It does so thoroughly: exact full name, user/ prefix, source via list_agent_capabilities, and rejection of bare slugs. This is far more useful than the raw schema alone.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Delete a capability previously saved with `save_agent_capability`'. It clearly scopes the tool to user-saved capabilities and distinguishes it from builtins, making it unambiguous among siblings like save_agent_capability and list_agent_capabilities.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives concrete input guidance: use the exact full name list_agent_capabilities reports, with the user/ prefix, not a bare slug. It also states an explicit when-not case: builtins cannot be deleted and will refuse with a clear message.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_pipelineDelete a saved pipelineA
DestructiveIdempotent

Delete exactly one pipeline YAML file previously written by save_pipeline, by the same name list_pipelines reports. Confined to the trusted pipelines directory (settings.PIPELINES_DIRECTORY) -- there is no directory argument, so this can never be pointed at an arbitrary path -- and a name that would resolve outside it is refused before anything is touched. Deleting an unknown name raises rather than silently no-op-ing, listing the names that do exist -- so a second delete of the same name also raises.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.8/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description directly contradicts the idempotentHint: true annotation: it explicitly states that deleting an unknown name raises and 'a second delete of the same name also raises,' meaning repeated calls with the same arguments produce an error rather than a harmless repeat. Per the rubric, a contradiction with annotations forces a score of 1. The safety-confinement and error-behavior details are otherwise excellent, but the idempotency conflict is a serious signal for an agent deciding whether re-invocation is safe.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, front-loaded with the core purpose, and every clause adds a distinct safety or behavioral fact (provenance, directory confinement, refusal-before-touch, error-on-unknown, re-delete raising). The heavy use of em-dash parentheticals makes the sentences dense and requires careful parsing, but there is no wasted wording.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With a single parameter, an output schema present, and annotations covering destructiveness, the description covers everything needed to invoke correctly: name provenance, path-confinement safety, and error behavior. The one gap is that the conflicting idempotency signal between the annotation and the description leaves the agent with ambiguous guidance about repeat calls, which keeps this from being fully complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the bare schema offers only a 'Name' string, so the description carries the full burden for the name parameter and meets it completely: it specifies that the name must have been written by save_pipeline, must match what list_pipelines reports, must not resolve outside the trusted directory, and that unknown names raise with a list of existing names. An agent can determine exactly what value to pass and what will happen if it is wrong.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a precise verb+resource statement: 'Delete exactly one pipeline YAML file previously written by save_pipeline, by the same name list_pipelines reports.' It names the exact resource, scopes it to previously saved pipelines, and connects to the two sibling tools that define valid names, which cleanly separates it from delete_capability and the run/preview pipeline tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description tells the agent where valid names come from (save_pipeline / list_pipelines) and frames the safety contract: there is no directory argument, so the tool is confined to settings.PIPELINES_DIRECTORY and refuses out-of-directory names before touching anything. It does not, however, explicitly state when not to use it or name an alternative tool for deleting other resources, so it stops just short of full routing guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

describe_data_sourceDescribe a data sourceA
Read-onlyIdempotent

Inspect a robotics, drone, or IoT log first: returns source metadata, available topics, and instructions for summarizing them. Use describe_topic next for field schemas before SQL. Does not return message rows or detect anomalies. Paths are resolved on the Bagel server; runtime and schema support depend on the selected service.

ParametersJSON Schema
NameRequiredDescriptionDefault
argsNo
pathYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false, so the safety profile is covered. The description adds meaningful behavioral context beyond annotations: it discloses that no message rows or anomalies are returned, and that path resolution and runtime/schema support depend on the selected service. This helps set expectations and avoid misinvocation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences with no filler. The primary action is front-loaded, the sibling alternative follows immediately, and the exclusions and service dependency are packed into the final sentence. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has an output schema, so return value details need not be in the description. Purpose, workflow ordering, exclusions, and server-side behavior are all covered. The only gap is parameter semantics (especially args), which prevents a perfect score, but the overall context is strong for a simple read-only inspection tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate for both parameters. It only mentions that 'Paths are resolved on the Bagel server', which hints at the path parameter but does not explain its format or semantics. The optional args parameter is entirely undocumented, leaving a clear gap for a required-parameter tool.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description opens with a specific verb and resource: 'Inspect a robotics, drone, or IoT log first'. It clearly states what the tool returns ('source metadata, available topics, and instructions') and differentiates from the sibling describe_topic by prescribing it as the next step for field schemas. No ambiguity remains about what this tool does.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly says to use this tool 'first' and directs the agent to 'Use describe_topic next for field schemas before SQL', giving an exact sequence and alternative. It also states exclusions ('Does not return message rows or detect anomalies') and notes server-side path resolution, making the invocation context unambiguous.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

describe_topicDescribe a topic in a data sourceA
Read-onlyIdempotent

Inspect one known topic before writing SQL or event predicates. Returns its DuckDB schema, original message definition, and query instructions, without message rows. Discover topic names with describe_data_source. Confirm units from the definition or user; do not infer units from field names.

ParametersJSON Schema
NameRequiredDescriptionDefault
argsNo
pathYes
topicYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.3/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already mark the operation read-only, idempotent, and non-destructive. The description adds non-obvious behavioral traits: it returns schema, definition, and query instructions but not message rows, and explicitly warns against inferring units from field names. No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, front-loaded with the action and purpose. Each sentence adds distinct value: when to use it, what it returns, and a critical data-interpretation warning. There is no filler or repetition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only inspection tool, the description covers when to call it, what it returns, and a critical unit-safety warning. An output schema exists to detail return values, but the input parameters are not fully defined, which is a minor completeness gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It clarifies that 'topic' is a known topic and references describe_data_source for discovery, but it never defines 'path' or the optional 'args' object, leaving required inputs under-specified.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific action—inspect a known topic—and the resource (a topic in a data source). It explicitly distinguishes itself from describe_data_source by pointing to that tool for discovering topic names, so an agent can select the right one.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly frames usage as a pre-step before writing SQL or event predicates, and directs topic discovery to describe_data_source. It does not explicitly name query_messages as the alternative for retrieving message rows, though the 'without message rows' clause implies that boundary.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

export_for_lerobotExport event windows as a LeRobot training dataset (beta)A
Idempotent

Write selected event windows as a LeRobotDataset v3.0: one episode per window, scalar signals grouped into feature vectors and resampled to a uniform fps using last observation carried forward. Returns the dataset directory, episode/frame counts, and loading instructions. Choose this for robot-learning dataset preparation, not interactive viewing or model training. Beta: load compatibility tested; actual training-run validation remains outstanding.

ParametersJSON Schema
NameRequiredDescriptionDefault
fpsYes
argsNo
nameNodataset
pathYes
taskYes
topicsYes
episodesYes
featuresYes
robot_typeNounknown

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.1/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Discloses behaviors beyond annotations: confirms it writes a new dataset (consistent with readOnlyHint false), implies idempotency (consistent with idempotentHint), describes LOCF resampling, mentions return values (dataset directory, counts, loading instructions), and notes beta status with a specific caveat. These details add substantial transparency not captured in structured fields.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, front-loaded with purpose and process, then returns, then usage guidance and beta caveat. Every sentence adds value without redundancy or irrelevant detail.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

While the description covers high-level behavior and return values, it omits any explanation of the 9 parameters, including nested structures like 'features' and 'episodes'. Given the complexity and zero schema coverage, the agent lacks essential input semantics. The beta note and usage guidance are helpful but do not compensate for this gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The description provides zero parameter-specific semantics. It mentions 'scalar signals grouped into feature vectors' and 'uniform fps' but does not explain how parameters like 'features', 'episodes', 'topics', 'task', or 'path' map to the schema. With 0% schema coverage, this is a critical gap for an agent to construct valid inputs.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Write') and resource ('event windows as a LeRobotDataset v3.0'), details the conversion process (episode per window, feature vector grouping, LOCF resampling), and differentiates from sibling export tools by naming the target format and use case. An agent can clearly distinguish this from export_for_plotjuggler, export_for_rerun, and export_for_lichtblick.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly instructs when to use ('Choose this for robot-learning dataset preparation') and when not ('not interactive viewing or model training'). While it doesn't name alternative tools, the guidance is unambiguous and sufficient to route the agent correctly.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

export_for_lichtblickExport an event window for Lichtblick / FoxgloveA
Idempotent

Write a selected time window as JSON-encoded MCAP plus a plot layout for Lichtblick or Foxglove; returns paths, curves, and opening instructions. Choose this for those viewers, not byte-preserving native ROS/CDR export. Inspect topic schemas and use source timestamps in seconds. Automatically selects up to eight numeric plot curves unless signals are specified. Requires a separate viewer; does not launch it.

ParametersJSON Schema
NameRequiredDescriptionDefault
argsNo
nameNoevent
pathYes
topicsYes
signalsNo
end_secondsYes
start_secondsYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already include idempotentHint=true and destructiveHint=false. The description adds meaningful behavioral context beyond this: it writes files, auto-selects up to eight numeric plot curves, returns paths/curves/instructions, and does not launch a viewer. No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences carry all key information without fluff: purpose, alternatives, usage constraint, parameter hints, and limitations. The most important action is front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers the core workflow, output nature, and viewer behavior, and the output schema covers return values. However, the meaning of path, topics, name, and args is not fully specified, so an agent may need to infer whether path is the output destination, a directory, or a file prefix.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It clarifies units for start_seconds/end_seconds and explains the optional signals behavior and default auto-selection, but it leaves path, topics, name, and args semantics mostly implicit, which is a significant gap for a 7-parameter tool.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific action ('Write a selected time window as JSON-encoded MCAP plus a plot layout') and clearly identifies the target viewers (Lichtblick or Foxglove). It distinguishes itself from byte-preserving native ROS/CDR export, which is meaningful given sibling export tools for other viewers.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly tells the agent when to choose this tool ('Choose this for those viewers') and when not to ('not byte-preserving native ROS/CDR export'). It also gives practical constraints: inspect topic schemas, use source timestamps in seconds, and remember that the viewer is not launched.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

export_for_plotjugglerExport an event window for PlotJugglerA
Idempotent

Write a selected time window as flattened scalar CSV plus a PlotJuggler XML layout; returns paths, plotted curves, and an opening command. Choose this for scalar plotting in PlotJuggler, not native bag preservation. Inspect topic schemas and use inclusive source timestamps in seconds. Automatically selects up to eight numeric curves unless signals are specified. Requires the separate viewer to open; does not launch it.

ParametersJSON Schema
NameRequiredDescriptionDefault
argsNo
nameNoevent
pathYes
topicsYes
signalsNo
end_secondsYes
start_secondsYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With annotations already marking this as non-read-only, non-destructive, and idempotent, the description adds useful behavioral details beyond them: it returns paths, plotted curves, and an opening command, automatically selects up to eight numeric curves unless signals are specified, and does not launch the viewer. These disclosures help an agent set expectations around side effects and output.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and front-loaded: the core output behavior and purpose come first, followed by selection guidance, parameter hints, and a launch caveat. Every sentence contributes information, and there is no redundant restating of the title or schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers the main workflow, output, and key behavioral constraints, and an output schema exists to define return values. Still, with 7 parameters and no schema-level parameter descriptions, the omission of path, name, and especially the free-form args object leaves the definition incomplete for an agent that needs to construct a fully correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It adds real meaning for start_seconds/end_seconds ('inclusive source timestamps in seconds'), signals ('unless signals are specified'), and topics ('Inspect topic schemas'). However, path, name, and args remain undocumented in both the schema and the description, which leaves meaningful gaps for a 7-parameter tool.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb and resource: writing a selected time window as flattened scalar CSV plus a PlotJuggler XML layout, with return values also stated. It also distinguishes itself by saying 'Choose this for scalar plotting in PlotJuggler, not native bag preservation', which sets it apart from sibling export tools like export_for_rerun and export_for_lichtblick.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly states when to use this tool ('Choose this for scalar plotting in PlotJuggler') and what it is not for ('not native bag preservation'). It also adds practical context such as inspecting topic schemas and using inclusive source timestamps in seconds, but it does not name a specific alternative tool for bag preservation.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

export_for_rerunExport an event window for the Rerun viewerA
Idempotent

Write a selected time window as scalar time series in a Rerun .rrd recording; returns its path, signals, and an opening command. Choose this when the user requests Rerun. Requires rerun-sdk (uv sync --group viz) and a separate viewer. Inspect topic schemas and use source timestamps in seconds. This exporter does not produce camera or 3D scene replay and does not launch the viewer.

ParametersJSON Schema
NameRequiredDescriptionDefault
argsNo
nameNoevent
pathYes
topicsYes
signalsNo
end_secondsYes
start_secondsYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare idempotentHint=true and destructiveHint=false, so the safety profile is covered. The description adds valuable behavioral context: it requires a separate viewer, does not launch the viewer, does not produce camera/3D replay, and instructs to inspect topic schemas and use source timestamps in seconds. This goes beyond the annotations without contradicting them.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences with no filler. The core action and return value are front-loaded, followed by usage conditions and exclusions. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers purpose, prerequisites, exclusions, and return value. With an output schema present, the return value doesn't need full enumeration. The only minor gap is that it doesn't explain the 'signals' parameter's relationship to 'topics' in detail, but the scalar time series framing and output schema compensate.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It explains the core parameters conceptually: 'selected time window' maps to start_seconds/end_seconds, 'scalar time series' maps to topics/signals, and 'returns its path' maps to path. It doesn't enumerate each parameter, but it gives enough semantic framing for an agent to understand the tool's purpose. The 'use source timestamps in seconds' note adds crucial unit semantics for the time parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Write'), a resource ('selected time window as scalar time series in a Rerun .rrd recording'), and the return value ('path, signals, and an opening command'). It also distinguishes itself from other exporters by naming the Rerun format and explicitly noting it does not produce camera/3D replay.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit when-to-use guidance ('Choose this when the user requests Rerun'), prerequisites ('Requires rerun-sdk (uv sync --group viz) and a separate viewer'), and exclusions ('does not produce camera or 3D scene replay and does not launch the viewer'). This clearly routes an agent away from sibling exporters like export_for_plotjuggler or export_for_lichtblick.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_pipelineRead a saved pipelineA
Read-onlyIdempotent

Return the full configuration of one pipeline saved by save_pipeline, by the name list_pipelines reports, so it can be inspected, edited and saved again or passed to run_pipeline. Confined to the trusted pipelines directory (settings.PIPELINES_DIRECTORY); a name resolving outside it is refused.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the annotations (readOnlyHint, idempotentHint, destructiveHint), the description adds meaningful behavior: the pipeline is confined to settings.PIPELINES_DIRECTORY and names resolving outside it are refused. This security-relevant boundary is valuable for an agent to know before invocation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences with no filler. The primary purpose is front-loaded, the workflow context follows, and the security constraint is stated last. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter read tool with an output schema and comprehensive annotations, the description is complete: it explains what is returned, how the parameter is sourced, and the filesystem safety constraint. Nothing essential is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema provides only 'name' as a string with no description, so the description carries the full burden. It clarifies that name must be one saved by save_pipeline and reported by list_pipelines, and that it must resolve within the trusted directory, adding critical meaning beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states an exact verb and resource: 'Return the full configuration of one pipeline saved by save_pipeline'. It also names the related tools list_pipelines and run_pipeline, making the tool's role in the pipeline lifecycle clear and differentiating it from siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear workflow context by explaining the returned configuration can be inspected, edited, saved again, or passed to run_pipeline, and identifies list_pipelines as the source of valid names. It does not explicitly state when not to use it, but the usage context is strongly implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_agent_capabilitiesList agent capabilitiesA
Read-onlyIdempotent

List every capability available to run: the predefined .poml capabilities shipped with Bagel, plus any user-saved capabilities (.poml or .md, named with a user/ prefix) discovered under the user-capabilities directory. Each entry has a name, a path to pass to run_poml_capability, and a one-line summary. Use this to discover available capabilities instead of guessing file paths.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already establish readOnly, idempotent, and non-destructive behavior. The description adds meaningful behavioral context beyond that: it names the sources of capabilities, the supported file extensions, the user/ prefix convention, and the fields returned, which helps the agent trust and use the result correctly.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is three sentences with no wasted text. It front-loads the core purpose, then adds only high-value details about scope, return fields, and usage intent.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter listing tool with an output schema and strong annotations, the description is complete. It tells the agent what will be listed, what each entry contains, and how the results connect to run_poml_capability, so no critical information is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, so there is nothing for the description to clarify at the parameter level. The baseline 4 applies, and the description compensates by explaining what the returned path is for, which is the most semantically relevant detail.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('List') and names the exact resource: every capability available to run, including predefined .poml capabilities and user-saved capabilities. It also distinguishes itself from related tools by tying entries to run_poml_capability and clarifying it is not about guessing file paths.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear usage context: use this to discover available capabilities instead of guessing file paths. It does not explicitly exclude alternatives like list_pipeline_capabilities, but the focus on agent capabilities and run_poml_capability makes the intended use clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_live_topicsList available live topicsA
Read-onlyIdempotent

Discover topic names from a live broker or ROS bridge before subscribing. Supported type_ values: mqtt, ros1.bridge, ros2.bridge. Requires a reachable service; specify host and port when defaults do not fit. Returns topic names without starting a persistent recording. Use describe_data_source for an existing recorded log.

ParametersJSON Schema
NameRequiredDescriptionDefault
argsNo
hostNo
portNo
type_Yes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and destructiveHint false. The description adds valuable behavioral context beyond that: it returns names 'without starting a persistent recording' and requires a reachable service, which informs the agent about network dependency and side-effect-free behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three concise sentences, front-loaded with the core purpose. Every sentence earns its place: purpose, supported types, prerequisites, side-effect note, and alternative. No fluff.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, return format isn't needed. The description covers purpose, type_ values, network requirements, and the alternative tool. It could elaborate on host/port defaults or pagination, but for a listing tool it's adequately complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It lists supported type_ values explicitly ('mqtt, ros1.bridge, ros2.bridge') and mentions host/port defaults. However, the 'args' parameter is completely unexplained, and host/port are only vaguely referenced without format or default values.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Discover') and resource ('topic names from a live broker or ROS bridge') and clearly differentiates from siblings by noting 'before subscribing' and explicitly routing to 'describe_data_source for an existing recorded log'. An agent can immediately tell what this tool does and how it differs from similar tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear context: use it before subscribing, requires a reachable service, and points to the alternative describe_data_source for recorded logs. It doesn't explicitly state when NOT to use it, but the alternative condition is clear enough.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_pipeline_capabilitiesList pipeline capabilitiesA
Read-onlyIdempotent

Discover available task and gate modules before authoring pipeline YAML. Returns module paths, constructor parameters, summaries, and availability. These are executable building blocks, not saved workflows: use list_pipelines for saved configs and list_agent_capabilities for reusable agent instructions.

ParametersJSON Schema
NameRequiredDescriptionDefault
include_unavailableNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is covered. The description adds useful context about return contents (module paths, constructor parameters, summaries, availability) and clarifies that these are executable building blocks rather than saved workflows.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences with no filler. The purpose is front-loaded, and the second sentence efficiently adds return details and sibling-tool routing.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description is largely complete for selection and invocation: it states purpose, timing, return contents, and alternatives. The only gap is the include_unavailable parameter, which is simple enough that the schema title and default mostly cover it, and an output schema exists for return details.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description carries the burden of explaining the include_unavailable parameter, but it never mentions it. The schema title and default provide some meaning, but the description does not add value beyond that.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'Discover available task and gate modules before authoring pipeline YAML.' It also explicitly distinguishes itself from list_pipelines and list_agent_capabilities, so an agent can tell them apart without inspecting schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives explicit timing ('before authoring pipeline YAML') and names alternatives with the conditions that select them: 'use list_pipelines for saved configs and list_agent_capabilities for reusable agent instructions.' This fully routes the agent to the correct tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_pipelinesList saved pipelinesA
Read-onlyIdempotent

List the pipeline YAML files saved by save_pipeline in the trusted pipelines directory (settings.PIPELINES_DIRECTORY): each entry's name, path, and a one-line summary (task count, site/asset, and cadence). Use this to discover what has already been saved before reusing, editing, or deleting it -- instead of guessing file names.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false, so the description does not need to restate safety. It adds value by explaining the exact content of each entry (name, path, one-line summary with task count, site/asset, cadence), and implicitly communicates that this is a non-mutating discovery operation. There is no contradiction between description and annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences, front-loaded with the core purpose and output details, then a usage hint. No waste, and every sentence contributes meaning. It is appropriately concise for a simple read-only listing tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given zero parameters, a rich output schema (implied), and annotations covering safety, the description fully covers what an agent needs: it specifies what is listed, where from, and why to use it. Nothing crucial is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters and schema description coverage is 100% (vacuously). The description clarifies that the list is scoped to the trusted pipelines directory (`settings.PIPELINES_DIRECTORY`), which adds meaning beyond the empty schema. Baseline for 0 parameters is 4, and the description meets it.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the specific action (list) and resource (pipeline YAML files saved by `save_pipeline`), and details the output fields (name, path, summary). It differentiates from sibling tools like save_pipeline and delete_pipeline by focusing on discovery and referencing the trusted pipelines directory.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear context on when to use the tool: before reusing, editing, or deleting a saved pipeline, and cautions against guessing file names. While it does not explicitly name alternative tools or when not to use it, the guidance is actionable and points to the intended workflow.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

preview_anomaliesPreview what the anomaly gate would flag (beta)A
Read-onlyIdempotent

Dry-run the on-robot anomaly screen over a recorded log without calling any decision model or writing artifacts. Learns a rolling baseline the way the src.pipeline.gates.anomaly gate does, then reports every window the screen would flag (mean shift, extreme sample, topic dropout) with the signal, value and z-score, flag counts per signal, and plain-language advice (signals that drift by design, warm-up longer than the log). Inspect topics first and pass signals as rates and errors (accelerations, angular rates, currents), never positions or orientations. Use it to choose signals and thresholds before saving an anomaly pipeline. It models screen mode with the decision model confirming every flag, every: N seconds cadences, and a gate that sees every fire (list the anomaly gate first).

ParametersJSON Schema
NameRequiredDescriptionDefault
argsNo
pathYes
topicsNo
signalsNo
max_signalsNo
z_thresholdNo
cadence_topicNo
warmup_minutesNo
window_secondsYes
cadence_secondsNo
dropout_secondsNo
baseline_window_minutesNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already provide readOnlyHint=true, idempotentHint=true, destructiveHint=false, so the description correctly aligns with these (no contradiction). It adds valuable behavioral context beyond annotations: it mentions learning a rolling baseline, modeling screen mode, every: N cadences, and gate ordering (list anomaly gate first). It also notes what it does NOT do (calling decision model, writing artifacts), which helps set expectations. Minor deduction for not explicitly stating output schema details, but the output schema is provided separately.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but efficiently packed, with multiple clauses per sentence. It front-loads the purpose and main output, then adds usage guidance. However, it is a long paragraph that might be better split into sections (e.g., purpose, usage, behavior) for readability. It's not overly verbose relative to the complexity, so it earns a 4.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (12 params, output schema rich, many siblings), the description covers the essential decision points: when to use (before saving an anomaly pipeline), what it returns (flag counts, z-scores, advice), what inputs to avoid (positions/orientations), and how it models the anomaly gate. The output schema likely lists the return structure, so return values aren't needed in the description. It is complete for an agent to decide and call correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema provides default values and types but zero descriptions for parameters (0% coverage). The description compensates by explaining the meaning of key parameters: 'signals' must be rates/errors, not positions/orientations; 'window_seconds' is part of the required params; mentions cadence and warm-up concepts, which map to cadence_seconds and warmup_minutes. It doesn't explicitly define every parameter (like max_signals, z_threshold), but the context is enough for an agent to infer their roles from the defaults and overall description. Given the high parameter count (12) and low schema coverage, this is a strong effort.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool does a dry-run of the anomaly gate over a recorded log, specifying the verb (preview/dry-run), resource (anomaly screen/gate), and scope (on a log, no decision model or artifacts). It distinguishes from siblings like preview_pipeline by emphasizing on-robot real-time screen behavior and the anomaly gate specifically, and it provides a concrete list of output types (flags, z-scores, advice).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly says 'Inspect topics first and pass signals as rates and errors... never positions or orientations' and 'Use it to choose signals and thresholds before saving an anomaly pipeline'. It also hints at when not to use it implicitly by requiring a recorded log and the specific screening mode. This gives clear prerequisites and a purpose that distinguishes it from other preview/run tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

preview_pipelinePreview an event-driven data reductionA
Read-onlyIdempotent

Preview an event-window reduction without writing artifacts. Inspect source and topic schemas first, then provide a SQL boolean predicate and nonnegative pre/post seconds. Detects false-to-true transitions, debounces nearby events, merges overlapping windows, and returns event timestamps, intervals, total seconds, and kept seconds/fraction. Report these results and obtain user confirmation before executing the reduction with run_pipeline or run_pipeline_batch. Does not predict output byte size.

ParametersJSON Schema
NameRequiredDescriptionDefault
argsNo
pathYes
predicateYes
event_topicYes
pre_secondsYes
post_secondsNo
debounce_secondsNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the readOnly/idempotent annotations, the description discloses concrete behavior: detecting false-to-true transitions, debouncing events, merging overlapping windows, and returning timestamps, intervals, and kept seconds/fraction. The explicit caveat that it does not predict output byte size adds useful limitation context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is three dense but purposeful sentences: purpose, input constraints and behavior, then workflow and limitation. Every sentence earns its place and the core purpose is front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

It covers behavior, outputs, input constraints, the required confirmation workflow, and a known limitation, while the output schema handles return-value detail. The remaining gap is the under-specified path, event_topic, and args semantics, which keeps this from being fully complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 0% schema description coverage, the description must compensate, and it partly does: predicate is a SQL boolean, pre/post seconds are nonnegative, and debounce_seconds maps to debouncing nearby events. However, path, event_topic, and especially the catch-all args parameter remain semantically unexplained beyond their schema titles.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Preview an event-window reduction without writing artifacts.' It also distinguishes itself from the run_pipeline siblings by explicitly framing this as a pre-execution preview step.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives clear usage context: inspect source and topic schemas, provide a SQL predicate and nonnegative pre/post seconds, then report results and get user confirmation before running run_pipeline or run_pipeline_batch. It does not provide explicit when-not-to-use exclusions beyond the byte-size caveat.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

query_messagesQuery topic messages with SQLA
Read-onlyIdempotent

Answer quantitative questions with read-only DuckDB SQL over one topic: filtering, aggregates, downsampling, and event evidence. Call describe_data_source and describe_topic first; use the returned schema and show the SQL to the user. Returns rows as dictionaries. Use read_loggings for textual diagnostics. Time bounds are inclusive source timestamps in seconds, not offsets from the first message.

ParametersJSON Schema
NameRequiredDescriptionDefault
argsNo
pathYes
topicYes
end_secondsNo
sql_statementYes
start_secondsNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare read-only, idempotent, and non-destructive behavior. The description adds valuable context beyond that: time bounds are 'inclusive source timestamps in seconds, not offsets from the first message,' and the return format is 'rows as dictionaries.' It also mandates showing the SQL to the user, which is a behavioral requirement not captured in annotations. No contradiction exists.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is three packed sentences, front-loaded with purpose, then usage prerequisites, then time semantics. Every sentence earns its place; there is no filler or redundant restatement of the title or annotations. It is concise without sacrificing critical information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (SQL over messages, six parameters) and that an output schema exists, the description covers the essential context: prerequisites, output format, usage pattern, and time semantics. It does not explain 'path' or 'args', which might be non-obvious, but the presence of an output schema reduces the need to describe return structure. Overall, it is fairly complete for a read-only query tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With schema description coverage at 0%, the description must compensate for parameter meaning. It clarifies start_seconds and end_seconds as inclusive source timestamps in seconds, and implies topic is the single topic being queried. However, it leaves 'path' and 'args' unexplained. This is partial compensation but not complete, so a mid-range score is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'Answer quantitative questions with read-only DuckDB SQL over one topic.' It clearly distinguishes from siblings like read_loggings (textual diagnostics) and describe_topic (schema description) by emphasizing quantitative SQL querying on a single topic. The mention of 'filtering, aggregates, downsampling, and event evidence' further specifies the tool's scope.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit, actionable usage guidance: 'Call describe_data_source and describe_topic first; use the returned schema and show the SQL to the user.' It also provides an explicit alternative: 'Use read_loggings for textual diagnostics.' This tells the agent exactly when to use this tool versus a specific sibling, and the prerequisite sequence is clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

read_loggingsRead logging messages from a data sourceA
Read-onlyIdempotent

Read textual INFO/WARN/ERROR diagnostics from a recorded source, optionally within inclusive source-time bounds in seconds. Returns logging records, or an empty list when no logging topics exist. Use query_messages for numerical signals, statistics, and threshold detection; empty diagnostics do not prove a healthy log.

ParametersJSON Schema
NameRequiredDescriptionDefault
argsNo
pathYes
end_secondsNo
start_secondsNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly, idempotent, and non-destructive behavior. The description adds useful context: returns logging records or an empty list when no topics exist, inclusive time bounds in seconds, and a semantic caveat about empty logs. No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences with no wasted wording: purpose first, then return behavior, then routing to the alternative tool. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Rich annotations, an output schema, and sibling context cover safety and return structure. The description covers purpose, time bounds, empty-list behavior, and alternatives. The only notable gap is the meaning of the optional args parameter.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It clarifies start_seconds and end_seconds as inclusive bounds in seconds, but leaves the required path and the opaque args parameter unexplained. This is only partial compensation for four undocumented parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description uses a specific verb ('Read') and resource ('textual INFO/WARN/ERROR diagnostics from a recorded source'), and explicitly contrasts with query_messages. An agent can distinguish this tool from siblings without opening schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly directs numerical signals, statistics, and threshold detection to query_messages, and warns that empty diagnostics do not prove a healthy log. This gives clear when-to-use and when-not-to-use guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

run_pipelineRun a pipelineA

Execute one pipeline config against its single config.path after configuration and input validation. For event reductions, first call preview_pipeline, report events and kept seconds, and obtain user confirmation. Returns pipeline status, run counts, and artifact paths; inspect status for failures. Tasks may write files or contact external services. Use save_pipeline to store without executing, or run_pipeline_batch for multiple sources.

ParametersJSON Schema
NameRequiredDescriptionDefault
configYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate non-read-only, open-world behavior, but the description adds concrete side-effect context: tasks may write files or contact external services, and status should be inspected for failures. It does not cover auth or rate limits, but with annotations carrying the broad safety profile, this is a minor gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four sentences deliver action, workflow prerequisite, return behavior, side effects, and alternatives with zero filler. The key execution semantics are front-loaded, and every sentence contributes.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a run tool with side effects, the description covers preconditions, confirmation flow, return inspection, and sibling routing. The main gap is the under-specified config object, which is more a parameter semantics issue; overall, an agent has enough context to call this tool safely and correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 0% schema description coverage, the description must compensate, and it partially does by identifying config as a pipeline config and naming config.path as the execution target. However, it does not document the expected type, format, or any other config fields, leaving an open object with minimal guidance.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool executes a pipeline config and even cites the specific config.path that drives execution. It differentiates from siblings by naming save_pipeline (store without executing) and run_pipeline_batch (multiple sources), so an agent can distinguish it without opening schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives explicit guidance for event reductions: call preview_pipeline first, report events and kept seconds, and get user confirmation. It also names alternatives and the conditions for choosing them, making when-to-use and when-not-to-use unambiguous.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

run_pipeline_batchRun a pipeline across many data sources (batch)A

Execute one pipeline config for multiple explicit source paths or globs, overriding config.path for each source. Each source runs independently; returns per-source statuses/errors and totals. For event reductions, preview each source to be reduced and obtain confirmation of the batch scope before executing. Tasks may write files or contact external services. Use run_pipeline for a single source. A glob matching nothing is treated as a literal path.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathsYes
configYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations indicate side effects (readOnlyHint=false), external interactions (openWorldHint=true), and non-idempotency (idempotentHint=false), but the description adds crucial context: 'Tasks may write files or contact external services' reinforces the side-effect risk. It also explains that each source runs independently and returns per-source statuses/errors and totals, which is behavior not in the annotations. The glob-literal-path fallback is a specific behavioral disclosure. This adds value beyond annotations, though it doesn't exhaustively cover all behavioral nuances (e.g., partial failure handling), so a 4 is appropriate.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is five sentences, each serving a distinct purpose: the core action, the independent execution and return shape, the confirmation requirement for event reductions, the side-effect warning, and the single-source alternative plus glob edge case. It is front-loaded with the primary function, and there is no filler or repetition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (batch execution, side effects, potential confirmation requirement) and the presence of an output schema (which presumably details the return structure), the description covers what an agent needs: the operation, how config.path is overridden, side-effect caveats, the guidance to preview for event reductions, the single-source alternative, and the glob edge case. Nothing critical for a correct invocation is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It does: 'paths' are described as 'explicit source paths or globs', and the config is 'one pipeline config' with the behavior that it overrides config.path for each source. This explains the meaning and relationship of both parameters. While it doesn't enumerate all possible config fields, it gives sufficient semantics for an agent to construct valid inputs.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific action: 'Execute one pipeline config for multiple explicit source paths or globs'. It clearly identifies the resource (pipeline config) and the batch nature, and explicitly differentiates from the sibling run_pipeline by noting 'Use run_pipeline for a single source.' An agent can unambiguously determine what this tool does and when to select it.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides explicit usage guidance: it tells the agent to use run_pipeline for a single source instead, and instructs to preview sources before execution for event reductions and obtain confirmation of batch scope. This directly addresses when to use and when not to use the tool, going beyond mere implication. It also warns about the glob-matching-nothing edge case, adding operational guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

run_poml_capabilityRun a capability defined in a POML fileA
Read-onlyIdempotent

Load reusable agent instructions from a POML or Markdown file; returns a prompt for the calling agent to follow. Discover paths with list_agent_capabilities. POML accepts template context; Markdown rejects nonempty context. Loading the prompt does not itself analyze data, execute a pipeline, or write reduction artifacts. Use run_pipeline for an approved executable pipeline config.

ParametersJSON Schema
NameRequiredDescriptionDefault
poml_pathYes
poml_contextNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already provide readOnlyHint=true and destructiveHint=false, but the description adds valuable behavioral context: it clarifies the tool returns a prompt, does not analyze data or write artifacts, and explains a critical difference between POML and Markdown regarding template context. This goes well beyond the structured annotation fields.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, each earning its place: purpose, discovery, context behavior, non-effects, and alternative. The purpose is front-loaded, and there is no redundant or vague wording.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple 2-parameter, read-only tool with an output schema, the description covers what it does, what it doesn't do, how to find inputs, and when to choose a sibling tool. No critical information an agent needs to call it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With schema description coverage at 0%, the description fully compensates by explaining poml_path as a path discoverable via list_agent_capabilities and poml_context as template context, including the rule that Markdown rejects nonempty context. Both parameters are meaningfully described beyond their raw schema definitions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('load') and resource ('reusable agent instructions from a POML or Markdown file') and explicitly states the return value ('returns a prompt for the calling agent to follow'). It distinguishes itself from siblings by referencing list_agent_capabilities for path discovery and run_pipeline as the alternative for executable pipeline configs.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides explicit when-to-use guidance: 'Discover paths with list_agent_capabilities' and 'Use run_pipeline for an approved executable pipeline config'. Also gives a clear exclusion—'Loading the prompt does not itself analyze data, execute a pipeline, or write reduction artifacts'—so an agent knows this is not for execution tasks.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

save_agent_capabilitySave a user capabilityA
DestructiveIdempotent

Save a reusable workflow as a named capability so it can be discovered with list_agent_capabilities and run with run_poml_capability in any future session. Content may be POML (validated before saving; supports context parameterization) or plain markdown instructions. Writes only to the user-capabilities directory; builtin capabilities cannot be modified.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYes
contentYes
overwriteNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate destructiveHint and idempotentHint, but the description adds valuable details: it validates POML before saving, supports context parameterization, and restricts writes to the user-capabilities directory. It also clarifies that builtin capabilities cannot be modified, which is a behavioral constraint not obvious from the schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise and well-structured. It front-loads the purpose, then details content formats, and ends with a scoping constraint. Every sentence adds value without redundancy, making it easy to parse quickly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description provides sufficient information for an agent to call the tool correctly: it explains the purpose, content format, and directory scope. Since an output schema exists, the return format does not need to be described. The only minor gap is the semantics of the overwrite parameter, which is not mentioned but is a boolean with a default in the schema, so it is discoverable.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 0% schema description coverage, the description must explain parameters. It does explain 'content' by specifying allowed formats (POML with validation and context parameterization, or plain markdown). It implicitly refers to 'name' as 'named capability', but it does not explain the 'overwrite' parameter at all, leaving its behavior unspecified. The description adds some meaning but does not fully compensate for the lack of schema documentation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb 'Save' and the resource 'a reusable workflow as a named capability'. It also differentiates from siblings by specifying the destination directory and that builtin capabilities cannot be modified, making it distinct from tools like delete_capability or save_pipeline.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explains when to use the tool: to save a reusable workflow as a capability for future sessions. It also gives context on content formats (POML or markdown) and a constraint (builtin cannot be modified), which implicitly guides when not to use it. However, it does not explicitly name alternative tools or provide direct comparisons with siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

save_pipelineSave a pipeline to a YAML fileB
Idempotent

Persist a pipeline configuration to a YAML file so it can be reused, edited, or run later with run.py. Returns the path to the written file.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYes
configYes
directoryNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description discloses that the operation writes to a YAML file and returns the file path, which adds useful behavioral context. It does not mention overwrite behavior or file-naming conventions, but the annotations already cover idempotency and non-destructiveness.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence that front-loads the action and resource, includes the key return value, and contains no filler. It is appropriately sized for the tool's purpose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with three parameters and zero schema descriptions, the description leaves too much unspecified—config shape, directory behavior, file extension, and overwrite semantics. The return path is stated, but an agent cannot reliably construct arguments from the description alone.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, and the description provides no meaningful detail about the 'config' object structure, the 'name' parameter, or the optional 'directory'. The generic phrase 'pipeline configuration' is the only hint and is insufficient for correct invocation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Persist') and clearly identifies the resource ('pipeline configuration') and output format ('YAML file'). It is clear about what the tool does, though it does not explicitly differentiate from sibling tools like run_pipeline or preview_pipeline.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'so it can be reused, edited, or run later with run.py' provides clear context for when to use the tool. However, it does not explicitly state when not to use it or name alternative tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

snap_hardwareSnapshot robot hardware into a WaffleForm (experimental beta)A

Auto-detect the robot's current hardware, firmware, and software using waffle-iron and return the resulting hardware state. Requires the waffle CLI on PATH (cargo install waffle-iron). The WaffleForm it writes is immediately queryable as a data source.

ParametersJSON Schema
NameRequiredDescriptionDefault
directoryNo.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

All annotations are false, so the description carries the full burden, and it performs well: it discloses auto-detection behavior, the side-effect of writing a WaffleForm, the dependency footprint, and the post-condition of data-source queryability. Could be stronger with failure modes (e.g., what happens if no robot is available, whether the directory is created). Not a contradiction, just an opportunity for more.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tightly-scoped sentences: purpose, prerequisite/installation context, and side-effect/composition note. There is zero filler, and the most important information (what it does) is front-loaded. Every sentence adds distinct value.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For the tool's complexity (1 optional param, no nested objects, output schema present), the description covers the essentials: core behavior, setup prerequisite, and downstream consumption model. Gaps include the role of the `directory` parameter and what happens on failure, but for a tool of this size these are minor. The description respects the line of what structured fields already convey and adds meaningful orchestration context.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 0% schema description coverage, the burden falls on the description to explain the `directory` parameter, but it's never mentioned. The schema itself only gives a name and default ('.'), so an agent must guess whether it's the output destination, the robot's config directory, or a scan root. Given the description does zero compensation for its single parameter, a 2 is appropriate here.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb-resource pair ('Auto-detect the robot's current hardware, firmware, and software') and clearly states the output ('return the resulting hardware state' and 'writes a WaffleForm'). It clearly distinguishes this from siblings like run_pipeline or query_messages by establishing a unique outcome (queryable data source) and the experimental beta caveat in the title adds useful maturity context.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Discloses a hard prerequisite ('Requires the waffle CLI on PATH (cargo install waffle-iron)') and implies when it's useful by noting the output is 'immediately queryable as a data source.' It stops short of explicitly naming alternatives or excluding contexts (e.g., 'don't use for X, use save_pipeline instead'), so it loses a point here, but the practical when-to-use context is well covered.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

subscribe_live_topicsSubscribe to live topic messagesA
Destructive

Start a background subscription to mqtt, ros1.bridge, or ros2.bridge and persist messages locally; returns the sink directory for subsequent analysis. Use list_live_topics first. An optional pipeline runs continuously on incoming messages and may write or upload artifacts. overwrite=True clears existing topic buffers. Stop it with unsubscribe_live_topics. Not for inspecting an existing file.

ParametersJSON Schema
NameRequiredDescriptionDefault
argsNo
hostNo
portNo
type_Yes
topicsNo
pipelineNo
overwriteNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the annotations (destructiveHint=true, readOnlyHint=false), the description discloses that it runs in the background, persists messages locally, may continuously write/upload artifacts via an optional pipeline, and that overwrite=True clears existing topic buffers. These are meaningful behavioral details the agent would not otherwise know.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and front-loaded with the core action and return value. Every sentence adds useful context: sequencing, pipeline caveat, overwrite behavior, lifecycle, and an exclusion. No filler or redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity and the presence of an output schema, the description covers the essential operational context: background execution, local persistence, returned sink directory, optional artifact-writing pipeline, destructive overwrite behavior, and how to stop the subscription. Nothing critical is missing for an agent to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It does explain type_ (mqtt, ros1.bridge, ros2.bridge), overwrite (clears buffers), and pipeline (continuous processing, may write/upload). However, it leaves host, port, topics, and particularly the opaque args object semantically unexplained, though some are inferable from their names.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Start a background subscription'), the resource types (mqtt, ros1.bridge, ros2.bridge), and the purpose (persist messages locally and return a sink directory). It clearly differentiates from siblings by saying to use list_live_topics first, to stop with unsubscribe_live_topics, and that it is not for inspecting an existing file.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit sequencing ('Use list_live_topics first'), the lifecycle counterpart ('Stop it with unsubscribe_live_topics'), and an exclusion ('Not for inspecting an existing file'). This is strong guidance on when and how to use the tool relative to alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

unsubscribe_live_topicsStop live topic subscriptionsA
Idempotent

Stop topics started by subscribe_live_topics on a running mqtt, ros1.bridge, or ros2.bridge subscription (same type_, host and port). Omit topics to stop all of them. A standing pipeline finishes its queued runs and its end-of-stream run first, which may write or upload artifacts. Recorded messages stay in the returned sink directory for analysis. Once nothing is subscribed the connection is closed. Never opens a connection: errors if there is no running subscription there.

ParametersJSON Schema
NameRequiredDescriptionDefault
hostNo
portNo
type_Yes
topicsNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description goes well beyond the annotations by disclosing that standing pipelines finish queued and end-of-stream runs first, may write or upload artifacts, that recorded messages remain in the sink directory, that the connection closes when nothing remains subscribed, and that the tool never opens a connection. These details meaningfully describe side effects and error behavior. No contradiction with the annotations is apparent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is information-dense with no wasted words and front-loads the core action. It is slightly long with several caveats, but each sentence contributes necessary behavioral context.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers the lifecycle of the operation: what is stopped, how to stop all, pipeline completion, artifact effects, retained recordings, connection closing, and the no-connection error case. This is sufficient for an agent to invoke the tool correctly without needing more context.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It does explain that type_, host, and port must match a running subscription and that omitting topics stops all topics. However, it does not clarify the meaning of null/default host and port values, nor the exact accepted values for type_, leaving some ambiguity for correct invocation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Stop') and a precise resource ('topics started by subscribe_live_topics'), and ties it to a running mqtt, ros1.bridge, or ros2.bridge subscription. It clearly distinguishes this tool from its inverse sibling, subscribe_live_topics, and from list_live_topics.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description makes the intended use clear: stop matching live topic subscriptions, with the explicit option to omit topics to stop all. It also gives an important boundary condition by stating that it errors when no running subscription exists. It does not explicitly contrast with other siblings, but the counterpart tool is named and the usage context is unambiguous.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 2 tool updatesv2.4.0
    • Addedget_pipeline
    • Addedunsubscribe_live_topics
  2. 6 tool updatesv2.3.1
    • Addeddelete_capability
    • Addeddelete_pipeline
    • Addedlist_pipelines
    • Addedpreview_anomalies
    • Addedsave_agent_capability
    • Changedsave_pipeline3 fields changed
      • addedInput schema / properties / directory / anyOf
        Added value: +[
        +  {
        +    "type": "string"
        +  },
        +  {
        +    "type": "null"
        +  }
        +]
      • changedInput schema / properties / directory / default
        Previous value: -"pipelines"New value: +null
      • removedInput schema / properties / directory / type
        Removed value: -"string"
  3. 2 tool updatesv2.1.0
    • Addedlist_agent_capabilities
    • Addedsnap_hardware
  4. 16 tool updatesv2.0.1
    • First observeddescribe_data_source
    • First observeddescribe_topic
    • First observedexport_for_lerobot
    • First observedexport_for_lichtblick
    • First observedexport_for_plotjuggler
    • First observedexport_for_rerun
    • First observedlist_live_topics
    • First observedlist_pipeline_capabilities
    • First observedpreview_pipeline
    • First observedquery_messages
    • First observedread_loggings
    • First observedrun_pipeline
    • First observedrun_pipeline_batch
    • First observedrun_poml_capability
    • First observedsave_pipeline
    • First observedsubscribe_live_topics

TDQS

A4/5.0

Scored across 25 tools

Disambiguation5/5

Each tool targets a distinct resource or action: capabilities, pipelines, data-source inspection, live-topic subscriptions, and exports are all clearly separated. Similar-sounding tools such as run_poml_capability versus run_pipeline, list_pipelines versus list_pipeline_capabilities, and query_messages versus read_loggings are explicitly cross-referenced and disambiguated in their descriptions.

Naming Consistency4/5

The set overwhelmingly follows a verb_noun snake_case pattern with consistent list/save/delete/run/preview/export verbs. Minor deviations include run_poml_capability supporting Markdown despite the POML-specific name, delete_capability breaking the save_agent_capability/list_agent_capabilities symmetry, and the somewhat vague snap_hardware verb.

Tool Count3/5

With 25 tools the server sits exactly in the heavy 16-25 range, making it borderline even though the domain is fairly broad. The pipeline and capability management clusters plus the four exporter tools add significant surface area, though none of the tools feel entirely redundant.

Completeness4/5

Core lifecycles are covered well: capabilities have list/save/run/delete, pipelines have save/get/list/delete/run/batch/preview, and data inspection covers schemas, SQL queries, and textual logs. Minor gaps remain, such as no tool to enumerate recorded data sources or produce byte-preserving native bag exports, but agents can work around these.

Maintenance

ActivityActive
ResponsivenessResponsive

Related MCP Connectors

Related MCP Servers

  • F
    license
    Not graded
    quality
    D
    maintenance
    Enables language models to perform hardware engineering tasks including CAD part design and heat transfer simulations. Provides tool calls for building mechanical components and running thermal analysis through natural language interactions.
    -
  • A
    license
    C
    quality
    A
    maintenance
    Provides unified control for both physical robots (ROS-based like Moorebot Scout, Unitree) and virtual robots in Unity3D/VRChat, enabling multi-robot coordination, environment generation, and automated 3D model creation.
    25
    11
    MIT
  • A
    license
    Not graded
    quality
    B
    maintenance
    Enables AI agents to control hardware devices like Arduino, Raspberry Pi, 3D printers, CNC machines, and custom robots via serial ports and HTTP. Provides tools for device discovery, command sending, sensor reading, servo control, G-code execution, and emergency stops with safety features.
    41 PyPI
    MIT