Skip to main content
Glama

TopicForge

PyPI version CI Python versions License: MIT Read-only

A read-only MCP (Model Context Protocol) server that lets an AI agent inspect a ROS2 graph, recorded bag files and the DDS layer underneath ROS. The code has no write path: it cannot publish to the bus or command a robot, and there is no permission system to configure.

It gives the agent twelve typed tools that return frozen Pydantic schemas, identical whether the server talks to a real robot or to its built-in mock fixtures. Ask why nav_planner gets no scan, and the agent reads the bus, finds the BEST_EFFORT writer facing a RELIABLE reader and names the incompatible policy (see examples/02-debug-qos-mismatch.md). It is meant for ROS2 developers, robotics ML/CV engineers and teams that cannot accept a write path into a production stack.

For DDS, TopicForge joins a domain as a read-only participant through one open-source binding (Eclipse CycloneDDS from PyPI) and reads the builtin discovery topics that the OMG DDS-RTPS protocol standardizes. So far the author has observed Cyclone DDS and Dust DDS participants on a live bus. RTI Connext, OpenDDS, CoreDX and Fast DDS announce themselves through the same standard discovery, but none of them has been observed yet. This covers discovery only: participants, readers, writers and their QoS. See docs/dds-interop-matrix.md.

Quickstart

No ROS2 is needed; the mock adapter serves deterministic fixtures for a small differential robot (LIDAR + RGB camera). Python 3.10 to 3.13.

pip install topicforge
TOPICFORGE_MODE=mock python -m topicforge
# Windows PowerShell: $env:TOPICFORGE_MODE="mock"; python -m topicforge

The server speaks MCP over stdio and waits for a client, so wire it into one. For Claude Desktop, add to claude_desktop_config.json:

{
  "mcpServers": {
    "topicforge": {
      "command": "python",
      "args": ["-m", "topicforge"],
      "env": { "TOPICFORGE_MODE": "auto" }
    }
  }
}

Then ask it to list the topics or to analyze /tmp/demo.mcap. For Claude Code: claude mcp add topicforge -- topicforge. Setup for a real ROS2 environment (WSL2, Linux, Docker, native Windows) is in docs/TESTING.md; recurring monitoring prompts and the privacy contract are in docs/TUTORIEL.md.

Related MCP server: ros2-mcp

Tools

Every response except health_check carries mode_effective ("live" or "mock"), so a caller can tell a real graph from fixtures.

Tool

Purpose

health_check

Environment and mode introspection. Always succeeds; reports mode next to requested_mode

list_topics

Discover the ROS2 graph

get_topic_info

Message type, publisher/subscriber counts and QoS for one topic

sample_messages

Peek recent messages on a ROS2 topic (count clamped to 50)

analyze_bag

Summarize a .mcap / .db3 recording or rosbag2_* directory (via ros2 bag info)

list_participants

DDS participants on the domain: vendor, name (EntityName QoS, Cyclone) and hostname

detect_qos_mismatches

Incompatible QoS pairs between DDS readers and writers

peek_dds_samples

Raw DDS samples; structured on the three builtin discovery topics, presence-only on user topics

participant_events

Timeline of participant discovered / lost events

topic_metrics

Frequency, sequence-gap and latency schema; data only for builtin discovery topics

peek_bag_samples

Decoded samples from a recorded bag, including ROS 1 .bag (needs pip install topicforge[bags])

list_endpoints

DDS writers and readers with structured QoS, per-topic roll-up that flags orphans (writer with no reader, reader with no writer)

Walkthroughs against the mock, each with the exact tool calls and payloads, are in examples/. To run the DDS tools against a real bus with several programs and vendors, see examples/dds/README.md (python examples/dds/run_all.py).

Modes

Mode

When to use

Backend

mock

Development, demos, CI, screencasts

Deterministic fixtures

live

ROS2 sourced and on PATH, and/or a DDS backend

ros2 CLI, DDS participant

auto

Detect what is available, else mock (default)

Best available

live and auto degrade instead of failing: if neither ros2 nor a DDS backend comes up, the server serves the mock fixtures and health_check reports mode: "mock" next to requested_mode. Only TOPICFORGE_MODE=mock forces fixtures unconditionally. The live ROS2 adapter shells out to the ros2 CLI, so rclpy does not need to be importable.

DDS backends

pip install topicforge[dds]                      # Eclipse CycloneDDS ([dds-cyclone] is the same thing)
TOPICFORGE_DDS_BACKEND=cyclone python -m topicforge

TOPICFORGE_DDS_BACKEND accepts mock (default), cyclone, fast and auto (fast, then cyclone, then mock, whichever binding imports). The default mock selects no DDS backend: installing the Cyclone binding is not enough, you must also set TOPICFORGE_DDS_BACKEND=cyclone. With TOPICFORGE_MODE=live and no backend selected, the DDS tools raise DDS module is not active: ... with the actual cause (backend not selected, binding not installed, or binding installed but the adapter failed to start), and health_check reports dds_backend: "none" plus dds_inactive_reason. An explicit value is honoured with or without ros2 on PATH, in any mode except mock. If the binding is missing or the participant cannot start, the server logs a warning naming the cause and falls back to the ROS2 CLI alone, or to the mock fixtures. When both ros2 and a DDS backend are up, a composite adapter routes the five ROS2 graph and bag tools to the CLI and the seven DDS tools to the DDS backend.

A Fast DDS adapter exists but has never run against a bus, and its fastdds Python binding is not on PyPI: build it from eProsima's sources and install it next to TopicForge. There is no [dds-fast] extra. opendds and dust are permanent stubs that never serve. rti, opensplice, coredx and intercom are rejected with a configuration error,; Cyclone already sees those vendors' participants through standard discovery. Full backend selection, the routing table and the QoS mismatch scenario are in docs/DDS_QUICKSTART.md; error messages are in docs/TROUBLESHOOTING.md.

Configuration reference

Variable

Default

Description

TOPICFORGE_MODE

auto

mock, live or auto

TOPICFORGE_LOG_LEVEL

INFO

DEBUG, INFO, WARNING, ERROR

TOPICFORGE_ROS2_BIN

ros2

Name or path of the ROS2 CLI binary

TOPICFORGE_TELEMETRY

off

Opt-in anonymous telemetry; an unrecognized value aborts startup. See Telemetry

TOPICFORGE_DDS_BACKEND

mock

mock (no DDS backend), cyclone, fast, auto (opendds and dust are stubs)

TOPICFORGE_DDS_DOMAIN_ID

0

DDS domain observed (0..232). Joined at startup; changing it needs a restart

TOPICFORGE_MAX_SAMPLE_BYTES

1048576

Size cap for one sampled message (1 KiB..64 MiB); a call returns at most 4 times that. Over-cap messages are dropped with a note

Samples with comments are in .env.example. Any invalid value stops the server with topicforge: configuration error: ... and exit code 2, so a typo cannot silently change behaviour.

Limitations

  • DDS validation is partial. The Cyclone adapter has run against a real bus, with Cyclone and Dust DDS participants, on Windows and in CI on Ubuntu and Windows (.github/workflows/demo.yml). The Fast DDS adapter has never run against a bus, and no RTI, OpenDDS, CoreDX or OpenSplice participant has been observed by this project. The multi-vendor claim rests on the RTPS protocol guarantee, not on a recorded cross-vendor run.

  • User-topic payloads are not decoded. peek_dds_samples on a user topic returns count 0 and a note that the topic is announced on the bus; no traffic is read. topic_metrics therefore has data only for the builtin discovery topics and says so in its status. It is a discovery-layer probe, not a publish-rate monitor.

  • Liveliness at runtime is not observed. A writer that is alive but silent (a hung process whose lease is still renewed) looks healthy, because TopicForge reads discovery, not data. An opt-in data probe is planned for 0.5.6. A crash and a clean leave cannot be told apart, and lost_ns is an upper bound of the death.

  • Cyclone vendor ids: participants that do not follow the RTPS vendor-id convention in their GUID prefix (Dust DDS, and RTI by default) are reported with vendor unknown.

  • Single domain: the server observes the domain it joined at startup; changing it needs a restart.

  • DDS Security is not handled. A participant without credentials sees an empty secure bus. detect_qos_mismatches checks Partition, type name, Reliability, Durability, Deadline, Liveliness, LatencyBudget, Ownership (kind), DestinationOrder and DataRepresentation (History as a risk); Presentation, XTypes assignability and runtime behavior are not checked, and the result lists them in policies_unchecked. It returns a MismatchScan envelope: read reports for the mismatches.

  • Fast DDS serves no list_endpoints.

  • sample_messages (live) runs ros2 topic echo --csv --once with a short timeout, so it returns at most one message, and a topic with no current publisher returns an empty sample. timestamp_ns is the message header.stamp for Header-stamped types and 0 for headerless ones. Arrays are cut at 128 elements by default; max_array_length (1..65536, or null for no cut) and arrays_summary_only change that, and a cut is listed under _truncated_after_columns.

  • analyze_bag (live) parses ros2 bag info text for the totals and counts; anomaly detection is mock-only. Per-topic times, rates ((n - 1) / span) and latched are added when the bag can be read locally (.db3 with the standard library, .mcap with rosbags), else rates fall back to count / bag duration (frequency_basis). peek_bag_samples reads the file itself, through rosbags, and is served only by the ROS2 CLI adapter or the mock; bags that embed no message definitions (Humble .db3) are decoded with the Humble definitions, or the distro the bag records, and note says so. Without ros2, bag tools return fixtures: check health_check for mode: "mock" before trusting bag output.

  • Synchronous handlers: the tools run on the MCP event loop; on Windows a hung ros2 launcher can block the server.

  • No streaming or push subscriptions: tools are strictly request/response.

Next: an opt-in probe to tell a hung writer from a healthy one, wider real-bus validation (Fast DDS, RTI, OpenDDS), and DDS Security. Open work is tracked in issues.

Telemetry

Opt-in, anonymous and off by default. When off, instrumentation returns the handler unchanged: no event is built, no transport is constructed, no network code runs (pinned by tests/test_telemetry.py::test_build_app_off_makes_no_transport_calls).

TOPICFORGE_TELEMETRY=on python -m topicforge

On-values: on, 1, true, yes, enabled. Off-values: unset, off, 0, false, no, disabled. Anything else is a configuration error rather than a silent "off".

When on, each tool call emits one event with exactly six fields:

Field

Example

Notes

tool_name

"list_topics"

One of the twelve tools, never argument values

latency_ms

12.34

Handler wall-clock duration, 2 decimals

mode

"mock"

Mode of the adapter actually serving: mock or live

version

"0.5.6"

TopicForge server version

session_id

"a1b2c3..."

Random UUID per process, never persisted

success

true

Whether the handler returned or raised

Never sent: topic names, message types or payloads, bag paths or contents, hostnames, usernames, IP addresses, environment variables, error messages. The field set is fenced by tests/test_telemetry.py::test_payload_contains_only_whitelisted_keys; adding a field requires updating this section. The default transport is a structured log line; there is no HTTP endpoint yet. The implementation is in src/topicforge/telemetry/.

Security model

TopicForge is designed for local trust: it runs as a subprocess of your MCP client on a machine you control and inspects your own ROS2 graph, DDS domain and bag files. It is not hardened for adversarial inputs.

  • TOPICFORGE_ROS2_BIN accepts an arbitrary path; treat it the way you treat PATH.

  • analyze_bag and peek_bag_samples open whatever path the client passes (no workspace isolation, no symlink restriction).

  • All ros2 invocations use subprocess.run with an argument list, never shell=True. ROS2 topic names are validated against ^/[A-Za-z0-9_/]+$ first.

  • The server loads no third-party code at startup.

  • No outbound network calls unless telemetry is turned on.

Before exposing TopicForge to untrusted MCP clients (hosted endpoints, shared environments), add path isolation and revisit the TOPICFORGE_ROS2_BIN policy. Vulnerability reports: see SECURITY.md.

Development

git clone https://github.com/yaniswav/TopicForge.git && cd TopicForge
python -m venv .venv && source .venv/bin/activate     # Windows: .venv\Scripts\Activate.ps1
pip install -e ".[dev]"
python -m ruff check src tests
python -m pytest

Tests run against the mock adapter, the live adapter's pure parsers and the binding-free DDS helpers; they never need a running ROS graph. Tests needing the cyclonedds or fastdds binding skip themselves when it is absent, and the integration marker (real-bus tests) is deselected by default. The Makefile (make check) uses POSIX shell syntax; on plain PowerShell run the commands above. See CONTRIBUTING.md.

Upgrading

TopicForge is pre-1.0; CHANGELOG.md lists every change, including the yanked releases and removed extras.

Layout

src/topicforge/
  server/      MCP bootstrap, build_app(settings)
  tools/       thin FastMCP handlers, no backend logic
  services/    input validation, orchestration, adapter factory
  adapters/    ros2_live, ros2_mock, dds_cyclone, dds_fast, common/ (binding-free logic)
  models/      frozen Pydantic schemas, the contract with MCP clients
  config/      settings and mode resolution
  telemetry/   opt-in, off by default
examples/      mock walkthroughs (*.md) and runnable live DDS examples (dds/)
scripts/       real-bus interop checks
docs/          guides, product plan
tests/         pytest suite, mock-only, no ROS2 or DDS required

Layers are strictly separated: handlers never call subprocess, adapters are the only code that talks to a backend, and new backends implement the MiddlewareAdapter protocol in adapters/base.py.

License

MIT, see LICENSE. Integration or support work for a specific ROS 2 / DDS setup: ethvignot.yanis@gmail.com.

Available Tools

12 tools
analyze_bagA

Summarize a ROS 2 bag at path. Returns a BagAnalysis with storage format, duration, message count, per-topic stats, detected anomalies and mode_effective (live or mock). Per topic, frequency_hz is (n - 1) / (last - first message time) of that topic (frequency_basis topic_span), with first_timestamp_ns, last_timestamp_ns and latched; a latched topic whose messages all fall within 1 second (e.g. /tf_static, a start-up burst) has a null frequency_hz, while a latched topic published over a longer span keeps its rate. When the bag cannot be read locally, or is a large .mcap (over 200 MiB), the rate falls back to count / bag duration (bag_duration) and note says why. Live mode runs ros2 bag info and accepts .mcap and .db3 files plus rosbag2_* directories (ROS 1 .bag files are not readable by ros2 bag info; use peek_bag_samples for those); mock mode returns fixture data for any path suffix except clearly non-bag ones. Raises an MCP error if the path is malformed, missing in live mode, or unparseable, or if no ros2 CLI is available. Anomaly detection is available in mock mode only. Read-only; no side effects.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathYesPath to a bag: a file ending in `.mcap` or `.db3`, or a `rosbag2_*` directory. `peek_bag_samples` also reads ROS 1 `.bag` files; `analyze_bag` does not (`ros2 bag info` cannot open them). Leading/trailing whitespace is stripped. Null bytes and otherwise malformed filesystem paths are rejected. Existence and bag format are validated by the live adapter (mock mode accepts any well-formed path).

Output Schema

ParametersJSON Schema
NameRequiredDescription
noteNoWhy the result is less detailed than usual, for example per-topic rates computed over the whole bag duration because the bag was too large to read per-topic message times. `None` when there is nothing to add.
pathYesPath to the analyzed bag, as supplied by the caller. May point to a file (`.mcap`, `.db3`, or ROS 1 `.bag` for `peek_bag_samples`) or to a `rosbag2_*` directory.
topicsYesPer-topic statistics for every topic present in the bag.
anomaliesNoHuman-readable notes about gaps, clock jumps, or other oddities. Populated in mock mode only; live mode does not detect anomalies.
bag_formatNoConcrete bag container format detected by the reader: `mcap` (Foxglove MCAP), `db3` (ROS2 rosbag2 SQLite), `bag` (ROS1 legacy chunked), or `unknown` when the reader could not classify. `None` when the bag was summarized from `ros2 bag info` text, which carries no format information.
message_countYesTotal number of messages across all recorded topics.
mode_effectiveYesRuntime mode the adapter served this response in: `live` (real ROS2 introspection) or `mock` (deterministic fixtures). Lets a caller tell a real graph from a demo one without calling `health_check`.
storage_formatNo`mcap`, `sqlite3`, or other storage identifier when known.
duration_secondsYesTotal bag duration, in seconds (wall clock between first and last message).
participants_recordedNoDDS participants recorded in the bag when the container format embeds participant metadata. MCAP can carry it via channel metadata records; ROS2 `.db3` and ROS1 `.bag` generally do not. Empty list when not available, which is the common case.
recording_duration_nsNoRecording duration in nanoseconds when readable from the bag's index. `None` when only `ros2 bag info` text was parsed; `duration_seconds` (float) is the always-populated fallback that downstream LLM consumers should prefer when this is `None`.
samples_decoded_countNoTotal decoded sample count across all topics produced by the bag reader. `0` when the reader only parsed metadata or when `rosbags` is not installed on the host. Use `peek_bag_samples` to pull the actual sample payloads for a specific topic.

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does so: it declares read-only/no side effects, enumerates the exact MCP error conditions (malformed path, missing in live mode, unparseable, no ros2 CLI), documents the large-.mcap fallback with a `note` field, and discloses that anomaly detection exists only in mock mode.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the one-line purpose, then layered detail; nearly every sentence carries distinct behavioral information (latched-topic rate rule, fallback, error list). It is dense to the point of being heavy, and the bolded mode blocks add visual weight without new semantics, so a 4 rather than 5.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with fallback paths, two execution modes, and error conditions, this is complete: return shape is named, edge cases (latched topics, oversized .mcap, unreadable bags) are handled, and an output schema exists so deeper return-value detail is unnecessary.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The single `path` parameter already has 100% schema description coverage covering accepted suffixes, whitespace stripping, and live-vs-mock validation. The description restates the format rules and the ROS 1 exclusion but adds little path syntax that the schema does not already carry, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Summarize a ROS 2 bag at path') and names the concrete artifact returned (a BagAnalysis with storage format, duration, per-topic stats). It also differentiates itself from the sibling peek_bag_samples by noting that ROS 1 .bag files must go there instead.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly routes the ROS 1 .bag case to peek_bag_samples and explains when the live adapter falls back to count/bag-duration. It does not explain how live vs mock mode is selected (no mode parameter exists), which leaves a small inference gap, so not a full 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

detect_qos_mismatchesA

Explain why DDS readers and writers on the same topic do not talk, and who will. Pairs every reader with every writer per topic and returns a MismatchScan: matched (pairs DDS will connect given the announced QoS; data flow is not observed), reports (incompatible or risky QoS pairs, each with participant names, requested vs offered values and the failed rule in details), not_matched (pairs DDS never matches: different partitions or type names; the QoS rules are NOT evaluated for them, so a partition split is not blamed on Reliability; latent_incompatible_policies lists the RxO policies that would ALSO be incompatible once the partition/type issue is fixed), hints (orphan topics with a near-identical name, i.e. probable typos, and type id notes), plus pairs_checked, topics_scanned, policies_checked and policies_unchecked. Checked: Partition (with * and ? wildcards), type name, Reliability, Durability, Deadline, Liveliness, LatencyBudget, Ownership, DestinationOrder, DataRepresentation, History (risky only, and only where announced: discovery does not carry it). Not checked: see policies_unchecked. An empty reports with a non-empty not_matched still means no data flows, and an all-empty result does not prove the bus healthy: discovery shows the QoS DECLARED, not runtime behavior (a reader logging 'deadline missed' with compatible QoS means the writer's real period exceeds the deadline, which TopicForge cannot observe). Pass topic to scope to one topic; omit to scan all. reports, matched and not_matched are capped at 200 entries each (truncated is true, the *_total fields keep the real counts, incompatible reports come first). A matched pair flagged late_joiner is a VOLATILE writer whose reader joined later on the same host: normal, not a fault. Read-only by architecture. Raises an MCP error when no DDS module is active; the mock backend returns fixtures.

ParametersJSON Schema
NameRequiredDescriptionDefault
topicNoOptional DDS topic name to scope the scan to: a bare name such as `scan` or a ROS 2 mangled name such as `rt/scan`. Omit to scan all topics.

Output Schema

ParametersJSON Schema
NameRequiredDescription
hintsYesLeads that are not findings: orphan topics with near-identical names (typos), type id differences, pairs that could not be fully checked.
matchedNoPairs that will be matched by DDS given the announced QoS: same topic and type name, overlapping partitions, no incompatible RxO policy (a pair with only a `risky` History finding still counts). Actual data flow is not observed.
reportsYesPairs that share a partition and a type but have incompatible or risky QoS.
truncatedNoTrue when `reports`, `matched` or `not_matched` was cut to its cap (200 entries each, incompatible reports first): see the `*_total` fields. Narrow the scan with `topic`.
not_matchedYesPairs separated by partition or type name. No data flows between them. An empty `reports` with a non-empty `not_matched` does not mean the bus is healthy.
matched_totalNoMatched pairs before the size cap.
pairs_checkedYesSame-topic (reader, writer) pairs examined, including those reported in `not_matched`.
reports_totalNoReports before the size cap.
mode_effectiveYesRuntime mode the adapter served this response in: `live` (real ROS2 introspection) or `mock` (deterministic fixtures). Lets a caller tell a real graph from a demo one without calling `health_check`.
topics_scannedYesTopics that had at least one endpoint in scope.
policies_checkedYesPolicies compared on every pair.
not_matched_totalNoNot-matched pairs before the size cap.
policies_uncheckedYesPolicies and facts this scan does not cover, each with a one-line reason. A clean result says nothing about them.

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden and does so thoroughly: read-only by architecture, raises an MCP error when no DDS module is active, mock backend returns fixtures, results capped at 200 with truncation flags, and explicit disclosure that discovery shows DECLARED QoS not runtime behavior. It also explains the late_joiner flag as normal rather than a fault.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is long and dense, but purpose is front-loaded in the first sentence and each subsequent sentence covers distinct output fields, caveats, or limits that a caller needs. The length is largely earned by the tool's complexity, though some field enumeration borders on restating the output schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Despite an output schema existing, the description goes further to explain what each field means and, critically, how to interpret ambiguous results (empty reports with non-empty not_matched still means no flow; all-empty does not prove health). Combined with error and mock-backend disclosure, nothing needed to invoke or interpret it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There is a single optional parameter and schema coverage is 100%, with the schema already documenting the bare vs ROS 2 mangled name formats and the omit-to-scan-all behavior. The description restates the scoping semantics but adds no syntax or meaning beyond what the schema provides, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The opening sentence states a specific verb (explain why) and resource (DDS readers/writers on a topic), and frames it as a diagnostic question distinct from siblings like topic_metrics or list_endpoints. An agent can tell exactly what it gets back without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear context for use ('explain why readers and writers do not talk') and scoping instructions ('Pass `topic` to scope to one topic; omit to scan all'), plus a when-not-to-trust caveat about empty results. It stops short of naming alternative sibling tools or an explicit exclusion, so it is clear but not exhaustive.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_topic_infoA

ROS 2 graph only; on a DDS-only setup use list_endpoints. Return info for a single ROS 2 topic. topic must be a fully qualified name, e.g. /cmd_vel. Returns a TopicInfo with mode_effective (live or mock) and, in live mode, the publishers' qos_reliability (reliable / best_effort / mixed) and qos_durability (volatile / transient_local / mixed; transient_local marks a latched topic such as /tf_static). Raises an MCP error if the topic name is malformed, the topic is unknown to the active graph, or no ros2 CLI is available. Read-only; no side effects.

ParametersJSON Schema
NameRequiredDescriptionDefault
topicYesFully qualified ROS2 topic name starting with `/`, e.g. `/cmd_vel` or `/camera/image_raw`. Each `/`-separated segment must start with a letter or underscore and contain only letters, digits, and underscores; everything else (whitespace, quotes, shell metacharacters, `//`, trailing `/`) is rejected before reaching the `ros2` CLI.

Output Schema

ParametersJSON Schema
NameRequiredDescription
nameYesFully qualified topic name, e.g. `/cmd_vel`.
qos_profileNoEffective DDS QoS profile for this topic when resolvable. `None` from the ROS2 CLI adapter or when the DDS module is inactive. The DDS module populates this on a best-effort basis (picks one representative endpoint if reader/writer QoS differ).
message_typeYesROS2 message type, e.g. `geometry_msgs/msg/Twist`.
reader_countNoDDS reader-endpoint count when the active backend can resolve endpoint-level info (Cyclone / Fast DDS). `None` from the ROS2 CLI adapter or when the DDS module is inactive.
writer_countNoDDS writer-endpoint count when the active backend can resolve endpoint-level info. `None` from the ROS2 CLI adapter or when the DDS module is inactive.
mode_effectiveYesRuntime mode the adapter served this response in: `live` (real ROS2 introspection) or `mock` (deterministic fixtures). Lets a caller tell a real graph from a demo one without calling `health_check`.
qos_durabilityNoDurability announced by the topic's publishers: `volatile`, `transient_local` (late subscribers receive the last samples; typical of latched topics such as `/tf_static`), or `mixed` when publishers disagree. `null` when unknown, with the same rules as `qos_reliability`.
publisher_countYesPublishers known to the graph.
qos_reliabilityNoReliability announced by the topic's publishers: `reliable`, `best_effort`, or `mixed` when publishers disagree. `null` when unknown: the topic has no publisher, or the value was not read (`list_topics` does not read QoS; `get_topic_info` does).
subscriber_countYesSubscribers known to the graph.

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full behavioral burden and does so: it discloses the live/mock mode distinction, the specific QoS fields returned, and the exact error conditions (malformed name, unknown topic, missing ros2 CLI). It also states 'Read-only; no side effects,' which is the safety profile an agent needs and is not contradicted by any annotation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Dense and front-loaded: the scope/routing constraint comes first, then the argument, then the return shape, then failure modes. Every clause adds information, though the detailed enumeration of TopicInfo fields partially duplicates the output schema and could be trimmed.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

All inputs, routing, failure modes, and the read-only contract are covered for a single-argument lookup tool. With an output schema present, the return-value prose is a bonus rather than a necessity, so nothing an agent needs is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the fully-qualified-name rule and character restrictions are already documented in the input schema. The description restates the format with the same `/cmd_vel` example rather than adding syntax or resolution behavior beyond it, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Return info for a single ROS 2 topic') and names the boundary condition against a sibling ('ROS 2 graph only; on a DDS-only setup use `list_endpoints`'). An agent can distinguish it from topic_metrics or list_topics without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly names the alternative tool and the condition that selects it (DDS-only setups -> list_endpoints), which is strong routing guidance. It doesn't cover adjacent siblings like topic_metrics, and the 'single topic' scope is only implicit in 'Return info for a single ROS 2 topic', so it falls just short of fully explicit when/when-not coverage.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

health_checkA

Report TopicForge environment state as a HealthReport: effective runtime mode (live or mock), ros_backend and dds_backend, ros_tools_available, ros2_available, ros2_distro (fed by the ROS_DISTRO env var), dds_domain_id and observed_domain_note (only the DDS domain joined at startup is observed; programs on other domains are invisible), the server version and the server-side sample cap. Reading mode: live with ros_backend none means the DDS tools are live and the ROS 2 tools are not available (a DDS-only setup: use list_endpoints for topics and wiring). With dds_backend none, dds_inactive_reason says why: backend not selected, binding not installed, or adapter failed to start. payload_decoding is disabled: DDS user-topic payloads are not decoded. dds_security is not_supported: on a secured domain participants show up but protected endpoints and data do not. For a live DDS backend it also reports dds_domain_id, observer_started_ns and now_ns (how long TopicForge has been watching: nothing before observer_started_ns was observed) and the discovery tracker status tracker_running / tracker_passes / tracker_errors / tracker_last_pass_ns / tracker_cache_evictions (errors or evictions above 0, or a stale last pass, mean the discovery data has gaps). Always succeeds: call it first when something looks wrong. Read-only; no side effects.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
modeYesRuntime mode of the adapter actually serving requests: `mock` or `live`. `live` with `ros_backend` `none` means the DDS tools are live and the ROS 2 tools are not available. Can differ from `requested_mode` when a live backend could not start (e.g. `live` requested without `ros2` installed falls back to `mock`).
now_nsNoServer wall-clock time (ns since epoch) when this report was built.
dds_backendNoDDS backend of the adapter actually serving requests. `none` when the DDS module is not active (default for ROS2-only installs). `mock` for synthetic fixtures. `cyclone` requires `pip install "topicforge[dds-cyclone]"` (Eclipse CycloneDDS); `fast` requires a Fast DDS Python binding built from eProsima sources (not on PyPI); `opendds` and `dust` are permanent stub adapters that never serve.
ros2_distroNoValue of `ROS_DISTRO` if set in the environment. **Env disclosure, by design**: under the local-trust threat model (see README 'Security model'), the MCP client is a trusted agent on a machine the user controls, and exposing the ROS2 distro lets it adapt to e.g. `humble`/`jazzy` differences. For a hosted multi-tenant TopicForge endpoint this field would be scrubbed .
ros_backendNoActive ROS2 backend. `ros2_cli` when the `ros2` CLI is on PATH and live mode resolves to a Ros2CliAdapter (alone or as the ROS half of a composite). `mock` when MockAdapter serves the ROS surface. `none` when no ROS2 backend is active (e.g. DDS-only live install with no `ros2` CLI). Together with `dds_backend` it tells the ROS2 and DDS halves of the runtime apart.
dds_securityNoDDS Security is not handled. On a secured domain TopicForge can show participants but not protected endpoints or data.
dds_domain_idNoDDS domain id observed when the DDS module is active.
requested_modeYesMode requested via configuration (may be `auto`).
ros2_availableYesWhether a `ros2` CLI is on PATH.
server_versionYesTopicForge server version (matches the PyPI release of the `topicforge` package).
tracker_errorsNoDiscovery tracker passes that raised (swallowed and logged). Non-zero means gaps.
tracker_passesNoCompleted discovery tracker passes since start.
tracker_runningNoWhether the continuous discovery tracker thread is alive (Cyclone). `None` when the backend has no tracker.
max_sample_countYesServer-side cap on the number of samples returned per `sample_messages` call. Requests above this limit are silently clamped; the value is exposed here so a client can size its requests proactively. Constant within a given server version.
payload_decodingNoWhether DDS user-topic payloads are decoded. `disabled` today: `peek_dds_samples` and `topic_metrics` do not return message content for user topics.
dds_inactive_reasonNoWhy `dds_backend` is `none` while the ROS 2 CLI serves: the backend was not selected (`TOPICFORGE_DDS_BACKEND` unset or `mock`), its Python binding is not installed, or the binding is installed but the adapter failed to start. `null` when a DDS backend is serving or the cause is not known.
observer_started_nsNoWall-clock time (ns since epoch) when the DDS observer joined the bus. Nothing earlier than this was watched: `now_ns` minus this is how long TopicForge has been observing. `None` without a live DDS observer.
ros_tools_availableNoTrue when the ROS 2 tools (`list_topics`, `get_topic_info`, `sample_messages`, `analyze_bag`, `peek_bag_samples`) can run, i.e. `ros_backend` is not `none`. False on a DDS-only setup: use `list_endpoints` for topics and wiring there.
middleware_availableNoTrue when a DDS backend is serving (`dds_backend` is not `none`). When the DDS module is inactive (`dds_backend == 'none'`), whether the *configured* backend's Python bindings are importable, so a missing binding is visible.
observed_domain_noteNoPlain statement of which DDS domain is observed, set when a DDS module is active: only the domain joined at startup is visible, a program on another domain is invisible.
tracker_last_pass_nsNoWall-clock time (ns since epoch) of the last completed tracker pass. A value far older than `now_ns` means lifecycle is stale.
payload_decoding_reasonNoOne-line reason for `payload_decoding`.
tracker_cache_evictionsNoDiscovery entries dropped because a tracker cache was full (4096 per cache). Non-zero means the bus is bigger than what is listed.

TDQS

A4.3/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and delivers: always succeeds, read-only, no side effects, the observation window (nothing before observer_started_ns was observed), domain visibility limits (only the DDS domain joined at startup), payload_decoding disabled, dds_security not_supported, and how to interpret tracker gaps.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the purpose, then layered interpretation guidance. It is dense and verbose for a no-arg tool, and some field enumeration overlaps the existing output schema, but the interpretation rules (mode, gap meanings) earn their space.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be re-explained, yet the description adds interpretation guidance on top. Combined with the always-succeeds and read-only guarantees, nothing an agent needs to call this correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters and has full schema description coverage, so the baseline of 4 applies. There is no parameter surface needing explanation in the description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource: reporting the TopicForge environment state as a HealthReport, enumerating the fields it covers. It is clearly the diagnostic/health tool among the data-inspection siblings, but it never names a sibling to contrast against, so differentiation is implicit.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

"Always succeeds: call it first when something looks wrong" gives a clear when-to-use trigger. It offers no explicit alternatives or exclusions (e.g., "use X to diagnose Y"), so it stops short of full routing guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_endpointsA

List every DDS endpoint (writer and reader) announced on the bus, one EndpointInfo per endpoint with role, topic, type_name, type_id, the owning participant_guid joined with its participant_name, and a structured qos (reliability, durability, history, deadline, liveliness kind and lease, ownership kind and strength, partitions, latency budget, destination order, data representation). Use it instead of parsing peek_dds_samples output and joining GUID prefixes. Spotting orphans: by_topic rolls the endpoints up per topic with writer_count, reader_count and orphan ("no_reader" = a writer nobody subscribes to, "no_writer" = a reader nobody publishes to), plus the union of partitions. Reading qos: a duration of None (deadline_ns, liveliness_lease_ns, latency_budget_ns) means infinite or not set; a policy field of None means the endpoint did not announce it. Ownership: among EXCLUSIVE writers the live one with the highest ownership_strength delivers to a reader; which writer currently owns an instance is reader-side runtime state TopicForge cannot observe. Departed endpoints: when a participant leaves, its endpoints are remembered (last 200, 1 h) and shown in by_topic as departed_writers / departed_readers (participant name and gone_ns), so a topic that lost its only writer is explained in one call; they are listed in endpoints only with include_departed. Topic filter: rt/scan and scan match each other (exact name first; note says which form matched), and a filter that matches nothing returns a note with the closest known topics. announced_ns is the discovery announcement's source timestamp on the announcing side's clock, which can differ from this host's clock. This lists discovery facts, not data flow: it shows which endpoints exist and how they are configured, not whether samples move. activity is always None (see activity_note): TopicForge holds no reader on user topics and cannot tell a silent or hung writer from a healthy one. Pair it with detect_qos_mismatches to see which pairs cannot match. TopicForge's own observer participant is excluded unless include_observer is true. Output is capped at 500 endpoints (truncated, total_discovered); by_topic still covers all matches. Read-only. Raises an MCP error when no DDS module is active. Mock mode returns a fixture matching the other mock DDS tools.

ParametersJSON Schema
NameRequiredDescriptionDefault
topicNoOnly endpoints on this DDS topic name: a bare name such as `scan` or a ROS 2 mangled name such as `rt/scan` (exact name first, then the alternate form). Omit to list every topic.
domain_idNoAccepted for compatibility (0..232). TopicForge observes the domain it joined at startup (TOPICFORGE_DDS_DOMAIN_ID); this argument does not switch domains, and the response `domain_id` says which one was observed.
include_departedNoAlso list endpoints whose participant left the bus (flagged with `gone_ns`). Defaults to false; `by_topic` reports them as `departed_writers` / `departed_readers` either way.
include_observerNoInclude TopicForge's own observer participant's endpoints. Defaults to false.
participant_guidNoOnly endpoints owned by this participant, as the guid reported by `list_participants` or by an earlier `list_endpoints`. Omit for all participants.

Output Schema

ParametersJSON Schema
NameRequiredDescription
noteNoHint when a `topic` filter matched nothing: names the closest known topics.
by_topicYesRoll-up over every matching endpoint.
returnedYesLength of `endpoints`.
domain_idYesDDS domain observed.
endpointsYesMatching endpoints, capped (see `truncated`).
truncatedYesTrue when matching endpoints exceeded the cap.
snapshot_nsYesWall-clock time of the snapshot, ns since epoch.
observer_guidNoGUID of TopicForge's own participant, `None` in mock.
mode_effectiveYesRuntime mode the adapter served this response in: `live` (real ROS2 introspection) or `mock` (deterministic fixtures). Lets a caller tell a real graph from a demo one without calling `health_check`.
total_discoveredYesEndpoints in the discovery cache before any filter.
departed_endpointsNoDeparted endpoints (their participant left) matching the filters. They are in `endpoints` only with `include_departed`; `by_topic` always carries them as `departed_writers` / `departed_readers`.
excluded_observer_endpointsNoEndpoints of TopicForge's own observer participant left out of `endpoints` (they are counted in `total_discovered`). Explains `total_discovered` vs `returned` together with the filters.

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden and does so thoroughly: read-only, raises an MCP error when no DDS module is active, mock mode returns a fixture, output capped at 500 endpoints with `truncated`/`total_discovered`, departed endpoints retained (last 200, 1 h), `activity` is always None and why, observer participant excluded unless `include_observer`, and how `None` duration/policy fields should be read. These are exactly the non-obvious traits an agent needs.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is long, but the bold-labeled sections (Spotting orphans, Reading qos, Ownership, Departed endpoints, Topic filter, This lists discovery facts) make it scannable and the core purpose is front-loaded in the first sentence. Nearly every sentence conveys a non-obvious behavior, though the volume is close to the upper bound of what is proportionate.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return-value documentation is not strictly required, yet the description still maps the fields, semantics of `None`, and the `by_topic` roll-up. Combined with the failure mode, cap behavior, and observer exclusion, an agent has everything needed to call and interpret this tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds real meaning: the topic-matching rule (`rt/scan` and `scan` match each other, exact name first, with a `note` indicating which form matched, and a closest-known-topics note on no match), the fact that `domain_id` does not switch domains, and that participant_guid composes from list_participants output. It goes beyond what the schema states.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The opening sentence gives a specific verb and resource ('List every DDS endpoint (writer and reader) announced on the bus') and enumerates the per-endpoint payload (role, topic, type_name, participant_guid joined with participant_name, structured qos). It explicitly separates itself from sibling `peek_dds_samples` by saying to use this 'instead of parsing `peek_dds_samples` output and joining GUID prefixes.' An agent can distinguish it from list_participants/list_topics without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It names an explicit alternative and the reason to prefer this tool ('instead of parsing `peek_dds_samples` output'), and points to a complementary tool ('Pair it with `detect_qos_mismatches` to see which pairs cannot match'). It also explains when the orphan and departed-endpoint views matter. There is no explicit 'do not use this when X' exclusions against other siblings, so it stops short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_participantsA

List DDS participants observed on the bus. Returns list[ParticipantInfo]: each entry carries guid, vendor (cyclone/fast/rti/rti_micro/opensplice/opendds/coredx/intercom/dust/mock/unknown) with vendor_source, optional name (announced EntityName QoS, e.g. lidar_driver), optional hostname, domain_id, is_observer and mode_effective (live/mock). Why vendor can be unknown: the vendor is read from the participant GUID prefix (vendor_source guid_prefix; none when unknown). Some vendors, e.g. Dust DDS and RTI Connext, do not put their vendor id there, and the Cyclone Python binding does not expose the RTPS header vendor id, so those participants are listed as unknown. is_observer is true for TopicForge's own read-only participant, which is listed like any other. Lifecycle fields: status (active/left), first_seen_ns / last_seen_ns (TopicForge's local clock), seen_count, announced_ns (DDS source timestamp of the announcement), and once left lost_ns + lost_time_source. lost_ns is an upper bound of when the participant died: exact after a clean shutdown, the lease expiry after a crash (the two cannot be told apart), so a crashed process died up to one lease before it (10 s Cyclone default, 20 s Fast DDS, 100 s RTI; the dead participant's lease, not ours). Cyclone tracks discovery continuously in the background, so these stay correct between calls; right after server start the call waits up to 3 s for discovery to warm up. Only the domain joined at startup is observed (see health_check dds_domain_id): a participant on another DDS domain is INVISIBLE here, so a missing participant may be on a different domain; domain_id does not switch domains (restart with TOPICFORGE_DDS_DOMAIN_ID). Works at the raw DDS layer beneath ROS, so it also sees non-ROS participants. Read-only by architecture: it cannot publish, modify QoS, or alter the bus. Raises an MCP error when no DDS module is active (install pip install topicforge[dds] and set TOPICFORGE_DDS_BACKEND=cyclone). The mock backend returns fixtures.

ParametersJSON Schema
NameRequiredDescriptionDefault
domain_idNoAccepted for compatibility (0..232). TopicForge observes the domain it joined at startup (TOPICFORGE_DDS_DOMAIN_ID); this argument does not switch domains, and the response `domain_id` says which one was observed.

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden and does so richly: read-only by architecture, raises an MCP error without a DDS module, waits up to 3 s for discovery warm-up, mock backend returns fixtures, and it discloses that `lost_ns` is only an upper bound with per-vendor lease timing. These are exactly the traits an agent cannot get from structured fields.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with purpose and return type, then bolded sub-sections keep the dense material navigable. It is long, and the exhaustive vendor enum restates what the output schema already lists, but nearly every sentence carries substantive semantic information rather than filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values needn't be explained, yet the description supplies the semantic layer the schema can't (why `vendor` can be `unknown`, what `is_observer` means, how lease expiry bounds `lost_ns`). Error conditions, domain limits, and warm-up behavior round out everything needed to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and the schema already explains that `domain_id` is accepted only for compatibility. The description echoes this and adds a small diagnostic hint (a missing participant may be on another domain), but does not add format or syntax meaning beyond the schema, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('List DDS participants observed on the bus') and immediately names the return type (`list[ParticipantInfo]`). An agent can distinguish it from siblings like list_endpoints, list_topics, and participant_events from the first sentence alone.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear context: it observes only the startup domain, participants on other domains are invisible, and it redirects to `health_check` for `dds_domain_id`. It also explains it operates beneath ROS so non-ROS participants appear. It stops short of explicitly contrasting with the closest sibling (participant_events for history vs. this for current state).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_topicsA

ROS 2 graph only; on a DDS-only setup use list_endpoints. List every ROS 2 topic on the current graph (or the mock graph in mock mode). Returns list[TopicInfo]: each entry carries name, message_type, publisher_count, subscriber_count, and mode_effective (live or mock) to tell a real graph from fixtures. Live mode leaves qos_reliability and qos_durability null here: call get_topic_info for a topic's QoS. Empty list when the graph has no topics or when live discovery times out. Raises an MCP error when no ros2 CLI is available (DDS-only setup). Read-only; no side effects.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and delivers: read-only/no side effects, mock vs live mode disclosure via mode_effective, the empty-list outcome on timeout or empty graph, the MCP-error outcome when no ros2 CLI exists, and the null-QoS-fields caveat. This is unusually rich failure-mode and state context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the routing constraint before the return description, and every clause adds information (return fields, null caveat, empty/error states). It is dense and slightly over-loaded with back-to-back bolded outcomes, but nothing is genuinely wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no annotations and a simple zero-param read tool with an output schema, the description still covers safety (read-only), both non-success outcomes (empty list, MCP error), mock-mode nuance, and pointers to siblings for QoS and DDS-only setups. Nothing an agent needs to invoke and interpret it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so there is nothing for the description to disambiguate and the baseline is 4. No parameter semantics are needed or missing.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('List every ROS 2 topic on the current graph') and immediately scopes it ('ROS 2 graph only'), naming the sibling to use otherwise (list_endpoints). An agent can distinguish it from neighbors like list_participants or list_endpoints without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit when-not routing: 'on a DDS-only setup use list_endpoints', plus a follow-up alternative for QoS ('call get_topic_info for a topic's QoS'). Both the primary alternative and the sub-case escalation path are named with their selecting conditions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

participant_eventsA

Return DDS participant lifecycle events (discovered / lost) from a recent window, e.g. 'who was on the bus 5 minutes ago and left?' or 'when did this participant first appear?'. Returns list[ParticipantEvent]: each entry carries guid, event_type, vendor, timestamp_ns (wall-clock ns since epoch), time_source, observed_ns, optional name (the participant's announced DDS name), optional hostname, domain_id, and mode_effective (live/mock). time_source says what timestamp_ns is: dds_source_timestamp (the DDS timestamp of the announcement or dispose) or observed_local (when TopicForge noticed, weakest). observed_ns is when TopicForge noticed, always at or after a DDS-derived timestamp_ns. Crash caveat: a lost timestamp is an upper bound of the death. After a clean shutdown it is exact; after a crash it is when the lease expired, so the process died between timestamp_ns minus the dead participant's lease and timestamp_ns (10 s Cyclone default, 20 s Fast DDS, 100 s RTI), and the two cases cannot be told apart. A restarted node is a new participant: expect one lost and one discovered per restart, with different guids and the same name. Sorted newest-first. Capped at 200 events, silently (reduce lookback_seconds if you hit it). TopicForge only knows what happened since it started watching (see health_check.observer_started_ns). Backend caveats: Fast DDS captures arrivals and removals through listener callbacks; Cyclone tracks discovery in the background (a pass every 0.5 s, independent of tool calls), so restarts and crashes are recorded as they happen, but a participant cycle faster than the discovery reader's history depth between two passes can be missed; mock returns a fixture timeline. Right after server start the call waits up to 3 s for discovery to warm up. Read-only by architecture. Raises an MCP error when no DDS module is active (install pip install topicforge[dds] and set TOPICFORGE_DDS_BACKEND=cyclone|fast).

ParametersJSON Schema
NameRequiredDescriptionDefault
domain_idNoAccepted for compatibility (0..232). TopicForge observes the domain it joined at startup (TOPICFORGE_DDS_DOMAIN_ID); this argument does not switch domains, and the response `domain_id` says which one was observed.
lookback_secondsNoWindow (in seconds) over which to return events. Defaults to 300 (5 minutes). Range: 1..86400 (1 second to 24 hours). Larger windows may hit the 200-event cap: narrow the window or filter on `domain_id` when that happens.

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations, so the description carries the full burden and does so richly: crash caveat with lease bounds per backend, restart semantics (one lost + one discovered per restart), 200-event silent cap, warm-up wait, backend-specific capture behavior, mock fixture mode, and the MCP error condition with required install/env settings.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Long but dense and front-loaded: purpose first, then return shape, then semantics, then caveats. Nearly every sentence carries a unique fact, though the field-by-field enumeration of the return type overlaps with the existing output schema and could be trimmed.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given output schema exists, return values needn't be explained, yet the description goes further with time_source semantics (dds_source_timestamp vs observed_local, weakest) and observed_ns ordering. Combined with crash/restart/backend caveats, an agent has everything needed to call and interpret it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents both parameters including the 200-event cap interaction with lookback_seconds and the domain_id compatibility note. The description's mention of reducing lookback_seconds largely repeats the schema, so baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Return DDS participant lifecycle events (`discovered`/`lost`)') with explicit scope (recent window). It is clearly distinguishable from siblings like list_participants (current snapshot) and health_check (observer metadata), which it cross-references.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides concrete motivating questions ('who was on the bus 5 minutes ago and left?', 'when did this participant first appear?') that map directly to invocation. However, it never explicitly contrasts with the nearest sibling, list_participants, nor states when-not to use it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

peek_bag_samplesA

Peek up to count samples from a recorded bag file. Unlike peek_dds_samples (live DDS) and sample_messages (live ROS 2 graph), this reads offline bag content for post-mortem analysis. Supported formats: MCAP (.mcap), ROS 2 rosbag2 SQLite (.db3), ROS 1 legacy chunked binary (.bag), detected from the file extension. Returns a SampleResult in the same shape as peek_dds_samples: each sample's payload carries a _decode_status annotation (full / partial / raw). count defaults to 5 and is silently clamped to 50. Requires the rosbags library (pip install topicforge[bags]) and the ROS 2 side of the runtime: on a DDS-only setup it raises an error. The mock backend returns fixture samples on canned bag paths. Bags that embed no message definitions (rosbag2 .db3 from Humble) are decoded with the type definitions of the bag's recorded distro, or Humble when it records none; note says which. Arrays over 4096 elements are cut and note lists the fields. Read-only by architecture: nothing writes to the bag file. Raises an MCP error when the bag path does not exist, the topic is not present in the bag, or rosbags is not installed.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathYesPath to a bag: a file ending in `.mcap` or `.db3`, or a `rosbag2_*` directory. `peek_bag_samples` also reads ROS 1 `.bag` files; `analyze_bag` does not (`ros2 bag info` cannot open them). Leading/trailing whitespace is stripped. Null bytes and otherwise malformed filesystem paths are rejected. Existence and bag format are validated by the live adapter (mock mode accepts any well-formed path).
countNoMaximum number of recent messages to return. Defaults to 5; silently clamped to 50 (the hard cap that keeps tool output bounded; read it from `health_check.max_sample_count`). Negative values raise an error. The returned `SampleResult.count` reflects the actual number of samples produced: it can be lower than the request (empty topic, timeout, mock fixture shorter than requested).
topicYesFully qualified ROS2 topic name starting with `/`, e.g. `/cmd_vel` or `/camera/image_raw`. Each `/`-separated segment must start with a letter or underscore and contain only letters, digits, and underscores; everything else (whitespace, quotes, shell metacharacters, `//`, trailing `/`) is rejected before reaching the `ros2` CLI.

Output Schema

ParametersJSON Schema
NameRequiredDescription
noteNoWhy `samples` is empty or limited, when the cause is not obvious (for example payload decoding is disabled for DDS user topics). `None` when there is nothing to add.
countYesNumber of samples actually returned. May be 0 (no publisher active in live mode, or empty mock fixture), less than the requested count (topic yielded fewer messages within the timeout), or capped by the the silent maximum of 50: request `count > 50` and you will receive at most 50 without warning.
topicYesTopic the samples were taken from, echoed from the request.
samplesYesThe sampled messages, ordered as received from the backend.
mode_effectiveYesRuntime mode the adapter served this response in: `live` (real ROS2 introspection) or `mock` (deterministic fixtures). Lets a caller tell a real graph from a demo one without calling `health_check`.

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and delivers: read-only by architecture, library dependency, explicit MCP error conditions (missing path, absent topic, missing library), silent clamping to 50, mock-backend behavior, decode-status annotation shape, and array truncation with `note`. This is the behavioral context an agent needs before calling.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads purpose and the sibling differentiation before drilling into formats and error behavior, and every sentence carries operational information. It is dense and long, but little of it is redundant.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given a rich output schema, the description need not explain return values, and it still covers the reader's prerequisites, error cases, decode-status semantics, mock behavior, and format support. Nothing an agent needs to call it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds cross-parameter error semantics not in the schema (raises an MCP error when the bag path does not exist or the topic is not present) and reinforces the format-detection-by-extension behavior tied to `path`.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Peek up to `count` samples from a recorded bag file') and immediately distinguishes itself from two named siblings, `peek_dds_samples` (live DDS) and `sample_messages` (live ROS 2 graph). An agent can tell exactly what this does without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly frames the use case ('reads offline bag content for post-mortem analysis') and names the alternatives it is not, with the condition that selects each. It also discloses the environment prerequisite (rosbags library, ROS 2 side of the runtime) and that it errors on a DDS-only setup.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

peek_dds_samplesA

Peek recent samples on a raw DDS topic. Unlike sample_messages (which uses the ros2 CLI), this reads the DDS layer directly and works without ROS 2. Returns a SampleResult {topic, count, samples, mode_effective, note}, the same shape as sample_messages. count defaults to 5 and is silently clamped to 50. Topic categories: (a) The 3 builtin discovery topics (DCPSParticipant, DCPSSubscription, DCPSPublication) return structured discovery payloads: on Cyclone the CURRENT discovery state (one record per live participant or endpoint), not a stream of recent events (use participant_events for history). DCPSPublication and DCPSSubscription are the raw writers and readers behind list_endpoints. The topic may be given as /scan, scan or rt/scan: all three resolve to the same topic. (b) User-defined topics: payload decoding is DISABLED on every backend. The call returns count 0, samples empty and a note saying so; that does NOT mean the topic is silent. Use list_endpoints for the topic's presence, writers, readers and QoS. A user topic that is not announced on the bus raises an error. Right after server start the call waits up to 3 s for discovery to warm up. Read-only by architecture: it cannot publish. Raises an MCP error when no DDS module is active or the topic is not announced on the bus.

ParametersJSON Schema
NameRequiredDescriptionDefault
countNoMaximum number of recent messages to return. Defaults to 5; silently clamped to 50 (the hard cap that keeps tool output bounded; read it from `health_check.max_sample_count`). Negative values raise an error. The returned `SampleResult.count` reflects the actual number of samples produced: it can be lower than the request (empty topic, timeout, mock fixture shorter than requested).
topicYesDDS topic name. Bare DDS names such as `scan` are valid, as are ROS 2 mangled names such as `rt/scan` and the builtin discovery topics `DCPSParticipant`, `DCPSSubscription`, `DCPSPublication`. Letters, digits, `_`, `/` and `::` are allowed; anything else is rejected.

Output Schema

ParametersJSON Schema
NameRequiredDescription
noteNoWhy `samples` is empty or limited, when the cause is not obvious (for example payload decoding is disabled for DDS user topics). `None` when there is nothing to add.
countYesNumber of samples actually returned. May be 0 (no publisher active in live mode, or empty mock fixture), less than the requested count (topic yielded fewer messages within the timeout), or capped by the the silent maximum of 50: request `count > 50` and you will receive at most 50 without warning.
topicYesTopic the samples were taken from, echoed from the request.
samplesYesThe sampled messages, ordered as received from the backend.
mode_effectiveYesRuntime mode the adapter served this response in: `live` (real ROS2 introspection) or `mock` (deterministic fixtures). Lets a caller tell a real graph from a demo one without calling `health_check`.

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden and does so richly: read-only by architecture (cannot publish), raises an MCP error in two specific situations, 3 s discovery warm-up, silent clamping of count, and payload decoding disabled on user topics returning count 0 with a note. This is exactly the behavioral context an agent needs.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Dense and front-loaded – the core distinction from `sample_messages` comes first, then topic categories, then failure modes. It is long, but nearly every sentence carries operational information; minor redundancy exists in restating the `SampleResult` shape that the output schema already covers.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a two-param tool with no annotations, the description covers selection rationale, topic-name handling, per-category return semantics, limits, and error conditions. With an output schema present, the brief restatement of the return shape is not a gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so baseline is 3, but the description adds real meaning: the three accepted topic name forms (`/scan`, `scan`, `rt/scan`) resolve to the same topic, and the builtin discovery topics have distinct return semantics. It also explains the clamping behavior for count, though that is largely mirrored in the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('peek recent samples on a raw DDS topic') and immediately differentiates from the sibling `sample_messages`, explaining that this reads the DDS layer directly and works without ROS 2. An agent can tell it apart from the other sampling/list tools without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly routes the agent: use `participant_events` for discovery history, `list_endpoints` for a user topic's presence/writers/readers/QoS, and `sample_messages` if you need the ros2 CLI. It also states the failure conditions (no DDS module, unannounced topic) and the warm-up wait.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

sample_messagesA

ROS 2 graph only; on a DDS-only setup use list_endpoints (topics and wiring) or peek_dds_samples on DCPSPublication / DCPSSubscription (raw discovery records). Peek up to count recent ROS 2 messages from topic. count defaults to 5 and is silently clamped to 50. Returns a SampleResult {topic, count, samples, mode_effective, note} where count is the actual number of samples returned (may be 0) and mode_effective is live or mock. Live mode runs ros2 topic echo --csv --once with a short timeout, so the result is empty when no publisher is active and at most one message comes back. Arrays: by default ros2 topic echo cuts arrays at 128 elements (a 541-beam LaserScan loses beams 128 and up); the cut is listed under _truncated_after_columns in the sample payload and in note (cut strings and bytes are listed under _truncated_columns). Raise max_array_length (up to 65536, or null for no cut) to read more, or set arrays_summary_only to see only the non-array fields. A message over the 1 MiB size cap (TOPICFORGE_MAX_SAMPLE_BYTES) is dropped and note says so. samples[i].timestamp_ns is the message's header.stamp (publish time) when the message is Header-stamped, and 0 for headerless types (e.g. std_msgs/String). The live parser exposes fields as positional CSV columns under samples[i].payload keys col_0, col_1, ..., with the verbatim CSV row under the reserved _raw_text key. Mock mode returns structured samples for the fictional demo robot. Raises an MCP error when no ros2 CLI is available. Read-only; never publishes. Distinct from peek_dds_samples, which reads the raw DDS layer.

ParametersJSON Schema
NameRequiredDescriptionDefault
countNoMaximum number of recent messages to return. Defaults to 5; silently clamped to 50 (the hard cap that keeps tool output bounded; read it from `health_check.max_sample_count`). Negative values raise an error. The returned `SampleResult.count` reflects the actual number of samples produced: it can be lower than the request (empty topic, timeout, mock fixture shorter than requested).
topicYesFully qualified ROS2 topic name starting with `/`, e.g. `/cmd_vel` or `/camera/image_raw`. Each `/`-separated segment must start with a letter or underscore and contain only letters, digits, and underscores; everything else (whitespace, quotes, shell metacharacters, `//`, trailing `/`) is rejected before reaching the `ros2` CLI.
max_array_lengthNoLongest array, string or bytes value to return in full, 1..65536; longer ones are cut. A cut array is listed in the sample's `_truncated_after_columns` (index of the last kept column); a cut string or bytes value is kept as its first N characters plus `...` and listed in `_truncated_columns`. Defaults to 128, the `ros2 topic echo` default, which cuts a 541-beam `LaserScan` after 128 ranges. Pass null to return everything in full (large for images and point clouds; a message over the server's size cap, 1 MiB by default, is dropped with a note, and a very large message may not print within the echo timeout).
arrays_summary_onlyNoWhen true, array fields are replaced by a short type and length summary instead of their elements. Use it to inspect the non-array fields of large messages (images, scans, point clouds). Defaults to false.

Output Schema

ParametersJSON Schema
NameRequiredDescription
noteNoWhy `samples` is empty or limited, when the cause is not obvious (for example payload decoding is disabled for DDS user topics). `None` when there is nothing to add.
countYesNumber of samples actually returned. May be 0 (no publisher active in live mode, or empty mock fixture), less than the requested count (topic yielded fewer messages within the timeout), or capped by the the silent maximum of 50: request `count > 50` and you will receive at most 50 without warning.
topicYesTopic the samples were taken from, echoed from the request.
samplesYesThe sampled messages, ordered as received from the backend.
mode_effectiveYesRuntime mode the adapter served this response in: `live` (real ROS2 introspection) or `mock` (deterministic fixtures). Lets a caller tell a real graph from a demo one without calling `health_check`.

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does so: read-only/never publishes, live mode runs `ros2 topic echo --csv --once` so results are empty with no active publisher and capped at one message, 1 MiB size cap drops messages with a note, array cut at 128 with `_truncated_after_columns`/`_truncated_columns` markers, and an MCP error is raised when no `ros2` CLI exists.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Dense but front-loaded: the DDS-only routing caveat leads, then the core action, then behavior. Every sentence carries information, though the parameter-behavior detail makes it longer than strictly necessary for selection.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex sampling tool with an output schema, the description covers failure modes, mode_effective semantics, timestamp provenance, CSV positional-column payload shape, and truncation markers. Nothing an agent needs to invoke it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds real semantics: `count` clamping to 50 and the fact that returned `count` may be lower than requested, and how `max_array_length`/`arrays_summary_only` change truncation behavior. It reinforces rather than merely repeats the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource+scope: 'Peek up to `count` recent ROS 2 messages from `topic`', and explicitly delimits itself from `peek_dds_samples` ('Distinct from `peek_dds_samples`, which reads the raw DDS layer'). An agent can select it without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Opens with the routing condition: ROS 2 graph only; on a DDS-only setup use `list_endpoints` or `peek_dds_samples` on `DCPSPublication`/`DCPSSubscription`. It names the alternatives and the precise condition that selects them.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

topic_metricsA

Return temporal metrics (frequency, sequence gaps, latency percentiles) for a DDS topic over a recent time window. Returns a TopicMetrics payload carrying status, samples_observed, frequency_hz_observed, frequency_hz_declared, sequence_gaps_count, latency_ns_p50/p95/p99, and boolean availability flags. Read status first: unsupported_user_topic means the topic is a user topic, whose payload is not decoded, so there are no metrics: every number is null or 0 and none of it is a measurement. no_samples_yet means a builtin topic with nothing buffered in the window. ok means metrics were computed. Limits: the buffer is filled only when peek_dds_samples runs on the topic, so frequency_hz_observed reflects how often it was called, not the real publish rate: treat it as a coarse presence signal. frequency_hz_declared is declared, not measured: 1 / deadline of the shortest QoS Deadline a writer on the topic announced in discovery, null when none announced one. Read-only by architecture. Raises an MCP error when no DDS module is active or window_seconds is out of range (1..3600). Right after server start the call waits up to 3 s for discovery to warm up.

ParametersJSON Schema
NameRequiredDescriptionDefault
topicYesDDS topic name. Bare DDS names such as `scan` are valid, as are ROS 2 mangled names such as `rt/scan` and the builtin discovery topics `DCPSParticipant`, `DCPSSubscription`, `DCPSPublication`. Letters, digits, `_`, `/` and `::` are allowed; anything else is rejected.
domain_idNoAccepted for compatibility (0..232). TopicForge observes the domain it joined at startup (TOPICFORGE_DDS_DOMAIN_ID); this argument does not switch domains, and the response `domain_id` says which one was observed.
window_secondsNoWindow in seconds over which to compute metrics (1..3600). Defaults to 60 seconds. Smaller windows reflect more recent state; larger windows smooth transient anomalies.

Output Schema

ParametersJSON Schema
NameRequiredDescription
topicYesTopic the metrics were computed for.
statusNoHow to read the numbers. `unsupported_user_topic`: the topic is a user topic, whose payload is not decoded, so no metric exists and the null fields are not a measurement. `no_samples_yet`: a supported topic with nothing buffered in the window. `ok`: metrics computed from buffered samples.
latency_ns_p50NoMedian publish-to-receive latency in nanoseconds, computed only when the sample type exposes a publish timestamp (typically via `header.stamp` on `Header`-stamped messages). `None` when `latency_available=False`.
latency_ns_p95No95th-percentile publish-to-receive latency (ns).
latency_ns_p99No99th-percentile publish-to-receive latency (ns).
mode_effectiveYesRuntime mode the adapter served this response in: `live` (real ROS2 introspection) or `mock` (deterministic fixtures). Lets a caller tell a real graph from a demo one without calling `health_check`.
window_secondsYesRequested window in seconds (1..3600). Echoed back from the tool call so the LLM can correlate the request.
samples_observedYesNumber of samples in the buffer matching `topic` within the window. `0` means TopicForge has not seen any sample on this topic recently: it does NOT mean the topic has no publisher, only that no `peek_dds_samples` call captured one in the window.
latency_availableNoTrue when at least one sample in the window carried both a publish timestamp and a receive timestamp. The percentile fields are `None` when this is False.
sequence_gaps_countNoNumber of missing sequence numbers detected in the buffered samples. `0` either means no gaps observed OR the sample type did not expose a sequence number (check `sequence_numbers_available` to disambiguate).
frequency_hz_declaredNoDeclared, not measured: `1 / deadline` for the shortest QoS Deadline period announced by a writer on this topic in discovery. `None` when no writer announced a finite Deadline or the topic is not announced. It is the rate the application promised, not the rate observed.
frequency_hz_observedNo`samples_observed / window_seconds_actual`. `None` when fewer than 2 samples were observed (a single sample does not define a frequency).
window_seconds_actualYesActual elapsed seconds within the window. May be smaller than `window_seconds` when the adapter buffered samples for less time than the requested window (e.g., the server just started). `0.0` when `samples_observed=0`.
sequence_numbers_availableNoTrue when the adapter successfully extracted sequence numbers from at least one sample. Sequence number support depends on the message type: `Header`-stamped messages with a `seq` field expose it; primitives like `std_msgs/String` do not.

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden, and it does so thoroughly: status semantics and their meaning for the numbers, the observability limitation of frequency_hz_observed, the declared-vs-measured distinction for frequency_hz_declared, the read-only architecture, the MCP error conditions (no DDS module, out-of-range window), and the 3 s discovery warm-up delay. This is exactly the behavioral context an agent needs to avoid misinterpreting results.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with purpose, then progressively adds interpretation rules (status), limits, error conditions, and startup caveats using bold markers for scannability. Despite its length, every sentence carries distinct decision-relevant information with no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, yet the description still names the payload fields and, more importantly, explains how to interpret them (status branches, null/0 caveats, declared-vs-measured). For a single-resource compute tool with three fully-documented params, this is complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all three parameters are already fully documented in the schema (including the domain_id no-op explanation and window_seconds range). The description adds only the error condition for an out-of-range window, so the schema does the heavy lifting and a baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Return) and resource (temporal metrics: frequency, sequence gaps, latency percentiles) scoped to a DDS topic over a time window. The named metric set distinguishes it from siblings like peek_dds_samples or get_topic_info, which don't compute derived temporal metrics.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explains the operational context clearly: the buffer is populated only when peek_dds_samples runs, so the metric is a coarse presence signal. It also tells the agent to read `status` first and enumerates the branches. It stops short of an explicit when-to-use-this-vs-X directive against siblings, so it's a strong 4 rather than a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 12 tool updatesv0.5.6
    • First observedanalyze_bag
    • First observeddetect_qos_mismatches
    • First observedget_topic_info
    • First observedhealth_check
    • First observedlist_endpoints
    • First observedlist_participants
    • First observedlist_topics
    • First observedparticipant_events
    • First observedpeek_bag_samples
    • First observedpeek_dds_samples
    • First observedsample_messages
    • First observedtopic_metrics

TDQS

A4.4/5.0

Scored across 12 tools

Disambiguation4/5

The three sampling tools (peek_dds_samples, sample_messages, peek_bag_samples) share the same SampleResult shape and verb, but each description explicitly states its source (raw DDS layer, ros2 CLI, offline bag) and cross-references the others. list_endpoints vs list_topics and peek_dds_samples on DCPSPublication vs list_endpoints also overlap slightly, but the descriptions resolve these with explicit 'use X instead' guidance.

Naming Consistency4/5

All names are snake_case, which is consistent, and most are verb_noun (list_endpoints, list_topics, get_topic_info, sample_messages, analyze_bag, detect_qos_mismatches). A few are noun_noun (topic_metrics, participant_events, health_check), a minor deviation that stays readable and predictable.

Tool Count5/5

12 tools is well within the ideal 3-15 range and each covers a distinct facet of DDS/ROS 2 introspection (discovery, sampling, bags, QoS, metrics, health). No tool looks redundant or filler.

Completeness4/5

The surface covers the topic-centric lifecycle well: participants and events, endpoints and QoS, ROS 2 topic listing/info, sampling from three sources, bag peek/analysis, metrics, and health. Gaps are minor for a topics-focused server (no ROS 2 node/service/action introspection, no time-window history queries beyond events), and read-only is a stated design choice rather than a missing operation.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    A
    quality
    D
    maintenance
    Read-only Modbus TCP monitoring server that exposes safe MCP tools for AI agents to read holding/input registers, coils, and device identity from industrial devices without write access.
    4
    Apache 2.0
  • A
    license
    Not graded
    quality
    C
    maintenance
    Read-only MCP server that exposes Autopsy digital forensics case data as tools, enabling LLM clients like Cline to browse filesystems, query artifacts, and search keywords without modifying the case.
    1
    MIT
  • A
    license
    Not graded
    quality
    C
    maintenance
    A read-only MCP server for bounded semantic inspection of ROS 2 perception systems, enabling discovery and metadata extraction from sensors such as cameras, depth sensors, and LiDARs without device configuration or actuation.
    Apache 2.0