Skip to main content
Glama
paramount-engineering

Roku Dev Studio MCP Server

Platform CI Version Electron License roku-dev-studio MCP server

Roku Developer Tools for macOS, Windows, and Linux — Remote Control, App Side-loading, ECP automation, RALE / App Connector, Network Inspector, Action Scripts, MCP server for AI agents (Cursor, Claude, VS Code), and a rds CLI. Supports both local network and internet-bridged devices.

A comprehensive cross-platform desktop application for controlling and developing on Roku devices over your local network or via remote server using the External Control Protocol (ECP).

Why Roku Dev Studio? (vs. official Roku tools)

Roku development is normally split across a pile of separate, single-purpose official tools — the Roku Remote Tool, the browser-based sideload installer, raw telnet, RALE, sca-cmd — that don't talk to each other. Roku Dev Studio doesn't replace Roku's own protocols (ECP, RALE, telnet, sca-cmd) — it wraps all of them in one GUI, one CLI (rds), and one MCP server:

Task

Without Roku Dev Studio

With Roku Dev Studio

Remote control

The official Roku Remote Tool, or raw ECP keypress calls via curl/Postman

One tab: full D-Pad, keyboard remote, and a floating mini-remote

Sideloading

The device's browser-based installer or a VS Code extension — one IP at a time

Sideload Relay — one push from your IDE installs, launches, and captures console on every targeted device

Debug console

telnet <ip> 8085 in a raw terminal or through an IDE — no search/filter/save either way

A structured console with search, filtering, and saved logs

BrightScript debugging

The socket debug protocol, usable mainly through a single IDE's extension

A standalone debugger: breakpoints, step execution, call stack, variables, watch

App inspection (RALE)

RALE alone only inspects SceneGraph nodes — no way to call into a channel or exchange data with it

App Connector — extends RALE with the ability to call your channel's own functions and pass data back and forth (GET/POST-style), unlocking automation that didn't exist before

Network traffic

A separately configured MITM proxy (Charles/mitmproxy/Fiddler) with manual device setup

Built-in local MITM proxy + optional hotspot packet capture

Static analysis

sca-cmd output cross-referenced by hand against Roku's cert docs

Runs sca-cmd for you, with cert-requirement links straight to Roku's docs

Remote locations / labs

Physical presence required — ECP only works on the local network

A bundled remote server bridges ECP over the internet

Repeatable testing

Hand-rolled scripts around ECP and RALE

Action Scripts — build a flow (keypresses, queries, conditionals, waits) from a GUI, or run it headless via rds

AI-agent access

Nothing official

A bundled MCP server lets Cursor, Claude Desktop, or VS Code drive a real device

Why have I built Roku Dev Studio?

This repository is an npm workspace monorepo. Run npm install and npm start from the repository root so workspaces link correctly. Installing runs a postinstall (npm run build:libs) that compiles the shared roku-dev-studio-platform and roku-dev-studio-api packages to their dist/ outputs, which the app and remote server import. Use npm run typecheck for a full TypeScript check across every workspace and npm test to run unit tests. CI runs these plus per-package build/syntax smoke checks on each push and pull request. Setup, scripts, and distributable builds are documented in INSTALLATION.md.

Related MCP server: brs-docs-mcp

Repository layout

Location

What it is

apps/roku-dev-studio/

Electron desktop app (main process, renderer, packaging). Dev and distributable builds: INSTALLATION.md.

packages/roku-dev-studio-api/

Shared Node library + rds CLI: discovery, ECP, screenshots, sideload, RALE, action-script runner, headless validator — package README.

packages/roku-dev-studio-mcp/

MCP server that lets AI agents (Cursor, Claude Desktop, VS Code) drive a Roku through this app — package README.

packages/roku-dev-studio-network-inspector/

Network Inspector engine: hotspot packet capture (DNS/SNI/HTTP) + local MITM proxy, transport-agnostic so it runs in both the desktop app and the remote server — package README.

packages/roku-dev-studio-rce/

Roku Cloud Emulator (RCE) client — Core API (accounts / devices / snapshots) and Device API (ECP proxy, ports-bridge sockets) for cloud-hosted virtual Rokus, used by the app's RCE locations — package README.

packages/roku-dev-studio-remote-server/

HTTP/WebSocket relay to control Rokus over the internet — package README.

packages/roku-dev-studio-platform/

Shared host-platform helpers (OS identity, modifier keys, path-safe, node-only filesystem helpers) used by the app and other packages so platform logic lives in one place. Built to dist/ on npm install — package README.

roku-components/

BrightScript-side artifacts: TrackerTask.xml (drop into your channel for App Connector / RALE), the fiddle/ SceneGraph scaffold, and demo/ (the bundled Roku Dev Studio Showcase channel behind Try Demo App) — components README.

Author: Hareendra Donapati

Glossary

Term

One-line meaning

ECP

External Control Protocol — Roku's HTTP API on port 8060 (KeyPress, Launch, Query, Deep-Link).

Telnet 8085 / 8080

The BrightScript debug console (8085) and dev system commands (8080) on a Developer-Mode Roku.

RALE

Roku Advanced Layout Editor — Roku's SceneGraph inspection protocol over a TCP socket (default port 49200), spoken by the TrackerTask component.

TrackerTask

The BrightScript component channel developers add to their app to make it reachable from RALE / App Connector — see roku-components/README.md.

App Connector

The Dev Studio tab that talks RALE: list / call your channel's GetExternalControlFunctions, plus built-ins (node lookup, registry editor, update node).

Network Inspector

The Dev Studio tab / engine that inspects a dev channel's HTTP(S) traffic through a local MITM proxy, with optional hotspot packet capture.

Sideload

Uploading and installing a .zip / .pkg dev channel onto a Developer-Mode Roku via its Dev Password.

Sideload Relay

RDS advertising itself as a Roku so one sideload from your IDE / browser fans out (install → launch → console) to many targeted devices.

Action Script

JSON-described automation that chains keypresses, queries, sideload, App Connector calls, screenshots, conditionals, waits, and variables. Built and run from the Action Scripts tab; also runnable headless via rds.

MCP server

Roku Dev Studio's Model Context Protocol server — lets Cursor / Claude Desktop / VS Code drive a real device through this app while it's open. Toggle clients in Settings → MCP Server.

Fiddle

The BrightScript scratch editor (Monaco + brighterscript lint) that wraps your snippet into a temporary channel and runs it on a selected device.

rds

The terminal CLI shipped by roku-dev-studio-api (rds discover, rds keypress, rds script run, rds rale repl, …).

Supported Platforms

Roku Dev Studio is available for:

Platform

Options

macOS

DMG installer, Portable ZIP archive

Windows

NSIS installer, Portable executable

Linux

DEB package, AppImage


Home

Home

Remote + Device Performance

App Connector (RALE)

Action Scripts Builder

Remote with Device Performance

App Connector

Action Scripts Builder

BrightScript Fiddle

MCP Server Settings

Dev App / Sideload

BrightScript Fiddle

Settings MCP Server

Dev App

More screenshots for every feature: FEATURES.md.

Features

See FEATURES.md for the full tour with screenshots. Quick index:

Remote Control (Floating Remote) · Device Performance · Device Discovery · App Launcher & Management · Device Queries · Dev App Management · Try Demo App · Sideload Relay · Console & Debugging · Ports Window · Console Monitor · BrightScript Debugger · App Connector (RALE) · Network Inspector · Network Session Viewer · Action Scripts · AI Agents (MCP Server) · BrightScript Fiddle · Log File Viewer · Static Channel Analysis · rds CLI · Remote Server Support · Settings · Language Switching · Crash Reporting · Developer Features

Remote Server Setup

Roku Dev Studio can control devices over the internet using a remote server bridge, so you can manage devices in Remote Locations without being on the same network as the desktop app. Run the relay (npm run remote-server from this repo, or npm install -g roku-dev-studio-remote-server), then add it via Add Remote Location in the device selector. The modal has two tabs: RDS Relay (Relay Server address + port) and RCE (a Roku Cloud Emulator account name + Personal Access Token). RCE devices list as shutdown / pending / running and must be started first (Start, with optional snapshot / firmware / Max Run Time options) — ECP, sideload and console only respond while a device is running.

Full setup (running the server as a service, network/firewall configuration, the HTTP/WebSocket API, and Swagger docs) lives in the remote server package README.

Project structure

.
├── apps/
│   └── roku-dev-studio/                 # Electron desktop app (see INSTALLATION.md)
├── packages/
│   ├── roku-dev-studio-api/             # Shared API + `rds` CLI (npm: roku-dev-studio-api)
│   ├── roku-dev-studio-mcp/             # MCP server bundled into the desktop app
│   ├── roku-dev-studio-network-inspector/ # Network capture + MITM proxy engine
│   ├── roku-dev-studio-rce/             # Roku Cloud Emulator client (accounts, devices, ECP proxy)
│   ├── roku-dev-studio-platform/        # Shared platform helpers (path-safe, OS identity)
│   └── roku-dev-studio-remote-server/   # HTTP/WS relay (npm: roku-dev-studio-remote-server)
├── roku-components/                     # TrackerTask + Fiddle SceneGraph assets
├── package.json                         # Workspace root (workspaces: apps/*, packages/*)
├── INSTALLATION.md
└── README.md

The Electron app’s own tree (TypeScript main.ts / preload.ts bundled to main.bundled.cjs / preload.bundled.cjs, renderer/, build assets) lives under apps/roku-dev-studio/.

Requirements

For Running the App:

  • Node.js 24.17+

  • npm (bundled with Node.js)

  • Roku device on local network (or remote server for remote access)

For Building:

  • All of the above

  • Platform-specific build tools:

    • macOS: Xcode Command Line Tools

    • Windows: Windows SDK (for NSIS installer)

    • Linux: Standard build tools (gcc, make, etc.)

See Installation for setup and build instructions.

License

This project is licensed under the MIT License.

Third-party components used in this software and their licences:

Library

Purpose

Licence

@tanstack/virtual-core

Virtualized list rendering (telnet console, large script results)

MIT

archiver

Building sideload .zip packages

MIT

brighterscript

BrightScript linting in the Fiddle editor

MIT

commander

rds CLI argument parsing

MIT

electron

Desktop app runtime

MIT

electron-builder

Packaging & installers

MIT

form-data

HTTP multipart uploads

MIT

modern-screenshot

DOM-to-image capture for chart cards / PDF export

MIT

monaco-editor

Code editor (Fiddle, action-script step editors)

MIT

pdf-lib

PDF generation

MIT

sharp

Image processing (icons/build)

Apache-2.0

solid-js

Reactive framework powering the new renderer

MIT

ws

WebSocket client

MIT

Their dependencies are used under the terms declared in package-lock.json and each package’s repository.

Available Tools

51 tools
app_connector_connectApp Connector: ConnectA
Idempotent

Open a RALE / App Connector session against the device's running Dev App. Mutates session state (establishes a connection), so it is not read-only; idempotent — reconnecting an open session is a no-op. You rarely need to call this explicitly: rale_command, app_function, and rale_get_node_by_id auto-connect on demand. Use it only to pre-warm the session or surface connection errors early.

ParametersJSON Schema
NameRequiredDescriptionDefault
deviceNoOptional target device (IP or serial). Omit to use the focused Dev Studio tab.

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description goes beyond the annotations by explaining that the tool mutates session state, is idempotent with reconnection being a no-op, and that other tools auto-connect on demand. It adds useful behavioral context without contradicting the provided readOnlyHint, openWorldHint, idempotentHint, or destructiveHint annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is three tightly written sentences with no filler: it front-loads the purpose, then covers side effects and idempotency, then gives usage guidance. Every sentence contributes distinct information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple one-parameter connect tool with rich annotations and no output schema, the description covers purpose, side effects, idempotency, auto-connect behavior, and explicit usage conditions. An agent has enough context to decide when to call it and what it does.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema already provides full coverage for the single optional 'device' parameter, including the behavior of omitting it (uses the focused Dev Studio tab). The description does not add param-level details, but none are needed given the schema's completeness.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Open a RALE / App Connector session against the device's running Dev App.' It clearly identifies what the tool does and implicitly distinguishes it from connection-disconnect and command-execution siblings such as app_connector_disconnect and rale_command.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage guidance is explicit and practical: it says the tool is rarely needed because rale_command, app_function, and rale_get_node_by_id auto-connect on demand, and it states the only recommended uses are pre-warming the session or surfacing connection errors early. This gives clear when-to-use and when-not-to-use direction.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

app_connector_disconnectApp Connector: DisconnectA
Idempotent

Close the RALE / App Connector session on the targeted device. Mutates session state (tears down the connection), so it is not read-only; idempotent — closing an already-closed session is a no-op. Use it to free the session or force a clean reconnect; normal RALE tools do not require you to disconnect between calls.

ParametersJSON Schema
NameRequiredDescriptionDefault
deviceNoOptional target device (IP or serial). Omit to use the focused Dev Studio tab.

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description adds behavior beyond the annotations by stating "Mutates session state (tears down the connection)" and clarifying idempotence: "closing an already-closed session is a no-op." These details complement the idempotentHint and readOnlyHint annotations rather than merely repeating them.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is three dense sentences with no filler. It front-loads the main action, then adds behavioral nuance, then gives usage context. Every sentence contributes distinct value.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple one-optional-parameter tool, the description combined with the annotations covers the purpose, side effects, idempotence, and when to call it. Nothing essential is missing for an agent to select and invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema already documents the only parameter, device, with 100% coverage, including that it is optional and what happens if omitted. The description does not add further parameter-specific meaning, so the baseline score of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: "Close the RALE / App Connector session on the targeted device." This clearly distinguishes it from sibling connect/disconnect tools such as app_connector_connect and telnet_disconnect by naming the exact session type.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives explicit triggers: "Use it to free the session or force a clean reconnect." It also provides useful when-not guidance: "normal RALE tools do not require you to disconnect between calls." However, it does not explicitly name alternative disconnect tools or state when to use telnet_disconnect instead, so it stops short of full routing guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

app_functionApp Connector: Call Channel FunctionA
Destructive

Invoke a single function on the sideloaded channel through the App Connector. Use this for any one-off function call exposed by the channel; only wrap it in an appFunction Action Script step when the call is part of a multi-step flow. The set of available functions is channel-specific — every sideloaded app exports its own. Always call list_app_connector_functions first to discover the exact name and the declared parameter list (params: [{ name, type }, …]) for the running channel before calling this tool. functionParams is a positional array with one entry per declared parameter, in declaration order. Each entry's value matches the declared type: String/Integer/Boolean/number types are primitives; roAssociativeArray is a JSON object (still wrapped in the outer array slot); roArray / roList is a JSON array (also wrapped). For a zero-arg function pass []. A named object ({ <paramName>: value }, keyed by names from list_app_connector_functions) is accepted for backward compatibility and rewritten to a positional array before the call is sent. Authors should still emit positional form: a typo in a key silently passes undefined for that slot. Auto-connects the App Connector session if needed; surfaces the call as a toast in Dev Studio. Invokes channel code, so it is not read-only and not assumed idempotent — a function may mutate app state.

ParametersJSON Schema
NameRequiredDescriptionDefault
deviceNoOptional target device (IP or serial). Omit to use the focused tab.
functionNameYesThe channel function name from list_app_connector_functions.
functionParamsNoPositional array of values, one per RALE-declared parameter. Use `[]` for zero-arg functions. A named object keyed by RALE param names is also accepted and will be normalized to positional before the call.

TDQS

A5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Matches annotations by stating 'not read-only and not assumed idempotent' and 'may mutate app state'. Adds extra context: auto-connects the App Connector session and surfaces a toast in Dev Studio, which are not in annotations. No contradictions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Every sentence serves a purpose – covering invocation, parameter format, backward compatibility, auto-connect, and side effects. Dense but not redundant; structure flows logically from usage to parameter details to behavior.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Covers all aspects needed to use the tool correctly: when to use, parameter format, defaults, prerequisites, side effects, and auto-connect behavior. No output schema exists, so no return explanation required. Fully complete for the tool's complexity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the description greatly enriches meaning: explains positional array format, zero-arg usage, named-object backward compatibility, normalization to positional, and the pitfall of key typos silently passing undefined. This goes far beyond the basic schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Clearly states 'Invoke a single function on the sideloaded channel through the App Connector' – a specific verb and resource. Distinguishes from siblings like list_app_connector_functions (lists) and app_connector_connect (connects).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly says 'Use this for any one-off function call' and contrasts with wrapping in an appFunction Action Script step for multi-step flows. Also instructs to always call list_app_connector_functions first, giving clear when-to-use guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

connect_deviceConnect to a DeviceA
Idempotent

Open (or focus, if already open) a Dev Studio device tab for the given Roku, making it the active target for renderer-routed tools (rale_command, telnet_*, app_function, get_telnet_log). Required device: Roku IP or serial from list_devices / scan_devices. Idempotent — a no-op if that device is already connected and focused. Not needed for main-direct ECP ops (keypress, launch_app, ecp_query, …) on local/remote devices, which accept a device argument directly; use test_connection to verify reachability without opening a tab. Only connects a device that's actually reachable right now — fails with a clear reason instead of guessing: a local device must be currently discoverable, a remote device's RDS Relay location must be online, and an rce (Roku Cloud Emulator) device must already be running — this tool will never start one; tell the user to start it in Roku Dev Studio first.

ParametersJSON Schema
NameRequiredDescriptionDefault
deviceYesRequired non-empty string: LAN IP (e.g. "192.168.1.68") or device serial exactly as shown by list_devices / scan_devices.

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already mark idempotentHint and openWorldHint, but the description adds crucial context beyond them: it is a no-op if already connected, it opens/focuses a tab with side effects, it will only connect reachable devices, and it never starts an rce emulator. This fully discloses behavioral traits that an agent must know to invoke and interpret the tool correctly.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but every sentence carries operational value: purpose, idempotency, exclusions, alternatives, and device-specific preconditions. The most important information is front-loaded, and the structure makes the routing decision and failure conditions easy to follow.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter tool with rich annotations and no output schema, the description covers all invocation-relevant context: what happens on success, when it is unnecessary, idempotency, reachability requirements, and failure behavior. No critical operational gap remains for an agent selecting or calling the tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and the input schema already defines `device` as a required non-empty string with LAN IP or serial examples. The description mostly restates that, adding only the source provenance ('from list_devices / scan_devices'), which is marginal value over the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a precise action: 'Open (or focus, if already open) a Dev Studio device tab for the given Roku' and clarifies the consequence of making it the active target for renderer-routed tools. It explicitly distinguishes the tool from direct-ECP alternatives like keypress and launch_app, so an agent can confidently tell it apart from siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly says when the tool is needed (for renderer-routed tools) and when it is not needed ('Not needed for main-direct ECP ops'), and it names test_connection as the alternative for reachability checks. It also gives per-device-type prerequisites (local discoverable, remote RDS Relay online, rce running), leaving no ambiguity about selection.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

console_monitor_findingsConsole Monitor: BrightScript FindingsA
Read-onlyIdempotent

Analyze the in-memory BrightScript debug console (port 8085) buffer and return the recognized BrightScript ISSUES and CRASHES — the same data the Console Monitor UI shows. Returns { connected, scannedLines, totalCaptured, totalIssues, issueTypeCount, byCategory, findings, crashes }, where each finding is { id, title, category, severity, meaning, cause, fix, docsUrl?, count, lines } and lines is that issue's unique console lines with per-line count and (when present) file/line. crashes are Micro Debugger dumps (uncaught runtime errors): each is { message, code?, file?, line?, backtrace[], count, exited?, app?, raw } where backtrace is the stack ({ depth, func, file?, line? }, innermost first) and exited marks a fatal EXIT_BRIGHTSCRIPT_CRASH. Only Roku/BrightScript-emitted diagnostics (BRIGHTSCRIPT: ERROR:/WARNING:, rendezvous, FormatJSON, roUrlEvent, …) are recognized — NOT arbitrary app log output. Data only accumulates while the console is connected: if connected is false call telnet_connect first. Read-only — it analyzes the buffer Dev Studio already holds and never touches the device.

ParametersJSON Schema
NameRequiredDescriptionDefault
deviceNoOptional target device (IP or serial). Omit to use the focused tab.

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly and idempotent, and the description strengthens this by adding 'never touches the device' and 'analyzes the buffer Dev Studio already holds'. It also discloses the connection-dependent accumulation behavior and the recognition filter, which are operational traits not present in annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long but densely informative, front-loading the core analysis action and then providing a complete return contract. Because there is no output schema, the detailed field documentation is necessary rather than redundant; every sentence adds operational value.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Despite having no output schema, the description fully documents the return shape for both findings and crashes, including nested fields and semantics. It covers the connected=false edge case and the recognition limitations, leaving no critical operational gap for an agent to call the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% for the only parameter (optional device, omit for focused tab), so the schema fully documents parameters. The description does not add parameter-level meaning, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Analyze the in-memory BrightScript debug console (port 8085) buffer and return the recognized BrightScript ISSUES and CRASHES'. It also delimits scope by stating it returns the same data the Console Monitor UI shows and excludes arbitrary app log output, which separates it from log-oriented sibling tools like get_telnet_log.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives an explicit precondition and action: 'if connected is false call telnet_connect first', plus a clear exclusion ('NOT arbitrary app log output'). This tells an agent exactly when the tool is appropriate and what to do when the prerequisite is not met.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

debugger_attachDebugger: AttachA
Idempotent

Open a BrightScript debug session to the Roku on control port 8081. REQUIRED FIRST — every other debugger_* tool needs an attached session. Prefer calling the read-only debugger_status before this one: if it already reports attached/running/stopped, skip this call entirely and go straight to the debugger_* operation you need. The port is only open when the channel was launched with debugging (sideload "with Debugging", or a STOP in the source auto-enables it); a plain sideload/relaunch does NOT open it, and attach returns an actionable error explaining that. Safe to call anyway even when already attached: if a healthy session for this device already exists (e.g. the user attached via the app's own debugger UI), this is a no-op that returns success without touching it — it only tears down and reconnects when there is no session, or the existing one is stale/errored (the control port is single-client, so a doomed reconnect would otherwise kill a working session for nothing). On success returns { ip, state }.

ParametersJSON Schema
NameRequiredDescriptionDefault
deviceNoOptional. Roku IP or serial; must match a connected Dev Studio tab. Omit to use the focused tab.

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well beyond the annotations by explaining the idempotent no-op behavior when a healthy session exists, the teardown/reconnect only for missing/stale sessions, and the single-client control port rationale. This gives the agent a solid model of side effects without contradicting any annotation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long but densely packed with actionable information, front-loaded with the required-first warning and the status-check shortcut. Every sentence earns its place by affecting whether or how the agent should invoke the tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's side effects, the lack of an output schema, and the large sibling family, the description is complete: it covers prerequisites, safe reuse, failure conditions, teardown behavior, and the success return shape. Nothing critical is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage for the one optional parameter is 100%, and the schema already explains the device field's meaning and default behavior. The description adds little about the parameter beyond calling it a Roku, which is consistent but not a meaningful supplement.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: it opens a BrightScript debug session to the Roku on control port 8081. It also clearly marks itself as the required prerequisite for the debugger_* family, differentiating it from siblings like debugger_status and debugger_detach.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides explicit when-to-use and when-not-to-use guidance: it says to call debugger_status first and skip attach entirely if already attached/running/stopped. It also explains the port-opening conditions and that attach returns an actionable error when the port isn't available.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

debugger_continueDebugger: ContinueA

Resume execution from a halted state (run until the next breakpoint / STOP / error). After calling this, use debugger_wait_for_stop to catch the next halt. Only meaningful while stopped.

ParametersJSON Schema
NameRequiredDescriptionDefault
deviceNoOptional. Roku IP or serial; must match a connected Dev Studio tab. Omit to use the focused tab.

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations (readOnlyHint=false, openWorldHint=true) already indicate mutation and side effects, but the description adds specific behavior: execution resumes until breakpoint/STOP/error, and it only works from a halted state. This goes beyond annotation basics and gives the agent a concrete expected behavior. No contradiction with the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with the main action, then the follow-up and precondition. No filler or redundant phrasing. Every word contributes to the agent's understanding.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description still tells the agent the immediate next step (use debugger_wait_for_stop) and the precondition (stopped). It does not mention the immediate return value, but the description implies the continue call doesn't provide the halt—so the follow-up is necessary. For a simple tool with one optional parameter, this is adequate.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The device parameter is fully described in the schema, including its type, role (Roku IP or serial), and default behavior (omit to use focused tab). The description adds nothing about parameters. Since schema description coverage is 100%, the baseline of 3 is appropriate; there is no extra value from the description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States the specific verb 'resume execution' and the resource (debugger). It clarifies the condition under which execution runs (until next breakpoint/STOP/error). It does not explicitly contrast with siblings like debugger_pause or debugger_step, but the action is distinct enough that an agent can infer its purpose.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit precondition: 'Only meaningful while `stopped`', which tells the agent when the tool should be used. It also names the follow-up tool debugger_wait_for_stop, guiding the agent on what to do next. It does not mention when to prefer it over debugger_pause/step, but the halted-state condition is a clear gating rule.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

debugger_detachDebugger: DetachA
Idempotent

Close the debug session for a device (releases the 8081 control socket). Idempotent — a no-op if not attached. The running channel keeps executing.

ParametersJSON Schema
NameRequiredDescriptionDefault
deviceNoOptional. Roku IP or serial; must match a connected Dev Studio tab. Omit to use the focused tab.

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond annotations, the description discloses the side effect of releasing the 8081 socket and that the running channel continues executing. It also restates idempotency in operational terms ('no-op if not attached'), reinforcing the idempotentHint annotation. No contradiction with annotations exists.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two compact sentences with no filler. The core action is front-loaded, and the key idempotency and execution-continuation details are packed efficiently into the remaining sentence.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-parameter tool with full schema coverage and informative annotations, the description covers the operation's side effects and no-op behavior. The lack of an output schema is not a concern at this complexity, though the description does not specify the return value on success.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The single optional device parameter is fully described in the schema (IP/serial, must match a connected Dev Studio tab, omit for focused tab), so the description carries little parameter burden. The description only ties the session to a device generically. Baseline 3 for high schema coverage is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: 'Close the debug session for a device', and adds the mechanism 'releases the 8081 control socket.' This clearly distinguishes it from sibling tools like debugger_attach and debugger_status by naming the exact action.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear context: it is a teardown/cleanup operation, idempotent, and safe to call when not attached. It does not explicitly name an alternative such as debugger_attach, but the action is unambiguous enough for an agent to infer when to invoke it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

debugger_evaluateDebugger: Evaluate (REPL)A
Destructive

Run a BrightScript expression/statement in the halted frame (the debug-console REPL) — e.g. print m.top.count or print type(node). Output streams to the device console; the result reports compile/runtime errors if any. Requires the target to be HALTED. Can have side effects (it executes code), so it is not read-only. For a plain variable read prefer debugger_get_variables. Optional stackFrameIndex / threadIndex select the scope.

ParametersJSON Schema
NameRequiredDescriptionDefault
deviceNoOptional. Roku IP or serial; must match a connected Dev Studio tab. Omit to use the focused tab.
expressionYesRequired. BrightScript to run in the halted frame (often `print <expr>`).
threadIndexNoOptional thread index (default the stopped/primary thread).
stackFrameIndexNoStack frame scope (default 0 = top frame).

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description explicitly states that execution can have side effects and is not read-only, reinforcing the annotations (readOnlyHint=false, destructiveHint=true). It also adds useful behavioral context not present in the annotations: output streams to the device console, and compile/runtime errors are reported.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and front-loaded, opening with the action, then giving examples, followed by the side-effect warning and alternative routing. Every sentence adds necessary information without filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with no output schema, the description adequately covers the critical context: halted-frame requirement, side-effect warning, output destination, and error reporting. It leaves default behavior to the schema, which already documents threadIndex and stackFrameIndex defaults, so the overall context is sufficient.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all parameters. The description adds useful examples for the expression parameter and notes that stackFrameIndex/threadIndex select the scope, but this largely restates what the schema already provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly identifies the operation as running a BrightScript expression/statement in the halted frame (the debug-console REPL), with concrete examples such as print m.top.count. This distinguishes it from read-only variable inspection tools like debugger_get_variables.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It states the core precondition that the target must be HALTED and explicitly routes plain variable reads to debugger_get_variables. It also warns that evaluation can have side effects, helping the agent decide when to use it versus safer alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

debugger_get_callstackDebugger: Get Call StackA
Read-onlyIdempotent

Return the call-stack frames (function, file, line — top frame first) for the halted thread. Requires the target to be HALTED. Optional threadIndex (default the stopped/primary thread). A frame index from here feeds stackFrameIndex in debugger_get_variables / debugger_evaluate.

ParametersJSON Schema
NameRequiredDescriptionDefault
deviceNoOptional. Roku IP or serial; must match a connected Dev Studio tab. Omit to use the focused tab.
threadIndexNoOptional thread index (default the stopped/primary thread).

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, covering the safety profile. The description adds behavioral value beyond those: the halted-thread precondition, the top-frame-first ordering, and the default thread behavior. No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences with no wasted words. The primary return information is front-loaded, the required condition is stated immediately after, and the cross-tool usage note is compact and useful.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description explains return shape, ordering, required precondition, and how the result connects to sibling tools, which is especially helpful given there is no output schema. It does not describe error behavior when the target is not halted, but the explicit HALTED requirement largely covers that gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description's threadIndex note repeats the schema's 'default the stopped/primary thread' almost verbatim)Skip, adding no new semantic value beyond the structured schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb ('Return') and resource ('call-stack frames') with concrete detail (function, file, line — top frame first). It is unambiguously distinct from sibling debugger tools and clearly scoped to the halted thread.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It states the key precondition ('Requires the target to be HALTED') and gives sequencing context by explaining that a returned frame index feeds stackFrameIndex into debugger_get_variables / debugger_evaluate. It does not enumerate exclusions or alternative tools, but no true alternative for call-stack retrieval exists among the siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

debugger_get_variablesDebugger: Get VariablesA
Read-onlyIdempotent

Return variables in scope at a stack frame while HALTED. With no variablePath, returns the frame's locals (incl. m); each entry has name, type, value, and for containers a childCount. To drill into a container, pass its variablePath (e.g. ["m","top"]; a quoted "key" segment forces a case-sensitive AA lookup, a bare number indexes an array) — the response is [container] whose .children is the next level. Optional stackFrameIndex (default 0 = top) and threadIndex.

ParametersJSON Schema
NameRequiredDescriptionDefault
deviceNoOptional. Roku IP or serial; must match a connected Dev Studio tab. Omit to use the focused tab.
threadIndexNoOptional thread index (default the stopped/primary thread).
variablePathNoOptional path segments to drill into a container (e.g. ["m","top","count"]). Omit for the frame's locals.
stackFrameIndexNoStack frame to read (default 0 = top frame).

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the readOnly/idempotent annotations, the description discloses the response shape (name, type, value, childCount), the locals default when variablePath is omitted, case-sensitive AA lookup behavior, array indexing rules, and the [container].children drill-down result.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but efficient: purpose and precondition first, then default behavior and return fields, then drill-down syntax and optional parameters. No filler or repetition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description supplies the essential return contract, including locals entries and container children. It covers defaults, path semantics, and the halted precondition; remaining details are either in the schema or not necessary to invoke the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Although schema coverage is already 100%, the description adds meaning: variablePath semantics (omitted default, quoted key vs bare number), stackFrameIndex default, and the shape of each returned variable. This materially enriches what the parameter schema alone provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Return variables in scope at a stack frame while HALTED.' It clearly focuses on variable inspection and drilling, which distinguishes it from sibling debugger tools like debugger_get_callstack or debugger_continue.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It clearly states the precondition (HALTED) and explains how to use variablePath, stackFrameIndex, and threadIndex to choose a frame or drill into containers. It does not explicitly name alternatives such as debugger_evaluate, so it stops short of full when-not-to-use guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

debugger_list_breakpointsDebugger: List BreakpointsA
Read-onlyIdempotent

List the breakpoints the debugger is tracking for a device: each with filePath, lineNumber, conditionalExpression?, hitCount?, verified (registered on the device), queued (waiting for the next halt to register), and breakpointId?. Read-only.

ParametersJSON Schema
NameRequiredDescriptionDefault
deviceNoOptional. Roku IP or serial; must match a connected Dev Studio tab. Omit to use the focused tab.

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, but the description adds value by explaining the meaning of 'verified' (registered on the device) and 'queued' (waiting for the next halt to register), as well as the optional fields. It does not contradict annotations and provides useful behavioral context beyond the structured data.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences: the first states the action and resource, the second enumerates the return fields with inline clarifications for the less-obvious ones. No filler, no redundancy, and the key purpose is front-loaded. Appropriate length for the information delivered.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description carries the burden of explaining the return shape, and it does so by listing each field with its meaning (verified, queued, optional fields). It covers the essential behavior of a read-only list tool; minor gaps like error cases or device-specific behavior are acceptable given the tool's simplicity and annotation coverage.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 100% description coverage for the single optional 'device' parameter, fully explaining its semantics (Roku IP/serial, must match connected tab, omit for focused tab). The description adds no parameter-specific information, so the baseline score of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('List') and resource ('breakpoints the debugger is tracking for a device'), clearly distinguishing it from sibling tools like debugger_set_breakpoints and debugger_remove_breakpoints. The field list further clarifies what is returned, leaving no ambiguity about the tool's function.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage for inspecting breakpoints (read-only) but does not explicitly say when to use this tool versus alternatives like debugger_set_breakpoints or debugger_status. It does not mention when not to use it or any preconditions beyond the device parameter. Usage guidance is only implicit through the tool's name and read-only hint.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

debugger_pauseDebugger: PauseA

Request a halt of a running channel (best-effort). Follow with debugger_wait_for_stop to get the snapshot once it stops. Only meaningful while running.

ParametersJSON Schema
NameRequiredDescriptionDefault
deviceNoOptional. Roku IP or serial; must match a connected Dev Studio tab. Omit to use the focused tab.

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already mark readOnlyHint=false大手, openWorldHint=true, idempotentHint=false, and destructiveHint=false, so the safety profile is known. The description adds valuable behavioral nuance by stating the action is 'best-effort' and asynchronous, implying the halt is not guaranteed and that the actual snapshot comes from a subsequent call. This goes beyond what annotations provide.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences with no filler. It front-loads the core action, then adds the essential follow-up and precondition. Every sentence contributes meaningful guidance, making it highly efficient.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple, one-parameter tool with 100% schema coverage and no output schema, the description covers the core behavioral contract: request the halt, use wait_for_stop for the snapshot, and only invoke while running. It could theoretically state what the immediate response is, but that is minor given the clear workflow and the annotations already covering non-destructiveness and side effects.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the single optional 'device' parameter is fully documented in the schema, including the fallback to the focused tab. The description itself does not add any additional parameter semantics or constraints beyond the schema, which warrants the baseline score of 3.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Request a halt') with a clear resource ('a running channel') and adds the qualifier 'best-effort,' which precisely defines the action. It is clearly distinguishable from sibling tools like debugger_continue, debugger_step, and debugger_wait_for_stop, which represent different debugger control operations.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit when-to-use context: 'Only meaningful while running.' It also provides a follow-up workflow, instructing the agent to use debugger_wait_for_stop to get the snapshot once stopping occurs. It does not explicitly name alternatives or when-not-to-use scenarios beyond the 'only meaningful' condition, but the guidance is sufficient for correct sequencing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

debugger_remove_breakpointsDebugger: Remove BreakpointsA
Idempotent

Remove breakpoints by location. locations: array of { filePath, lineNumber } (matching what debugger_list_breakpoints reports). Removal is by file:line so it also clears a still-queued breakpoint that has no device id yet. Returns { removed }.

ParametersJSON Schema
NameRequiredDescriptionDefault
deviceNoOptional. Roku IP or serial; must match a connected Dev Studio tab. Omit to use the focused tab.
locationsYesBreakpoint locations to remove.

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the annotations (idempotentHint, readOnlyHint false, etc.), the description adds meaningful behavior: removal is by file:line and clears still-queued breakpoints without a device ID, and it returns '{ removed }'. This context is not present in structured data and helps the agent predict side effects and results.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, zero fluff. The core action is front-loaded, followed by parameter semantics and return value. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a two-parameter tool with a fully documented schema and annotations, the description covers the essential operational behavior and output. The return value is stated, the input format is clarified, and the queued-breakpoint nuance removes ambiguity. Nothing critical is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so baseline is 3. The description adds value by explaining that the locations array must match debugger_list_breakpoints output and that removal by file:line also covers queued breakpoints, which is not obvious from the schema alone. This raises it to a 4.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'Remove breakpoints by location.' It also clarifies that locations match what debugger_list_breakpoints reports, which distinguishes it from sibling tools like debugger_set_breakpoints and debugger_list_breakpoints. The purpose is immediately unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear context by specifying the locations format as matching debugger_list_breakpoints and notes the queued-breakpoint behavior, which guides the agent on when and how to invoke the tool. It does not explicitly name alternatives or when-not-to-use conditions, but the reference to list_breakpoints implies the correct input source and the tool name makes its role obvious.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

debugger_set_breakpointsDebugger: Set BreakpointsA
Idempotent

Add breakpoints. breakpoints: array of { path, line, condition?, hitCount? } — path is a pkg:/… source path (a bare path is prefixed with pkg:/), condition is an optional BrightScript expression (Roku OS 11.5+), hitCount skips that many hits first. IMPORTANT: the device only registers breakpoints while HALTED — one added while the channel is running comes back pending:true and is queued to register at the next stop. Each result carries a breakpointId (registered) or an error. Existing conditions are replaced on re-add.

ParametersJSON Schema
NameRequiredDescriptionDefault
deviceNoOptional. Roku IP or serial; must match a connected Dev Studio tab. Omit to use the focused tab.
breakpointsYesBreakpoints to set.

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the annotations, the description discloses the pending/queued behavior when a breakpoint is added while the channel is running, the re-add replacement of existing conditions, and that each result returns a breakpointId or an error. This is meaningful operational detail the structured annotations do not provide.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with the action, then provides the parameter shape, then highlights the critical runtime behavior in a clear IMPORTANT clause. Each sentence serves a purpose and there is no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description still explains what each result carries, covers the pending state, and describes re-add semantics. Combined with the schema's full parameter documentation, an agent has enough to call the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and the description largely restates what the schema already documents: path prefixing, condition as a BrightScript expression, and hitCount behavior. It adds little genuinely new parameter meaning beyond a compact summary, so the baseline of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with the exact action and resource: 'Add breakpoints.' It also names the data shape and the device context, and the sibling tools include debugger_remove_breakpoints and debugger_list_breakpoints, so the tool is immediately distinguishable as the write/set operation for breakpoints.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives critical usage context: the device must be HALTED for breakpoints to register, and breakpoints added while running are queued with pending:true. It does not explicitly name alternatives like debugger_remove_breakpoints or debugger_list_breakpoints, but the use case is clear enough that an agent can select this tool confidently.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

debugger_statusDebugger: StatusA
Read-onlyIdempotent

Return the session state for a device WITHOUT blocking: one of disconnected (not attached), connecting, attached, running, stopped (HALTED — safe to inspect), or error. Call this before debugger_attach — if it already reports attached/running/stopped, a session is already up (maybe from the app's own debugger UI) and you can skip straight to the operation you need. Also poll this to decide whether inspection tools will work; to block until the next halt use debugger_wait_for_stop instead.

ParametersJSON Schema
NameRequiredDescriptionDefault
deviceNoOptional. Roku IP or serial; must match a connected Dev Studio tab. Omit to use the focused tab.

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is covered. The description adds valuable behavioral context beyond annotations: it explicitly states the tool is non-blocking, explains the meaning of the 'stopped' state (HALTED — safe to inspect), and clarifies that a session may already be up from the app's own debugger UI. It doesn't describe polling frequency or error behavior, but the non-blocking disclosure is a meaningful addition.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, each earning its place: the first defines the return values and non-blocking behavior, the second gives a concrete pre-attach usage pattern, and the third routes to the blocking alternative. The most critical information (non-blocking, state values) is front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only status tool with one optional parameter, no output schema, and rich annotations, the description is complete. It covers what the tool returns, when to use it, how to interpret the key state, and which sibling to use for blocking behavior. Nothing an agent needs to call it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents the single optional 'device' parameter. The description adds context about the parameter's purpose ('must match a connected Dev Studio tab') and the omit-to-use-focused-tab behavior, which is slightly beyond the schema. However, the schema already covers the core semantics, so the description's added value is marginal.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Return'), a specific resource ('session state for a device'), and enumerates all possible return values. It also distinguishes itself from debugger_wait_for_stop by explicitly noting it is non-blocking, which differentiates it from a sibling tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit when-to-use guidance: call before debugger_attach to check if a session is already up, and poll to decide whether inspection tools will work. It also names the alternative (debugger_wait_for_stop) for blocking behavior, providing clear routing between siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

debugger_stepDebugger: StepA

Single-step the halted thread. kind: "over" (default — next line, skipping calls), "in" (into the call), or "out" (finish the current function). Requires the target to be HALTED. Follow with debugger_wait_for_stop to get the new location.

ParametersJSON Schema
NameRequiredDescriptionDefault
kindNoStep kind (default "over").
deviceNoOptional. Roku IP or serial; must match a connected Dev Studio tab. Omit to use the focused tab.
threadIndexNoOptional thread to step (default the stopped/primary thread).

TDQS

A4.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate the tool is not read-only (readOnlyHint=false) and not idempotent (idempotentHint=false), so the mutation behavior is implicit. The description adds the halted requirement and the fact that it moves execution, which is useful context beyond the annotations. However, it does not describe what happens if the thread is invalid or if stepping fails, leaving some behavioral gaps.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences with no wasted words. The primary action is front-loaded, followed by the kind options, then the critical precondition and next-step guidance. Every clause serves a purpose, and the structure is logical and easy to parse.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's moderate complexity (3 optional params, no output schema) and annotations that already cover mutation, the description is fairly complete. It provides the essential precondition, parameter semantics, and a follow-up for retrieving results. It lacks explicit error-handling details, but that is minor given that the schema covers all parameters and the annotations cover safety. An agent can call this tool correctly with the information provided.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description adds significant value for the 'kind' parameter by explaining each enum value in plain language ('over' = next line skipping calls, 'in' = into the call, 'out' = finish current function), which goes beyond the schema's minimal 'Step kind (default "over").' The device and threadIndex parameters are adequately described in the schema, and the description does not repeat them.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('Single-step the halted thread') with a specific verb and resource. It also enumerates the three step kinds, making the tool's function unambiguous. It distinguishes itself from siblings like debugger_continue and debugger_pause by focusing on single-stepping.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly states a critical prerequisite ('Requires the target to be HALTED') and provides a direct follow-up instruction ('Follow with debugger_wait_for_stop to get the new location'). This gives the agent clear guidance on when and how to use the tool, including a sequence that avoids wasted calls.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

debugger_wait_for_stopDebugger: Wait for StopA
Read-onlyIdempotent

Block (server-side poll) until the target HALTS at a breakpoint / STOP / step-completion / runtime error, then return { stopped: true, stop: { reason, detail, threads, stackFrames, variables } } — the top-frame snapshot. Returns { stopped: false, timedOut: true } if it is still running at the deadline, or { stopped:false, state } if the session ended. Call this right after debugger_continue / debugger_step, or after triggering the app, to know when you can inspect. Optional timeoutMs (default 15000, max 30000).

ParametersJSON Schema
NameRequiredDescriptionDefault
deviceNoOptional. Roku IP or serial; must match a connected Dev Studio tab. Omit to use the focused tab.
timeoutMsNoMax ms to wait (default 15000, capped at 30000).

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description discloses important behavioral traits beyond the annotations: server-side blocking/polling, timeout handling with distinct return variants, and the exact state returned when the session ends. Annotations already declare readOnlyHint and idempotentHint, but the description adds substantial behavioral context without contradicting the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four sentences, each earning its place: core blocking behavior and return value, alternative timeout/session returns, when to call it, and timeout parameter details. The most important information is front-loaded, and there is no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, but the description fully supplies the return contract for all three outcomes: stopped, timed out, or session ended. It also covers invocation timing and the optional timeout parameter, making the tool complete and callable without further research.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% for both parameters, so the schema already documents device and timeoutMs. The description adds no new parameter-level meaning beyond restating the timeout default and cap, so the baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific action and resource: it blocks until the target halts at a breakpoint/STOP/step-completion/runtime error, then returns a structured stop snapshot. This clearly distinguishes it from sibling debugger tools like debugger_continue or debugger_status, which do different things.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly tells the agent when to call this tool: right after debugger_continue, debugger_step, or after triggering the app. It doesn't explicitly state when not to use it or name alternatives, but the invocation context is clear and concrete.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_sideloadDelete Sideloaded ChannelA
DestructiveIdempotent

Remove the currently sideloaded Dev App from the device. Password optional when Dev Studio has remembered it for this device. Destructive; idempotent (deleting when nothing is sideloaded still ends with no Dev App). To install/replace a Dev App use sideload — you do not need to delete first, since sideload overwrites.

ParametersJSON Schema
NameRequiredDescriptionDefault
deviceNoTarget device — IP (e.g. "192.168.1.137") or serial (e.g. "X0004EX9Q7M"). Omit to use the focused device.
passwordNoOmit if Roku Dev Studio has saved the Dev Password for this device (Remember on the device tab).

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint and idempotentHint; the description adds human-readable detail that deletion when nothing is sideloaded still results in no Dev App, and notes the password-optional behavior tied to Dev Studio's memory. This adds context beyond the structured hints without contradicting them.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences: first states the main action, second covers edge behavior and password, third routes to the alternative. No filler, front-loaded with the core operation.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive, idempotent two-parameter tool with no output schema, the description covers what is removed, the empty-state behavior, password handling, and the alternative install path. An agent has everything needed to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Input schema provides 100% coverage with clear descriptions for both 'device' and 'password'; the description restates password optionality but does not add new meaning beyond the schema. Therefore the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific action ('Remove the currently sideloaded Dev App') and explicitly distinguishes it from the sibling 'sideload' by noting that sideload overwrites rather than requiring a delete first. This makes the tool's purpose immediately distinguishable from its closest alternative.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly names the alternative tool (sideload) and the condition under which it should be used (install/replace), and clarifies that delete_sideload is not a prerequisite. Also explains when password is optional, giving the agent clear decision criteria.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

device_performance_metricsDevice Performance Metrics (CPU / Memory / Objects)A
Read-onlyIdempotent

Time-series Device Performance metrics — the same chanperf/r2d2-bitmaps/app-object-counts data the Remote tab's CPU/Memory/BrightScript Objects quad charts poll and plot, returned as compact per-timestamp entries with a decoding legend. Requires "Show Device Performance" (quad layout) to have been turned on for this device tab at some point this session, with the sideloaded Dev channel as the foreground app — if it never was, devicePerformanceEnabled is false and samples is empty (never an error). charts selects which of cpu/memory/objects to include (default: all three) — each requested type appears as its own key per sample (c=cpu, m=memory, o=objects; see legend for field meanings). windowSec (default 60) sets how far back from now to report; if the device's retained history is shorter, actualWindowSec/sampleCount reflect what was actually available. A window whose natural sample count exceeds maxSamples (default 120, max 500) is evenly downsampled across the window (not truncated from one end) and downsampled is set true. cpuProcessSnapshot (only present when cpu is requested) is a single latest-value object (process state, channel uptime, CPU time, cumulative fault counts) — Roku's <proc-stat> block has no historical series of its own, only the fault-rate numbers inside each cpu sample do. Read-only.

ParametersJSON Schema
NameRequiredDescriptionDefault
chartsNoWhich chart types to include. Omit (or pass an empty array) for all three.
deviceNoOptional target device (IP or serial). Omit to use the focused tab.
windowSecNoHow far back from now to report, in seconds (e.g. 60 for the last minute, 3600 for the last hour). Default 60.
maxSamplesNoCap on returned samples; the window is evenly downsampled if it would exceed this. Default 120, hard cap 500.

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the annotations (readOnly, idempotent, non-destructive), the description discloses important behavioral traits: the quiet-failure mode when the feature was never enabled, uniform downsampling across the window rather than truncation, the `actualWindowSec`/`sampleCount` fallback when history is short, and the special single-value `cpuProcessSnapshot` that is not a historical series. This goes well beyond the annotations and helps the agent set expectations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but every clause earns its place. It front-loads the core purpose and data source, then logically covers prerequisites, parameter effects, edge cases, and one special output. No filler or repeated schema information—only clarifying detail.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, so the description must cover return shape and behavior. It does so thoroughly: compact per-timestamp entries, a decoding legend, `devicePerformanceEnabled`, `samples`, `actualWindowSec`, `sampleCount`, `downsampled`, and `cpuProcessSnapshot`. It also covers defaults, caps, and the only meaningful failure mode, making the tool fully invocable without additional lookups.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Even though schema coverage is 100%, the description adds substantial meaning: it explains the per-sample key convention (`c`/`m`/`o`) and the role of `legend`, how `windowSec` interacts with retained history, how `maxSamples` causes even downsampling plus a `downsampled` flag, and why `cpuProcessSnapshot` appears only with cpu. This is significant value beyond the schema's parameter descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: it returns time-series device performance metrics for CPU, memory, and objects, tied to a known data source (chanperf/r2d2-bitmaps/app-object-counts). This clearly differentiates it from all sibling tools, which target debugging, networking, or app control, not performance telemetry.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear context and a concrete prerequisite: the 'Show Device Performance' quad layout must have been enabled earlier in the session, and the behavior when it was not is explicitly defined (empty samples, never an error). It does not name explicit alternatives or when-not-to-use versus a specific sibling, but the distinct purpose makes alternatives unnecessary.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ecp_postECP POST (raw)A
Destructive

POST to an arbitrary ECP endpoint (e.g. /sgrendezvous/track). Side-effecting — agents should use list_post_presets for safe defaults. For read-only lookups use ecp_query instead.

ParametersJSON Schema
NameRequiredDescriptionDefault
deviceNoTarget device — IP (e.g. "192.168.1.137") or serial (e.g. "X0004EX9Q7M"). Omit to use the focused device.
endpointYesECP path to POST to (e.g. "/sgrendezvous/track", "/input/12345"). Prefer a value from list_post_presets; arbitrary paths are sent verbatim.

TDQS

A4.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The annotations already declare destructiveHint=true, readOnlyHint=false, and idempotentHint=false, so the description's 'Side-effecting' statement mostly restates what the annotations provide. It adds the safety-focused guidance to prefer presets, but does not reveal additional behavioral details such as what state may change, auth requirements, or response behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two concise sentences with no filler. It front-loads the core purpose, then immediately provides safety and alternative routing guidance.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a raw, side-effecting POST tool, the description covers the key concerns: arbitrary endpoint, side effects, safe defaults, and read-only alternatives. It could mention what kind of response to expect, but given the arbitrary nature and lack of output schema, the description is sufficiently complete for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the endpoint parameter description already includes examples and the warning that arbitrary paths are sent verbatim. The tool description adds context about the endpoint being arbitrary, but does not materially enhance parameter understanding beyond the schema, so the baseline of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific action ('POST to an arbitrary ECP endpoint') with a concrete example path and differentiates itself from sibling tools by naming list_post_presets and ecp_query. An agent immediately understands what this tool does and what it is not for.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly tells agents when not to use this tool: use list_post_presets for safe defaults and ecp_query for read-only lookups. This is clear routing guidance that prevents unsafe or incorrect invocations.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ecp_queryECP Query (read-only)A
Read-onlyIdempotent

Run a read-only ECP GET against a device (device info, installed apps, active app, media player state, …). Pick an endpoint from list_query_presets or pass any /query/* path. Read-only — does not change device state. This is the go-to inspection tool; for state-changing POSTs use ecp_post, and to enumerate app ids call this with "/query/apps".

ParametersJSON Schema
NameRequiredDescriptionDefault
deviceNoTarget device — IP (e.g. "192.168.1.137") or serial (e.g. "X0004EX9Q7M"). Omit to use the focused device.
endpointYesECP path (e.g. /query/active-app) or telnet preset (e.g. telnet:plugins).

TDQS

A4.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, openWorldHint, idempotentHint, and destructiveHint=false. The description reinforces the read-only nature and adds behavioral context by listing example return categories and advising to use list_query_presets for endpoint discovery. It doesn't contradict annotations and provides additional context beyond the safety flags.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact yet information-dense. The core action is front-loaded, followed by examples, endpoint guidance, explicit read-only note, and final routing to alternatives. Every sentence adds value with no filler or repetition, making it easy to parse quickly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple GET tool with 2 parameters and no output schema, the description covers purpose, usage, endpoint selection, alternatives, and the read-only guarantee. It's fully sufficient for an agent to decide when to call it and how to construct the request without additional lookup.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, but the description enriches meaning: it clarifies the endpoint parameter by specifying the two acceptable formats (preset names or /query/* paths) and provides a concrete usage example ('/query/apps'). This is beyond what the schema offers, making the tool easier to invoke correctly.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Run') and resource ('read-only ECP GET'), lists concrete data examples (device info, installed apps, active app, media player state), and explicitly differentiates from the sibling ecp_post by mentioning state-changing POSTs. The agent can immediately tell what this tool does and how it differs from alternatives.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly names when to use ('go-to inspection tool'), when not to (state-changing POSTs use ecp_post), and provides an actionable pattern: pick an endpoint from list_query_presets or pass any /query/* path. Also gives a concrete example for enumerating app ids ('/query/apps'). This is fully actionable guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_action_schemaGet Action SchemaA
Read-onlyIdempotent

Return the authoring schema (label, description, required and optional fields) for ONE Action Script step type. Read-only. Call this after list_action_types (which enumerates every type) when you are about to author or fix a specific step and need its exact field names before running validate_script. Required argument type — one of the values from list_action_types (also enumerated in this tool's inputSchema). For the whole authoring contract at once, prefer get_capability_bundle.

ParametersJSON Schema
NameRequiredDescriptionDefault
typeYesExact step type key from list_action_types (e.g. appFunction, wait, keypress).

TDQS

A4.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is fully covered. The description adds that the result is an authoring schema scoped to one step type, but otherwise restates 'Read-only' and does not disclose error behavior for invalid types; the enum parameter largely mitigates that gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four sentences cover purpose, usage, alternative, and required argument with minimal waste. 'Read-only' and the parenthetical 'also enumerated in this tool's inputSchema' are slightly redundant with annotations and schema, but the description remains compact and front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-parameter read-only lookup tool, the description is fully adequate: it states what the return value is, when to call it, how it relates to siblings, and which argument is required. No output schema exists, but the description's mention of 'label, description, required and optional fields' covers the expected return content.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the single `type` parameter has an enum plus a descriptive comment, so the schema carries most of the semantic load. The description adds the cross-reference to list_action_types and clarifies that the argument is required, but it mostly reinforces what the schema already states.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description opens with a specific verb and resource: 'Return the authoring schema ... for ONE Action Script step `type`'. It distinguishes itself from list_action_types (enumerates every type) and get_capability_bundle (whole contract at once), so an agent knows exactly what scope this tool covers.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit invocation context: call after list_action_types when authoring/fixing a specific step and needing exact field names before validate_script. Also names the alternative get_capability_bundle for the whole contract, so the agent can route between siblings without guessing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_app_iconGet App IconA
Read-onlyIdempotent

Fetch the 336x210 app icon for one installed channel on the device, returned as base64 / data URL (ECP /query/icon/). Read-only. Discover valid app ids with ecp_query "/query/apps" (or launch_app's notes); "dev" is the sideloaded Dev App. Use this to preview a channel's branding — for a picture of the current screen use screenshot instead.

ParametersJSON Schema
NameRequiredDescriptionDefault
appIdYesChannel id whose icon to fetch (e.g. "837" for YouTube, "dev" for the sideloaded Dev App). Get ids from ecp_query "/query/apps".
deviceNoTarget device — IP (e.g. "192.168.1.137") or serial (e.g. "X0004EX9Q7M"). Omit to use the focused device.

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, openWorldHint, idempotentHint, and destructiveHint=false. The description adds meaningful behavioral context on top: the exact 336x210 dimensions, the ECP endpoint, the base64/data URL return format, and that the app must be installed. This is more than the minimal bar, though it doesn't describe error behavior for invalid app ids.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three concise sentences with the key facts front-loaded. The 'Read-only' phrase and the appId discovery sentence are slightly redundant with annotations and schema, but every sentence still adds useful invocation context.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple 2-parameter tool with no output schema, the description supplies the return format, dimensions, endpoint, valid-id discovery, and a clear alternative for a different use case. The optional device parameter is already fully documented in the schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Input schema coverage is 100%, so the baseline is 3. The description repeats the appId examples and the /query/apps source already present in the schema, and adds only the 'installed' qualifier. It doesn't materially extend the schema's parameter descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb and resource: 'Fetch the 336x210 app icon for one installed channel' and specifies the return format (base64 / data URL). It distinguishes itself from the screenshot tool by stating that screenshot should be used for a picture of the current screen.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly says when to use this tool ('to preview a channel's branding') and names the alternative for screen captures ('use screenshot instead'). It also tells the agent how to discover valid app ids via ecp_query /query/apps, which is practical guidance for invocation.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_capability_bundleGet Capability BundleA
Read-onlyIdempotent

Single payload of every static capability (actions, vocabularies, RALE built-ins, presets, authoring rules, op directory, actionScriptAgentContract). Load once before authoring scripts, then cache. Same JSON is also available as resource roku-dev-studio://capability-bundle.json.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and destructiveHint. The description adds value beyond those by explaining the payload is static, that it should be loaded once and cached, and that the same JSON is exposed as a resource. No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences with no filler. The core purpose is front-loaded, the payload contents are enumerated compactly, and the usage directive ('load once, then cache') is placed immediately after the purpose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a parameterless tool with no output schema, the description is complete: it identifies what is returned, enumerates its contents, states the expected usage pattern, and gives an alternative resource path. An agent has enough information to call it correctly and cache the result.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, so schema coverage is trivially 100% and the baseline is 4. The description adds meaningful context about what the returned payload contains, which is the closest analogue to parameter semantics for a parameterless tool.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb and resource: it returns a single payload of every static capability, and it enumerates the major content categories (actions, vocabularies, RALE built-ins, presets, authoring rules, op directory, actionScriptAgentContract). This clearly distinguishes it from siblings like list_action_types or get_action_schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit when-to-use guidance: load once before authoring scripts, then cache. It also notes the same JSON is available as a resource URL, which helps an agent decide between fetching and using a cached resource. It doesn't contrast with specific sibling tools, but for a zero-parameter static bundle this is sufficient.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_selected_deviceGet Selected DeviceA
Read-onlyIdempotent

Return the single device tab the user currently has focused in Dev Studio (ip, serial, modelName, friendlyDeviceName, …), or an empty/null result when no tab is focused. Read-only. Call this to resolve the implicit target before a device op when the user says "this device" / "the current one" and gave no IP. For the full inventory (all connected / discovered / remembered devices) use list_devices instead; to change the focus use connect_device.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, openWorldHint, idempotentHint, and destructiveHint=false, so the safety profile is clear. The description adds that it returns empty/null when no tab is focused, which is valuable behavioral context not covered by annotations. It also reinforces the read-only nature without contradiction.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise, front-loaded with the main purpose, and includes key details without fluff. It efficiently covers return value, null case, read-only note, and usage guidance in a few sentences. Every sentence adds value.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter read-only tool with clear annotations and no output schema, the description is complete. It explains what the tool returns, when to use it, and alternatives, which is all an agent needs to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, so the schema provides no parameter semantics. The description compensates by explaining what the return value includes (ip, serial, etc.) and when it might be empty. However, it doesn't need to explain parameters since there are none, so a high score is warranted.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool returns the focused device tab and lists example fields. It distinguishes itself from sibling tools list_devices (full inventory) and connect_device (change focus), making its purpose unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly states when to call this tool: when resolving the implicit target for device ops when the user says 'this device' and gave no IP. It also provides alternatives: use list_devices for full inventory, connect_device to change focus. This is excellent guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_telnet_logGet Telnet / BrightScript Console LogA
Read-onlyIdempotent

Read lines from the BrightScript debug console (port 8085) buffer that Dev Studio holds in memory. Returns { lines, cursor, totalLines, connected }. Pass afterCursor (the cursor from a previous call) to get only new lines — use this for polling. maxLines caps the response (default 500, max 2000). Lines only accumulate while the console is connected: if connected is false call telnet_connect first, then re-run this tool. The Roku 8085 telnet socket only allows one client at a time — telnet_connect will close any existing telnet session held by another tool/IDE before attaching. Read-only — it drains the buffer Dev Studio already holds and never touches the device.

ParametersJSON Schema
NameRequiredDescriptionDefault
deviceNoOptional target device (IP or serial).
maxLinesNoMax lines to return (default 500, max 2000).
afterCursorNoCursor returned by a previous call. Omit (or pass 0) for the full buffer.

TDQS

A5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description goes beyond the readOnlyHint and idempotentHint annotations by explicitly stating the tool is read-only and never touches the device. It also reveals important behavioral traits: lines accumulate only while connected, and the tool drains an existing buffer, with a caveat about the single-client socket.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise and well-structured, using about four sentences to cover the return shape, parameters, prerequisites, and caveats. Each sentence adds essential information without redundancy or unnecessary detail.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the sibling tools include telnet_connect and console_monitor_findings, the description clearly places this tool in context by referencing the connection requirement and the buffer mechanism. It also describes the return structure, making it sufficiently complete even without an output schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema covers 100% of parameters with descriptions, and the description further enriches semantics by explaining that afterCursor is from a previous call, omitting it returns the full buffer, and maxLines has a default and maximum. This provides complete context for each parameter's purpose.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool reads lines from the BrightScript debug console buffer, specifies the resource (port 8085) and the return format, and distinguishes itself from sibling tools like telnet_connect and console_monitor_findings by focusing on reading the buffered output.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides explicit usage guidance: it is for polling (using afterCursor), explains how maxLines caps the response, and instructs to call telnet_connect first if the console is not connected. It also warns about the single-client limitation, giving clear conditions for correct use.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

input_textSend Text InputA

Type a literal text string into whatever input field is currently focused on the device (ECP /input endpoint). Requires a text field to already be focused — use keypress to navigate into one first. Mutates the focused field: repeated calls append, so this is NOT read-only or idempotent. Use this instead of sending characters as individual keypress keys.

ParametersJSON Schema
NameRequiredDescriptionDefault
textYesText to send.
deviceNoTarget device — IP (e.g. "192.168.1.137") or serial (e.g. "X0004EX9Q7M"). Omit to use the focused device.

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already flag readOnlyHint=false and idempotentHint=false; the description goes beyond those flags by explaining the mechanism — repeated calls append to the focused field — which is precisely the behavioral nuance structured hints cannot convey. The focus prerequisite is also a real runtime condition worth stating. Minor omission: the failure mode when no field is focused is not described.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four sentences, each with a distinct job: core action, precondition, mutation semantics, and alternative routing. No filler or repetition of schema content, and the most decision-relevant facts are front-loaded in the first sentence.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a two-parameter tool with no output schema, the description covers the action, the focus precondition, the append/mutation behavior, and when to choose it over keypress. The only gap is error behavior (e.g., what happens if nothing is focused or input is rejected), which is a minor omission for a simple input-send tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both text and device are already documented with examples and default behavior. The description adds only the 'literal' nuance to the text parameter, which is marginal beyond the schema's 'Text to send.' The baseline 3 for high schema coverage applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource — 'Type a literal text string into whatever input field is currently focused on the device' — anchored to the ECP /input endpoint. It also preemptively distinguishes itself from keypress, its nearest sibling, so an agent can route correctly without opening the schema. The domain is completely unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The precondition is stated explicitly ('Requires a text field to already be focused'), the alternative tool to satisfy it is named ('use keypress to navigate into one first'), and the substitution rule is given ('Use this instead of sending characters as individual keypress keys'). Both when-to-use and when-not-to-use are covered with zero inference required.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

keypressSend Remote KeyA

Send one ECP remote key (e.g. "Home", "Up", "Select", "Play") to a Roku device — mirrors a physical remote press. Mutates on-screen state: each key advances the UI, so repeated calls are NOT a no-op. Use keypress for navigation/transport keys; to type characters into a focused text field use input_text (far faster than sending "Lit_" keys one at a time).

ParametersJSON Schema
NameRequiredDescriptionDefault
keyYesECP key name. See list_keypress_options for the full set.
deviceNoTarget device — IP (e.g. "192.168.1.137") or serial (e.g. "X0004EX9Q7M"). Omit to use the focused device.

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description goes beyond the annotations by explicitly stating that each key advances the UI and repeated calls are not a no-op, which aligns with idempotentHint=false. It also conveys the side-effect profile ('mutates on-screen state') beyond what the annotations alone imply.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences with no filler: the core action, the behavioral caution, and the routing alternative are each compact and front-loaded. Every sentence contributes to correct usage.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple two-parameter tool with full schema coverage and no output schema, the description is complete. It explains what the tool does, its side effect profile, and how to choose it versus the closest sibling, leaving no critical gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, and the schema already documents the key enum and device semantics. The description adds examples of valid keys but does not materially extend meaning beyond the schema, matching the baseline for high coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description states a specific verb-resource pair: send one ECP remote key to a Roku device, with concrete examples. It is clearly distinguished from sibling tools like input_text and ecp_post by describing its exact role as a remote-press mirror.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly names the alternative input_text and gives the condition for choosing it ('to type characters into a focused text field'), while reserving keypress for navigation/transport keys. This is direct when-not guidance that enables correct tool selection.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

launch_appLaunch Roku AppA

Launch a channel / app on the device to its home screen by app id. Discover ids with ecp_query "/query/apps"; "dev" is the sideloaded Dev App. Changes device state (foregrounds the app). To open the app directly on a specific title use deep_link instead; optional launch params are passed through as ECP query params.

ParametersJSON Schema
NameRequiredDescriptionDefault
appIdYesChannel id (e.g. "837" for YouTube, "dev" for sideloaded).
deviceNoTarget device — IP (e.g. "192.168.1.137") or serial (e.g. "X0004EX9Q7M"). Omit to use the focused device.
paramsNoOptional URL-encoded launch params.

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description goes beyond the annotations by stating 'Changes device state (foregrounds the app)', which clarifies the real-world side effect. It also reveals that optional launch params are passed through as ECP query params, adding behavioral detail that annotations alone do not provide.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and front-loaded: the core action is stated first, followed by id discovery, side effect, alternative, and params. Each sentence earns its place with no redundant filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the lack of an output schema, the description still fully covers what an agent needs: what the tool does, how to get valid appId values, the state-change behavior, the sibling alternative, and how params are handled. Nothing critical for correct invocation is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description adds semantic value by explaining how to discover appId values via ecp_query and that params are passed through as ECP query params, which goes beyond the schema's bare field descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb 'Launch' and the resource 'channel / app on the device to its home screen by app id.' It also distinguishes this from the sibling deep_link by explicitly noting deep_link opens a specific title, making the tool's purpose unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly explains how to discover app ids via ecp_query "/query/apps", notes the special 'dev' id for sideloaded apps, and provides a clear alternative: use deep_link when opening a specific title. This gives the agent concrete decision criteria.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_action_typesList Action TypesA
Read-onlyIdempotent

Return every supported Action Script step type (with label, description, required / optional fields). Read-only. Start here when authoring a script, then call get_action_schema for one type's exact fields, and validate_script before send_script_to_builder. For the full authoring contract in one call use get_capability_bundle or read resource roku-dev-studio://action-script-contract.md.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, covering the safety profile. The description adds 'Read-only' (redundant) but also contextual behavior like 'Start here' which frames the tool as a non-mutating entry point. It doesn't contradict annotations and adds mild workflow context, so a 4 is appropriate given the lower bar set by the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, front-loaded with the core purpose, then a concise workflow sequence, then an alternative. Every sentence serves a function: stating output, giving usage order, and offering a fallback. No fluff, and the structure is easy to scan.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter, read-only, idempotent tool with no output schema, the description provides everything needed: what it returns, the recommended order of operations, and alternatives. It's complete for an agent to invoke it correctly without further investigation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, and the schema has no properties. Baseline for 0 params is 4. The description doesn't need to explain any parameters, and the schema is fully covered. Nothing is missing.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'Return every supported Action Script step type' with metadata like label, description, and fields. It clearly distinguishes this from siblings by naming the exact next step (get_action_schema) and positioning itself as the entry point. An agent can immediately understand what this tool does and why it exists.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly says 'Start here when authoring a script' and provides a recommended sequence: call get_action_schema for one type, then validate_script before send_script_to_builder. It also offers an alternative for a full contract (get_capability_bundle or a resource URI). This gives unambiguous when-to-use guidance and alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_app_connector_functionsList App Connector FunctionsA
Read-onlyIdempotent

Live functionName + parameter metadata from RALE getExternalControlFunctions. Each entry has name, params: [{ name, type }, …], and an optional description string when the channel includes one in its payload — surface that description verbatim to the user when explaining what a function does. Call before authoring appFunction steps so names and param keys/order match. Optional device (IP or serial).

ParametersJSON Schema
NameRequiredDescriptionDefault
deviceNoOptional. Roku IP or serial; must match a connected Dev Studio tab when set.

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and non-destructive. The description adds value by explaining the output structure (name, params array, optional description) and instructs to surface the description verbatim to the user, which is behavioral guidance beyond the annotations. It does not contradict annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is multi-sentence but each clause earns its place: the source, the output shape, the verbatim-handling instruction, the usage timing, and the parameter. It is front-loaded with the core purpose and efficiently packed. No redundancy or filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Without an output schema, the description compensates by detailing the return structure (name, params, optional description) and the source. It also explains the usage relationship with appFunction. Given the tool is simple and annotations cover safety, this is sufficiently complete for an agent to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and the device parameter is described in the schema as 'Optional. Roku IP or serial; must match a connected Dev Studio tab when set.' The description simply restates 'Optional device (IP or serial)' without adding new meaning. Baseline 3 applies since the schema already carries the full semantics.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: it lists app connector functions, sourced from RALE getExternalControlFunctions. It clearly distinguishes itself from sibling tools by focusing on function metadata for appFunction authoring, and names the sibling appFunction as the downstream consumer. The purpose is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly says 'Call before authoring appFunction steps so names and param keys/order match', giving a concrete trigger condition. It also notes the optional device parameter, providing context for when to include it. This directly guides the agent on when to invoke this tool relative to appFunction.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_devicesList All Known DevicesA
Read-onlyIdempotent

Return every device Dev Studio already knows about — connected, discovered, remembered, remote, or RCE — without running a network scan. Read-only. Each entry: ip, serial, modelName, friendlyDeviceName, softwareVersion, source, isTabOpen, isTabFocused, isReachable. isTabOpen only means a Dev Studio tab/session exists for this device — it does NOT mean the device will respond right now (it could be powered off or off-network). Check isReachable before relying on a device to answer a live command; other tools that need to reach the device (keypress, ecp_query, rale_command, …) will themselves fail with a clear "not responding" error if it's unreachable — treat isReachable: false as a signal to tell the user to check the device rather than retrying blindly. source is local (physical, this LAN), remote (physical, via an RDS Relay location), or rce (Roku Cloud Emulator — no real IP, ip is a serial stand-in; most main-direct ops work against it directly, but sideload/delete_sideload don't — use connect_device + Dev Studio's Sideload Relay/Dev App tab for those). Use this as the first step to resolve a device argument (IP or serial) for other tools. Related tools: get_selected_device returns only the one focused device; scan_devices actively probes the network for NEW devices not yet known; connect_device opens/focuses a tab for one of these entries.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, openWorldHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is covered. The description adds substantial behavioral context beyond annotations: the meaning of isTabOpen vs isReachable, the source field semantics (local/remote/rce), the fact that rce devices have no real IP, and the caveat that sideload/delete_sideload don't work against rce devices. It also discloses that other tools will fail with a clear 'not responding' error, which sets expectations for downstream behavior. It doesn't describe pagination or ordering, but for a zero-parameter list tool this is minor.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long but every sentence earns its place: it covers scope, field semantics, source types, usage guidance, and sibling differentiation. It is front-loaded with the core purpose and read-only nature, then expands into necessary caveats. It could be slightly tightened, but the density of useful information justifies the length.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter, read-only list tool with no output schema, the description is remarkably complete. It explains the return fields, the meaning of each source type, the isReachable caveat, how to use it as a first step, and how it relates to siblings. An agent has everything it needs to call this tool correctly and interpret its results.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters and schema coverage is 100% (empty object), so there are no parameter semantics to document. The description compensates by thoroughly explaining the return fields and their meanings, which is the closest equivalent to parameter semantics for a no-input tool. Baseline 4 is appropriate for a zero-parameter tool.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource ('Return every device Dev Studio already knows about') and immediately distinguishes itself from scanning by stating 'without running a network scan.' It enumerates the exact fields returned and explicitly contrasts itself with related tools (get_selected_device, scan_devices, connect_device), so an agent can tell it apart from siblings without opening schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly says to use this as the first step to resolve a `device` argument for other tools, and names the alternatives with their conditions: get_selected_device for the one focused device, scan_devices for actively probing new devices, connect_device for opening/focusing a tab. It also gives behavioral guidance on how to interpret isReachable and when to tell the user to check the device rather than retrying.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

network_inspector_analyzeNetwork Inspector: AnalyzeA
Read-onlyIdempotent

Aggregate the captured buffer into hotspots and rollups in one call — counts by event type, by HTTP status class (2xx/3xx/4xx/5xx), top hosts (with error counts), top content types, total HTTP/MITM transactions, error count, and the largest responses. Use this to orient on a session before drilling into individual events. Accepts the same optional filters as network_inspector_list_events (device, host, method, type, status, statusClass, contentType, errorsOnly, mitmOnly). Requires Network Inspector enabled.

ParametersJSON Schema
NameRequiredDescriptionDefault
hostNoOptional case-insensitive substring matched against hostname, TLS SNI, or request URL.
typeNoOptional event type filter.
deviceNoOptional Roku IP or serial. Omit to include every Roku on the hotspot.
methodNoOptional HTTP method filter.
statusNoOne or more values to match (OR).
mitmOnlyNoOnly count decrypted-HTTPS transactions.
errorsOnlyNoOnly count HTTP transactions with a response status >= 400.
contentTypeNoOne or more values to match (OR).
statusClassNoOne or more values to match (OR).

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, openWorldHint, idempotentHint, and destructiveHint false, so the safety profile is covered. The description adds a meaningful behavioral prerequisite ('Requires Network Inspector enabled') and clarifies that the tool aggregates existing captured data rather than capturing new data ('Aggregate the captured buffer'). This goes beyond the annotations without contradicting them.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and front-loaded: it starts with the core action and output list, then the use case, then the shared filter reference, then the requirement. Three sentences, no filler, and every sentence contributes actionable information for tool selection and invocations.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description takes on the burden of describing what the tool returns, and it does so concretely by listing the aggregations (counts by event type, status class, top hosts, content types, transaction totals, error counts, largest responses). It also mentions the enabled prerequisite and the accepted filters. It does not explain the shape of 'hotspots' or pagination, but for an aggregate summary tool this is a minor omission given the output list provided.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the parameters are already fully documented in the input schema. The description adds only that these are 'the same optional filters as network_inspector_list_events', which gives a useful cross-tool consistency hint but does not explain individual parameter semantics beyond the schema. Baseline 3 is appropriate, as the schema carries the detail.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Aggregate the captured buffer into hotspots and rollups in one call' and enumerates the exact computed outputs (counts by event type, HTTP status class, top hosts, content types, etc.). It further distinguishes itself from siblings by stating its intended position in a workflow: 'Use this to orient on a session before drilling into individual events.' An agent can immediately tell this is the aggregation/summary tool, not a list/detail tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description clearly states when to use the tool ('orient on a session before drilling into individual events') and names the sibling whose filters it mirrors (network_inspector_list_events). It also adds a prerequisite ('Requires Network Inspector enabled'). However, it does not explicitly name an alternative for the drill-down step or state a 'when not to use' condition, so it stops just short of fully explicit when/when-not guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

network_inspector_findNetwork Inspector: Find in ContentA
Read-onlyIdempotent

Search the FULL content of captured transactions — request/response URL, headers, and bodies — for query, unlike network_inspector_list_events' host filter which only matches hostname/SNI/URL. This is the tool for "which request(s) contain X" (a session id, an error string, a specific JSON field/value) across the whole buffer, without paging through every event with get_event_detail. Each result carries total (match count), scopes (per-scope breakdown: url/reqHeaders/reqBody/respHeaders/respBody), and the matching event's summary (host/url/method/status) inline. query is required; scopes optionally narrows which parts are searched (omit for all); caseSensitive (default false); regex treats query as a JS regex (a dangerous/over-long pattern safely degrades to a literal search rather than erroring). device optional — omit to search every Roku with captured traffic. limit caps results (default 50, max 500). Requires Network Inspector enabled (see network_inspector_status).

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMax matching events to return (default 50, max 500).
queryYesRequired. Text (or regex, with `regex: true`) to search for.
regexNoTreat `query` as a JS regular expression (default false).
deviceNoOptional Roku IP or serial. Omit to search every Roku with captured traffic.
scopesNoOptional. Which parts to search: url, reqHeaders, reqBody, respHeaders, respBody. Omit for all.
caseSensitiveNoCase-sensitive match (default false).

TDQS

A5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, openWorldHint, idempotentHint, and destructiveHint=false, and the description adds valuable behavioral details: the result structure (total, scopes, summary), the safe degradation of dangerous/over-long regex patterns to a literal search, and the prerequisite that Network Inspector must be enabled. This goes well beyond the annotations and gives the agent a complete picture of expected behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, well-organized paragraph with zero filler. The main purpose is front-loaded, followed by usage distinction, result structure, parameter details, and a prerequisite. Every sentence carries essential information, and the structure makes it easy to scan for the key facts.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (6 parameters, 1 required, no output schema), the description is exceptionally complete. It covers the tool's purpose, when to use it, all parameters with their defaults and interactions, the result structure, and a prerequisite. Nothing an agent needs to call it correctly is missing, and it even explains edge-case behavior (regex degradation).

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% with each parameter described, but the description adds semantic context beyond the schema: it explains that scopes can be omitted to search all parts, that regex treats query as a JS regex with graceful degradation, that device defaults to all Rokus, and that limit caps results. It also clarifies the meaning of the result's scopes breakdown, which is not in the schema. This fully compensates for any gaps and enhances parameter understanding.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Search') and a well-defined resource ('FULL content of captured transactions' including URL, headers, bodies), and explicitly contrasts with sibling network_inspector_list_events' host filter. This makes the tool's purpose unmistakable and differentiates it from alternatives.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly says this is the tool for 'which request(s) contain X' across the whole buffer, and contrasts with using network_inspector_list_events for host-only filtering. It also names an alternative (get_event_detail) and mentions the prerequisite that Network Inspector must be enabled, providing clear when-to-use and when-not-to-use guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

network_inspector_get_ca_infoNetwork Inspector: HTTPS CA InfoA
Read-onlyIdempotent

Return the Dev Studio MITM CA fingerprint, proxy host:port, and the BrightScript snippet needed to trust the proxy so HTTPS request/response bodies become visible to the Network Inspector. Use when network_inspector_list_events shows TLS handshakes but no decrypted HTTP bodies, to guide the user through enabling HTTPS decryption for their sideloaded dev channel.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and idempotentHint=true, but the description adds meaningful context beyond them: what is returned, the proxy trust purpose, and the practical effect of making HTTPS bodies visible. It also notes the 'sideloaded dev channel' prerequisite without contradicting the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, with returned items front-loaded and the use case following immediately. Every clause adds useful information; there is no filler, tautology, or restatement.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a parameterless, read-only, idempotent helper, the description fully covers what the agent needs: what the tool returns, why it is useful, and exactly when to call it. The named return items are sufficiently clear even without an output schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema has zero parameters and schema_description_coverage is 100%, so there is nothing for the description to compensate for. The description appropriately reinforces that only fixed, parameterless information is returned.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'Return the Dev Studio MITM CA fingerprint, proxy host:port, and the BrightScript snippet.' This clearly distinguishes it from sibling network_inspector tools, and the reference to network_inspector_list_events establishes how it fits in the broader workflow.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives an explicit trigger: 'Use when network_inspector_list_events shows TLS handshakes but no decrypted HTTP bodies.' This clearly identifies the situation and user goal, though it does not enumerate when-not-to-use conditions or explicitly name alternative tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

network_inspector_get_event_detailNetwork Inspector: Event DetailA
Read-onlyIdempotent

Fetch the full headers and body for one captured event by id (from network_inspector_list_events). Bodies are capped at maxBodyChars (default 4096) and the response lists warnings when truncated; pass includeFullBody: true to override. DNS/TLS/TCP events have no body and may return 404. Requires Network Inspector enabled.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesRequired. Event id from network_inspector_list_events.
maxBodyCharsNoPer-side body character cap when includeFullBody is not set (default 4096).
includeFullBodyNoReturn untruncated request/response bodies (can be large).

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Adds significant behavior beyond annotations: truncation at maxBodyChars, warnings when truncated, use of includeFullBody to override, and the no-body/404 edge case for DNS/TLS/TCP events. This is consistent with readOnly, idempotent, and destructiveHint false annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, front-loaded with the core purpose and id provenance, followed by behavior caveats and prerequisite. Every sentence contributes actionable information with no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Covers the critical edge cases, prerequisites, truncation, override, and possible 404. It does not fully describe the response envelope beyond 'headers and body' and 'warnings', and there is no output schema, but the description is strong enough for an agent to call it correctly in most cases.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so baseline is 3. The description adds real meaning by tying id to list_events output, providing the default 4096 cap, warning that includeFullBody can return very large data, and explaining per-event body constraints.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Fetch'), a clear resource ('full headers and body for one captured event'), and the id source ('from network_inspector_list_events'). This distinguishes it from list_events and other network inspector sibling tools without ambiguity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides concrete context: use with an id from network_inspector_list_events, requires Network Inspector enabled, and warns that DNS/TLS/TCP events may yield 404. It does not explicitly contrast with alternative sibling tools, but the id-source instruction and clear purpose make appropriate use obvious.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

network_inspector_list_eventsNetwork Inspector: List EventsA
Read-onlyIdempotent

List captured network events as lightweight summaries (no full headers/body — drill down with network_inspector_get_event_detail using an event id). Summary-first by design to protect context. All filters optional and AND-combined across fields; give a field an array to OR within it (e.g. status: [404, 500]): device (IP or serial; omit for all Rokus on the hotspot), host (case-insensitive substring of hostname/SNI/URL), method (GET/POST/…), type (one of the network event types), status (exact HTTP response status code(s)), statusClass ('2xx'|'3xx'|'4xx'|'5xx'), contentType (case-insensitive substring against the response Content-Type, e.g. "json"), errorsOnly (HTTP status >= 400 — a shortcut for statusClass 4xx+5xx), mitmOnly (decrypted-HTTPS transactions only), limit (default 200, max 2000). Returns most-recent events. Requires Network Inspector enabled (see network_inspector_status).

ParametersJSON Schema
NameRequiredDescriptionDefault
hostNoOptional case-insensitive substring matched against hostname, TLS SNI, or request URL.
typeNoOptional event type filter.
limitNoMax events to return (default 200, max 2000).
deviceNoOptional Roku IP or serial. Omit to include every Roku on the hotspot.
methodNoOptional HTTP method filter (e.g. "GET", "POST").
statusNoOne or more values to match (OR).
mitmOnlyNoOnly decrypted-HTTPS transactions captured via the MITM proxy.
errorsOnlyNoOnly HTTP transactions with a response status >= 400.
contentTypeNoOne or more values to match (OR).
statusClassNoOne or more values to match (OR).

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the annotations' readOnly/idempotent hints, the description adds genuinely useful behavioral context: 'Summary-first by design to protect context' (explains the deliberately truncated output), 'Returns most-recent events' (ordering contract), AND/OR filter combinatorics, and the prerequisite state dependency. No contradiction with the annotations exists.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but well-ordered: purpose, drill-down pointer, design rationale, filter semantics, ordering, prerequisite. The long parameter glossary is justified by 10 parameters, though it partially restates schema content and could be tightened by trimming redundant per-field phrasing.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 10-parameter, no-output-schema read tool, the description covers prerequisites, filtering semantics, defaults, ordering, and drill-down routing. The one gap is that it tells the agent what summaries lack ('no full headers/body') but never what fields they contain, leaving the return shape to inference.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Although schema coverage is 100%, the description adds meaning the schema lacks: cross-field AND-combination with OR-within-arrays, the errorsOnly shortcut mapping to statusClass 4xx+5xx, the device omit-for-all-Rokus semantics, and the limit default/max. These semantics materially change how an agent composes a correct call and go well beyond the per-field schema text.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource ('List captured network events as lightweight summaries') and adds a scope qualifier ('no full headers/body'), then explicitly names the sibling it is not (network_inspector_get_event_detail). An agent can distinguish this tool from the other network_inspector_* siblings without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives explicit routing to network_inspector_get_event_detail for drill-down and states the prerequisite ('Requires Network Inspector enabled (see network_inspector_status)'). It does not, however, address the overlapping siblings network_inspector_find and network_inspector_analyze, leaving some ambiguity about when those should be preferred, so it falls just short of full when/not-when coverage.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

network_inspector_statusNetwork Inspector: StatusA
Read-onlyIdempotent

Report whether Dev Studio's Network Inspector is enabled and actively capturing, plus connected Roku clients, packet/event counts, MITM (HTTPS decryption) state, and prerequisites[] remediation. Call this first before the other network_inspector_* tools — if ready is false, relay notice / remediation to the user (enable the feature, grant capture access, connect the Roku to the hotspot). Reads return nothing useful until ready is true.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already signal readOnly/openWorld/idempotent/non-destructive. The description adds important behavioral context beyond those hints: the gating 'ready' condition, the need to relay notice/remediation when not ready, and the fact that reads return nothing useful until ready is true.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two well-structured sentences front-load the purpose and then give actionable gating instructions. The slight repetition that reads are not useful until ready is true reinforces a critical operational constraint rather than being wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description enumerates the key status dimensions an agent needs: ready state, connected clients, packet/event counts, MITM state, and prerequisites remediation. It also gives the conditional workflow for acting on ready=false, making the tool fully understandable for a zero-parameter status probe.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool accepts zero parameters, and the schema covers 100% of them trivially. With no parameters, the description has no parameter meaning to add and instead focuses on status output semantics, which is the appropriate 0-parameter baseline behavior.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb-resource pair ('Report whether Dev Studio's Network Inspector is enabled') and enumerates what it covers: Roku clients, packet/event counts, MITM state, and prerequisites remediation. It also differentiates itself by explicitly instructing 'Call this first' before other network_inspector_* tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit usage directive: 'Call this first before the other network_inspector_* tools'. It also tells the agent exactly how to behave when ready is false (relay notice/remediation) and clarifies that reads are not useful until ready is true.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

probe_bridgeProbe Dev Studio BridgeA
Read-onlyIdempotent

Returns { live, port, pid, startedAt } or { live: false, reason }. Call once per session before the first bridge-dependent tool; once live=true, call direct ops (keypress, launch_app, ecp_query, rale_command, …) and send_script_to_builder freely without re-probing.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The annotations already declare readOnlyHint, idempotentHint, and destructiveHint false, so the safety profile is covered. The description adds valuable behavioral context beyond those annotations: the session-level contract, the 'call once before first bridge-dependent tool' rule, and the `live:false` failure shape with a reason field. This is meaningful supplementary detail rather than a restatement of annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, focused sentence with the return contract front-loaded and the usage rule immediately after. Every clause contributes either to what the tool returns or when to call it. There is no filler or repetition of schema or annotation data.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter, no-output-schema status probe, the description fully equips an agent to invoke it correctly: it explains the return shape, the failure case, and the session-level usage policy. It even names the sibling operations that depend on a successful probe. Nothing essential is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters and schema description coverage is 100%, so there is no parameter meaning the description needs to add. The baseline for a no-parameter tool is 4, and the description appropriately focuses on return-value semantics instead. No parameter-related gap exists.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly defines the tool as a bridge liveness probe by specifying the exact return shape: `{ live, port, pid, startedAt }` or `{ live: false, reason }`. It distinguishes itself from sibling status tools like `debugger_status` or `test_connection` by framing itself as a one-time precondition for bridge-dependent operations. It lacks an explicit verb phrase like 'checks whether', but the return contract makes the purpose unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives an explicit usage rule: call once per session before the first bridge-dependent tool. It also states when re-probing is unnecessary once `live=true`, and it names the affected direct operations such as keypress, launch_app, ecp_query, rale_command, and send_script_to_builder. This is strong, actionable guidance for an agent deciding when to invoke this tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

rale_commandRALE Command (full; read + write)A
Destructive

Run any built-in RALE command against the active App Connector session — including destructive ones (addRegistryField, removeRegistrySection, clearRegistry, …). Use list_rale_builtins for the catalog. Every call surfaces as a toast in Dev Studio. Some commands read (getNodeById, getRegistry) and some write; the tool as a whole is not read-only — for a plain read prefer rale_get_node_by_id.

ParametersJSON Schema
NameRequiredDescriptionDefault
argsNoCommand-specific argument object; the shape depends on `command` (see the `args` template for that entry in list_rale_builtins). Omit for commands that take none.
deviceNoOptional target device (IP or serial). Omit to use the focused tab.
commandYesRALE built-in command name exactly as listed by list_rale_builtins (e.g. "getNodeById", "getRegistry", "clearRegistry").

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true and readOnlyHint=false, and the description adds value beyond them: every call surfaces as a toast in Dev Studio, the mixed read/write nature is spelled out ('the tool as a whole is not read-only'), and concrete destructive command examples are given. No contradiction with annotations; the description reinforces and extends them.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four sentences with no wasted words: purpose and destructive scope are front-loaded, followed by catalog pointer, toast behavior, and read/write routing advice. Every sentence earns its place and the structure moves from most critical (what it does) to supporting detail.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a high-complexity dispatch tool with open-world semantics, no output schema, and a dynamic args object, the description provides strong coverage: destructive warning, catalog discovery path, side-effect visibility (toast), and sibling routing. It could note that output shape varies by command or describe error behavior, but the list_rale_builtins pointer substantially mitigates that gap for a generic runner.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all three parameters including the shape-dependent args object. The description adds a few illustrative command names (getNodeById, getRegistry, clearRegistry) and points to the args template in list_rale_builtins, but it doesn't add substantive parameter semantics beyond the schema — baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource ('Run any built-in RALE command against the active App Connector session') and explicitly flags destructive capability. It differentiates from siblings by naming rale_get_node_by_id as the plain-read alternative and list_rale_builtins as the catalog source, so an agent can select it correctly without opening schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear context: use this for any built-in command including destructive ones, consult list_rale_builtins for the catalog, and prefer rale_get_node_by_id for plain reads. It names an explicit exclusion (plain reads) and the alternative, though it doesn't enumerate when-not-to-use cases beyond reads or address overlap with other write-capable siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

rale_get_node_by_idRALE: Get Node by IDA
Read-onlyIdempotent

Read-only convenience wrapper over rale_command getNodeById: fetch one SceneGraph node (its fields / children) by its id from the running Dev App via the App Connector. Requires a connected App Connector session (auto-connects if needed). Use this — not the general rale_command — for the common "inspect one node" case; drop to rale_command only for other RALE built-ins (registry, focus, other queries). Required id (the node's id field as authored in XML/BrightScript). Optional path (array of child indices/ids to disambiguate when the id is not globally unique; omit or [] for a global lookup) and device (IP or serial; omit for the focused tab).

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesRequired. Node id string from the scene / registry.
pathNoOptional. Scene graph path segments; use [] or omit for root.
deviceNoOptional. IP or serial for a specific Dev Studio device tab.

TDQS

A4.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, openWorldHint, idempotentHint, and non-destructive behavior, so the safety profile is covered. The description adds useful behavioral context: it requires a connected App Connector session, can auto-connect, and fetches node fields/children from the running Dev App. It does not mention return format or error behavior, but annotations lower the burden here.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and front-loaded with the core purpose, then adds usage guidance and parameter semantics in a structured way. Every sentence carries necessary information with no fluff or repetition of schema details.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple read-only fetch tool with three well-documented parameters and strong annotations, the description is complete. It covers the tool's purpose, connection requirements, when to use it versus rale_command, and enough parameter semantics for correct invocation. The expected result (node fields/children) is also stated.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so each parameter is already documented. The description adds meaningful clarifications beyond the schema: id is the authored XML/BrightScript id, path is for disambiguation when id is not globally unique, and device can be an IP/serial or omitted for the focused tab. These details help the agent invoke parameters correctly.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool fetches one SceneGraph node by id and identifies it as a read-only convenience wrapper over rale_command getNodeById. It names the exact resource and operation, and differentiates it from the general rale_command sibling.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly says when to use this tool ('common inspect one node' case), when not to use it, and directs the agent to rale_command for other RALE built-ins. It also states the connectivity requirement and that the tool auto-connects if needed, leaving no ambiguity about prerequisites.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

scan_devicesScan Network for Roku DevicesA
Read-onlyIdempotent

Discover Roku devices on the local network via SSDP (multicast) and, optionally, a subnet HTTP sweep. Read-only; does not connect devices — follow up with connect_device to open a tab. Use this to FIND unknown devices; to list devices Dev Studio already knows (connected / remembered) without a network scan, use list_devices instead.

ParametersJSON Schema
NameRequiredDescriptionDefault
timeoutMsNoOverall SSDP discovery window in milliseconds (default 4000). Does not affect the per-host subnet sweep timeout.
includeSubnetScanNoAlso sweep the local /24 subnet over HTTP to catch devices that did not answer SSDP multicast (slower, more thorough). Default false.

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and idempotentHint=true, so the safety profile is covered. The description adds value by clarifying that the tool 'does not connect devices' and explaining the network scan nature, which is relevant context beyond the structured hints. No contradictions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences with zero waste. The purpose is front-loaded, the scope is stated, and the sibling distinction is given in the second sentence. Every word contributes to agent decision-making.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is simple with two optional parameters and no output schema, but the description covers purpose, side effects (none destructive), usage alternatives, and follow-up actions. The return value (discovered devices) is implicitly clear from 'Discover Roku devices'. An agent has all necessary information to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, meaning both parameters (timeoutMs and includeSubnetScan) are already well-documented in the schema. The description mentions SSDP and subnet HTTP sweep, which loosely maps to the parameters, but it does not add meaningful detail beyond what the schema provides. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb 'Discover' with the resource 'Roku devices on the local network' and specifies the method (SSDP multicast and optional HTTP sweep). It explicitly distinguishes itself from siblings by naming connect_device (for opening a tab) and list_devices (for known devices), so an agent can unambiguously select the right tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides explicit when-to-use and when-not-to-use guidance: 'Use this to FIND unknown devices' and directs to list_devices for already-known devices. It also notes the follow-up action with connect_device, giving a clear workflow. No inference is required.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

screenshotCapture ScreenshotA
Idempotent

Capture a screenshot of the current device screen and return it inline as an MCP image content block (JPEG, base64). Hosts (Cursor, Claude Desktop, etc.) render this image to the user, so for any human-facing capture let returnImageBase64 default to true (or omit it). Set returnImageBase64: false ONLY for batch / metadata-only flows where no one will view the screenshot; in that case the response is just { success, filename, bytes } and the image will not appear in the chat. Password is optional when Dev Studio has remembered it for this device.

ParametersJSON Schema
NameRequiredDescriptionDefault
deviceNoTarget device — IP (e.g. "192.168.1.137") or serial (e.g. "X0004EX9Q7M"). Omit to use the focused device.
passwordNoOmit if Roku Dev Studio has saved the Dev Password for this device (Remember on the device tab).
returnImageBase64NoDefault true. Keep true (or omit) for any user-facing capture so the screenshot is rendered inline in the chat. Set false ONLY for batch / metadata-only flows where no one will view the image; when false, the user will see only the JSON metadata and nothing visible.
waitAfterTriggerMsNoDelay in milliseconds after triggering the capture before reading the image, to let the screen settle after a navigation (default 0). Increase if the screenshot catches a mid-transition frame.

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare idempotentHint=true and destructiveHint=false. The description adds the return behavior: inline image when true, and a JSON object with { success, filename, bytes } when false, plus the note that hosts render the image. It also discloses password optionality. This goes beyond the annotations without contradicting them, though it doesn't mention any side effects (none expected given idempotence).

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is four sentences, each with a purpose: core action, host rendering context, parameter usage guidance, and password note. It is front-loaded with the primary behavior and then provides conditional details. Slightly longer than necessary but every sentence adds value, so it earns a 4.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with 4 parameters, no output schema, and annotations present, the description explains the return format in both modes, when to use the boolean, and password handling. It does not cover error cases or what happens if the device is unspecified, but those are covered by schema and the focused-device concept. The description is sufficient for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description adds meaningful context for `returnImageBase64` (when to set false) and clarifies password behavior ('optional when Dev Studio has remembered it'), which reinforces the schema. For `device` and `waitAfterTriggerMs`, the schema already explains them, so the description adds no extra value there, but the added emphasis on the boolean is valuable.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('Capture a screenshot of the current device screen') and the resource (the device screen). It also specifies the return format ('inline as an MCP image content block (JPEG, base64)'), which distinguishes it from other tools. No sibling performs the same function, so differentiation is inherent, but the description is precise and unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit guidance on when to use the `returnImageBase64` parameter: 'for any human-facing capture let returnImageBase64 default to true (or omit it)' and 'Set returnImageBase64: false ONLY for batch / metadata-only flows'. It also notes the password is optional when remembered, covering usage prerequisites. This is clear, actionable usage direction.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

send_script_to_builderSend Script to BuilderA
Idempotent

Drop a validated Action Script into Dev Studio Builder for human review (does not auto-run). Runs the same validation as validate_script. Arguments: script (object or JSON string), optional device. Use only for multi-step / conditional / saved-or-reviewed flows — if the task is a single deterministic action (one keypress, one launch, one RALE command, one ECP query/POST, one screenshot), call the matching direct op (keypress, launch_app, rale_command, ecp_query, ecp_post, screenshot, …) directly instead of wrapping it in a one-step script.

ParametersJSON Schema
NameRequiredDescriptionDefault
deviceNoOptional. Target Roku IP (e.g. 192.168.1.68) or serial. Must match an open Dev Studio device tab when provided.
scriptYesRequired. Same shape as for validate_script: object with `steps` array, or a JSON string that parses to that object.

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already give readOnlyHint=false, idempotentHint=true, and destructiveHint=false. The description adds meaningful behavioral context beyond annotations: it does not auto-run, performs the same validation as validate_script, and is intended for human review. While it could clarify what happens after validation (e.g., what 'sending to builder' changes in the UI), the added behavioral info is significant.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with the core purpose and key behavioral fact, then covers arguments, then gives decisive usage routing. Every sentence adds value, and the guidance is compact despite covering both when-to-use and when-not-to-use.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description is complete enough for a moderately complex tool with 2 parameters, no output schema, and decent annotations. It covers what the tool does, its validation behavior, parameter types, and the exact conditions for use versus direct alternatives. A minor gap is that it assumes the agent understands what 'Dev Studio Builder' is and what 'human review' entails, but for the intended audience this is sufficient.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%: both `script` and `device` are already documented in the schema with shape and constraints. The description repeats the argument names and basic types, but does not add meaning beyond the schema, except for referencing validate_script's script shape, which the schema itself already mentions. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a clear specific action: sending a validated Action Script into Dev Studio Builder for human review, and explicitly distinguishes it from validate_script and direct operations. The verb 'drop' is informal but unambiguous, and the resource ('Dev Studio Builder') and behavior ('does not auto-run') are concrete.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly states when to use the tool: only for multi-step, conditional, or saved/reviewed flows. It also provides clear when-not-to-use guidance by listing direct operations like keypress, launch_app, rale_command, ecp_query, and screenshot for single deterministic actions. This is exemplary routing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

sideloadSideload Channel PackageA
Destructive

Upload and install a .zip channel package on the device. Destructive: replaces any currently sideloaded Dev App (readOnly false; not safe for autonomous use without user intent). Provide the zip via filePath — an absolute path to a .zip on the SAME machine that runs Roku Dev Studio. Do NOT use this when running in a remote agent sandbox (Claude.ai, ChatGPT web) where files only exist inside the agent's container — the path will not resolve on the user's machine. OR via contentBase64 + filename — the .zip bytes inline; this server writes them to a temp file on the user's machine, sideloads, and cleans up. Use this whenever the agent has file content but no shared filesystem with Roku Dev Studio. Password is optional when Dev Studio has remembered it for this device.

ParametersJSON Schema
NameRequiredDescriptionDefault
deviceNoTarget device — IP (e.g. "192.168.1.137") or serial (e.g. "X0004EX9Q7M"). Omit to use the focused device.
filePathNoAbsolute path to a .zip on the same machine that runs Roku Dev Studio. Mutually exclusive with `contentBase64`. Will fail with an actionable error if the path looks like an agent sandbox path (e.g. /mnt/user-data/...).
filenameNoSuggested filename for the temp file when using `contentBase64` (e.g. "my-app.zip"). Optional but recommended; if omitted, "agent-upload.zip" is used.
passwordNoDeveloper password. Omit if Roku Dev Studio has saved it for this device (Remember on the device tab).
contentBase64NoZip bytes encoded as base64. The server writes them to a temp file on the user's machine, sideloads, then deletes the temp file. Use this when running in a remote agent sandbox so file content travels through MCP rather than relying on a shared filesystem. Provide `filename` alongside.

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already include destructiveHint: true and readOnlyHint: false, but the description adds specific behavioral detail: it replaces the currently sideloaded Dev App, writes base64 content to a temp file, sideloads, then cleans upaint. This meaningfully exceeds the structured annotation hints.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is longer than average, but its length is justified by two mutually exclusive input modes and a meaningful destructive warning. The flow is logical: core purpose, destructive warning, filePath pathway, contentBase64 pathway, then password nuance. Every sentence earns its place, though it could be trimmed slightly without losing meaning.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has no output schema)Skip, but the description covers prerequisites, both input modes, cleanup behavior, destructive effects, and edge-case path failures. It omits what the success response looks like, which an agent might want, but the invocation guidance is thorough enough that this is a minor gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description adds beyond the schema by explaining the real-world constraint that filePath only works on the same machine as Roku Dev Studio)Skip, and by giving a decision rule for choosing contentBase64 over filePath. This adds semantic value, but some sentences mostly restate the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Upload and install a .zip channel package on the device.' It further distinguishes the operation by noting it replaces the currently sideloaded Dev App, which separates it from sibling tools like delete_sideload or launch_app.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly says when not to use this tool ('Do NOT use this when running in a remote agent sandbox') and when to use it ('Use this whenever the agent has file content but no shared filesystem'). It also warns that the operation is destructive and not safe for autonomous use without user intent, giving clear usage boundaries.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

telnet_connectTelnet Console: ConnectA
Idempotent

Open the BrightScript debug console (TCP 8085) for the targeted device, exactly as if the user had clicked the Connect button on the Telnet Console tab. Idempotent: returns { connected: true, already: true } when already attached. Lines do not accumulate until this is called. After it returns successfully, poll the buffer with get_telnet_log({ afterCursor }). Roku's 8085 socket is single-client: connecting here will displace another tool (e.g. an IDE telnet session) that may currently hold it. Opening the socket is a side effect (displaces other clients), so this is not read-only.

ParametersJSON Schema
NameRequiredDescriptionDefault
deviceNoOptional target device (IP or serial). Omit to use the focused tab.

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description explicitly discloses side effects ('Opening the socket is a side effect (displaces other clients), so this is not read-only'), which reinforces the readOnlyHint=false annotation. It also details the idempotent return shape and the requirement to call before logs accumulate, adding behavioral context beyond the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but efficient. It front-loads the core purpose, then covers idempotency, sequencing, and side effects in a logical order. Every sentence adds necessary information without redundancy, achieving high information density without verbosity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description is complete for a connection tool with no output schema. It explains the return value for the idempotent case, the subsequent polling step, and the side effect of displacing other clients. Given the tool's complexity and the lack of an output schema, the description covers all the agent needs to know to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema already describes the optional 'device' parameter with a clear description ('Optional target device (IP or serial). Omit to use the focused tab.'). The description does not add further semantic detail about the parameter beyond what the schema provides, so the baseline 3 applies with 100% schema coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states the exact verb and resource: 'Open the BrightScript debug console (TCP 8085) for the targeted device', explicitly likening it to the Connect button. It distinguishes from siblings like telnet_disconnect and get_telnet_log by its role and sequencing.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit usage guidance: it must be called before lines accumulate, and after successful return the agent should poll with get_telnet_log. It also warns about the single-client nature that can displace other clients, implying when to avoid use (e.g., if another session is needed) and clarifies idempotent behavior.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

telnet_disconnectTelnet Console: DisconnectA
Idempotent

Close the BrightScript debug console (TCP 8085) for the targeted device, mirroring the Disconnect button. Idempotent: returns { connected: false, already: true } when no session is open. Use this to release the 8085 socket so another tool can attach, or to stop log accumulation. Closing the socket is a side effect, so this is not read-only.

ParametersJSON Schema
NameRequiredDescriptionDefault
deviceNoOptional target device (IP or serial). Omit to use the focused tab.

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate readOnlyHint=false and idempotentHint=true. The description adds valuable details: the exact idempotent return value `{ connected: false, already: true }` when no session is open, and explicitly warns that closing the socket is a side effect, reinforcing that it is not read-only. This goes beyond the annotations and fully discloses behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise (three sentences), front-loaded with the primary action, and includes essential details like idempotency and side effects without any fluff. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple tool with one optional parameter and no output schema, the description covers everything an agent needs: the action, the idempotent behavior with a specific return value, the use cases, and the side-effect warning. The annotations complement it well, and nothing essential is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% since the only parameter 'device' has a clear description in the schema ('Optional target device (IP or serial). Omit to use the focused tab.'). The tool description does not add any additional meaning beyond the schema, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the specific verb 'Close' and the resource 'BrightScript debug console (TCP 8085)', and it explicitly mirrors the Disconnect button. It distinguishes itself from siblings like telnet_connect (opposite operation) and get_telnet_log (reading logs) by specifying the exact scope.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides explicit use cases: 'Use this to release the 8085 socket so another tool can attach, or to stop log accumulation.' It clearly implies when to use it, though it does not explicitly name alternative tools or state when not to use it. The context is sufficient for an agent to decide.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

test_connectionTest Device ConnectionA
Read-onlyIdempotent

Probe one known device IP for ECP reachability and return basic device info. Does not require a Dev Studio tab to be open — use it to confirm a specific IP is a reachable Roku before connect_device. Read-only. Differs from its siblings: probe_bridge checks whether Dev Studio itself is running (not a device); scan_devices discovers unknown devices on the network; test_connection verifies one address you already have.

ParametersJSON Schema
NameRequiredDescriptionDefault
deviceNoTarget device — IP (e.g. "192.168.1.137") or serial (e.g. "X0004EX9Q7M"). Omit to use the focused device.

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false. The description adds valuable behavioral context: no Dev Studio tab requirement and that it verifies a single known address. This goes beyond the annotations without contradicting them.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, each earning its place: purpose, a key usage condition, and sibling differentiation. Front-loaded with the core action, no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Complete for a simple read-only probe with one optional parameter. The description covers what it does, when to use it, how it differs from siblings, and the safety profile is already in annotations. No critical information is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The sole parameter 'device' is fully described in the schema (IP or serial, optional, with fallback to focused device). The description does not add new parameter semantics, but schema coverage is 100%, so baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Probe'), a clear resource ('one known device IP for ECP reachability'), and the output ('return basic device info'). It explicitly contrasts with siblings probe_bridge and scan_devices, making its purpose unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit when-to-use: 'confirm a specific IP is a reachable Roku before connect_device.' It also clarifies it does not require a Dev Studio tab, and directly names alternatives with their different purposes, so the agent knows exactly when to pick this tool over siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

validate_scriptValidate Action ScriptA
Read-onlyIdempotent

Validate an Action Script before send_script_to_builder. Argument script: JSON object or JSON string. Response: ok, errors[] (path, code, message, expected?), stepCounts, humanSummary, referenceTools. ok=false is returned as isError. Contract: resource roku-dev-studio://action-script-contract.md. Only author a script for multi-step / conditional / polling / saved-or-reviewed flows — for a single action use the matching direct op (keypress, launch_app, rale_command, ecp_query, ecp_post, screenshot, …).

ParametersJSON Schema
NameRequiredDescriptionDefault
scriptYesRequired. The script root object: at minimum `{ "steps": [ ... ] }`, optionally `version`, `name`, `description`. Pass as a native JSON object, or as a single JSON **string** that parses to that object (not double-encoded).

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already mark the operation as read-only, idempotent, and non-destructive. The description adds meaningful behavioral context: the response structure, that ok=false is surfaced as isError, and the reference to an external contract resource. This goes beyond the annotations without contradicting them.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and information-dense. It front-loads the purpose, then covers the parameter, response contract, error behavior, and usage boundaries in a few sentences with no filler or redundant elaboration.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Despite having no output schema, the description fully explains the response fields and error semantics. Combined with the single well-documented parameter and the explicit usage guidance, the agent has everything needed to invoke and interpret this tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema already provides 100% parameter coverage with detailed explanation of the script object's minimum shape and the two acceptable formats. The description only briefly restates 'script: JSON object or JSON string', adding little beyond the schema. Baseline of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'Validate an Action Script before send_script_to_builder'. It clearly differentiates validation from the subsequent build/send step and from the direct operations listed as alternatives.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly tells when to use this tool ('before send_script_to_builder') and when not to: for single actions, use direct operations like keypress, launch_app, rale_command, ecp_query, ecp_post, screenshot. It also specifies the flow types that warrant a script: multi-step, conditional, polling, or saved-or-reviewed.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 13 tool updatesv1.3.0
    • Changedconnect_device1 field changed
      • changedInput schema / properties / device / description
        Previous value: -"Required non-empty string: LAN IP (e.g. \"192.168.1.75\") or device serial exactly as shown by list_devices / scan_devices."New value: +"Required non-empty string: LAN IP (e.g. \"192.168.1.68\") or device serial exactly as shown by list_devices / scan_devices."
    • Changeddeep_link1 field changed
      • changedInput schema / properties / device / description
        Previous value: -"Target device — IP (e.g. \"192.168.1.154\") or serial (e.g. \"X00046N6S6F\"). Omit to use the focused device."New value: +"Target device — IP (e.g. \"192.168.1.137\") or serial (e.g. \"X0004EX9Q7M\"). Omit to use the focused device."
    • Changeddelete_sideload1 field changed
      • changedInput schema / properties / device / description
        Previous value: -"Target device — IP (e.g. \"192.168.1.154\") or serial (e.g. \"X00046N6S6F\"). Omit to use the focused device."New value: +"Target device — IP (e.g. \"192.168.1.137\") or serial (e.g. \"X0004EX9Q7M\"). Omit to use the focused device."
    • Changedecp_post1 field changed
      • changedInput schema / properties / device / description
        Previous value: -"Target device — IP (e.g. \"192.168.1.154\") or serial (e.g. \"X00046N6S6F\"). Omit to use the focused device."New value: +"Target device — IP (e.g. \"192.168.1.137\") or serial (e.g. \"X0004EX9Q7M\"). Omit to use the focused device."
    • Changedecp_query1 field changed
      • changedInput schema / properties / device / description
        Previous value: -"Target device — IP (e.g. \"192.168.1.154\") or serial (e.g. \"X00046N6S6F\"). Omit to use the focused device."New value: +"Target device — IP (e.g. \"192.168.1.137\") or serial (e.g. \"X0004EX9Q7M\"). Omit to use the focused device."
    • Changedget_app_icon1 field changed
      • changedInput schema / properties / device / description
        Previous value: -"Target device — IP (e.g. \"192.168.1.154\") or serial (e.g. \"X00046N6S6F\"). Omit to use the focused device."New value: +"Target device — IP (e.g. \"192.168.1.137\") or serial (e.g. \"X0004EX9Q7M\"). Omit to use the focused device."
    • Changedinput_text1 field changed
      • changedInput schema / properties / device / description
        Previous value: -"Target device — IP (e.g. \"192.168.1.154\") or serial (e.g. \"X00046N6S6F\"). Omit to use the focused device."New value: +"Target device — IP (e.g. \"192.168.1.137\") or serial (e.g. \"X0004EX9Q7M\"). Omit to use the focused device."
    • Changedkeypress1 field changed
      • changedInput schema / properties / device / description
        Previous value: -"Target device — IP (e.g. \"192.168.1.154\") or serial (e.g. \"X00046N6S6F\"). Omit to use the focused device."New value: +"Target device — IP (e.g. \"192.168.1.137\") or serial (e.g. \"X0004EX9Q7M\"). Omit to use the focused device."
    • Changedlaunch_app1 field changed
      • changedInput schema / properties / device / description
        Previous value: -"Target device — IP (e.g. \"192.168.1.154\") or serial (e.g. \"X00046N6S6F\"). Omit to use the focused device."New value: +"Target device — IP (e.g. \"192.168.1.137\") or serial (e.g. \"X0004EX9Q7M\"). Omit to use the focused device."
    • Changedscreenshot1 field changed
      • changedInput schema / properties / device / description
        Previous value: -"Target device — IP (e.g. \"192.168.1.154\") or serial (e.g. \"X00046N6S6F\"). Omit to use the focused device."New value: +"Target device — IP (e.g. \"192.168.1.137\") or serial (e.g. \"X0004EX9Q7M\"). Omit to use the focused device."
    • Changedsend_script_to_builder1 field changed
      • changedInput schema / properties / device / description
        Previous value: -"Optional. Target Roku IP (e.g. 192.168.1.75) or serial. Must match an open Dev Studio device tab when provided."New value: +"Optional. Target Roku IP (e.g. 192.168.1.68) or serial. Must match an open Dev Studio device tab when provided."
    • Changedsideload2 fields changed
      • addedInput schema / oneOf
        Added value: +[
        +  {
        +    "required": [
        +      "filePath"
        +    ]
        +  },
        +  {
        +    "required": [
        +      "contentBase64"
        +    ]
        +  }
        +]
      • changedInput schema / properties / device / description
        Previous value: -"Target device — IP (e.g. \"192.168.1.154\") or serial (e.g. \"X00046N6S6F\"). Omit to use the focused device."New value: +"Target device — IP (e.g. \"192.168.1.137\") or serial (e.g. \"X0004EX9Q7M\"). Omit to use the focused device."
    • Changedtest_connection1 field changed
      • changedInput schema / properties / device / description
        Previous value: -"Target device — IP (e.g. \"192.168.1.154\") or serial (e.g. \"X00046N6S6F\"). Omit to use the focused device."New value: +"Target device — IP (e.g. \"192.168.1.137\") or serial (e.g. \"X0004EX9Q7M\"). Omit to use the focused device."
  2. 36 tool updatesv1.0.2
    • Addedapp_connector_connect
    • Addedapp_connector_disconnect
    • Addedapp_function
    • Addedconnect_device
    • Addedconsole_monitor_findings
    • Addeddebugger_attach
    • Addeddebugger_continue
    • Addeddebugger_detach
    • Addeddebugger_evaluate
    • Addeddebugger_get_callstack
    • Addeddebugger_get_variables
    • Addeddebugger_list_breakpoints
    • Addeddebugger_pause
    • Addeddebugger_remove_breakpoints
    • Addeddebugger_set_breakpoints
    • Addeddebugger_status
    • Addeddebugger_step
    • Addeddebugger_wait_for_stop
    • Changeddeep_link4 fields changed
      • addedInput schema / properties / appId / description
        Added value: +"Channel id to launch (e.g. \"837\" for YouTube, \"dev\" for the sideloaded Dev App). Discover ids with ecp_query \"/query/apps\"."
      • addedInput schema / properties / contentId / description
        Added value: +"App-specific content identifier to deep-link to (the value the channel expects for this title/episode). Omit for a plain launch."
      • changedInput schema / properties / mediaType / description
        Previous value: -"e.g. \"movie\", \"episode\", \"series\"."New value: +"Content kind, e.g. \"movie\", \"episode\", \"series\", \"season\", \"short-form\"."
      • addedInput schema / properties / params
        Added value: +{
        +  "additionalProperties": {
        +    "type": "string"
        +  },
        +  "description": "Extra key/value query params beyond contentId/mediaType, for channels that expect additional launch args (e.g. { \"season\": \"2\" }).",
        +  "type": "object"
        +}
    • Addeddevice_performance_metrics
    • Changedecp_post1 field changed
      • addedInput schema / properties / endpoint / description
        Added value: +"ECP path to POST to (e.g. \"/sgrendezvous/track\", \"/input/12345\"). Prefer a value from list_post_presets; arbitrary paths are sent verbatim."
    • Addedecp_query
    • Changedget_app_icon1 field changed
      • addedInput schema / properties / appId / description
        Added value: +"Channel id whose icon to fetch (e.g. \"837\" for YouTube, \"dev\" for the sideloaded Dev App). Get ids from ecp_query \"/query/apps\"."
    • Addedget_telnet_log
    • Addedlist_action_types
    • Addedlist_app_connector_functions
    • Addedlist_devices
    • Changednetwork_inspector_analyze3 fields changed
      • addedInput schema / properties / contentType
        Added value: +{
        +  "description": "One or more values to match (OR).",
        +  "items": {
        +    "description": "Case-insensitive substring against the response Content-Type (falls back to the request's), e.g. \"json\" or \"image\".",
        +    "type": "string"
        +  },
        +  "type": "array"
        +}
      • addedInput schema / properties / status
        Added value: +{
        +  "description": "One or more values to match (OR).",
        +  "items": {
        +    "description": "Exact HTTP response status code, e.g. 404.",
        +    "type": "number"
        +  },
        +  "type": "array"
        +}
      • addedInput schema / properties / statusClass
        Added value: +{
        +  "description": "One or more values to match (OR).",
        +  "items": {
        +    "description": "Response status class.",
        +    "enum": [
        +      "2xx",
        +      "3xx",
        +      "4xx",
        +      "5xx"
        +    ],
        +    "type": "string"
        +  },
        +  "type": "array"
        +}
    • Addednetwork_inspector_find
    • Addednetwork_inspector_get_ca_info
    • Addednetwork_inspector_list_events
    • Changedrale_command3 fields changed
      • addedInput schema / properties / args / description
        Added value: +"Command-specific argument object; the shape depends on `command` (see the `args` template for that entry in list_rale_builtins). Omit for commands that take none."
      • changedInput schema / properties / command / description
        Previous value: -"RALE built-in command name."New value: +"RALE built-in command name exactly as listed by list_rale_builtins (e.g. \"getNodeById\", \"getRegistry\", \"clearRegistry\")."
      • changedInput schema / properties / device / description
        Previous value: -"Optional target device (IP or serial)."New value: +"Optional target device (IP or serial). Omit to use the focused tab."
    • Addedscan_devices
    • Addedscreenshot
    • Addedtelnet_connect
    • Addedvalidate_script
  3. 20 tool updatesv1.0.1
    • First observeddeep_link
    • First observeddelete_sideload
    • First observedecp_post
    • First observedget_action_schema
    • First observedget_app_icon
    • First observedget_capability_bundle
    • First observedget_selected_device
    • First observedinput_text
    • First observedkeypress
    • First observedlaunch_app
    • First observednetwork_inspector_analyze
    • First observednetwork_inspector_get_event_detail
    • First observednetwork_inspector_status
    • First observedprobe_bridge
    • First observedrale_command
    • First observedrale_get_node_by_id
    • First observedsend_script_to_builder
    • First observedsideload
    • First observedtelnet_disconnect
    • First observedtest_connection

TDQS

A4/5.0

Scored across 51 tools

Disambiguation4/5

Tools are grouped into clear functional families (action scripting, debugger, network inspector, ECP, RALE/telnet, device discovery), and each has an explicit, well-scoped purpose. A few near-neighbor pairs like probe_bridge vs test_connection or rale_command vs app_function could be confused, but their descriptions draw clear boundaries and cross-reference each other.

Naming Consistency3/5

Many tools follow a consistent verb_noun or domain-prefix_verb pattern, but the set mixes conventions: keypress, launch_app, app_function, rale_command, network_inspector_status, debugger_status, and device_performance_metrics don't fit a single style. The domain prefixes like debugger_, telnet_, and network_inspector_ help readability, but the naming is not fully predictable across all 51 tools.

Tool Count2/5

At 51 tools this is well beyond the 25+ threshold that the rubric treats as too many, even though the scope spans multiple substantial subsystems. The count is heavy for agents to navigate and is only partially mitigated by clear grouping.

Completeness4/5

The surface is very complete across device discovery, ECP control, debugging lifecycle, network inspection, telnet, sideloading, RALE, and action scripting, with full read/read-write and lifecycle coverage. The main gap is that descriptions reference helper catalog tools like list_query_presets, list_post_presets, and list_rale_builtins that aren't actually exposed in the tool list, though get_capability_bundle partially compensates.

Maintenance

ActivityActive
ResponsivenessUnresponsive

Related MCP Connectors

Related MCP Servers

  • F
    license
    Not graded
    quality
    D
    maintenance
    Enables AI agents to develop, test, and certify Roku applications by providing direct control over device functions like app deployment, remote input, and SceneGraph inspection. It supports automated workflows including real-time log collection, media monitoring, and certification verification.
    1
    -
  • F
    license
    Not graded
    quality
    B
    maintenance
    MCP server for Roku BrightScript documentation and device control, enabling doc search, device introspection, keypress/keysequence input, app launch, and sideloading.
    -
  • A
    license
    Not graded
    quality
    D
    maintenance
    Enables AI agents to inspect and control Roku devices—query UI elements, send remote input, launch channels, and run tests—using the Model Context Protocol or a CLI.
    18 npm
    4
    MIT
  • A
    license
    A
    quality
    B
    maintenance
    An MCP server that scaffolds runnable Roku channels from a validated AppSpec, zips them, and optionally sideloads to a Roku device.
    5
    11 npm
    MIT