Skip to main content
Glama

mcp-vm-relay

The Model Context Protocol front end of the VM relay, packaged as a Claude Code plugin and as a pi package. pi-vm-relay, which gave pi a native relay tool directly, is retired; mcp-vm-relay replaces it for both Claude Code and pi, the latter loaded through pi-mcp-adapter. It is an MCP server: nineteen relay_* tools, each with its own schema, title and annotations, and the server's instructions. The two user commands, status and trajectory, ship as one small command file per host (see "Commands"). There is no skill and no hook. Together they offer recorded, snapshot-evidenced interruptive computer-use and browser-use in a fresh, dedicated VM. The agent judges interruption; non-disruptive work stays with the agent's local tools. Nothing here executes local computer-use, spawns subagents, targets physical machines, records video, or attaches to existing browsers.

How it is built

The relay implementation lives in this repository: the manager (src/manager.ts), the vm-service client, the registry with its OS-level lock, the guest transfer and transport, the evidence package and its state merge, the selected-environment profile, the console and saved-image contracts, the strict tool contract (src/schema.ts), the action dispatch and bounded result rendering (src/surface.ts), and the two guest programs (src/guest/receiver.ts, src/guest/mcp-host.ts). It was inherited from pi-vm-relay at commit 8990123, synchronized with its implementation through commit 0d69fc7 before that project's retirement, and is maintained here as mcp-vm-relay's own implementation from then on. src/server.ts binds the core to MCP over standard input and output.

The committed dist/ holds server.mjs (the MCP server with the core bundled in), the guest bundles the manager stages (receiver.mjs, mcp-host.mjs) and the doctor. Compiled relay-driver code, the recorded-execution and evidence substrate, is bundled with hashes and provenance in dist/build-info.json. Consumers load dist/; no build, SDK checkout or network fetch happens at install time.

The Claude Code plugin and pi both reach this same server as plain MCP: pi through pi-mcp-adapter (see "Use with pi" below), Claude Code by starting it directly. Every tool's identity, schema and contract are shared; the plugin adds nothing Claude-Code-specific beyond the marketplace packaging and the vm-relay-operator agent definition.

In pi-vm-relay (retired, native pi extension)

Here (MCP server)

a registered relay tool with thirteen actions

nineteen relay_* MCP tools from the server relay, each with its own schema, title and annotations

doctrine injected before each agent turn

the server's MCP instructions, plus a compact version in the relay_probe and relay_acquire descriptions for clients that do not surface instructions

/relay-status, /relay-trajectory commands

/mcp-vm-relay-status and /mcp-vm-relay-trajectory in pi, /mcp-vm-relay:status and /mcp-vm-relay:trajectory in Claude Code; each asks the assistant to call relay_status or relay_trajectory

prompt-composed enclosure

the vm-relay-operator agent definition (Claude Code only), limited to the relay's tools and read-only file tools

typed image blocks in a run result, with pi's tool_result hook keeping the error flag

MCP image content blocks in the tool result, with the MCP isError flag set beside them

session shutdown pauses lease renewal

the same when the server's stdio closes or it is signalled: renewal pauses, the recording detaches, the VM is retained for an explicit relay_finish or relay_release, and the vm-service TTL is the backstop

the settled agent pauses renewal

none: MCP has no such hook, so the instructions and each run tool's description say "finish or release before you return"

Related MCP server: screenbox

Install

Use with Claude Code

Try it from a checkout:

claude --plugin-dir /path/to/mcp-vm-relay

Or add this repository as a marketplace and install the plugin from it:

/plugin marketplace add WeZZard/mcp-vm-relay
/plugin install mcp-vm-relay@mcp-vm-relay

Once loaded, the model sees the nineteen tools as mcp__plugin_mcp-vm-relay_relay__relay_search, mcp__plugin_mcp-vm-relay_relay__relay_probe, and so on through mcp__plugin_mcp-vm-relay_relay__relay_trajectory (see "The tools" below for the full table). The vm-relay-operator agent definition allows them by those names. The server can also be configured directly in a project's .mcp.json with node /path/to/mcp-vm-relay/dist/server.mjs, in which case the tools are mcp__relay__relay_search and so on.

Use with pi

pi install npm:@wezzard/mcp-vm-relay

This needs pi-mcp-adapter installed. The tools appear as relay_search, relay_probe, and so on through relay_trajectory, with no host-specific prefix (pi-mcp.json sets toolPrefix: "none"). The package also ships pi prompt templates for the two commands, /mcp-vm-relay-status and /mcp-vm-relay-trajectory.

Use with npx

npx -y @wezzard/mcp-vm-relay

This runs the MCP server directly on stdio, for a client that speaks MCP without a Claude Code plugin or pi package wrapper.

Prerequisites

Node 22+, Python 3 with fcntl for the registry lock, and for VM execution an Apple-silicon Mac with Tart, a running vm-service and prepared guest images. Check readiness with node dist/doctor.mjs. It does not acquire a VM or establish guest UI readiness. Guest execution needs Node and CuaDriver with capture/input permissions; Linux requires native X11 and serve --no-overlay. The playwright and chrome-devtools targets of relay_run need a browser in the guest (installed Google Chrome, or a path given as browserExecutable); their pinned servers are downloaded once on the host and staged, so the guest needs no network or npm. Console viewing needs the vm-service guest-sharing backend and its guest preparation; see docs/console.md here.

The tools

Each tool's inputSchema is a plain object with no root anyOf/oneOf/allOf and no action field, derived directly from the strict per-action contract in src/schema.ts (nested unions inside a property, such as relay_image's target, are unaffected). A call is mapped back to that contract's shape and dispatched exactly as the retired single relay tool was, so behaviour, results, image blocks and isError semantics are unchanged from before this split. relay_run forwards the cua-driver, Playwright MCP and Chrome DevTools MCP tool calls a model already knows; relay_exec, relay_script and relay_code share the optional reason, step, snapshots and timeoutMs. Evidence is automatic for all four: see "relay_run" below.

Tool

Title

Replaces (action, kind)

Purpose

relay_search

Search installed applications

search

Find installed applications from image inventories by name and optional OS.

relay_probe

Probe host and guest readiness

probe

Default scope: "host" reports host facts, service availability and owned state. scope: "guest" checks the Node and CuaDriver executables and each relay_run target's availability on the owned guest without installing, starting or claiming capture readiness.

relay_acquisition_capabilities

Read acquisition capabilities

acquisition-capabilities

Read versioned acquisition options (VNC backends) without allocation or ownership recovery.

relay_acquire

Acquire a VM

acquire

Register the task, acquire a fresh VM and start its heartbeat; declare outputs before work. Optional vnc prepares sharing without opening a viewer.

relay_stage

Stage the guest runtime

stage

Push and hash-check the runtime, support files and an opt-in workspace; browserExecutable names the guest browser for the browser targets. Failed staging can be retried; a staged runtime accepts corrected executable paths; resetRecording: true archives the recording and starts a new one on the same VM.

relay_exec

Run a guest command

run, exec

One recorded guest command. diagnostic: true records command diagnosis or repair without screenshot evidence, including before staging.

relay_script

Run a guest script

run, script

One recorded guest script file (localPath, language).

relay_code

Run guest code

run, code

One recorded inline guest code snippet (code, language).

relay_run

Run an MCP tool call in the VM

run, mcp

One recorded tool call (target: cua, playwright or chrome-devtools; that server's own tool and args), forwarded unchanged to the server inside the VM. Replaces relay_cua and relay_browser.

relay_tools

List a target's tools

tools

The target's real tools/list from its in-guest server: names, descriptions and input schemas; tool narrows to one.

relay_image

Retrieve a saved image

image

Retrieve one saved display image, declared application image or immutable image reference without input, capture, directory export or acquisition.

relay_extract

Extract declared outputs

extract

Pull only declared files or directories with source and host checksum verification.

relay_finish

Finish and deliver evidence

finish

Extract declared outputs, deliver and verify the snapshot package, destroy the VM and unregister.

relay_release

Release the VM

release

Retain available evidence and abandon or destroy the VM.

relay_console_resolve

Resolve console status

console-resolve

Resolve non-secret console status for the owned lease and environment.

relay_console_open

Open console viewing

console-open

Open explicitly user-requested viewing on the service host, with console_id, attempt_id, userRequested: true, reason and expected.

relay_console_cancel

Cancel console viewing

console-cancel

Cancel the identified viewing attempt without releasing the VM.

relay_status

Show relay status

(unchanged)

This session's owned lease: backend binding, guest state, renewal state, console observation, last error, staging state and evidence path, plus the project directory, the VM service origin and the selected environment; {"active": false} when nothing is owned.

relay_trajectory

Open the trajectory viewer

(unchanged)

Verify a delivered evidence package (every artifact, hash and reference) and open its trajectory viewer in the local human-facing browser; human review remains pending.

readOnlyHint/idempotentHint are true for relay_search, relay_probe, relay_acquisition_capabilities, relay_console_resolve, relay_image and relay_status; destructiveHint is true only for relay_finish and relay_release, which destroy the VM; openWorldHint is true only for the four run tools, whose guest code may reach the network.

A run's result carries the execution identity and outcome, the derived step record, and, when the snapshot plan captured the after phase and delivery succeeded, the saved after-image as an MCP image content block. A command's result also carries the guest's bounded standard output and error; a relay_run result carries the target tool's own text and image blocks. The full receipt stays in the evidence package.

Commands

Each command asks the assistant to call one relay tool and report the answer. It changes nothing itself.

Command in pi

Command in Claude Code

Arguments

Purpose

/mcp-vm-relay-status

/mcp-vm-relay:status

none

Calls relay_status and reports this session's owned lease as is. Read-only.

/mcp-vm-relay-trajectory

/mcp-vm-relay:trajectory

the package directory

Calls relay_trajectory, which verifies the package and opens its viewer. Human review remains pending.

They are not MCP prompts. pi's MCP adapter can only name a prompt /mcp__<package>__<server>__<prompt>, so each host gets its own command file instead: pi prompt templates in pi-prompts/ (listed in package.json under pi.prompts) and Claude Code plugin commands in commands/. Both are generated by npm run build from scripts/host-commands.mjs; do not edit them by hand. CI and the release workflow check the packed tarball with scripts/verify-package.mjs: the generated files must match their source, pi's glob must select exactly the two templates, and the packed server must start without the prompts capability. The release publishes and attaches the same tarball it checked.

Ownership, failure and recovery

  • One server session owns at most one VM. Another task needs a separate acquisition.

  • An operation failure retains the VM so the agent can inspect, repair and submit a new operation. Failed and uncertain operations are never automatically replayed.

  • Use finish to deliver evidence and release, or release to abandon explicitly. A failed delivery retains the VM; a release failure retains ownership until destruction is verified.

  • When the session ends (the server's stdio closes or it is signalled), lease renewal pauses and the recording detaches; the VM is not destroyed. The backend TTL and grace period handle abandoned leases. Status reports the last confirmed expiration time.

  • A restarted server with the same MCP_VM_RELAY_SESSION reconciles its durable ownership and reattaches the recording session without replaying prior work. Context compaction does not reset VM state; probe reports the owned state without relying on earlier messages.

  • A run tool's timeoutMs defaults to 120,000 ms and accepts integers up to 3,600,000 ms. It bounds command execution, not the snapshot delay or the lease lifetime. A timeout reports confirmed termination or uncertainty and keeps the VM available.

  • relay_exec with diagnostic: true records command diagnosis or repair without screenshots, for explicitly requested diagnosis when capture is unavailable. Before staging it runs through vm-service in the guest's default directory; after staging in the recording workspace. It cannot join a snapshot group and is not visual verification.

  • After diagnosing damaged recording state, relay_stage with resetRecording: true archives the old recording and starts a new recording session in the same VM. An existing receiver lock refuses the reset until its operation is reconciled. Prior evidence paths remain in the owned status.

relay_run: MCP tool calls in the VM

relay_run takes the same tool calls a model already sends to cua-driver, Playwright MCP or Chrome DevTools MCP and runs them against that server inside the VM:

{"target":"playwright","tool":"browser_navigate","args":{"url":"https://example.com"}}
{"target":"cua","tool":"click","args":{"pid":812,"window_id":3,"x":40,"y":90}}
{"target":"chrome-devtools","tool":"take_snapshot","args":{}}

relay_tools {"target":"playwright"} returns the server's real tools/list, so the exact names and schemas are one call away. The arguments are checked against the tool's input schema (a JSON Schema validator, Ajv) before anything is sent; a mismatch is refused with the correct schema in the error.

  • Evidence is automatic. Every call is admitted by the guest receiver, journaled, and bracketed by a before and an after display snapshot. The step record is derived from the call: the title is <target>.<tool> (for example playwright.browser_click), the expected result is "returns without a tool error", and cua-driver's accessibility forms are labelled accessibility. reason, expected and afterIntervalMs are optional overrides. The same rule now applies to relay_exec, relay_script and relay_code: their reason, step and snapshots are accepted but no longer required.

  • Default waits before the after-snapshot: 300 ms for playwright and chrome-devtools (both servers already wait for the page to settle before they answer), 500 ms for cua and for commands.

  • Sessions persist. A small guest-resident MCP host keeps one MCP SDK client per target, starting each server on first use, so a browser page or a native session survives between calls. relay_finish and relay_release stop it.

  • Results pass through. The target's text blocks follow the relay's own result text, bounded like every result (50 KiB, 2000 lines); its images (up to four) are delivered inline under the usual image rules, before the relay's after-snapshot. Each call's full result, images and the servers' logs land in workspace/relay-run, the extraction relay-run, and come home with relay_finish.

  • Outcomes. A call the relay proves was never sent is refused; a call that may have reached the server and has no answer (timeout, crash, a malformed answer) is uncertain, and the server is stopped so nothing it still holds can act later; the target's own isError: true is completed-with-tool-error (exit status 3). All three are error results and keep the VM. Nothing is ever replayed.

The launch commands, pinned versions and the full outcome mapping are in docs/relay-run.md.

Saved images

A run tool returns its saved after-image as a typed image block whenever the snapshot plan captures that phase: a standalone event or a text group's last event. A first or intermediate group event and a diagnostic command do not invent an image. Execution and image delivery are independent outcomes: a completed command can have a failed delivery, and a delivered image does not turn a failed command into a successful one. Either failure sets the MCP isError flag while the content, image included, is kept. The delivery identity (imageDelivery) leads the result text so it survives truncation.

If inline delivery fails, relay_image retrieves the same saved image through one closed selector; it never repeats input, creates a capture, exports the consumer directory or acquires a VM:

{"target":{"source":"display","sessionId":"<recording-session>","executionId":"<saved-execution>","phase":"after"}}
{"target":{"source":"application","name":"relay-run","path":"playwright/<call>/image-1.png"}}
{"target":{"source":"reference","imageId":"image-<64 lowercase hex digits>"}}
  • Display selection needs before or after; a phase the plan did not request returns not-requested. Application selection needs an acquisition-time declaration: a directory declaration takes one relative file path, a file declaration omits path. A reference resolves to the same original bytes or an explicit failure.

  • Originals are limited to 64 MiB, 40,000,000 decoded pixels and 32,768 pixels per dimension; PNG, JPEG and WebP are accepted. The preview is at most 2,000 × 2,000 pixels and 4 MiB of base64. Presentation differs from pi here: pi resizes through its own image helper, while this core has no codec dependency. A PNG original is decoded and, when it exceeds the preview bounds, resampled in-process (area averaging, node:zlib); a JPEG or WebP original has its container validated and passes through unchanged when within bounds, and is reported presentation-unavailable otherwise. No EXIF orientation is applied; dimensions are those stored. The presentation receipt records the policy, the original and preview dimensions, the MIME type, hash, size and whether bytes were transformed, beside the untouched original.

  • Each delivery has a 90-second deadline and a shared budget of three byte-transfer attempts. Only transient transfer failures are retried; authorization, unsafe-path, capture, format and integrity failures are not. Recommend at most two explicit recovery calls for the same reference; if a required image still cannot be inspected, stop exploratory input and finish or release.

  • Immutable catalog entries bind each reference to the owner, enclosure, backend and original identity. Closed or sealed enclosures return stale-reference; delivered originals remain readable with the host's file tools. Finalization materializes verified originals at canonical snapshot paths and preserves prior state when merging guest evidence.

  • Attachment means the block was included in the result; it does not prove the model inspected it or that a human reviewed it.

Live console viewing

Acquisition never opens a viewer. relay_acquisition_capabilities reads the backend's VNC options; relay_acquire with vnc: true checks OS availability and requires a ready console and lease identity in the response. relay_console_open is permitted only following an explicit user request, requires userRequested: true, console_id, attempt_id, reason and expected, and opens on the declared service host, not on a remote client. relay_console_resolve refreshes the non-secret observation; relay_console_cancel closes managed viewing resources and never releases the VM. Console status ready is guest preflight, not launch success; transport connection, authentication, displayed pixels and human confirmation are separate observations that console actions never establish. On macOS the human must choose Standard sharing of the existing console, not a new Log In session or a High Performance display; the relay never confirms that selection, and server_enforced_view_only: false with session_binding: viewer-selection-unverified remain limitations after transport connects. Failed or uncertain launches retain ownership; resolve or cancel the attempt rather than replaying it. Console tests use mocked backends; no live viewer, installation or platform acceptance is claimed here.

Selected environments

Set VM_ENVIRONMENT_FILE to a profile that binds a loopback vm-service endpoint, an image repository, a Tart store and separate service, image and relay state directories together. The profile is validated before use, is authoritative over the individual MCP_VM_RELAY_* variables, and an invalid selection fails rather than falling back. Leases are bound to their backend and store: an owned lease restored under a different environment fails before any backend operation, and destruction is verified with the selected Tart binary and TART_HOME. docs/selected-environments.md documents the schema. The bundle the profile exports to subprocesses is the canonical one vm-service defines, so its relay entries keep the VM_RELAY_STATE_DIR and VM_RELAY_URL names and the three runtimes agree; this server reads its own settings from the profile.

Evidence and cleanup

Default output is relay-evidence/<unique-task>/ under the project, with state/ (the guest journal, action records, receipts and snapshot PNGs), host/ (reasons, submissions, transfer facts, receipts, diagnostics, image deliveries and lifecycle events), extractions/ (declared files, including the page captures), and manifest.json, summary.json, trajectory.json, index.html and OPENING.txt. Open index.html directly: no server, network, VM or external assets are needed. Diagnostic commands appear in the trajectory as command-only steps without screenshots. finish verifies the package before destroying the VM; a failed delivery removes only the derived files it created and retains the VM for a corrected attempt. Cleanup checks the read-only host Tart inventory (the selected one, when an environment is selected) before claiming destruction. Leases default to 4 hours with a heartbeat; the vm-service reaper is the final backstop for process death.

Configuration

  • MCP_VM_RELAY_PROJECT: the project directory (the plugin passes Claude Code's). Evidence lands under relay-evidence/<task>/ there.

  • MCP_VM_RELAY_SESSION: an explicit session identity. By default each server process takes a fresh random one; a stable identity lets a restarted server reconcile and reattach the enclosure of the same identity.

  • MCP_VM_RELAY_URL: the loopback vm-service origin, default http://localhost:6240. Non-loopback servers, redirects and physical targets are refused.

  • MCP_VM_RELAY_STATE_DIR: host state, otherwise $XDG_STATE_HOME/mcp-vm-relay or ~/.local/state/mcp-vm-relay, one subdirectory per session.

  • MCP_VM_RELAY_REGISTRY: an explicit task registry file; otherwise an existing compatible ~/AGENTS.md VM table, or a private managed registry.md in the state directory.

  • MCP_VM_RELAY_PYTHON: the Python 3 used for the registry lock; default python3 on PATH.

  • VM_ENVIRONMENT_FILE: a selected environment profile; it overrides the URL, state directory and registry above and binds the Tart store.

  • Guest files go to /var/tmp/<UTC-timestamp>-mcp-vm-relay-<task>/, with the runtime in receiver.mjs, support in support/ and work in workspace/.

Development

Build relay-driver using its own instructions, then in this checkout:

npm ci
npm run setup:dev -- /absolute/path/to/built/relay-driver
npm run check                                            # build, typecheck, tests

Three tests need a sibling checkout beside this repository and skip with a message when it is absent: the console fixture regeneration and the environment parity test need ../vm-service, and the shared catalog text corpus needs ../pilot-images. The JPEG and WebP presentation fixtures in tests/fixtures/ were encoded once from the test's spatial PNG with a codec outside this repository and are committed, because the core carries none.

End-to-end with headless Claude Code

npm run test:e2e

This runs claude --print with the plugin loaded from this checkout (--plugin-dir), the model limited to the relay's tools, and the server pointed at a loopback stand-in for vm-service (tests/e2e/fixture-service.ts), which keeps leases as records, runs the relay's own Node and transfer commands locally, and fakes the desktop capture driver. Four cases run, each a separate headless session: the status call (which also proves the server is pointed at the stand-in before anything is acquired), an invalid call refused by the contract, a full acquire, stage, exec, finish lifecycle with a verified package, and a lease abandoned when the session ends, which is retained with its renewal paused rather than released. What is checked is what happened on disk and at the service, not what the model said. The suite needs claude on PATH with an account and spends a few model turns per case, so it is not part of npm test.

The relay-driver links are development-only. The tests run the manager, the guest programs and the shipped server against fixture HTTP services, fixture MCP servers built with the MCP SDK and a real MCP client over stdio; no VM, desktop or user registry is touched. Commit the regenerated dist/ with source changes.

Releasing

Cut a release with:

node scripts/bump.mjs X.Y.Z && git push --follow-tags

scripts/bump.mjs sets X.Y.Z as the version in package.json, .claude-plugin/plugin.json and the pinned pi-mcp.json arg, refreshes package-lock.json, rebuilds dist/, commits Release vX.Y.Z and creates the annotated tag vX.Y.Z. It refuses to run against a dirty working tree or a malformed version, and it never pushes — git push --follow-tags is a separate, explicit step.

Pushing the tag triggers .github/workflows/release.yml, which rebuilds vmctl, checks the working build against a fresh one, and publishes to npm using trusted publishing (OIDC — no NPM_TOKEN secret involved), then creates the matching GitHub release.

Trusted publishing has to be configured on npmjs.com, and npm requires the package to already exist before you can do that. The first publish must therefore be done by hand (npm publish --access public from a maintainer's machine) before its trusted publisher can be configured for this workflow.

License

MIT — see LICENSE.

Available Tools

19 tools
relay_acquireAcquire a VMA

Acquire one fresh VM for this task and start its ownership heartbeat. Requires task (a short slug), image (a key from relay_probe or relay_search) and extractions: every output you intend to bring home, declared before any work ([] is allowed; nothing is extracted automatically). Optional ttlHours (default 4), env (a credential pack, never baked into images), fullWorkspace (opt in to deliver the whole workspace) and vnc (default false; only prepares console-sharing capacity, never opens a viewer). One task per enclosure: call this once per task, not again for another task. Continue with relay_stage, relay_run (or a command tool), relay_image/relay_extract, then relay_finish or relay_release explicitly before you return; a failed step keeps the VM for repair rather than replaying uncertain input.

ParametersJSON Schema
NameRequiredDescriptionDefault
envNoCredential pack, default none; never baked into images.
vncNoPrepare guest console sharing, default false. Never opens a viewer automatically.
taskYes
imageYesImage key from relay action=probe, e.g. ubuntu2404 or macos26.
ttlHoursNo
extractionsYes
fullWorkspaceNoExplicit opt-in to deliver the entire workspace.

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With all annotations false, the description carries the full behavioral burden and does so richly: it discloses the ownership heartbeat, that nothing is extracted automatically, that env is never baked into images, that vnc only prepares console-sharing and never opens a viewer, and that a failed step keeps the VM for repair rather than replaying uncertain input.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but mostly front-loaded, leading with the core action and required parameters. The final lifecycle sentence packs many directives and there is slight redundancy between 'One task per enclosure' and 'call this once per task,' so it is not maximally tight.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 7-parameter lifecycle tool with no output schema and unhelpful annotations, the description gives the agent everything needed to call correctly: prerequisites, optional-parameter semantics, a one-call-per-task constraint, next steps, and failure behavior. The lack of an explicit return-value description is mitigated by the stated handoff to relay_stage/relay_run.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 57%, but the description compensates by explaining task as a short slug, image as a key from relay_probe/relay_search, extractions as pre-declared outputs ([] allowed), ttlHours default 4, env as a credential pack, fullWorkspace as opt-in, and vnc's exact console-sharing semantics.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The opening sentence names a specific action and resource: 'Acquire one fresh VM for this task and start its ownership heartbeat.' This clearly positions it as the acquisition step and separates it from image-lookup tools like relay_probe/relay_search and later stage/run tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives explicit preconditions ('Requires task, image, extractions') and lifecycle direction ('Continue with relay_stage, relay_run ... relay_finish or relay_release'), plus an exclusion: 'call this once per task, not again for another task.' It does not name an alternative tool for the same purpose, but the pipeline context makes when-to-use clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

relay_acquisition_capabilitiesRead acquisition capabilitiesA
Read-onlyIdempotent

Read the backend's versioned VNC console options for the images it serves. Read-only: it neither acquires a VM nor recovers ownership of one.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false. The description adds valuable context by explicitly stating it 'neither acquires a VM nor recovers ownership of one,' clarifying behaviors beyond the generic annotation hints. This is useful behavioral disclosure about side-effect boundaries without contradicting the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two concise sentences with no filler. The core purpose is front-loaded, and the clarifying non-behavior sentence adds meaningful distinction without redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a parameterless, read-only capability inspection tool with strong annotations, the description covers purpose and side-effect boundaries sufficiently. It does not describe the return format, but this is a minor gap given the simple nature of the tool and the absence of an output schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters and schema description coverage is 100%, so the schema fully documents that no inputs are needed. The description correctly focuses on the tool's purpose rather than parameter details. The baseline of 4 for zero-parameter tools applies here.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'Read the backend's versioned VNC console options for the images it serves.' It also explicitly differentiates the tool from acquisition and ownership-recovery actions by saying it 'neither acquires a VM nor recovers ownership of one,' which distinguishes it from sibling tools like relay_acquire and relay_console_resolve.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies this tool is for inspecting capabilities without side effects, and the read-only clarification helps an agent understand not to expect acquisition. However, it does not explicitly state when to prefer this tool over alternatives such as relay_acquire or relay_console_resolve, nor when not to use it. Usage context is present but indirect.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

relay_codeRun guest codeB

Run inline guest code as this call's single recorded operation. Evidence is automatic, as for relay_exec, whose optional reason, step, snapshots and timeoutMs it shares. Requires code (up to 1 MiB) and language (javascript, typescript or python).

ParametersJSON Schema
NameRequiredDescriptionDefault
codeYes
stepNoOptional step record overrides; any field left out is derived from the call.
reasonNoOptional intent, retained as evidence. Default: derived from the call.
languageYes
snapshotsNoOptional. Consecutive keystrokes may form an explicit text group (first/member/last); only first gets a before-snapshot and only last an after-snapshot.
timeoutMsNoExecution timeout in milliseconds, default 120000, maximum 3600000. Independent of snapshot delay and lease TTL.

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already signal readOnlyHint=false and openWorldHint=true; the description adds that evidence is automatic and that optional params are shared with relay_exec. However, for an arbitrary code runner it does not disclose side effects, result/output behavior, failure modes, or why openWorld semantics apply, leaving significant behavior implicit.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences front-load the core purpose and add the only required-input details; no filler. The relay_exec reference is compact but slightly dense, so it is not quite a perfect example of concision.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For an execution tool with no output schema, high complexity, and six parameters, the description omits what the tool returns or surfaces on success/failure, and offers no routing guidance against similar siblings. It is enough to attempt a call, but not enough to invoke correctly and interpret the result confidently.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 67%, so the schema already documents reason, step, snapshots, and timeoutMs. The description adds that code is required up to 1 MiB and enumerates languages, but those values are largely mirrored by schema constraints (maxLength, enum); it does not explain the semantics of code beyond the name and title.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The first sentence states a concrete action and object ('Run inline guest code') and frames it as 'this call's single recorded operation', which is distinctive. It references relay_exec but does not explicitly differentiate from siblings like relay_script or relay_run, so full sibling separation is not achieved.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies a use case—inline guest code in a single recorded operation—and points to relay_exec as a behavioral reference. It never states when to prefer relay_code over relay_exec, relay_script, or relay_run, and gives no exclusions or prerequisites beyond the required parameters.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

relay_console_cancelCancel console viewingA

Cancel an open or uncertain console-viewing attempt by console_id and attempt_id. Closes managed viewing resources only; it never releases or destroys the VM.

ParametersJSON Schema
NameRequiredDescriptionDefault
attempt_idYes
console_idYes

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate this is a non-read-only, non-idempotent, non-destructive operation. The description adds that it only closes managed viewing resources and never releases or destroys the VM, providing a clear boundary for side effects. It doesn't discuss failure states or prerequisites, but the added scoping is valuable.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is exactly two sentences, front-loads the core action, and every sentence adds value—first the operation, then the boundary. No wasted words or redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple two-parameter cancel operation, the description covers the action's effect and its non-destructive boundary. It does not explain return values or error conditions, but given this is a small tool with no output schema, that omission is minor and unlikely to prevent correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It mentions both `console_id` and `attempt_id` as the identifiers of the attempt, but provides no additional detail about their relationship, format, or roles beyond the parameter names and regex patterns in the schema. This is minimal but adequate for a simple cancellation action.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific action ('Cancel an open or uncertain console-viewing attempt') and clearly identifies the resource and scope ('Closes managed viewing resources only; it never releases or destroys the VM'). It distinguishes itself from related siblings like relay_console_open and relay_release by framing what it does and does not affect.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives clear context for when to use the tool (canceling an open or uncertain console-viewing attempt) and an explicit exclusion ('never releases or destroys the VM'), which helps an agent decide when not to use it. However, it does not explicitly name alternative sibling tools that may be considered, leaving some inference to the agent.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

relay_console_openOpen console viewingA

Open console viewing for the owned lease, only following an explicit user request to watch. Requires console_id and attempt_id from relay_console_resolve, userRequested: true, reason (intent) and expected (what the human should expect to see); it opens on the declared service host, never a remote client, and is never triggered automatically by relay_acquire's vnc option. See relay_console_resolve for what a ready status does and does not prove. A failed or uncertain open keeps the attempt: resolve or cancel it rather than opening again with a new attempt.

ParametersJSON Schema
NameRequiredDescriptionDefault
reasonYesIntent of this execution. Retained as evidence, never identity, authorization or a retry key.
expectedYes
attempt_idYes
console_idYes
userRequestedYes

TDQS

A5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only indicate it is not read-only, not idempotent, not destructive, and not open-world. The description goes far beyond this by disclosing that it opens on the declared service host, never a remote client, and is never triggered automatically. It also reveals that a failed or uncertain open keeps the attempt, which is critical behavioral context for the agent to handle retries correctly.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single dense paragraph that front-loads the core purpose and then adds necessary constraints and cross-references. Every sentence carries essential information—no filler. It is structured logically: purpose, prerequisites, behavioral constraints, and failure handling.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with 5 required parameters, no output schema, and only false annotations, the description covers all essential aspects: what it does, when to use it, where it operates, what it requires, and how to handle failures. It appropriately defers to relay_console_resolve for deeper status semantics, and nothing critical for correct invocation is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 20% (only 'reason' has a description). The description compensates fully by explaining that console_id and attempt_id come from relay_console_resolve, that userRequested must be true (which it enforces as const), that reason captures intent, and that expected is what the human should expect to see. This gives the agent the semantic meaning needed to construct valid calls.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Open console viewing for the owned lease', and immediately constrains it to explicit user requests. It distinguishes itself from siblings by referencing relay_console_resolve and relay_acquire, clarifying it is never automatically triggered by relay_acquire's vnc option. The purpose is unambiguous and clearly differentiated.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides explicit when-to-use guidance: only after an explicit user request to watch. It also gives what-not-to-do: never triggered automatically by relay_acquire's vnc option. It instructs that a failed or uncertain open should be resolved or canceled rather than re-opened, effectively routing the agent to the correct alternative actions (relay_console_resolve/relay_console_cancel).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

relay_console_resolveResolve console statusA
Read-onlyIdempotent

Refresh this session's non-secret console-viewing status for the owned lease. Read-only, and it never recovers ownership. A ready status means guest preflight only, not that a viewer launched: console operations carry no screenshots and never establish authentication, displayed pixels or a human's confirmation. On macOS, viewing requires the human to choose Standard sharing of the existing console, not a new Log In session or a High Performance display, and the relay never auto-confirms that choice; server_enforced_view_only: false and an unverified viewer selection remain limitations even once transport connects. Never replay an uncertain attempt: resolve it here, or close it with relay_console_cancel.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well beyond the readOnlyHint, idempotentHint, and destructiveHint annotations. It discloses that the tool never recovers ownership, never establishes authentication or human confirmation, never auto-confirms the macOS sharing choice, and that view-only limitations persist even after transport connects. No contradiction with the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The core action is front-loaded in the first sentence, and every subsequent sentence adds meaningful caveats or routing guidance. The description is dense and somewhat long for a zero-parameter refresh action, but the platform-specific warnings justify most of the length.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description explains the meaning of a 'ready' result, the operation's limitations, the macOS prerequisites, and the cancellation alternative, which is strong context given zero parameters and rich annotations. It does not spell out the exact response shape or possible status values beyond 'ready,' but that is a minor gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, so the schema is trivially complete and there are no parameter descriptions needed. The description adds useful contextual semantics by referring to 'this session' and 'the owned lease,' which frames the scope of the operation. This is the appropriate baseline for a no-parameter tool.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States the specific action: 'Refresh this session's non-secret console-viewing status for the owned lease.' It clarifies this is read-only, never recovers ownership, and contrasts with viewer-launching behavior by noting it carries no screenshots and never establishes authentication or displayed pixels. The 'ready' status caveat further pins down exactly what the tool does and does not confirm.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides explicit when-to-use guidance: 'Never replay an uncertain attempt: resolve it here, or close it with relay_console_cancel.' It also names the alternative tool and gives platform-specific prerequisites for macOS (Standard sharing of the existing console, not a new Log In session or High Performance display), so an agent knows how to interpret a successful resolution.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

relay_execRun a guest commandA

Run one guest command (argv) as one recorded operation; never hide several interactions in one call. Evidence is automatic: the relay snapshots the display before and after and records a step derived from the command. Optional reason (intent, never authorization), step (id, title, expected, inputMode) and snapshots.afterIntervalMs (default 500 ms) enrich or tune the record; consecutive keystrokes may form an explicit text group with snapshots.group (first/member/last). timeoutMs bounds execution (default 120000, up to 3600000). The result carries the outcome, bounded stdout/stderr and the after-snapshot inline. A refused, uncertain or nonzero result keeps the VM; never replay input whose effect is uncertain. diagnostic: true records a diagnosis or repair without snapshots, even before relay_stage; it is never visual verification, and its result has one output field with stdout and stderr combined.

ParametersJSON Schema
NameRequiredDescriptionDefault
argvYes
stepNoOptional step record overrides; any field left out is derived from the call.
reasonNoOptional intent, retained as evidence. Default: derived from the call.
snapshotsNoOptional. Consecutive keystrokes may form an explicit text group (first/member/last); only first gets a before-snapshot and only last an after-snapshot.
timeoutMsNoExecution timeout in milliseconds, default 120000, maximum 3600000. Independent of snapshot delay and lease TTL.
diagnosticNoExplicit command diagnosis or repair without screenshot evidence, including before staging. Commands and outcomes remain recorded. Never claim visual verification.

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description discloses many non-obvious behaviors beyond annotations: automatic before/after display snapshots, derived step recording, bounded stdout/stderr in results, VM retention on refused/uncertain/nonzero results, and the diagnostic mode's different snapshot/output behavior. This is exactly the kind of context annotations cannot convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but every sentence earns its place. It front-loads the core operation, then clusters optional parameters and safety/behavioral notes without fluff. Long paragraphs are justified by the tool's complexity and the high-value details they convey.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complex nested schema and no output schema, the description compensates fully: it explains result contents, snapshot grouping behavior, timeout bounds, diagnostic mode semantics, and safety guarantees. An agent has enough to invoke the tool correctly and interpret its return value.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Even with 83% schema coverage, the description adds meaning beyond the schema: `reason` is intent and 'never authorization', snapshots.group is for consecutive keystrokes with explicit first/member/last phases, afterIntervalMs defaults to 500 ms, and diagnostic output is a single combined `output` field. These clarifications help an agent tune calls correctly.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Run one guest command (`argv`) as one recorded operation'. It also adds a crucial scoping rule, 'never hide several interactions in one call', which distinguishes this executor from bulk or multi-step siblings. The title reinforces the same message, so an agent can select it confidently.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear usage context: it is for single guest commands, evidence is recorded automatically, and `diagnostic: true` is for diagnosis/repair even before relay_stage. It also warns not to replay input whose effect is uncertain. It does not explicitly name sibling tools for comparison, but the guidance is sufficient for correct invocation.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

relay_extractExtract declared outputsA

Pull one or more already-declared extraction outputs home early, by names. Only outputs declared on relay_acquire may be pulled; each pull verifies source and host hashes and rejects path traversal, symlinks, or a source that changed underneath it.

ParametersJSON Schema
NameRequiredDescriptionDefault
namesYes

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only indicate readOnlyHint=false and destructiveHint=false, leaving safety to the description. The description adds valuable behavioral context: verification of source and host hashes, rejection of path traversal, symlinks, and changed sources. It does not contradict annotations; the pull operation is consistent with readOnlyHint=false.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences, front-loaded with the core action and then adding essential constraints. Every word contributes; there is no filler or redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with one parameter and no output schema, the description covers what it does, the precondition (declared on relay_acquire), and the safety checks. It does not specify return values or post-pull effects, but those are likely implied by the operation name and context. Minor gap, not critical.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% (no parameter descriptions). The description mentions 'by `names`', indicating the parameter is a list of output names, and ties it to 'already-declared extraction outputs'. This adds some meaning beyond the bare array-of-strings schema, but it does not explain format, constraints, or examples, so it only partially compensates for the schema gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('pull'), resource ('already-declared extraction outputs'), and mechanism ('by names'). Differentiates from siblings by noting the constraint 'Only outputs declared on relay_acquire may be pulled', which clearly separates it from other relay tools like relay_acquire or relay_stage.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explains the intended use: pulling already-declared outputs early. It provides a clear context (outputs must be declared on relay_acquire) but does not explicitly name alternatives or give when-not-to-use guidance. Still, the context is specific enough for an agent to infer when this tool is appropriate.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

relay_finishFinish and deliver evidenceA
Destructive

Complete the task: extract every declared output, deliver and verify a portable evidence package, then destroy the VM and unregister it. The result reports delivery, snapshot completeness, execution outcome and human review as separate facts; a verified package is not a passing test, and a delivered package is not itself human approval. Call this, or relay_release, explicitly before you return, since ending the session does not do it for you. A failed delivery keeps the VM for a corrected attempt.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations say destructiveHint=true, and the description goes further by disclosing VM destruction, unregistration, and the delivery/verification workflow. It also clarifies important result semantics: a verified package is not a passing test, and a delivered package is not human approval, which is valuable beyond the structured annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and front-loaded with the main workflow, then covers result semantics, call timing, and failure behavior. Every sentence adds distinct operational value with no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter finalization tool with no output schema, the description covers what the tool does, what the result reports, when to call it, and what happens on failure. Nothing critical is missing for an agent to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There are zero parameters and the schema is fully empty, so the baseline is 4. The description adds no parameter-level ambiguity and needs no parameter documentation; its mention of 'every declared output' is behavioral context rather than a parameter definition.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb-resource pair: 'Complete the task' by extracting declared outputs, delivering/verifying evidence, and destroying/unregistering the VM. This clearly separates it from most siblings, but it does not distinguish relay_finish from the closely related relay_release beyond saying 'Call this, or relay_release'.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives explicit timing guidance: call before returning because ending the session does not finalize. It also names relay_release as an alternative and describes failure behavior (failed delivery keeps the VM), but it does not specify when to prefer this tool over relay_release.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

relay_imageRetrieve a saved imageA
Read-onlyIdempotent

Retrieve one already-saved image; never a new capture, input or directory export, and it never acquires a VM. target selects a display phase (sessionId/executionId/phase: before/after), a declared application file (name, plus a relative path for a directory declaration), or an immutable reference (imageId). PNG, JPEG and WebP only; originals up to 64 MiB and 40,000,000 decoded pixels; the delivered preview is at most 2000x2000 px and 4 MiB of base64 (PNG originals are resampled in-process; JPEG/WebP pass through only within bounds). Each delivery has a 90-second deadline and up to three transfer attempts. A relay_run tool image is the application file name: "relay-run" at the path its result gives. If a run's inline image did not arrive, recover it here with the same imageId or display selector, never by repeating the input or capturing again, and stop after at most two such recovery calls if it still cannot be inspected. A closed enclosure returns stale-reference for a display or reference target; its delivered originals stay readable with host file tools. An attached image proves only that the block was included, not that anyone inspected or reviewed it.

ParametersJSON Schema
NameRequiredDescriptionDefault
targetYes

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and idempotentHint=true, but the description adds extensive behavioral context: it confirms the tool never acquires a VM, specifies format and size limits (PNG/JPEG/WebP, up to 64 MiB, 40M pixels, preview at most 2000x2000 px and 4 MiB base64), describes delivery behavior (90-second deadline, up to three attempts), and clarifies the meaning of 'attached image proves only inclusion, not inspection'. No contradictions with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but well-structured, with the core purpose front-loaded and technical details organized logically. While it is longer than typical tool descriptions, the complexity of the tool justifies the length. It could be slightly more concise by trimming redundant explanations, but every sentence carries value.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (nested anyOf target variants, no output schema, no explicit parameter documentation), the description covers all necessary aspects: usage scenarios, constraints (size, format, deadlines), error handling (`stale-reference`), and recovery workflow. An agent would have sufficient information to call the tool correctly and interpret results.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 0% description coverage, but the description compensates fully by explaining the semantics of the `target` parameter: it distinguishes the three source types (display, application, reference), defines the fields for each (e.g., `sessionId`/`executionId`/`phase` for display; `name` and `path` for application; `imageId` for reference), and clarifies the purpose of `relay-run`. This adds meaning far beyond the raw schema structure.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description explicitly states 'Retrieve one already-saved image' and enumerates what it is not ('never a new capture, input or directory export, and it never acquires a VM'), clearly distinguishing it from siblings. The three target forms (display, application, reference) are defined with precise semantics, leaving no ambiguity about the tool's function.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It provides explicit when-to-use instructions, including recovery of relay_run images ('If a run's inline image did not arrive, recover it here...'), and when-not-to-use guidance ('never by repeating the input or capturing again'), with a stop condition ('stop after at most two such recovery calls'). It also mentions alternative tools for closed enclosures (host file tools). This is exemplary usage guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

relay_probeProbe host and guest readinessA
Read-onlyIdempotent

Read-only host and service facts, used to judge whether a task is interruptive and to pick an image; it reports current owned state on its own, so call it again after context compaction rather than trusting earlier messages. Default scope: "host" reports host permissions and activity, VM service availability, and this session's owned lifecycle state, not guest readiness; unknown is not idle. scope: "guest" (on an already-owned VM) reports whether Node, CuaDriver and each relay_run target (cua, playwright, chrome-devtools) are available, without installing or starting anything, and never claims capture readiness. Relay only interruptive work judged from these facts; non-disruptive or headless work stays with local tools. Continue in order: relay_acquire, relay_stage, relay_run (or a command tool), relay_image/relay_extract, then relay_finish or relay_release, called explicitly before you return.

ParametersJSON Schema
NameRequiredDescriptionDefault
scopeNo

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already mark read-only, idempotent, and non-destructive; the description adds that it reports current owned state and must be re-called after compaction, that guest scope installs or starts nothing, and that it never claims capture readiness. It also clarifies 'unknown is not idle.' These are meaningful behavioral details beyond the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Every sentence carries distinct information: purpose, state-freshness caveat, host scope semantics, guest scope semantics, usage boundary, and workflow order. It is dense but not wasteful, and it front-loads the read-only purpose and the key re-call behavior.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the absence of an output schema, the description summarizes what each scope reports and warns about known pitfalls (unknown vs idle, capture readiness). It also embeds the tool in the relay sequence so an agent knows when and in what order to invoke it. Some exact response structure is not specified, but the essential context is present.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% and the schema only lists an enum, but the description thoroughly explains both `scope: "host"` and `scope: "guest"` including what each reports and their caveats. This gives an agent the full semantic meaning of the only parameter, even without additional schema descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: it probes read-only host/service facts to judge interruptiveness and pick an image. It explicitly distinguishes the two scopes and notes that `host` is not guest readiness, which separates it from a generic probe. The tool's role in the relay workflow is immediately clear.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly says when to use it: before interruptive relay work, and after context compaction rather than trusting earlier messages. It names the alternative (local tools for non-disruptive/headless work) and provides the full ordered relay lifecycle. No ambiguity about where this tool fits.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

relay_releaseRelease the VMA
Destructive

Abandon the task: destroy the owned VM and unregister it, retaining whatever evidence already exists, without claiming the task succeeded. Use this instead of relay_finish when the task is not being completed. Safe to retry if a previous release attempt failed; a failed release keeps ownership until destruction is verified.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already mark destructiveHint=true, but the description adds valuable context about evidence retention, success status, and retry behavior. It goes beyond the structured metadata to describe side effects and failure semantics.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three concise sentences, front-loaded with the primary action, then the usage distinction, then retry safety. Every sentence adds necessary information with no redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter destructive action with no output schema, the description covers purpose, usage, side effects, and retry behavior. Nothing an agent needs to decide to call this tool is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, so there is nothing to document. The description correctly omits parameter details, and the empty schema fully covers the interface. Baseline of 4 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description explicitly states the tool destroys and unregisters the owned VM, retains evidence, and does not claim success. It clearly distinguishes itself from relay_finish by naming the alternative and the condition for use.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It directly instructs to use this instead of relay_finish when the task is not being completed, and clarifies retry semantics (safe to retry, failed release keeps ownership). This gives explicit when-to-use and when-not-to-use guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

relay_runRun an MCP tool call in the VMA

Send one tool call to an MCP server inside the VM. target is cua (cua-driver), playwright (Playwright MCP) or chrome-devtools (Chrome DevTools MCP); tool and args are that server's own tool name and arguments, forwarded unchanged. Use relay_tools to see a target's exact tools. Evidence is automatic: snapshots before and after, and a step record; reason, expected and afterIntervalMs are optional overrides. Never replay an uncertain call.

ParametersJSON Schema
NameRequiredDescriptionDefault
argsNoThe tool's own arguments, forwarded unchanged.
toolYesThe target server's own tool name.
reasonNoOptional intent, retained as evidence. Default: derived from the call.
targetYescua (cua-driver), playwright (Playwright MCP) or chrome-devtools (Chrome DevTools MCP).
expectedNoOptional expected result for the step record.
timeoutMsNoExecution timeout in milliseconds, default 120000, maximum 3600000. Independent of snapshot delay and lease TTL.
afterIntervalMsNoWait in milliseconds from the end of the call to the after-snapshot. Optional; the relay has a default.

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the annotations (readOnlyHint=false, openWorldHint=true, idempotentHint=false), the description discloses a concrete behavioral trait: 'Evidence is automatic: snapshots before and after, and a step record,' and that reason/expected/afterIntervalMs only override that evidence. The caution 'Never replay an uncertain call' aligns with the non-idempotent, open-world hints. It stops short of a 5 because it does not address failure/error propagation or side-effect expectations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four sentences, roughly 70 words, each earning its place: purpose, target enumeration, discovery routing, evidence model, and a safety caution. The core verb-resource pair is front-loaded in the first sentence, and there is no fluff or redundant restatement of schema descriptions.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex 7-parameter open-world tool with no output schema, the description covers the essential operational knowledge: target choices, forwarding semantics, the discovery prerequisite, automatic evidence, and optional overrides. Minor gaps remain — the exact return/step-record shape and the lease prerequisite implied by siblings relay_acquire/relay_release — but these do not prevent a correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description adds meaning by grouping parameters into two semantic classes — tool/args are 'forwarded unchanged' (the actual remote call) versus reason/expected/afterIntervalMs as 'optional overrides' of the evidence record — a relationship the per-property schema entries only imply. This framing tells the agent which parameters affect the remote server and which only affect the step record.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb-resource pair, 'Send one tool call to an MCP server inside the VM,' and enumerates the three valid targets with parenthetical mappings (cua, playwright, chrome-devtools). It also clarifies the forwarding semantics for tool and args, and the relay_tools reference helps distinguish this tool from its discovery sibling without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly routes the discovery step to the relevant sibling — 'Use relay_tools to see a target's exact tools' — which is the natural confusion point. It also states a when-not condition with 'Never replay an uncertain call' and clarifies which parameters are optional evidence overrides, so an agent knows what to supply and what to leave out.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

relay_scriptRun a guest scriptA

Run one guest script file as this call's single recorded operation. Evidence is automatic, as for relay_exec, whose optional reason, step, snapshots and timeoutMs it shares. Requires localPath (a host file path) and language (javascript, typescript or python); it runs under the staged workspace.

ParametersJSON Schema
NameRequiredDescriptionDefault
stepNoOptional step record overrides; any field left out is derived from the call.
reasonNoOptional intent, retained as evidence. Default: derived from the call.
languageYes
localPathYes
snapshotsNoOptional. Consecutive keystrokes may form an explicit text group (first/member/last); only first gets a before-snapshot and only last an after-snapshot.
timeoutMsNoExecution timeout in milliseconds, default 120000, maximum 3600000. Independent of snapshot delay and lease TTL.

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description adds meaningful context beyond the annotations: evidence is automatic, this is a single recorded operation, and execution happens under the staged workspace. It does not contradict the annotations. However, for a tool that executes arbitrary guest scripts, it does not disclose potential side effects, sandboxing limits, or what happens on failure, leaving the agent to infer behavioral risk beyond the basic openWorldHint and readOnlyHint flags.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences with no filler. It front-loads the core purpose ('Run one guest script file'), then packs the essential constraints and cross-reference to relay_exec into the second sentence. Every clause contributes useful information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with no output schema, the description is missing a critical piece: what the caller receives after execution (stdout, exit code, recorded evidence reference, error behavior). It covers invocation inputs, shared optional parameters, and workspace context, but an agent cannot fully predict how to consume the result. The nested snapshots and step objects are explained by the schema, so the biggest gap is the unspecified return/result semantics.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The description adds meaning for the two required parameters that the schema leaves undocumented: localPath is 'a host file path' and language is restricted to 'javascript', 'typescript' or 'python'. Since schema coverage is 67%, the description compensates for the biggest gap. It does not need to re-explain step, reason, snapshots, and timeoutMs because those already have descriptions in the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'Run one guest script file as this call's single recorded operation.' It adds useful scope by calling out that this is a single recorded operation, which separates it from broader multi-step tools. However, it does not explicitly contrast it with sibling tools like relay_exec, relay_code, or relay_run, instead only referencing relay_exec for shared optional parameter semantics.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies the primary use case: run a guest script file under the staged workspace. It also states required inputs (localPath and language), which gives the agent a baseline for when the tool is applicable. But it never explicitly says when to prefer this tool over relay_exec, relay_code, or relay_run, nor does it provide exclusions or alternative routing guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

relay_stageStage the guest runtimeA

Push and hash-check the guest runtime, plus an optional workspace and support files (files, landing under support/), onto an acquired VM. The guest Node and CuaDriver executables must already exist (nodePath, cuaDriver); Linux guests need native X11 and cua-driver serve --no-overlay. browserExecutable names the guest browser the playwright and chrome-devtools targets launch (default: installed Google Chrome); their pinned servers are staged on first use. A failed stage can be retried; once staged, only corrected executable paths may be resubmitted, and staging does not prove capture readiness. resetRecording: true archives the current recording's evidence and starts a fresh recording on the same VM; it refuses while a receiver lock is held.

ParametersJSON Schema
NameRequiredDescriptionDefault
filesNo
nodePathNoGuest Node executable, default node; use image nvm path if needed.
cuaDriverNoGuest CUA executable; OS default when omitted.
workspaceNoHost directory copied to the guest workspace, opt-in.
resetRecordingNoExplicitly archive the current recording and start a new session on the same VM. Retains prior evidence; refuses while the receiver lock exists. Use after diagnosing recording damage.
browserExecutableNoGuest browser executable for the playwright and chrome-devtools targets; default is installed Google Chrome.

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With all annotation hints false or unhelpful, the description carries the behavioral burden and does so richly. It discloses hash-checking, staging of support files, first-use server staging, retry semantics, the limitation that 'staging does not prove capture readiness,' and reset/resubmission behavior including the receiver-lock refusal.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single dense paragraph with no filler. Every clause earns its place: scope, prerequisites, Linux-specific requirements, retry rules, reset behavior, and the readiness caveat are all packed in without redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given six parameters, no output schema, and no meaningful annotations, the description is unusually complete. It covers prerequisites, retry and resubmission constraints, browser default and target behavior, reset semantics, and the warning that staging does not prove capture readiness—enough for an agent to invoke it correctly after acquisition.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is high (83%), so the baseline is 3, but the description adds meaningful context: `files` lands under `support/`, `nodePath`/`cuaDriver` must already exist, `browserExecutable` selects the launched guest browser for specific targets, and `resetRecording` archives evidence and starts fresh. It does not add meaning to `files[].local`, whose schema description is also missing.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific action: 'Push and hash-check the guest runtime, plus an optional workspace and support files... onto an acquired VM.' This clearly identifies the resource and distinguishes staging from siblings like relay_acquire or relay_exec/run.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Prerequisites are explicit: guest Node and CuaDriver executables must already exist, and Linux guests need native X11 plus `cua-driver serve --no-overlay`. It also defines retry/resubmission constraints and when `resetRecording` is appropriate, though it does not explicitly name alternatives or say 'use relay_stage instead of X.'

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

relay_statusShow relay statusA
Read-onlyIdempotent

Show this session's owned VM lease (backend binding, guest state, renewal, console observation, last error), staging state and evidence path, plus the project directory, the VM service origin and the selected environment in use; active:false when nothing is owned. Read-only; it does not touch the VM.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, destructiveHint=false, and idempotentHint=true, so the safety profile is covered. The description adds the active:false behavior when nothing is owned and repeats 'Read-only; it does not touch the VM,' which is redundant with annotations. The added context about the active flag is useful, but overall the description provides only marginal behavioral detail beyond the annotations, earning a 3.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single long sentence but front-loaded with the primary output (owned VM lease) and then lists the rest. It is efficient, avoids fluff, and every phrase adds information. It could be slightly more structured, but it is appropriately sized and informative.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has no parameters, no output schema, and annotations covering the safety profile, the description is quite complete. It enumerates all the fields the agent can expect (lease details, staging state, evidence path, project directory, service origin, environment, and active flag). The only missing element is the exact return format (e.g., JSON structure), but that is not critical for a status tool and the description suffices for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, and the schema coverage is 100% (empty schema). The description does not need to explain parameters, and the baseline for zero-parameter tools is 4. No additional parameter information is required, so this dimension is satisfied.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool shows the session's owned VM lease, staging state, evidence path, project directory, VM service origin, and selected environment, plus an active:false indicator. It uses a specific verb ('show') and identifies a distinct resource (relay status) with a detailed list of contents, making its purpose unambiguous. However, it does not explicitly differentiate from sibling tools like relay_probe or relay_console_resolve, though the read-only inspection nature is evident from context.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies the tool is for inspecting the current session state before or after operations, but it does not provide explicit guidance on when to use it versus alternatives. No exclusions or conditions are stated, leaving the agent to infer from the read-only nature and the listed contents. This is adequate but lacks the explicit 'when to use' framing that would merit a 4.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

relay_toolsList a target's toolsA
Idempotent

List a relay_run target's real tools from its server inside the VM: names, descriptions and input schemas; tool returns just one. Starts the server if needed, without sending it any tool call. Requires relay_stage.

ParametersJSON Schema
NameRequiredDescriptionDefault
toolNoOptional: return only this tool.
targetYescua (cua-driver), playwright (Playwright MCP) or chrome-devtools (Chrome DevTools MCP).

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare idempotentHint=true and destructiveHint=false; the description adds valuable context that it starts the server if needed without sending a tool call. This goes beyond what annotations provide, with no contradiction.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two concise sentences with the core purpose front-loaded and no filler. The second sentence adds side effects and prerequisites without redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, but the description covers what is returned (names, descriptions, input schemas) and the side effect of starting the server. It also notes the prerequisite. Missing error handling is not essential for a listing tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% with descriptive parameter text for both `tool` and `target`. The description adds only a minor clarification that `tool` returns just one, which is already implied by the schema. No significant extra meaning over the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific action ('List') and a clear resource ('a relay_run target's real tools'), and spells out what is returned (names, descriptions, input schemas) plus the `tool` option to return one. This clearly distinguishes it from siblings like relay_search or relay_probe.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides a prerequisite ('Requires relay_stage') but does not explicitly compare with sibling tools or state when to use this instead of alternatives. The context is clear enough to infer use, but no exclusions or alternatives are named.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

relay_trajectoryOpen the trajectory viewerA
Idempotent

Verify a delivered relay evidence package (every artifact, hash and reference) and open its trajectory viewer in the local browser. Human review remains pending.

ParametersJSON Schema
NameRequiredDescriptionDefault
directoryYesThe package directory, absolute or relative to the project.

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare idempotentHint=true and destructiveHint=false. The description adds that it opens a local browser (side effect), verifies every artifact/hash/reference, and leaves human review pending. This contextual behavior goes beyond the annotations without contradicting them.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with the core action and scope, followed by a state note. No wasted words; the description is efficient and well-structured.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with one parameter and no output schema, the description covers purpose, verification scope, side effect (browser open), and pending human review. It does not explain failure behavior or return value, but these are minor given the simplicity and annotations.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The only parameter 'directory' is fully described in the schema (100% coverage). The description does not add extra parameter details beyond referring to the package directory. Baseline 3 is appropriate since the schema handles parameter meaning.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('verify') and resource ('delivered relay evidence package') followed by an action ('open its trajectory viewer'). It clearly identifies what the tool does, and the trajectory viewer is distinct from sibling console tools, though it doesn't name an alternative explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage: when a relay evidence package has been delivered and needs verification and viewing. It gives clear context ('delivered relay evidence package') but does not explicitly state when not to use it or name alternatives like relay_console_open. This is adequate without exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 20 tool updatesv0.5.1
    • Removedrelay
    • Addedrelay_acquire
    • Addedrelay_acquisition_capabilities
    • Addedrelay_code
    • Addedrelay_console_cancel
    • Addedrelay_console_open
    • Addedrelay_console_resolve
    • Addedrelay_exec
    • Addedrelay_extract
    • Addedrelay_finish
    • Addedrelay_image
    • Addedrelay_probe
    • Addedrelay_release
    • Removedrelay_review
    • Addedrelay_run
    • Addedrelay_script
    • Addedrelay_search
    • Addedrelay_stage
    • Addedrelay_tools
    • Addedrelay_trajectory
  2. 3 tool updatesv0.3.1
    • First observedrelay
    • First observedrelay_review
    • First observedrelay_status

TDQS

A4/5.0

Scored across 19 tools

Disambiguation4/5

Most tools have distinct purposes: acquire, stage, execute, extract, finish, release are lifecycle steps; probe, status, search, and capabilities cover discovery/state; console operations are separate; image and trajectory are retrieval/verification. Some overlap exists between relay_exec, relay_script, and relay_code (all run guest code) but inputs differ (command vs file vs inline), and relay_probe vs relay_status could be confused but descriptions clarify host/guest facts vs lease state.

Naming Consistency3/5

All tools share the 'relay_' prefix, but the second part mixes verbs (exec, search, acquire, stage, extract, finish, release, resolve, open, cancel) with nouns (status, trajectory, script, code, run, tools, image) and even a noun phrase (acquisition_capabilities). While readable, the pattern is inconsistent—some names describe actions, others describe resources—making it less predictable than a uniform verb_noun convention.

Tool Count4/5

19 tools is on the higher end but appropriate for a VM relay server that manages a full lifecycle (acquire, stage, run variants, extract, finish/release) plus supporting operations (status, probe, search, console, image, trajectory). Each tool serves a distinct purpose in the workflow, though the count could be trimmed by merging some run variants, but the complexity justifies it.

Completeness5/5

The tool surface covers the entire VM workflow: discovery (search, probe, capabilities), acquisition (acquire), setup (stage), execution (exec, script, code, run), monitoring (status, console resolve/open/cancel), output retrieval (extract, image), and completion (finish, release). There are no obvious gaps—even edge cases like early extraction and evidence verification (trajectory) are handled. The lifecycle is closed with explicit finish/release steps.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • F
    license
    Not graded
    quality
    C
    maintenance
    Enables autonomous desktop automation by delegating tasks to vision-based agents operating within cloud-based virtual machine sandboxes. It allows users to manage VMs, execute complex computer tasks, and receive text-based screen summaries across Linux, Windows, and macOS environments.
    2
    -
  • A
    license
    Not graded
    quality
    B
    maintenance
    Provides AI agents with isolated virtual desktops containing a real Chromium browser, enabling them to see, click, type, and navigate like a human, with features like snapshots and remote control.
    23
    AGPL 3.0
  • A
    license
    Not graded
    quality
    C
    maintenance
    Enables a Copilot Studio agent to operate a dedicated, disposable Windows VM by running commands, PowerShell scripts, file operations, and background jobs over authenticated HTTPS.
    MIT