vm-relay
This server provides an MCP interface for interruptive computer-use and browser-use inside a dedicated VM, with recorded, snapshot-evidenced execution and explicit lifecycle management.
Search installed application images by name and OS (
relay_search).Probe host and guest readiness, including owned VM state and run-target availability (
relay_probe).Read VNC/console acquisition capabilities (
relay_acquisition_capabilities).Acquire a fresh VM with a task slug, image, declared outputs, TTL, env pack, optional VNC preparation, and full-workspace opt-in (
relay_acquire).Stage the guest runtime, support files, workspace, browser executable, and reset recording (
relay_stage).Run guest commands, scripts, or inline code with automatic before/after snapshots, optional step/reason/timeout, and diagnostic mode (
relay_exec,relay_script,relay_code).Forward MCP tool calls to in-guest cua-driver, Playwright MCP, or Chrome DevTools MCP servers (
relay_run).List a target server's real tools or inspect one tool's schema (
relay_tools).Retrieve saved display, application, or immutable-reference images (
relay_image).Extract declared outputs early with hash verification (
relay_extract).Finish a task: deliver verified evidence, destroy the VM, and unregister (
relay_finish).Abandon a task and destroy the VM while retaining existing evidence (
relay_release).Resolve, open, or cancel live console viewing, only on explicit user request (
relay_console_resolve,relay_console_open,relay_console_cancel).Show current session lease, staging state, evidence path, and environment (
relay_status).Verify a delivered evidence package and open its trajectory viewer in the local browser (
relay_trajectory).
Provides browser-use automation inside a dedicated VM, controlling a fresh persistent Google Chrome (Chromium) instance to navigate, click, type, press keys, read page content, and capture snapshots with settle-waiting.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@vm-relayUse a fresh VM to open our staging site and capture evidence of any interruptions during checkout."
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
mcp-vm-relay
The Model Context Protocol front end of the VM relay, packaged as a Claude Code
plugin and as a pi package. pi-vm-relay,
which gave pi a native relay tool directly, is retired; mcp-vm-relay replaces
it for both Claude Code and pi, the latter loaded through pi-mcp-adapter. It
is an MCP server: nineteen relay_* tools, each with its own schema, title
and annotations, and the server's instructions. The two user commands,
status and trajectory, ship as one small command file per host (see
"Commands"). There is no skill and no hook. Together they offer recorded,
snapshot-evidenced interruptive computer-use and browser-use in a fresh,
dedicated VM. The agent judges interruption; non-disruptive work stays with the
agent's local tools. Nothing here executes local computer-use, spawns
subagents, targets physical machines, records video, or attaches to existing
browsers.
How it is built
The relay implementation lives in this repository: the manager
(src/manager.ts), the vm-service client, the registry with its OS-level lock,
the guest transfer and transport, the evidence package and its state merge, the
selected-environment profile, the console and saved-image contracts, the strict
tool contract (src/schema.ts), the action dispatch and bounded result
rendering (src/surface.ts), and the two guest programs
(src/guest/receiver.ts, src/guest/mcp-host.ts). It was inherited from
pi-vm-relay at commit 8990123,
synchronized with its implementation through commit 0d69fc7 before that
project's retirement, and is maintained here as mcp-vm-relay's own
implementation from then on. src/server.ts binds the core to MCP over
standard input and output.
The committed dist/ holds server.mjs (the MCP server with the core bundled
in), the guest bundles the manager stages (receiver.mjs, mcp-host.mjs) and
the doctor. Compiled relay-driver
code, the recorded-execution and evidence substrate, is bundled with hashes and
provenance in dist/build-info.json. Consumers load dist/; no build, SDK
checkout or network fetch happens at install time.
The Claude Code plugin and pi both reach this same server as plain MCP: pi
through pi-mcp-adapter (see "Use with pi" below), Claude Code by starting it
directly. Every tool's identity, schema and contract are shared; the plugin
adds nothing Claude-Code-specific beyond the marketplace packaging and the
vm-relay-operator agent definition.
In pi-vm-relay (retired, native pi extension) | Here (MCP server) |
a registered | nineteen |
doctrine injected before each agent turn | the server's MCP |
|
|
prompt-composed enclosure | the |
typed image blocks in a run result, with pi's | MCP image content blocks in the tool result, with the MCP |
session shutdown pauses lease renewal | the same when the server's stdio closes or it is signalled: renewal pauses, the recording detaches, the VM is retained for an explicit |
the settled agent pauses renewal | none: MCP has no such hook, so the instructions and each run tool's description say "finish or release before you return" |
Related MCP server: screenbox
Install
Use with Claude Code
Try it from a checkout:
claude --plugin-dir /path/to/mcp-vm-relayOr add this repository as a marketplace and install the plugin from it:
/plugin marketplace add WeZZard/mcp-vm-relay
/plugin install mcp-vm-relay@mcp-vm-relayOnce loaded, the model sees the nineteen tools as
mcp__plugin_mcp-vm-relay_relay__relay_search,
mcp__plugin_mcp-vm-relay_relay__relay_probe, and so on through
mcp__plugin_mcp-vm-relay_relay__relay_trajectory (see "The tools" below for
the full table). The vm-relay-operator agent definition allows them by those
names. The server can also be configured directly in a project's .mcp.json
with node /path/to/mcp-vm-relay/dist/server.mjs, in which case the tools are
mcp__relay__relay_search and so on.
Use with pi
pi install npm:@wezzard/mcp-vm-relayThis needs pi-mcp-adapter installed. The tools appear as relay_search,
relay_probe, and so on through relay_trajectory, with no host-specific
prefix (pi-mcp.json sets toolPrefix: "none"). The package also
ships pi prompt templates for the two commands, /mcp-vm-relay-status and
/mcp-vm-relay-trajectory.
Use with npx
npx -y @wezzard/mcp-vm-relayThis runs the MCP server directly on stdio, for a client that speaks MCP without a Claude Code plugin or pi package wrapper.
Prerequisites
Node 22+, Python 3 with fcntl for the registry lock, and for VM execution an
Apple-silicon Mac with Tart, a running
vm-service and prepared guest images.
Check readiness with node dist/doctor.mjs. It does not acquire a VM or
establish guest UI readiness. Guest execution needs Node and CuaDriver with
capture/input permissions; Linux requires native X11 and serve --no-overlay.
The playwright and chrome-devtools targets of relay_run need a browser in
the guest (installed Google Chrome, or a path given as browserExecutable);
their pinned servers are downloaded once on the host and staged, so the guest
needs no network or npm. Console viewing needs the vm-service guest-sharing backend and its
guest preparation; see docs/console.md here.
The tools
Each tool's inputSchema is a plain object with no root anyOf/oneOf/allOf
and no action field, derived directly from the strict per-action contract in
src/schema.ts (nested unions inside a property, such as relay_image's
target, are unaffected). A call is mapped back to that contract's shape and
dispatched exactly as the retired single relay tool was, so behaviour,
results, image blocks and isError semantics are unchanged from before this
split. relay_run forwards the cua-driver, Playwright MCP and Chrome DevTools
MCP tool calls a model already knows; relay_exec, relay_script and
relay_code share the optional reason, step, snapshots and timeoutMs.
Evidence is automatic for all four: see "relay_run" below.
Tool | Title | Replaces ( | Purpose |
| Search installed applications |
| Find installed applications from image inventories by name and optional OS. |
| Probe host and guest readiness |
| Default |
| Read acquisition capabilities |
| Read versioned acquisition options (VNC backends) without allocation or ownership recovery. |
| Acquire a VM |
| Register the task, acquire a fresh VM and start its heartbeat; declare outputs before work. Optional |
| Stage the guest runtime |
| Push and hash-check the runtime, support files and an opt-in workspace; |
| Run a guest command |
| One recorded guest command. |
| Run a guest script |
| One recorded guest script file ( |
| Run guest code |
| One recorded inline guest code snippet ( |
| Run an MCP tool call in the VM |
| One recorded tool call ( |
| List a target's tools |
| The target's real |
| Retrieve a saved image |
| Retrieve one saved display image, declared application image or immutable image reference without input, capture, directory export or acquisition. |
| Extract declared outputs |
| Pull only declared files or directories with source and host checksum verification. |
| Finish and deliver evidence |
| Extract declared outputs, deliver and verify the snapshot package, destroy the VM and unregister. |
| Release the VM |
| Retain available evidence and abandon or destroy the VM. |
| Resolve console status |
| Resolve non-secret console status for the owned lease and environment. |
| Open console viewing |
| Open explicitly user-requested viewing on the service host, with |
| Cancel console viewing |
| Cancel the identified viewing attempt without releasing the VM. |
| Show relay status | (unchanged) | This session's owned lease: backend binding, guest state, renewal state, console observation, last error, staging state and evidence path, plus the project directory, the VM service origin and the selected environment; |
| Open the trajectory viewer | (unchanged) | Verify a delivered evidence package (every artifact, hash and reference) and open its trajectory viewer in the local human-facing browser; human review remains pending. |
readOnlyHint/idempotentHint are true for relay_search, relay_probe,
relay_acquisition_capabilities, relay_console_resolve, relay_image and
relay_status; destructiveHint is true only for relay_finish and
relay_release, which destroy the VM; openWorldHint is true only for the
four run tools, whose guest code may reach the network.
A run's result carries the execution identity and outcome, the derived step
record, and, when the snapshot plan captured the after phase and delivery
succeeded, the saved after-image as an MCP image content block. A command's
result also carries the guest's bounded standard output and error; a
relay_run result carries the target tool's own text and image blocks. The
full receipt stays in the evidence package.
Commands
Each command asks the assistant to call one relay tool and report the answer. It changes nothing itself.
Command in pi | Command in Claude Code | Arguments | Purpose |
|
| none | Calls |
|
| the package directory | Calls |
They are not MCP prompts. pi's MCP adapter can only name a prompt
/mcp__<package>__<server>__<prompt>, so each host gets its own command file
instead: pi prompt templates in pi-prompts/ (listed in package.json under
pi.prompts) and Claude Code plugin commands in commands/. Both are
generated by npm run build from scripts/host-commands.mjs; do not edit them
by hand. CI and the release workflow check the packed tarball with
scripts/verify-package.mjs: the generated files must match their source, pi's
glob must select exactly the two templates, and the packed server must start
without the prompts capability. The release publishes and attaches the same
tarball it checked.
Ownership, failure and recovery
One server session owns at most one VM. Another task needs a separate acquisition.
An operation failure retains the VM so the agent can inspect, repair and submit a new operation. Failed and uncertain operations are never automatically replayed.
Use
finishto deliver evidence and release, orreleaseto abandon explicitly. A failed delivery retains the VM; a release failure retains ownership until destruction is verified.When the session ends (the server's stdio closes or it is signalled), lease renewal pauses and the recording detaches; the VM is not destroyed. The backend TTL and grace period handle abandoned leases. Status reports the last confirmed expiration time.
A restarted server with the same
MCP_VM_RELAY_SESSIONreconciles its durable ownership and reattaches the recording session without replaying prior work. Context compaction does not reset VM state;probereports the owned state without relying on earlier messages.A run tool's
timeoutMsdefaults to 120,000 ms and accepts integers up to 3,600,000 ms. It bounds command execution, not the snapshot delay or the lease lifetime. A timeout reports confirmed termination or uncertainty and keeps the VM available.relay_execwithdiagnostic: truerecords command diagnosis or repair without screenshots, for explicitly requested diagnosis when capture is unavailable. Before staging it runs through vm-service in the guest's default directory; after staging in the recording workspace. It cannot join a snapshot group and is not visual verification.After diagnosing damaged recording state,
relay_stagewithresetRecording: truearchives the old recording and starts a new recording session in the same VM. An existing receiver lock refuses the reset until its operation is reconciled. Prior evidence paths remain in the owned status.
relay_run: MCP tool calls in the VM
relay_run takes the same tool calls a model already sends to cua-driver,
Playwright MCP or Chrome DevTools MCP and runs them against that server inside
the VM:
{"target":"playwright","tool":"browser_navigate","args":{"url":"https://example.com"}}
{"target":"cua","tool":"click","args":{"pid":812,"window_id":3,"x":40,"y":90}}
{"target":"chrome-devtools","tool":"take_snapshot","args":{}}relay_tools {"target":"playwright"} returns the server's real tools/list,
so the exact names and schemas are one call away. The arguments are checked
against the tool's input schema (a JSON Schema validator, Ajv) before anything
is sent; a mismatch is refused with the correct schema in the error.
Evidence is automatic. Every call is admitted by the guest receiver, journaled, and bracketed by a before and an after display snapshot. The step record is derived from the call: the title is
<target>.<tool>(for exampleplaywright.browser_click), the expected result is "returns without a tool error", and cua-driver's accessibility forms are labelledaccessibility.reason,expectedandafterIntervalMsare optional overrides. The same rule now applies torelay_exec,relay_scriptandrelay_code: theirreason,stepandsnapshotsare accepted but no longer required.Default waits before the after-snapshot: 300 ms for
playwrightandchrome-devtools(both servers already wait for the page to settle before they answer), 500 ms forcuaand for commands.Sessions persist. A small guest-resident MCP host keeps one MCP SDK client per target, starting each server on first use, so a browser page or a native session survives between calls.
relay_finishandrelay_releasestop it.Results pass through. The target's text blocks follow the relay's own result text, bounded like every result (50 KiB, 2000 lines); its images (up to four) are delivered inline under the usual image rules, before the relay's after-snapshot. Each call's full result, images and the servers' logs land in
workspace/relay-run, the extractionrelay-run, and come home withrelay_finish.Outcomes. A call the relay proves was never sent is
refused; a call that may have reached the server and has no answer (timeout, crash, a malformed answer) isuncertain, and the server is stopped so nothing it still holds can act later; the target's ownisError: trueiscompleted-with-tool-error(exit status 3). All three are error results and keep the VM. Nothing is ever replayed.
The launch commands, pinned versions and the full outcome mapping are in
docs/relay-run.md.
Saved images
A run tool returns its saved after-image as a typed image block whenever the
snapshot plan captures that phase: a standalone event or a text group's last
event. A first or intermediate group event and a diagnostic command do not
invent an image. Execution and image delivery are independent outcomes: a
completed command can have a failed delivery, and a delivered image does not
turn a failed command into a successful one. Either failure sets the MCP
isError flag while the content, image included, is kept. The delivery
identity (imageDelivery) leads the result text so it survives truncation.
If inline delivery fails, relay_image retrieves the same saved image through
one closed selector; it never repeats input, creates a capture, exports the
consumer directory or acquires a VM:
{"target":{"source":"display","sessionId":"<recording-session>","executionId":"<saved-execution>","phase":"after"}}
{"target":{"source":"application","name":"relay-run","path":"playwright/<call>/image-1.png"}}
{"target":{"source":"reference","imageId":"image-<64 lowercase hex digits>"}}Display selection needs
beforeorafter; a phase the plan did not request returnsnot-requested. Application selection needs an acquisition-time declaration: a directory declaration takes one relative file path, a file declaration omitspath. A reference resolves to the same original bytes or an explicit failure.Originals are limited to 64 MiB, 40,000,000 decoded pixels and 32,768 pixels per dimension; PNG, JPEG and WebP are accepted. The preview is at most 2,000 × 2,000 pixels and 4 MiB of base64. Presentation differs from pi here: pi resizes through its own image helper, while this core has no codec dependency. A PNG original is decoded and, when it exceeds the preview bounds, resampled in-process (area averaging,
node:zlib); a JPEG or WebP original has its container validated and passes through unchanged when within bounds, and is reportedpresentation-unavailableotherwise. No EXIF orientation is applied; dimensions are those stored. The presentation receipt records the policy, the original and preview dimensions, the MIME type, hash, size and whether bytes were transformed, beside the untouched original.Each delivery has a 90-second deadline and a shared budget of three byte-transfer attempts. Only transient transfer failures are retried; authorization, unsafe-path, capture, format and integrity failures are not. Recommend at most two explicit recovery calls for the same reference; if a required image still cannot be inspected, stop exploratory input and finish or release.
Immutable catalog entries bind each reference to the owner, enclosure, backend and original identity. Closed or sealed enclosures return
stale-reference; delivered originals remain readable with the host's file tools. Finalization materializes verified originals at canonical snapshot paths and preserves prior state when merging guest evidence.Attachment means the block was included in the result; it does not prove the model inspected it or that a human reviewed it.
Live console viewing
Acquisition never opens a viewer. relay_acquisition_capabilities reads the
backend's VNC options; relay_acquire with vnc: true checks OS availability
and requires a ready console and lease identity in the response.
relay_console_open is permitted only following an explicit user request,
requires userRequested: true, console_id, attempt_id, reason and
expected, and opens on the declared service host, not on a remote client.
relay_console_resolve refreshes the non-secret observation;
relay_console_cancel closes managed viewing resources and never releases the
VM. Console status ready is guest preflight, not launch
success; transport connection, authentication, displayed pixels and human
confirmation are separate observations that console actions never establish.
On macOS the human must choose Standard sharing of the existing console, not a
new Log In session or a High Performance display; the relay never confirms
that selection, and server_enforced_view_only: false with session_binding: viewer-selection-unverified remain limitations after transport connects.
Failed or uncertain launches retain ownership; resolve or cancel the attempt
rather than replaying it. Console tests use mocked backends; no live viewer,
installation or platform acceptance is claimed here.
Selected environments
Set VM_ENVIRONMENT_FILE to a profile that binds a loopback vm-service
endpoint, an image repository, a Tart store and separate service, image and
relay state directories together. The profile is validated before use, is
authoritative over the individual MCP_VM_RELAY_* variables, and an invalid
selection fails rather than falling back. Leases are bound to their backend and
store: an owned lease restored under a different environment fails before any
backend operation, and destruction is verified with the selected Tart binary
and TART_HOME. docs/selected-environments.md documents the schema. The
bundle the profile exports to subprocesses is the canonical one vm-service
defines, so its relay entries keep the VM_RELAY_STATE_DIR and
VM_RELAY_URL names and the three runtimes agree; this server reads its own
settings from the profile.
Evidence and cleanup
Default output is relay-evidence/<unique-task>/ under the project, with
state/ (the guest journal, action records, receipts and snapshot PNGs),
host/ (reasons, submissions, transfer facts, receipts, diagnostics, image
deliveries and lifecycle events), extractions/ (declared files, including the
page captures), and manifest.json, summary.json, trajectory.json,
index.html and OPENING.txt. Open index.html directly: no server, network,
VM or external assets are needed. Diagnostic commands appear in the trajectory
as command-only steps without screenshots. finish verifies the package before
destroying the VM; a failed delivery removes only the derived files it created
and retains the VM for a corrected attempt. Cleanup checks the read-only host
Tart inventory (the selected one, when an environment is selected) before
claiming destruction. Leases default to 4 hours with a heartbeat; the
vm-service reaper is the final backstop for process death.
Configuration
MCP_VM_RELAY_PROJECT: the project directory (the plugin passes Claude Code's). Evidence lands underrelay-evidence/<task>/there.MCP_VM_RELAY_SESSION: an explicit session identity. By default each server process takes a fresh random one; a stable identity lets a restarted server reconcile and reattach the enclosure of the same identity.MCP_VM_RELAY_URL: the loopback vm-service origin, defaulthttp://localhost:6240. Non-loopback servers, redirects and physical targets are refused.MCP_VM_RELAY_STATE_DIR: host state, otherwise$XDG_STATE_HOME/mcp-vm-relayor~/.local/state/mcp-vm-relay, one subdirectory per session.MCP_VM_RELAY_REGISTRY: an explicit task registry file; otherwise an existing compatible~/AGENTS.mdVM table, or a private managedregistry.mdin the state directory.MCP_VM_RELAY_PYTHON: the Python 3 used for the registry lock; defaultpython3on PATH.VM_ENVIRONMENT_FILE: a selected environment profile; it overrides the URL, state directory and registry above and binds the Tart store.Guest files go to
/var/tmp/<UTC-timestamp>-mcp-vm-relay-<task>/, with the runtime inreceiver.mjs, support insupport/and work inworkspace/.
Development
Build relay-driver using its own instructions, then in this checkout:
npm ci
npm run setup:dev -- /absolute/path/to/built/relay-driver
npm run check # build, typecheck, testsThree tests need a sibling checkout beside this repository and skip with a
message when it is absent: the console fixture regeneration and the
environment parity test need ../vm-service, and the shared catalog text
corpus needs ../pilot-images. The JPEG and WebP presentation fixtures in
tests/fixtures/ were encoded once from the test's spatial PNG with a codec
outside this repository and are committed, because the core carries none.
End-to-end with headless Claude Code
npm run test:e2eThis runs claude --print with the plugin loaded from this checkout
(--plugin-dir), the model limited to the relay's tools, and the server pointed
at a loopback stand-in for vm-service (tests/e2e/fixture-service.ts), which
keeps leases as records, runs the relay's own Node and transfer commands
locally, and fakes the desktop capture driver. Four cases run, each a separate
headless session: the status call (which also proves the server is pointed at
the stand-in before anything is acquired), an invalid call refused by the
contract, a full acquire, stage, exec, finish lifecycle with a verified package,
and a lease abandoned when the session ends, which is retained with its renewal
paused rather than released. What is checked is what happened on disk and at
the service, not what the model said. The suite needs claude on PATH with an
account and spends a few model turns per case, so it is not part of npm test.
The relay-driver links are development-only. The tests run the manager, the
guest programs and the shipped server against fixture HTTP services, fixture
MCP servers built with the MCP SDK and a real MCP client over stdio; no VM, desktop or user
registry is touched. Commit the regenerated dist/ with source changes.
Releasing
Cut a release with:
node scripts/bump.mjs X.Y.Z && git push --follow-tagsscripts/bump.mjs sets X.Y.Z as the version in package.json,
.claude-plugin/plugin.json and the pinned pi-mcp.json arg, refreshes
package-lock.json, rebuilds dist/, commits Release vX.Y.Z and creates the
annotated tag vX.Y.Z. It refuses to run against a dirty working tree or a
malformed version, and it never pushes — git push --follow-tags is a
separate, explicit step.
Pushing the tag triggers
.github/workflows/release.yml, which
rebuilds vmctl, checks the working build against a fresh one, and publishes
to npm using trusted publishing
(OIDC — no NPM_TOKEN secret involved), then creates the matching GitHub
release.
Trusted publishing has to be configured on npmjs.com, and npm requires the
package to already exist before you can do that. The first publish must
therefore be done by hand (npm publish --access public from a maintainer's
machine) before its trusted publisher can be configured for this workflow.
License
MIT — see LICENSE.
Available Tools
19 toolsrelay_acquireAcquire a VMA
Acquire one fresh VM for this task and start its ownership heartbeat. Requires task (a short slug), image (a key from relay_probe or relay_search) and extractions: every output you intend to bring home, declared before any work ([] is allowed; nothing is extracted automatically). Optional ttlHours (default 4), env (a credential pack, never baked into images), fullWorkspace (opt in to deliver the whole workspace) and vnc (default false; only prepares console-sharing capacity, never opens a viewer). One task per enclosure: call this once per task, not again for another task. Continue with relay_stage, relay_run (or a command tool), relay_image/relay_extract, then relay_finish or relay_release explicitly before you return; a failed step keeps the VM for repair rather than replaying uncertain input.
| Name | Required | Description | Default |
|---|---|---|---|
| env | No | Credential pack, default none; never baked into images. | |
| vnc | No | Prepare guest console sharing, default false. Never opens a viewer automatically. | |
| task | Yes | ||
| image | Yes | Image key from relay action=probe, e.g. ubuntu2404 or macos26. | |
| ttlHours | No | ||
| extractions | Yes | ||
| fullWorkspace | No | Explicit opt-in to deliver the entire workspace. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With all annotations false, the description carries the full behavioral burden and does so richly: it discloses the ownership heartbeat, that nothing is extracted automatically, that env is never baked into images, that vnc only prepares console-sharing and never opens a viewer, and that a failed step keeps the VM for repair rather than replaying uncertain input.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but mostly front-loaded, leading with the core action and required parameters. The final lifecycle sentence packs many directives and there is slight redundancy between 'One task per enclosure' and 'call this once per task,' so it is not maximally tight.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 7-parameter lifecycle tool with no output schema and unhelpful annotations, the description gives the agent everything needed to call correctly: prerequisites, optional-parameter semantics, a one-call-per-task constraint, next steps, and failure behavior. The lack of an explicit return-value description is mitigated by the stated handoff to relay_stage/relay_run.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 57%, but the description compensates by explaining task as a short slug, image as a key from relay_probe/relay_search, extractions as pre-declared outputs ([] allowed), ttlHours default 4, env as a credential pack, fullWorkspace as opt-in, and vnc's exact console-sharing semantics.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening sentence names a specific action and resource: 'Acquire one fresh VM for this task and start its ownership heartbeat.' This clearly positions it as the acquisition step and separates it from image-lookup tools like relay_probe/relay_search and later stage/run tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives explicit preconditions ('Requires task, image, extractions') and lifecycle direction ('Continue with relay_stage, relay_run ... relay_finish or relay_release'), plus an exclusion: 'call this once per task, not again for another task.' It does not name an alternative tool for the same purpose, but the pipeline context makes when-to-use clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
relay_acquisition_capabilitiesRead acquisition capabilitiesARead-onlyIdempotent
Read the backend's versioned VNC console options for the images it serves. Read-only: it neither acquires a VM nor recovers ownership of one.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false. The description adds valuable context by explicitly stating it 'neither acquires a VM nor recovers ownership of one,' clarifying behaviors beyond the generic annotation hints. This is useful behavioral disclosure about side-effect boundaries without contradicting the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two concise sentences with no filler. The core purpose is front-loaded, and the clarifying non-behavior sentence adds meaningful distinction without redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a parameterless, read-only capability inspection tool with strong annotations, the description covers purpose and side-effect boundaries sufficiently. It does not describe the return format, but this is a minor gap given the simple nature of the tool and the absence of an output schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters and schema description coverage is 100%, so the schema fully documents that no inputs are needed. The description correctly focuses on the tool's purpose rather than parameter details. The baseline of 4 for zero-parameter tools applies here.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'Read the backend's versioned VNC console options for the images it serves.' It also explicitly differentiates the tool from acquisition and ownership-recovery actions by saying it 'neither acquires a VM nor recovers ownership of one,' which distinguishes it from sibling tools like relay_acquire and relay_console_resolve.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies this tool is for inspecting capabilities without side effects, and the read-only clarification helps an agent understand not to expect acquisition. However, it does not explicitly state when to prefer this tool over alternatives such as relay_acquire or relay_console_resolve, nor when not to use it. Usage context is present but indirect.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
relay_codeRun guest codeB
Run inline guest code as this call's single recorded operation. Evidence is automatic, as for relay_exec, whose optional reason, step, snapshots and timeoutMs it shares. Requires code (up to 1 MiB) and language (javascript, typescript or python).
| Name | Required | Description | Default |
|---|---|---|---|
| code | Yes | ||
| step | No | Optional step record overrides; any field left out is derived from the call. | |
| reason | No | Optional intent, retained as evidence. Default: derived from the call. | |
| language | Yes | ||
| snapshots | No | Optional. Consecutive keystrokes may form an explicit text group (first/member/last); only first gets a before-snapshot and only last an after-snapshot. | |
| timeoutMs | No | Execution timeout in milliseconds, default 120000, maximum 3600000. Independent of snapshot delay and lease TTL. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already signal readOnlyHint=false and openWorldHint=true; the description adds that evidence is automatic and that optional params are shared with relay_exec. However, for an arbitrary code runner it does not disclose side effects, result/output behavior, failure modes, or why openWorld semantics apply, leaving significant behavior implicit.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences front-load the core purpose and add the only required-input details; no filler. The relay_exec reference is compact but slightly dense, so it is not quite a perfect example of concision.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an execution tool with no output schema, high complexity, and six parameters, the description omits what the tool returns or surfaces on success/failure, and offers no routing guidance against similar siblings. It is enough to attempt a call, but not enough to invoke correctly and interpret the result confidently.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 67%, so the schema already documents reason, step, snapshots, and timeoutMs. The description adds that code is required up to 1 MiB and enumerates languages, but those values are largely mirrored by schema constraints (maxLength, enum); it does not explain the semantics of code beyond the name and title.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The first sentence states a concrete action and object ('Run inline guest code') and frames it as 'this call's single recorded operation', which is distinctive. It references relay_exec but does not explicitly differentiate from siblings like relay_script or relay_run, so full sibling separation is not achieved.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies a use case—inline guest code in a single recorded operation—and points to relay_exec as a behavioral reference. It never states when to prefer relay_code over relay_exec, relay_script, or relay_run, and gives no exclusions or prerequisites beyond the required parameters.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
relay_console_cancelCancel console viewingA
Cancel an open or uncertain console-viewing attempt by console_id and attempt_id. Closes managed viewing resources only; it never releases or destroys the VM.
| Name | Required | Description | Default |
|---|---|---|---|
| attempt_id | Yes | ||
| console_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already indicate this is a non-read-only, non-idempotent, non-destructive operation. The description adds that it only closes managed viewing resources and never releases or destroys the VM, providing a clear boundary for side effects. It doesn't discuss failure states or prerequisites, but the added scoping is valuable.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is exactly two sentences, front-loads the core action, and every sentence adds value—first the operation, then the boundary. No wasted words or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple two-parameter cancel operation, the description covers the action's effect and its non-destructive boundary. It does not explain return values or error conditions, but given this is a small tool with no output schema, that omission is minor and unlikely to prevent correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It mentions both `console_id` and `attempt_id` as the identifiers of the attempt, but provides no additional detail about their relationship, format, or roles beyond the parameter names and regex patterns in the schema. This is minimal but adequate for a simple cancellation action.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action ('Cancel an open or uncertain console-viewing attempt') and clearly identifies the resource and scope ('Closes managed viewing resources only; it never releases or destroys the VM'). It distinguishes itself from related siblings like relay_console_open and relay_release by framing what it does and does not affect.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives clear context for when to use the tool (canceling an open or uncertain console-viewing attempt) and an explicit exclusion ('never releases or destroys the VM'), which helps an agent decide when not to use it. However, it does not explicitly name alternative sibling tools that may be considered, leaving some inference to the agent.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
relay_console_openOpen console viewingA
Open console viewing for the owned lease, only following an explicit user request to watch. Requires console_id and attempt_id from relay_console_resolve, userRequested: true, reason (intent) and expected (what the human should expect to see); it opens on the declared service host, never a remote client, and is never triggered automatically by relay_acquire's vnc option. See relay_console_resolve for what a ready status does and does not prove. A failed or uncertain open keeps the attempt: resolve or cancel it rather than opening again with a new attempt.
| Name | Required | Description | Default |
|---|---|---|---|
| reason | Yes | Intent of this execution. Retained as evidence, never identity, authorization or a retry key. | |
| expected | Yes | ||
| attempt_id | Yes | ||
| console_id | Yes | ||
| userRequested | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only indicate it is not read-only, not idempotent, not destructive, and not open-world. The description goes far beyond this by disclosing that it opens on the declared service host, never a remote client, and is never triggered automatically. It also reveals that a failed or uncertain open keeps the attempt, which is critical behavioral context for the agent to handle retries correctly.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single dense paragraph that front-loads the core purpose and then adds necessary constraints and cross-references. Every sentence carries essential information—no filler. It is structured logically: purpose, prerequisites, behavioral constraints, and failure handling.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with 5 required parameters, no output schema, and only false annotations, the description covers all essential aspects: what it does, when to use it, where it operates, what it requires, and how to handle failures. It appropriately defers to relay_console_resolve for deeper status semantics, and nothing critical for correct invocation is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 20% (only 'reason' has a description). The description compensates fully by explaining that console_id and attempt_id come from relay_console_resolve, that userRequested must be true (which it enforces as const), that reason captures intent, and that expected is what the human should expect to see. This gives the agent the semantic meaning needed to construct valid calls.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Open console viewing for the owned lease', and immediately constrains it to explicit user requests. It distinguishes itself from siblings by referencing relay_console_resolve and relay_acquire, clarifying it is never automatically triggered by relay_acquire's vnc option. The purpose is unambiguous and clearly differentiated.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides explicit when-to-use guidance: only after an explicit user request to watch. It also gives what-not-to-do: never triggered automatically by relay_acquire's vnc option. It instructs that a failed or uncertain open should be resolved or canceled rather than re-opened, effectively routing the agent to the correct alternative actions (relay_console_resolve/relay_console_cancel).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
relay_console_resolveResolve console statusARead-onlyIdempotent
Refresh this session's non-secret console-viewing status for the owned lease. Read-only, and it never recovers ownership. A ready status means guest preflight only, not that a viewer launched: console operations carry no screenshots and never establish authentication, displayed pixels or a human's confirmation. On macOS, viewing requires the human to choose Standard sharing of the existing console, not a new Log In session or a High Performance display, and the relay never auto-confirms that choice; server_enforced_view_only: false and an unverified viewer selection remain limitations even once transport connects. Never replay an uncertain attempt: resolve it here, or close it with relay_console_cancel.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Goes well beyond the readOnlyHint, idempotentHint, and destructiveHint annotations. It discloses that the tool never recovers ownership, never establishes authentication or human confirmation, never auto-confirms the macOS sharing choice, and that view-only limitations persist even after transport connects. No contradiction with the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The core action is front-loaded in the first sentence, and every subsequent sentence adds meaningful caveats or routing guidance. The description is dense and somewhat long for a zero-parameter refresh action, but the platform-specific warnings justify most of the length.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description explains the meaning of a 'ready' result, the operation's limitations, the macOS prerequisites, and the cancellation alternative, which is strong context given zero parameters and rich annotations. It does not spell out the exact response shape or possible status values beyond 'ready,' but that is a minor gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so the schema is trivially complete and there are no parameter descriptions needed. The description adds useful contextual semantics by referring to 'this session' and 'the owned lease,' which frames the scope of the operation. This is the appropriate baseline for a no-parameter tool.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States the specific action: 'Refresh this session's non-secret console-viewing status for the owned lease.' It clarifies this is read-only, never recovers ownership, and contrasts with viewer-launching behavior by noting it carries no screenshots and never establishes authentication or displayed pixels. The 'ready' status caveat further pins down exactly what the tool does and does not confirm.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides explicit when-to-use guidance: 'Never replay an uncertain attempt: resolve it here, or close it with relay_console_cancel.' It also names the alternative tool and gives platform-specific prerequisites for macOS (Standard sharing of the existing console, not a new Log In session or High Performance display), so an agent knows how to interpret a successful resolution.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
relay_execRun a guest commandA
Run one guest command (argv) as one recorded operation; never hide several interactions in one call. Evidence is automatic: the relay snapshots the display before and after and records a step derived from the command. Optional reason (intent, never authorization), step (id, title, expected, inputMode) and snapshots.afterIntervalMs (default 500 ms) enrich or tune the record; consecutive keystrokes may form an explicit text group with snapshots.group (first/member/last). timeoutMs bounds execution (default 120000, up to 3600000). The result carries the outcome, bounded stdout/stderr and the after-snapshot inline. A refused, uncertain or nonzero result keeps the VM; never replay input whose effect is uncertain. diagnostic: true records a diagnosis or repair without snapshots, even before relay_stage; it is never visual verification, and its result has one output field with stdout and stderr combined.
| Name | Required | Description | Default |
|---|---|---|---|
| argv | Yes | ||
| step | No | Optional step record overrides; any field left out is derived from the call. | |
| reason | No | Optional intent, retained as evidence. Default: derived from the call. | |
| snapshots | No | Optional. Consecutive keystrokes may form an explicit text group (first/member/last); only first gets a before-snapshot and only last an after-snapshot. | |
| timeoutMs | No | Execution timeout in milliseconds, default 120000, maximum 3600000. Independent of snapshot delay and lease TTL. | |
| diagnostic | No | Explicit command diagnosis or repair without screenshot evidence, including before staging. Commands and outcomes remain recorded. Never claim visual verification. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description discloses many non-obvious behaviors beyond annotations: automatic before/after display snapshots, derived step recording, bounded stdout/stderr in results, VM retention on refused/uncertain/nonzero results, and the diagnostic mode's different snapshot/output behavior. This is exactly the kind of context annotations cannot convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but every sentence earns its place. It front-loads the core operation, then clusters optional parameters and safety/behavioral notes without fluff. Long paragraphs are justified by the tool's complexity and the high-value details they convey.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the complex nested schema and no output schema, the description compensates fully: it explains result contents, snapshot grouping behavior, timeout bounds, diagnostic mode semantics, and safety guarantees. An agent has enough to invoke the tool correctly and interpret its return value.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Even with 83% schema coverage, the description adds meaning beyond the schema: `reason` is intent and 'never authorization', snapshots.group is for consecutive keystrokes with explicit first/member/last phases, afterIntervalMs defaults to 500 ms, and diagnostic output is a single combined `output` field. These clarifications help an agent tune calls correctly.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Run one guest command (`argv`) as one recorded operation'. It also adds a crucial scoping rule, 'never hide several interactions in one call', which distinguishes this executor from bulk or multi-step siblings. The title reinforces the same message, so an agent can select it confidently.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear usage context: it is for single guest commands, evidence is recorded automatically, and `diagnostic: true` is for diagnosis/repair even before relay_stage. It also warns not to replay input whose effect is uncertain. It does not explicitly name sibling tools for comparison, but the guidance is sufficient for correct invocation.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
relay_extractExtract declared outputsA
Pull one or more already-declared extraction outputs home early, by names. Only outputs declared on relay_acquire may be pulled; each pull verifies source and host hashes and rejects path traversal, symlinks, or a source that changed underneath it.
| Name | Required | Description | Default |
|---|---|---|---|
| names | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only indicate readOnlyHint=false and destructiveHint=false, leaving safety to the description. The description adds valuable behavioral context: verification of source and host hashes, rejection of path traversal, symlinks, and changed sources. It does not contradict annotations; the pull operation is consistent with readOnlyHint=false.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences, front-loaded with the core action and then adding essential constraints. Every word contributes; there is no filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with one parameter and no output schema, the description covers what it does, the precondition (declared on relay_acquire), and the safety checks. It does not specify return values or post-pull effects, but those are likely implied by the operation name and context. Minor gap, not critical.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% (no parameter descriptions). The description mentions 'by `names`', indicating the parameter is a list of output names, and ties it to 'already-declared extraction outputs'. This adds some meaning beyond the bare array-of-strings schema, but it does not explain format, constraints, or examples, so it only partially compensates for the schema gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('pull'), resource ('already-declared extraction outputs'), and mechanism ('by names'). Differentiates from siblings by noting the constraint 'Only outputs declared on relay_acquire may be pulled', which clearly separates it from other relay tools like relay_acquire or relay_stage.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains the intended use: pulling already-declared outputs early. It provides a clear context (outputs must be declared on relay_acquire) but does not explicitly name alternatives or give when-not-to-use guidance. Still, the context is specific enough for an agent to infer when this tool is appropriate.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
relay_finishFinish and deliver evidenceADestructive
Complete the task: extract every declared output, deliver and verify a portable evidence package, then destroy the VM and unregister it. The result reports delivery, snapshot completeness, execution outcome and human review as separate facts; a verified package is not a passing test, and a delivered package is not itself human approval. Call this, or relay_release, explicitly before you return, since ending the session does not do it for you. A failed delivery keeps the VM for a corrected attempt.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations say destructiveHint=true, and the description goes further by disclosing VM destruction, unregistration, and the delivery/verification workflow. It also clarifies important result semantics: a verified package is not a passing test, and a delivered package is not human approval, which is valuable beyond the structured annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded with the main workflow, then covers result semantics, call timing, and failure behavior. Every sentence adds distinct operational value with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter finalization tool with no output schema, the description covers what the tool does, what the result reports, when to call it, and what happens on failure. Nothing critical is missing for an agent to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There are zero parameters and the schema is fully empty, so the baseline is 4. The description adds no parameter-level ambiguity and needs no parameter documentation; its mention of 'every declared output' is behavioral context rather than a parameter definition.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb-resource pair: 'Complete the task' by extracting declared outputs, delivering/verifying evidence, and destroying/unregistering the VM. This clearly separates it from most siblings, but it does not distinguish relay_finish from the closely related relay_release beyond saying 'Call this, or relay_release'.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives explicit timing guidance: call before returning because ending the session does not finalize. It also names relay_release as an alternative and describes failure behavior (failed delivery keeps the VM), but it does not specify when to prefer this tool over relay_release.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
relay_imageRetrieve a saved imageARead-onlyIdempotent
Retrieve one already-saved image; never a new capture, input or directory export, and it never acquires a VM. target selects a display phase (sessionId/executionId/phase: before/after), a declared application file (name, plus a relative path for a directory declaration), or an immutable reference (imageId). PNG, JPEG and WebP only; originals up to 64 MiB and 40,000,000 decoded pixels; the delivered preview is at most 2000x2000 px and 4 MiB of base64 (PNG originals are resampled in-process; JPEG/WebP pass through only within bounds). Each delivery has a 90-second deadline and up to three transfer attempts. A relay_run tool image is the application file name: "relay-run" at the path its result gives. If a run's inline image did not arrive, recover it here with the same imageId or display selector, never by repeating the input or capturing again, and stop after at most two such recovery calls if it still cannot be inspected. A closed enclosure returns stale-reference for a display or reference target; its delivered originals stay readable with host file tools. An attached image proves only that the block was included, not that anyone inspected or reviewed it.
| Name | Required | Description | Default |
|---|---|---|---|
| target | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and idempotentHint=true, but the description adds extensive behavioral context: it confirms the tool never acquires a VM, specifies format and size limits (PNG/JPEG/WebP, up to 64 MiB, 40M pixels, preview at most 2000x2000 px and 4 MiB base64), describes delivery behavior (90-second deadline, up to three attempts), and clarifies the meaning of 'attached image proves only inclusion, not inspection'. No contradictions with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but well-structured, with the core purpose front-loaded and technical details organized logically. While it is longer than typical tool descriptions, the complexity of the tool justifies the length. It could be slightly more concise by trimming redundant explanations, but every sentence carries value.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (nested anyOf target variants, no output schema, no explicit parameter documentation), the description covers all necessary aspects: usage scenarios, constraints (size, format, deadlines), error handling (`stale-reference`), and recovery workflow. An agent would have sufficient information to call the tool correctly and interpret results.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 0% description coverage, but the description compensates fully by explaining the semantics of the `target` parameter: it distinguishes the three source types (display, application, reference), defines the fields for each (e.g., `sessionId`/`executionId`/`phase` for display; `name` and `path` for application; `imageId` for reference), and clarifies the purpose of `relay-run`. This adds meaning far beyond the raw schema structure.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description explicitly states 'Retrieve one already-saved image' and enumerates what it is not ('never a new capture, input or directory export, and it never acquires a VM'), clearly distinguishing it from siblings. The three target forms (display, application, reference) are defined with precise semantics, leaving no ambiguity about the tool's function.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It provides explicit when-to-use instructions, including recovery of relay_run images ('If a run's inline image did not arrive, recover it here...'), and when-not-to-use guidance ('never by repeating the input or capturing again'), with a stop condition ('stop after at most two such recovery calls'). It also mentions alternative tools for closed enclosures (host file tools). This is exemplary usage guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
relay_probeProbe host and guest readinessARead-onlyIdempotent
Read-only host and service facts, used to judge whether a task is interruptive and to pick an image; it reports current owned state on its own, so call it again after context compaction rather than trusting earlier messages. Default scope: "host" reports host permissions and activity, VM service availability, and this session's owned lifecycle state, not guest readiness; unknown is not idle. scope: "guest" (on an already-owned VM) reports whether Node, CuaDriver and each relay_run target (cua, playwright, chrome-devtools) are available, without installing or starting anything, and never claims capture readiness. Relay only interruptive work judged from these facts; non-disruptive or headless work stays with local tools. Continue in order: relay_acquire, relay_stage, relay_run (or a command tool), relay_image/relay_extract, then relay_finish or relay_release, called explicitly before you return.
| Name | Required | Description | Default |
|---|---|---|---|
| scope | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already mark read-only, idempotent, and non-destructive; the description adds that it reports current owned state and must be re-called after compaction, that guest scope installs or starts nothing, and that it never claims capture readiness. It also clarifies 'unknown is not idle.' These are meaningful behavioral details beyond the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Every sentence carries distinct information: purpose, state-freshness caveat, host scope semantics, guest scope semantics, usage boundary, and workflow order. It is dense but not wasteful, and it front-loads the read-only purpose and the key re-call behavior.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the absence of an output schema, the description summarizes what each scope reports and warns about known pitfalls (unknown vs idle, capture readiness). It also embeds the tool in the relay sequence so an agent knows when and in what order to invoke it. Some exact response structure is not specified, but the essential context is present.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% and the schema only lists an enum, but the description thoroughly explains both `scope: "host"` and `scope: "guest"` including what each reports and their caveats. This gives an agent the full semantic meaning of the only parameter, even without additional schema descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: it probes read-only host/service facts to judge interruptiveness and pick an image. It explicitly distinguishes the two scopes and notes that `host` is not guest readiness, which separates it from a generic probe. The tool's role in the relay workflow is immediately clear.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says when to use it: before interruptive relay work, and after context compaction rather than trusting earlier messages. It names the alternative (local tools for non-disruptive/headless work) and provides the full ordered relay lifecycle. No ambiguity about where this tool fits.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
relay_releaseRelease the VMADestructive
Abandon the task: destroy the owned VM and unregister it, retaining whatever evidence already exists, without claiming the task succeeded. Use this instead of relay_finish when the task is not being completed. Safe to retry if a previous release attempt failed; a failed release keeps ownership until destruction is verified.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already mark destructiveHint=true, but the description adds valuable context about evidence retention, success status, and retry behavior. It goes beyond the structured metadata to describe side effects and failure semantics.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three concise sentences, front-loaded with the primary action, then the usage distinction, then retry safety. Every sentence adds necessary information with no redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter destructive action with no output schema, the description covers purpose, usage, side effects, and retry behavior. Nothing an agent needs to decide to call this tool is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so there is nothing to document. The description correctly omits parameter details, and the empty schema fully covers the interface. Baseline of 4 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description explicitly states the tool destroys and unregisters the owned VM, retains evidence, and does not claim success. It clearly distinguishes itself from relay_finish by naming the alternative and the condition for use.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It directly instructs to use this instead of relay_finish when the task is not being completed, and clarifies retry semantics (safe to retry, failed release keeps ownership). This gives explicit when-to-use and when-not-to-use guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
relay_runRun an MCP tool call in the VMA
Send one tool call to an MCP server inside the VM. target is cua (cua-driver), playwright (Playwright MCP) or chrome-devtools (Chrome DevTools MCP); tool and args are that server's own tool name and arguments, forwarded unchanged. Use relay_tools to see a target's exact tools. Evidence is automatic: snapshots before and after, and a step record; reason, expected and afterIntervalMs are optional overrides. Never replay an uncertain call.
| Name | Required | Description | Default |
|---|---|---|---|
| args | No | The tool's own arguments, forwarded unchanged. | |
| tool | Yes | The target server's own tool name. | |
| reason | No | Optional intent, retained as evidence. Default: derived from the call. | |
| target | Yes | cua (cua-driver), playwright (Playwright MCP) or chrome-devtools (Chrome DevTools MCP). | |
| expected | No | Optional expected result for the step record. | |
| timeoutMs | No | Execution timeout in milliseconds, default 120000, maximum 3600000. Independent of snapshot delay and lease TTL. | |
| afterIntervalMs | No | Wait in milliseconds from the end of the call to the after-snapshot. Optional; the relay has a default. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the annotations (readOnlyHint=false, openWorldHint=true, idempotentHint=false), the description discloses a concrete behavioral trait: 'Evidence is automatic: snapshots before and after, and a step record,' and that reason/expected/afterIntervalMs only override that evidence. The caution 'Never replay an uncertain call' aligns with the non-idempotent, open-world hints. It stops short of a 5 because it does not address failure/error propagation or side-effect expectations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four sentences, roughly 70 words, each earning its place: purpose, target enumeration, discovery routing, evidence model, and a safety caution. The core verb-resource pair is front-loaded in the first sentence, and there is no fluff or redundant restatement of schema descriptions.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex 7-parameter open-world tool with no output schema, the description covers the essential operational knowledge: target choices, forwarding semantics, the discovery prerequisite, automatic evidence, and optional overrides. Minor gaps remain — the exact return/step-record shape and the lease prerequisite implied by siblings relay_acquire/relay_release — but these do not prevent a correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description adds meaning by grouping parameters into two semantic classes — tool/args are 'forwarded unchanged' (the actual remote call) versus reason/expected/afterIntervalMs as 'optional overrides' of the evidence record — a relationship the per-property schema entries only imply. This framing tells the agent which parameters affect the remote server and which only affect the step record.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb-resource pair, 'Send one tool call to an MCP server inside the VM,' and enumerates the three valid targets with parenthetical mappings (cua, playwright, chrome-devtools). It also clarifies the forwarding semantics for tool and args, and the relay_tools reference helps distinguish this tool from its discovery sibling without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly routes the discovery step to the relevant sibling — 'Use relay_tools to see a target's exact tools' — which is the natural confusion point. It also states a when-not condition with 'Never replay an uncertain call' and clarifies which parameters are optional evidence overrides, so an agent knows what to supply and what to leave out.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
relay_scriptRun a guest scriptA
Run one guest script file as this call's single recorded operation. Evidence is automatic, as for relay_exec, whose optional reason, step, snapshots and timeoutMs it shares. Requires localPath (a host file path) and language (javascript, typescript or python); it runs under the staged workspace.
| Name | Required | Description | Default |
|---|---|---|---|
| step | No | Optional step record overrides; any field left out is derived from the call. | |
| reason | No | Optional intent, retained as evidence. Default: derived from the call. | |
| language | Yes | ||
| localPath | Yes | ||
| snapshots | No | Optional. Consecutive keystrokes may form an explicit text group (first/member/last); only first gets a before-snapshot and only last an after-snapshot. | |
| timeoutMs | No | Execution timeout in milliseconds, default 120000, maximum 3600000. Independent of snapshot delay and lease TTL. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description adds meaningful context beyond the annotations: evidence is automatic, this is a single recorded operation, and execution happens under the staged workspace. It does not contradict the annotations. However, for a tool that executes arbitrary guest scripts, it does not disclose potential side effects, sandboxing limits, or what happens on failure, leaving the agent to infer behavioral risk beyond the basic openWorldHint and readOnlyHint flags.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences with no filler. It front-loads the core purpose ('Run one guest script file'), then packs the essential constraints and cross-reference to relay_exec into the second sentence. Every clause contributes useful information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with no output schema, the description is missing a critical piece: what the caller receives after execution (stdout, exit code, recorded evidence reference, error behavior). It covers invocation inputs, shared optional parameters, and workspace context, but an agent cannot fully predict how to consume the result. The nested snapshots and step objects are explained by the schema, so the biggest gap is the unspecified return/result semantics.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description adds meaning for the two required parameters that the schema leaves undocumented: localPath is 'a host file path' and language is restricted to 'javascript', 'typescript' or 'python'. Since schema coverage is 67%, the description compensates for the biggest gap. It does not need to re-explain step, reason, snapshots, and timeoutMs because those already have descriptions in the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'Run one guest script file as this call's single recorded operation.' It adds useful scope by calling out that this is a single recorded operation, which separates it from broader multi-step tools. However, it does not explicitly contrast it with sibling tools like relay_exec, relay_code, or relay_run, instead only referencing relay_exec for shared optional parameter semantics.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the primary use case: run a guest script file under the staged workspace. It also states required inputs (localPath and language), which gives the agent a baseline for when the tool is applicable. But it never explicitly says when to prefer this tool over relay_exec, relay_code, or relay_run, nor does it provide exclusions or alternative routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
relay_searchSearch installed applicationsARead-onlyIdempotent
Find an installed application in the image inventories by name, then choose an image yourself; optional before relay_acquire. Matches canonical names and aliases by exact, prefix and substring comparison, with an optional hard os filter (linux/macos). Returns each match's installed versions (which may be null), image keys, OS and architecture; a truncated result means whole installations were omitted, and an empty result is distinct from a catalog error. Never allocates, boots, installs, reserves capacity or chooses an image for you.
| Name | Required | Description | Default |
|---|---|---|---|
| os | No | ||
| name | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint and idempotentHint, but the description adds behavioral details: matching logic (exact/prefix/substring), optional os filter, return shape (installed versions possibly null, image keys, OS, architecture), and the distinction between truncated results and empty results versus catalog errors. It also clarifies it never allocates, which adds safety context beyond the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is comprehensive yet efficient, leading with the core purpose, then matching behavior, then return semantics, and ending with what it does not do. Every sentence adds distinct value with no redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no output schema, the description explains return values (installed versions, image keys, OS, architecture) and handles edge cases (null versions, truncated results, empty vs error). It also specifies the optional filter. For a search tool, this covers everything an agent needs to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description fully compensates by explaining the 'name' parameter as matching by exact, prefix, and substring, and the 'os' parameter as an optional hard filter with enum values linux/macos. This gives meaning beyond the schema's bare property definitions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a clear verb 'Find' and resource 'installed application in the image inventories', and it explicitly differentiates from relay_acquire by stating it is optional before acquiring and does not choose an image. This distinguishes it from sibling tools like relay_acquire and relay_probe.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It states it is optional before relay_acquire, and clarifies what it never does (allocates, boots, installs, reserves capacity, chooses an image), guiding when to use it. It does not explicitly name alternatives, but the context of being a search step is clear. The 'optional before relay_acquire' provides a clear usage slot.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
relay_stageStage the guest runtimeA
Push and hash-check the guest runtime, plus an optional workspace and support files (files, landing under support/), onto an acquired VM. The guest Node and CuaDriver executables must already exist (nodePath, cuaDriver); Linux guests need native X11 and cua-driver serve --no-overlay. browserExecutable names the guest browser the playwright and chrome-devtools targets launch (default: installed Google Chrome); their pinned servers are staged on first use. A failed stage can be retried; once staged, only corrected executable paths may be resubmitted, and staging does not prove capture readiness. resetRecording: true archives the current recording's evidence and starts a fresh recording on the same VM; it refuses while a receiver lock is held.
| Name | Required | Description | Default |
|---|---|---|---|
| files | No | ||
| nodePath | No | Guest Node executable, default node; use image nvm path if needed. | |
| cuaDriver | No | Guest CUA executable; OS default when omitted. | |
| workspace | No | Host directory copied to the guest workspace, opt-in. | |
| resetRecording | No | Explicitly archive the current recording and start a new session on the same VM. Retains prior evidence; refuses while the receiver lock exists. Use after diagnosing recording damage. | |
| browserExecutable | No | Guest browser executable for the playwright and chrome-devtools targets; default is installed Google Chrome. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With all annotation hints false or unhelpful, the description carries the behavioral burden and does so richly. It discloses hash-checking, staging of support files, first-use server staging, retry semantics, the limitation that 'staging does not prove capture readiness,' and reset/resubmission behavior including the receiver-lock refusal.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single dense paragraph with no filler. Every clause earns its place: scope, prerequisites, Linux-specific requirements, retry rules, reset behavior, and the readiness caveat are all packed in without redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given six parameters, no output schema, and no meaningful annotations, the description is unusually complete. It covers prerequisites, retry and resubmission constraints, browser default and target behavior, reset semantics, and the warning that staging does not prove capture readiness—enough for an agent to invoke it correctly after acquisition.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is high (83%), so the baseline is 3, but the description adds meaningful context: `files` lands under `support/`, `nodePath`/`cuaDriver` must already exist, `browserExecutable` selects the launched guest browser for specific targets, and `resetRecording` archives evidence and starts fresh. It does not add meaning to `files[].local`, whose schema description is also missing.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific action: 'Push and hash-check the guest runtime, plus an optional workspace and support files... onto an acquired VM.' This clearly identifies the resource and distinguishes staging from siblings like relay_acquire or relay_exec/run.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Prerequisites are explicit: guest Node and CuaDriver executables must already exist, and Linux guests need native X11 plus `cua-driver serve --no-overlay`. It also defines retry/resubmission constraints and when `resetRecording` is appropriate, though it does not explicitly name alternatives or say 'use relay_stage instead of X.'
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
relay_statusShow relay statusARead-onlyIdempotent
Show this session's owned VM lease (backend binding, guest state, renewal, console observation, last error), staging state and evidence path, plus the project directory, the VM service origin and the selected environment in use; active:false when nothing is owned. Read-only; it does not touch the VM.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, destructiveHint=false, and idempotentHint=true, so the safety profile is covered. The description adds the active:false behavior when nothing is owned and repeats 'Read-only; it does not touch the VM,' which is redundant with annotations. The added context about the active flag is useful, but overall the description provides only marginal behavioral detail beyond the annotations, earning a 3.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single long sentence but front-loaded with the primary output (owned VM lease) and then lists the rest. It is efficient, avoids fluff, and every phrase adds information. It could be slightly more structured, but it is appropriately sized and informative.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has no parameters, no output schema, and annotations covering the safety profile, the description is quite complete. It enumerates all the fields the agent can expect (lease details, staging state, evidence path, project directory, service origin, environment, and active flag). The only missing element is the exact return format (e.g., JSON structure), but that is not critical for a status tool and the description suffices for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, and the schema coverage is 100% (empty schema). The description does not need to explain parameters, and the baseline for zero-parameter tools is 4. No additional parameter information is required, so this dimension is satisfied.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool shows the session's owned VM lease, staging state, evidence path, project directory, VM service origin, and selected environment, plus an active:false indicator. It uses a specific verb ('show') and identifies a distinct resource (relay status) with a detailed list of contents, making its purpose unambiguous. However, it does not explicitly differentiate from sibling tools like relay_probe or relay_console_resolve, though the read-only inspection nature is evident from context.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the tool is for inspecting the current session state before or after operations, but it does not provide explicit guidance on when to use it versus alternatives. No exclusions or conditions are stated, leaving the agent to infer from the read-only nature and the listed contents. This is adequate but lacks the explicit 'when to use' framing that would merit a 4.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
relay_toolsList a target's toolsAIdempotent
List a relay_run target's real tools from its server inside the VM: names, descriptions and input schemas; tool returns just one. Starts the server if needed, without sending it any tool call. Requires relay_stage.
| Name | Required | Description | Default |
|---|---|---|---|
| tool | No | Optional: return only this tool. | |
| target | Yes | cua (cua-driver), playwright (Playwright MCP) or chrome-devtools (Chrome DevTools MCP). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare idempotentHint=true and destructiveHint=false; the description adds valuable context that it starts the server if needed without sending a tool call. This goes beyond what annotations provide, with no contradiction.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two concise sentences with the core purpose front-loaded and no filler. The second sentence adds side effects and prerequisites without redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, but the description covers what is returned (names, descriptions, input schemas) and the side effect of starting the server. It also notes the prerequisite. Missing error handling is not essential for a listing tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% with descriptive parameter text for both `tool` and `target`. The description adds only a minor clarification that `tool` returns just one, which is already implied by the schema. No significant extra meaning over the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific action ('List') and a clear resource ('a relay_run target's real tools'), and spells out what is returned (names, descriptions, input schemas) plus the `tool` option to return one. This clearly distinguishes it from siblings like relay_search or relay_probe.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides a prerequisite ('Requires relay_stage') but does not explicitly compare with sibling tools or state when to use this instead of alternatives. The context is clear enough to infer use, but no exclusions or alternatives are named.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
relay_trajectoryOpen the trajectory viewerAIdempotent
Verify a delivered relay evidence package (every artifact, hash and reference) and open its trajectory viewer in the local browser. Human review remains pending.
| Name | Required | Description | Default |
|---|---|---|---|
| directory | Yes | The package directory, absolute or relative to the project. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare idempotentHint=true and destructiveHint=false. The description adds that it opens a local browser (side effect), verifies every artifact/hash/reference, and leaves human review pending. This contextual behavior goes beyond the annotations without contradicting them.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the core action and scope, followed by a state note. No wasted words; the description is efficient and well-structured.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with one parameter and no output schema, the description covers purpose, verification scope, side effect (browser open), and pending human review. It does not explain failure behavior or return value, but these are minor given the simplicity and annotations.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The only parameter 'directory' is fully described in the schema (100% coverage). The description does not add extra parameter details beyond referring to the package directory. Baseline 3 is appropriate since the schema handles parameter meaning.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('verify') and resource ('delivered relay evidence package') followed by an action ('open its trajectory viewer'). It clearly identifies what the tool does, and the trajectory viewer is distinct from sibling console tools, though it doesn't name an alternative explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage: when a relay evidence package has been delivered and needs verification and viewing. It gives clear context ('delivered relay evidence package') but does not explicitly state when not to use it or name alternatives like relay_console_open. This is adequate without exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
20 tool updates
v0.5.1- Removed
relay - Added
relay_acquire - Added
relay_acquisition_capabilities - Added
relay_code - Added
relay_console_cancel - Added
relay_console_open - Added
relay_console_resolve - Added
relay_exec - Added
relay_extract - Added
relay_finish - Added
relay_image - Added
relay_probe - Added
relay_release - Removed
relay_review - Added
relay_run - Added
relay_script - Added
relay_search - Added
relay_stage - Added
relay_tools - Added
relay_trajectory
3 tool updates
v0.3.1- First observed
relay - First observed
relay_review - First observed
relay_status
TDQS
Scored across 19 tools
Most tools have distinct purposes: acquire, stage, execute, extract, finish, release are lifecycle steps; probe, status, search, and capabilities cover discovery/state; console operations are separate; image and trajectory are retrieval/verification. Some overlap exists between relay_exec, relay_script, and relay_code (all run guest code) but inputs differ (command vs file vs inline), and relay_probe vs relay_status could be confused but descriptions clarify host/guest facts vs lease state.
All tools share the 'relay_' prefix, but the second part mixes verbs (exec, search, acquire, stage, extract, finish, release, resolve, open, cancel) with nouns (status, trajectory, script, code, run, tools, image) and even a noun phrase (acquisition_capabilities). While readable, the pattern is inconsistent—some names describe actions, others describe resources—making it less predictable than a uniform verb_noun convention.
19 tools is on the higher end but appropriate for a VM relay server that manages a full lifecycle (acquire, stage, run variants, extract, finish/release) plus supporting operations (status, probe, search, console, image, trajectory). Each tool serves a distinct purpose in the workflow, though the count could be trimmed by merging some run variants, but the complexity justifies it.
The tool surface covers the entire VM workflow: discovery (search, probe, capabilities), acquisition (acquire), setup (stage), execution (exec, script, code, run), monitoring (status, console resolve/open/cancel), output retrieval (extract, image), and completion (finish, release). There are no obvious gaps—even edge cases like early extraction and evidence verification (trajectory) are handled. The lifecycle is closed with explicit finish/release steps.
Maintenance
Related MCP Connectors
- TabfleetOAuthcom.tabfleet
Launch, inspect, control, and share isolated cloud browsers for your agents.
Run in-product voice interviews with AI agents and analyze source-linked evidence.
Run multi-step tasks in a real Chrome browser: persistent environments, live view, human takeover.
Hosted MCP catalog with 30 tenant-isolated browser, RAG, AI, mail and media tools.
Related MCP Servers
- FlicenseNot gradedqualityCmaintenanceEnables autonomous desktop automation by delegating tasks to vision-based agents operating within cloud-based virtual machine sandboxes. It allows users to manage VMs, execute complex computer tasks, and receive text-based screen summaries across Linux, Windows, and macOS environments.2-
- AlicenseNot gradedqualityBmaintenanceProvides AI agents with isolated virtual desktops containing a real Chromium browser, enabling them to see, click, type, and navigate like a human, with features like snapshots and remote control.23AGPL 3.0
- AlicenseNot gradedqualityCmaintenanceEnables a Copilot Studio agent to operate a dedicated, disposable Windows VM by running commands, PowerShell scripts, file operations, and background jobs over authenticated HTTPS.MIT
- AlicenseNot gradedqualityBmaintenanceEnables an agent to drive a real browser while enforcing provenance-based gating so URLs and form values must trace to the user or an allowlist, never to untrusted page content.13 PyPIAGPL 3.0