PS5 MCP
PS5 MCP gives an agent eyes and hands on a jailbroken PS5: capture the screen, send virtual controller input, manage titles and files, and automate visual or recorded flows.
See the console:
snapshot(JPEG frame, up to 1920x1080),status(payload link/version, counters, capture state), andshow_viewerto pop the live window for a human.Send controller input:
press,hold(multiple buttons),stick(analog deflection),trigger(L2/R2),touchpad,sequence(timed multi-step states, capped at 60 s), andrelease_all. Most input tools can return a frame viasnapshot_after_ms.Control apps and the system:
home(suspend/PS-button stand-in),close_app,launcha title by ID,list_apps, anduninstall_apps(destructive).Install software:
installa.pkgpackage or.elfpayload from the Mac (optionally running the payload on install).Wait for things to happen:
wait_for_change(screen changed) andwait_for/save_template/list_templatesfor template-matching on a screen region, present or gone.Record and replay:
record_start,record_stop,list_recordings,play_recording— captures agent and keyboard input, replays with original timing.Share the console:
claim_console/release_console(referenced in the README; claims queue other agents and block their input, console commands, and installs).
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@PS5 MCPtake a snapshot of my PS5 screen and press X"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
PS5 MCP
See and control a jailbroken PS5 from an MCP agent or your keyboard, with live HDMI video and audio. macOS only for now, Windows and Linux support is planned.
Requirements
A jailbroken PS5, currently firmware 13.60. Payload Manager and PS5 Web File Manager are recommended (see Console services).
An HDMI capture card connected to the Mac, with HDCP disabled on the PS5 (Settings → System → HDMI → Enable HDCP).
macOS 14 or later, with network access to the console.
To build: Xcode with a signing account, Homebrew, Python 3.12+, and the tools installed below. The payload uses PS5 Payload SDK v0.43 and LLVM 18.1.8.
Console services
Controlling the console needs only padd running, however it was loaded. These payloads add padd Start/Stop, the
game/app list, Install, and file transfers:
Service | Port | Used for |
Payload Manager (v0.5.2 tested) | 8084 | Deploying |
PS5 Web File Manager (v1.9 tested; Payload Manager can load it) | 8888 | Verifying the |
zftpd (v1.6.0 FTP-only | 2120 |
|
Without them, load padd.elf with any ELF loader.
PS5 MCP expects Web File Manager on port 8888; it moves to the next free port when 8888 is taken.
To add zftpd, download zftpd-ps5-v1.6.0.elf from its
releases, rename it to zftpd.elf, and add it to Payload Manager's
library with uv run ps5mcp install zftpd.elf (or Payload Manager's web page). Do not run it yourself: push and
pull start it when they first need it. Without zftpd, push is slower and pull copies single files only.
Related MCP server: vvd
How it works
paddis a payload running on the PS5. It creates a virtual DualSense and accepts controller input over TCP port 9305.The PS5 MCP app captures HDMI video and audio, owns the connection to
padd, and provides keyboard controls, payload Start/Stop, status, and install/uninstall. It talks to Payload Manager and Web File Manager only for the tasks listed above.The MCP server connects agents to the app's local API. Multiple agent sessions and the keyboard can share the console.
Setup
brew install llvm@18 ffmpeg mpv uv xcodegen
make bootstrap # pinned SDK and Python environment
make check # lint, tests, payloads, and the macOS release appThe release app is built at build/PS5 MCP.app. It is signed ad hoc, so macOS asks for its permissions again after
each rebuild. To keep them, sign with your Apple team: make app DEVELOPMENT_TEAM=… APP_BUNDLE_ID=… or
app/Configuration/LocalSigning.xcconfig (copy the .example). Quit the app before rebuilding.
No make target contacts the console.
Usage
Point it at the console:
export PS5_HOST=192.168.1.20(your PS5's IP address). There is no default. The CLI commands also take--host. In the app, set the address in Settings (⌘,). The app saves it and reconnects to it at once. When the app starts,--hostandPS5_HOSTwin over the saved address.Open the app:
uv run ps5mcp vieworopen --env PS5_HOST=$PS5_HOST "build/PS5 MCP.app". Allow camera, microphone, local network, and Documents access when prompted.Start
paddwith Start oruv run ps5mcp padd start, once per console boot. The app answers the user-selection dialog automatically. If your DualSense turns off, press its PS button and select your user again.Control the console from the app window or connect an MCP client (see MCP clients).
Stop
paddwith Stop oruv run ps5mcp padd stopbefore turning off the console. Killing it through Payload Manager can leave a phantom controller until reboot.
The MCP server starts the app when needed, headless by default (PS5MCP_VIEW=1 shows the window).
Closing an MCP session leaves the app running; quitting the app releases input and leaves padd running.
MCP clients
Every client runs the same server from this checkout: uv run --project <repo> ps5mcp-server, with PS5_HOST set.
Use the full path to uv (command -v uv) if the client does not have Homebrew on its PATH.
Claude Code:
claude mcp add ps5 -e PS5_HOST=192.168.1.20 -- uv run --project /path/to/ps5-mcp ps5mcp-serverCodex: add this to ~/.codex/config.toml (or run
codex mcp add ps5 --env PS5_HOST=192.168.1.20 -- uv run --project /path/to/ps5-mcp ps5mcp-server and add the
timeouts after). Installs and file transfers can take minutes, longer than Codex's default tool timeout.
[mcp_servers.ps5]
command = "/opt/homebrew/bin/uv"
args = ["run", "--project", "/path/to/ps5-mcp", "ps5mcp-server"]
startup_timeout_sec = 60 # the first start can launch the app
tool_timeout_sec = 600 # install, push, and pull
[mcp_servers.ps5.env]
PS5_HOST = "192.168.1.20"Other clients that read a JSON config: see examples/mcp.json.
Check the setup with claude mcp list or codex mcp list, then /mcp in a session. Agents from all clients share
the console through the same claim and queue (see Sharing the console). A running session
keeps the server code it started with; restart it after updating this checkout.
Live window and keyboard
The window shows live video, audio, and connection status. The sidebar has these controls:
padd: Start and Stop.
Console: Home, close the running game/app, Release all input (⌘.), and Force release agent… to take the console from an agent that claimed it and will not let go (the next agent in the queue gets it).
Launch: pick an installed game/app and launch it (⌘L). The list comes from PS5 Web File Manager and is cached in
titles/; ↻ reloads it.Capture: save a snapshot as PNG (⌘S), and record the picture and sound to an MP4 (⇧⌘R). Stop asks where to save the video; quitting while recording saves it to
~/Movies.Recordings: record your button presses (⌘R), then pick a saved recording and play it. These are the same files that agents use with
play_recording.Install: install a
.pkgpackage or add an.elfpayload to Payload Manager (and run it), and uninstall checked games/apps (padd 1.3). For titles ShadowMountPlus manages, uninstalling also deletes their source folder or image through ShadowMountPlus, so its next scan does not install them again.
Holding a key holds its button and gives you priority over agents; releasing it returns control.
Toggle the sidebar with ⌥⌘S, video with ⌥⌘V, and audio with ⌥⌘M. Video and audio preferences
persist. Hiding video keeps capture running; closing the window silences audio but keeps the app and API running.
Quit with ⌘Q or uv run ps5mcp capture stop.
Keys | Pad |
Arrows | D-pad |
Enter / Space | Cross |
Backspace / Esc | Circle |
| Square / Triangle |
Q / E, Z / C | L1 / R1, L2 / R2 |
WASD, IJKL | Left stick, right stick |
Tab, T | Options, touchpad click |
H | Home screen |
Features
MCP tools
Feature | Tools |
Capture and status |
|
Controller input |
|
Console apps |
|
File transfer |
|
Visual automation |
|
Recording |
|
Sharing |
|
Input tools can return a frame with snapshot_after_ms. home suspends the game like the PS button;
use press("circle") to back out of system screens. A game takes input only from the controller that launched it,
so launch games with launch to control them through the MCP.
Sharing the console
Several agents can use the console. An agent calls claim_console(reason) so that other agents cannot press
buttons, run console commands, or install until it calls release_console(). Another agent that calls
claim_console joins a queue. With wait_s, the call returns when the console is free for that agent.
A claim ends when its session closes, or after 10 minutes without a tool call. The sidebar shows the agent that has
the console and the number of agents in the queue. The keyboard and the window's buttons always work. The CLI (ps5mcp pad, ps5mcp padd) is blocked like any agent.
File transfer
push(local_path, remote_path) copies a file or folder from this Mac to the console, and
pull(remote_path, local_path) copies one back. Console paths are absolute. A path names the destination itself;
end it with / to copy into that folder. Folders are copied recursively. A file with the same size and a
destination that is not older is skipped (force=True copies it anyway). A shorter, newer destination resumes
when its last 64 KB match the source; otherwise the file is copied again.
Transfers use FTP through zftpd on port 2120. If zftpd is not running, the tool starts it from Payload Manager's
library. Starting it counts as a console action, so it is refused while another agent has claimed the console.
zftpd cannot be stopped remotely: it runs until the console restarts, and loading it again replaces the running
copy. Without zftpd, push sends files one at a time through Web File Manager, and pull copies single files only.
Native uploads set eboot.bin, .prx and .sprx files to 0755 and read back the remote
execute bits. Other files with source execute bits are also restored when the upload contains native
binaries; ordinary assets are left alone. This check runs even for skipped and resumed files. A rejected
SITE CHMOD or an unreadable/incorrect remote mode fails the upload. Native executable uploads require
zftpd: Web File Manager cannot provide this permission guarantee.
For folders containing eboot.bin and a PS5 sce_sys/param.json with titleId, push then polls
ShadowMount's /api/v1/games until that title is installed, managed, and available from the uploaded
source path (up to 90 seconds). Its API must be reachable from this Mac: allow LAN access in ShadowMount's
settings, and use PS5MCP_SMP_PORT if its port differs from 10101. The default console-local listener
cannot be reached by this check. Failure reports that files may already be uploaded and readiness is
unverified; retrying repairs permissions and checks registration again. Within this MCP server session,
launch blocks titles whose upload is still running or failed its checks. Other sessions and the console's
own launcher do not share this guard. Uploads without a recognized manifest report registration as
unverified. These checks prevent the reported execute-permission failure; they do not prove a title will boot.
The .pkg installer follows a separate Web File Manager upload and install-task path; it does not upload
unpacked title executables through FTP. Package installs have not been validated for this reported issue.
Existing titles are not changed automatically. A read-only permission audit can use FTP MLSD's
unix.mode or Unix LIST for eboot.bin and native modules, then scope any repairs to the identified files.
Variable | Effect |
| Remove the |
| Never start zftpd (for example, during experiments that must not load extra payloads) |
| zftpd's port (default 2120) |
| ShadowMount's API port for native-title registration checks (default 10101) |
Browser viewing
uv run ps5mcp stream serves live MJPEG video at http://127.0.0.1:8090/ (about 10 fps).
Add --bind 0.0.0.0 to view from another machine, or set PS5MCP_STREAM_PORT=8090 to start it with the MCP server.
Recording
Record keyboard input with uv run ps5mcp control --record menu-dance (Ctrl-C to finish), then replay it
with uv run ps5mcp play menu-dance. Agents can use the recording tools above.
You can also record and play from the app window. Recordings and templates live in ~/.local/state/ps5-mcp/.
Other commands
uv run ps5mcp snapshot -o frame.png # capture a frame
uv run ps5mcp pad press right # send a button press
uv run ps5mcp pad home suspend # return to the home screen
uv run ps5mcp pad close # close the running game
uv run ps5mcp pad launch PPSA06654 # launch a title
uv run ps5mcp pad uninstall PPSA06654 # uninstall a title (padd 1.3)
uv run ps5mcp install game.pkg # install a package (or an .elf payload; --run starts it)
uv run ps5mcp push build/title /data/homebrew/FAKE00001 # copy a folder to the console (zftpd)
uv run ps5mcp pull /data/playden/logs ./logs/ # copy a folder from the console (--no-start, --force)
uv run ps5mcp padd status # payload version, uptime, and counters
uv run ps5mcp latency # measure input-to-frame latencyFor development without a capture card, use --video synthetic or --video file:PATH when launching the app,
or set PS5MCP_VIDEO for the server. Close QuickTime/OBS before capturing: only one process can own the card.
The ffmpeg fallback is available with uv run ps5mcp capture start --backend ffmpeg, followed by
uv run ps5mcp view --backend ffmpeg. Run it from a regular terminal with camera and microphone permission.
Hardware runs
These commands contact the console. Quit the app before running soak or killtest.
uv run ps5mcp padd soak --minutes 10
uv run ps5mcp padd killtestRuns are archived in results/ with firmware, loader version, payload hash, logs, and snapshots.
Unfinished runs stay marked incomplete.
Development
Findings · App API · Virtual-pad ABI · Wire protocol · PS button and app control
License
MIT, see LICENSE. The virtual-pad interface values come from FGG-XSense and the launch parameter layout from zftpd; this code is a separate implementation.
100% vibecoded
This project is 100% vibecoded with AI coding agents. It's experimental, may contain bugs, and comes with no guarantees. Use at your own risk.
Available Tools
24 toolsclose_appA
Close the running game or app (unsaved progress is lost). The console returns to the home screen.
| Name | Required | Description | Default |
|---|---|---|---|
| snapshot_after_ms | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the burden and does well by disclosing data loss ('unsaved progress is lost') and the resulting state ('console returns to home screen'). It does not mention permissions, failure modes (e.g., what if no app is running), or whether the close is graceful vs. forced.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short clauses with zero waste. The critical warning (unsaved progress is lost) is placed early, and the resulting state is stated succinctly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple single-parameter tool with no annotations or output schema, the description covers the primary action and its side effect, which is adequate. However, it omits any parameter guidance and does not address failure conditions or return behavior, leaving some gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description does not mention the single parameter 'snapshot_after_ms' at all. Schema description coverage is 0%, so the description should compensate, but it provides no explanation of the parameter's purpose (e.g., delay before taking a snapshot) or its default behavior.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Close the running game or app') with clear scope, and distinguishes itself from siblings like launch, home, and uninstall_apps by being the terminator of the current app rather than a launcher or remover.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The parenthetical '(unsaved progress is lost)' and 'console returns to the home screen' make the context of use clear: terminate the current running app. There is no explicit comparison to alternatives like press/home, but the implication is strong enough that an agent can select this over others.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
holdC
Hold several buttons together for duration_ms, then release them.
| Name | Required | Description | Default |
|---|---|---|---|
| buttons | Yes | ||
| duration_ms | Yes | ||
| snapshot_after_ms | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It does disclose one genuinely useful trait beyond the schema: the buttons are automatically released after duration_ms. It omits blocking/wait semantics, whether it errors if a button is already held, and how it interacts with release_all.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single tight sentence, front-loaded with the operation and ending with the release behavior. It earns its place, though at this length it leaves significant gaps rather than being optimally sized for a 3-parameter tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no annotations, no output schema, 0% schema coverage, and three parameters, the description should do far more. It leaves the agent without valid button values, snapshot_after_ms semantics, or any behavioral caveats needed to invoke the tool reliably.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and there are three parameters. The description touches buttons and duration_ms implicitly, but explains nothing about valid button identifiers (no enum), units beyond the name, or the purpose of snapshot_after_ms, which goes completely unmentioned.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb (hold), the resource (several buttons), and the temporal parameter (duration_ms), making the core operation unambiguous. It does not, however, distinguish itself from the many related siblings such as press, trigger, stick, or release_all, so the agent gets no help disambiguating.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to prefer hold over press, sequence, or stick, nor any note about prerequisites (e.g. whether buttons must be free, whether pressing is equivalent plus wait). The description states what happens but never when to choose it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
homeA
Go to the PS5 home screen (stands in for the PS button, which the pad cannot send).
method: "suspend" (suspend the running game, as the PS button does), "system" or "shellcore". From system screens such as Settings, press circle instead.
| Name | Required | Description | Default |
|---|---|---|---|
| method | No | suspend | |
| snapshot_after_ms | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It usefully discloses that 'suspend' suspends the running game (like the PS button) and that system screens need circle instead, but says nothing about the safety profile, side effects on running apps in the 'system'/'shellcore' modes, or what snapshot_after_ms does.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the purpose in the first sentence and adds only the exception and method semantics. The mid-sentence line break before 'method:' is slightly awkward formatting, but the content is dense and waste-free.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 2-parameter tool with no annotations and no output schema, the definition is only partially complete: it omits snapshot_after_ms entirely and gives no side-effect or safety context. The method semantics and system-screen exception are present, so it is adequate but not fully self-sufficient.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% and 'method' has no enum, so the description is the only source for its values ('suspend', 'system', 'shellcore') and their semantics — real added value. However, snapshot_after_ms is never mentioned anywhere, leaving half the parameters undocumented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific action and resource — navigate to the PS5 home screen — and explains the reason it exists ('stands in for the PS button, which the pad cannot send'). That rationale implicitly separates it from press/launch/close_app, though no sibling is named directly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit when-not guidance: 'From system screens such as Settings, press circle instead.' It also explains what the method values mean. There is no broader statement of when to prefer home over close_app or launch, but the key exception is covered.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
installA
Install a file from this Mac onto the console.
path: a .pkg package (uploaded through Web File Manager to /data/ps5-mcp/pkg, then installed; can take minutes) or an .elf payload (added to Payload Manager's library). run=True also starts the .elf once.
| Name | Required | Description | Default |
|---|---|---|---|
| run | No | ||
| path | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It usefully discloses that .pkg installation 'can take minutes' and that run only affects .elf, but says nothing about permissions, failure modes, overwrite behavior, or whether installation is reversible.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Short, front-loaded with the core action, and the path/run behavior is presented as a compact list with no filler. The wording is slightly telegraphic ('run=True also starts the .elf once'), but every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter mutation tool with no annotations and no output schema, the description covers inputs and latency but omits return/result behavior and error handling. An agent could call it correctly, but the picture is incomplete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, and it does for both parameters: it defines the accepted path formats (.pkg at /data/ps5-mcp/pkg, or an .elf in the Payload Manager library) and the effect of run=True. It does not clarify path syntax beyond these examples, but the added meaning is substantial.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource ('Install a file from this Mac onto the console') and even names the source platform, so an agent can tell it apart from sibling write tools like launch or uninstall_apps. It stops short of explicitly contrasting itself with those siblings, so it is clear but not fully differentiated.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explains the two accepted input forms (.pkg vs .elf) and that run=True starts the .elf, which implies usage. However, it never states when to prefer install over launch/uninstall_apps, nor any prerequisite such as the file already being present on the console.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
launchA
Launch an installed title by id (e.g. PPSA01325). list_apps() shows what is installed.
A game takes input only from the controller that launched it: launch it here to control it, because a game the user started with their own controller ignores this pad (and one launched here ignores theirs).
| Name | Required | Description | Default |
|---|---|---|---|
| title_id | Yes | ||
| snapshot_after_ms | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, and it discloses a genuinely non-obvious behavioral trait (a game only accepts input from the controller that launched it, so launching here is required to control it). It omits error behavior for uninstalled/running titles and how launch blocks or returns.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action and id format, and the second paragraph earns its place by explaining a non-obvious control-routing constraint. Wording is slightly convoluted ('ignores this pad ... ignores theirs') but not wasteful.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema or annotations, so the description should do more: the unexplained snapshot_after_ms and the absence of any failure modes leave gaps. Its coverage of the discovery workflow and input-routing behavior is solid but partial for a 2-parameter act-on-device tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate. It adds a useful id format example for title_id, but snapshot_after_ms — an opaque parameter with a 5000ms default — is never mentioned or explained in either place.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (launch) and resource (an installed title by id), with a concrete id example (PPSA01325). It is clearly distinguishable from siblings like close_app, install, and list_apps, which it explicitly routes the agent to.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear context: call list_apps() to discover installed titles, and explains the condition that selects this tool over the user's own controller (input routing follows the launching controller). It lacks an explicit when-not clause, e.g. what happens if the title is already running.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_appsC
Installed title ids (from /user/app on the console).
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description must carry full behavioral disclosure. It implies a safe read ('installed title ids') but omits return format, ordering, pagination, and whether it requires an active console connection.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single short phrase that is front-loaded and has no filler. It is arguably too terse rather than verbose, which is a content gap, not a structure problem.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter, no-annotation, no-output-schema tool, the description should at minimum say what form the result takes and any connection assumptions. It leaves an agent unable to predict the return shape or when the call is valid.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so per the rubric this is a baseline 4. The description adds useful context on what is returned (installed title ids) and their source.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States it lists installed title ids, a specific verb+resource, but the parenthetical '/user/app on the console' is opaque and it doesn't distinguish from siblings like status or list_templates beyond the resource name.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance, no mention of alternatives such as status, and no context on when an agent would need app listings versus snapshot or launch.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_recordingsC
Saved recordings with their length.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It implies a read-only listing of persisted recordings, but does not state side effects, permissions, ordering, or return shape beyond 'their length'.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single short fragment, so it is concise, but it is under-specified rather than optimally structured. It lacks a clear verb and does not front-load the action before the resource.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter list tool with no output schema, the description gives the core resource and one returned field (length). However, it omits whether all saved recordings are returned, their ordering, and any other relevant context, leaving gaps for an agent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so there are no parameter semantics to document. The baseline score for zero parameters is 4.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names the resource ('Saved recordings') and adds that length is included, but it uses a fragment with no verb and does not distinguish the tool from siblings like play_recording or record_stop. An agent can infer it lists recordings, but the purpose is not stated with a specific action.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance, prerequisites, or alternatives are provided. The description does not say whether this lists all recordings, only saved ones, or how it relates to record_start, record_stop, or play_recording.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_templatesC
Saved template names.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden, but it only states what the tool returns in vague terms. It does not disclose whether this is a read-only operation, whether results are paginated, or how many names are returned. For a listing tool with zero annotation coverage, this is a significant gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is extremely short and contains no filler, so it is concise. However, three words is arguably under-specified rather than optimally concise for a tool definition, and there is no structure beyond a fragment.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is simple with no parameters and no output schema, so the description need not document arguments. Still, without an output schema it should say more about the return shape, such as whether it lists all templates or only names, and whether there are pagination or scope caveats.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so the baseline score is 4. The description does not need to explain parameter semantics because there are none.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description is just a noun phrase, 'Saved template names,' with no verb. It largely restates the tool name rather than describing the action. It does not distinguish this tool clearly from the sibling save_template beyond the obvious name contrast.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this tool versus alternatives. It does not mention save_template, play_recording, or any other sibling. The agent receives no context for selecting it over related tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
play_recordingC
Replay a saved recording with its original timing (up to 5 minutes).
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | ||
| snapshot_after_ms | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full behavioral burden. It discloses original timing and a 5-minute cap, but says nothing about side effects, blocking behavior, whether it replays input events, error handling, or what snapshot_after_ms does.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single front-loaded sentence with no wasted words. It is appropriately concise, though it could benefit from a small amount of structured detail for such a behavior-heavy tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is likely complex and has no annotations, no output schema, and 0% schema description coverage. The description covers only replay timing, leaving key details about parameters, prerequisites, side effects, and return behavior missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must explain both parameters. It indirectly implies that 'name' refers to a saved recording, but snapshot_after_ms is never mentioned and its purpose remains opaque.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb and resource: 'Replay a saved recording.' It also adds a timing constraint, but it does not distinguish this tool from related siblings such as list_recordings, record_start, or snapshot.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description says what the tool does but gives no explicit guidance on when to use it versus alternatives. There is no mention of prerequisites, when not to use it, or how it relates to recording or template tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
pressB
Press and release one button (e.g. cross, circle, up, options). Optionally return a frame afterwards.
| Name | Required | Description | Default |
|---|---|---|---|
| button | Yes | ||
| hold_ms | No | ||
| snapshot_after_ms | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description must carry the behavioral burden. It discloses the core behavior of pressing and releasing, and that a frame may optionally be returned after the action. However, it does not address timing precision, key event semantics, side effects, permissions, or success behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single tight sentence, front-loaded with the core operation and examples. Every phrase earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Without annotations or an output schema, the description should compensate more. It does not clarify parameter semantics (hold_ms, snapshot_after_ms), return behavior, or when to prefer this over hold or sequence, leaving significant gaps for an input tool with three parameters.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% and none of the three parameters are described in the schema. The description does not explain hold_ms (the press duration) or snapshot_after_ms (delay before optional frame capture), leaving most parameter meaning undocumented. Only the button parameter is touched via examples.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (press and release) and resource (one button), with examples of valid button values. It is distinguishable from sibling tool hold (press-hold), release_all (release only), and sequence, but does not explicitly name these distinctions.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Implies use for momentary button presses by the name and description, but gives no explicit when/when-not guidance or direction to alternatives like hold, release_all, or sequence.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
record_startA
Start recording every input (agent tools and the keyboard in the viewer) as timed steps.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It usefully discloses the capture scope (every input: agent tools and keyboard in viewer) and output format (timed steps), but does not state persistence, how recording is stopped, or any side effects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, front-loaded sentence with no redundant or filler text. It efficiently communicates the core action and scope.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter start action with no output schema, the description covers what is recorded, but it omits lifecycle context such as how to stop recording or where recordings are stored. Given the presence of sibling tools like record_stop and list_recordings, a bit more routing context would improve completeness.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters and schema coverage is 100%, so there are no parameter semantics to document. The description does not need to compensate for any schema gaps; baseline 4 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Start') and resource ('recording'), and specifies what is captured (agent tools and the keyboard in the viewer) as timed steps. It is clearly distinct from record_stop by verb, though it does not explicitly name the sibling alternative.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by the imperative 'Start recording', but the description provides no explicit when-to-use, prerequisites, or guidance on when not to record or how this relates to record_stop.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
record_stopB
Stop recording and save it under name (letters, digits, - and _).
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It discloses the save action and name character constraints, but does not explain permissions, whether an existing recording is overwritten, what happens if no recording is active, or the return behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single sentence with no wasted words, front-loading the action and the save parameter. It is appropriately sized for a simple one-parameter tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple stop-and-save command the description is minimally adequate, but it omits important context such as the need for an active recording and the behavior when the name already exists. The lack of output schema and annotations means these gaps matter.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It usefully specifies that `name` accepts letters, digits, hyphens, and underscores, adding meaningful format constraints beyond the bare string type in the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: stop recording and save under a given name. It inherently distinguishes itself from record_start and play_recording by action, but it does not explicitly reference any sibling tool or scope condition.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this tool versus alternatives such as record_start, play_recording, or list_recordings. It also omits the prerequisite that a recording must already be in progress.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
release_allA
Release every button and centre the sticks (agent and keyboard).
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description discloses scope ('every button and centre the sticks', agent and keyboard), which is real behavioral content. With no annotations, it carries the full burden but doesn't state whether this is safe/idempotent, whether it affects active recordings or sequences, or what happens to in-flight holds.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single short sentence, fully front-loaded, with no wasted words. Optimal for a zero-argument action.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-param, no-output action with no annotations, the description adequately conveys what is released and reset. Minor gap: no mention of interaction with recordings/sequences or idempotency, which would help in this gesture-control family.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Zero parameters, so the baseline of 4 applies. The description correctly implies the tool takes no inputs by describing a blanket operation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb-and-resource action ('Release every button and centre the sticks') with clear effect. Siblings like press, hold, stick, trigger are distinctly action-specific, and release_all reads as the inverse/reset operation among them, though it doesn't explicitly differentiate itself by name.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No indication of when to use this versus alternatives such as individually releasing via press/hold inverses or stopping a sequence. An agent must infer that this is a mass reset after press/hold/stick operations.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
save_templateB
Save a region of the current frame (full 1920x1080 coordinates) as a template for wait_for.
| Name | Required | Description | Default |
|---|---|---|---|
| x | Yes | ||
| y | Yes | ||
| name | Yes | ||
| width | Yes | ||
| height | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden of behavioral disclosure. It states that a region of the current frame is saved and clarifies the coordinate space, but omits critical details such as whether existing templates are overwritten, what permissions are needed, and what the tool returns or errors on.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, well-structured sentence with no wasted words. The core action and coordinate context are front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given five required parameters, no annotations, no output schema, and zero schema description coverage, the description is too thin. It does not explain what the tool returns, how errors are handled, or any behavioral expectations, leaving significant gaps for an agent to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It clarifies that x, y, width, and height use full 1920x1080 coordinates, and implies name is the template identifier, but it does not define coordinate origin, whether width/height are inclusive, or any constraints on name.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Save') and resource ('region of the current frame ... as a template for wait_for'), making the tool's function immediately clear. It does not explicitly differentiate itself from siblings like list_templates, so it stops short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'as a template for wait_for' implies the tool is used when preparing a template for later matching, but there is no explicit guidance on when to use it versus alternatives or any prerequisites. Usage is only implied, not stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
sequenceA
Run timed states in order, then release everything.
Each step: {"buttons": [...], "lx"/"ly"/"rx"/"ry": -1..1, "l2"/"r2": 0..1, "duration_ms": N}. A step with no inputs is a pause. Total duration is capped at 60 s.
| Name | Required | Description | Default |
|---|---|---|---|
| steps | Yes | ||
| snapshot_after_ms | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden. It discloses the total duration cap of 60 seconds, that a step with no inputs acts as a pause, and that everything is released at the end. This is strong behavioral context, though it doesn't mention error handling or whether steps are atomic.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Extremely concise and well-structured: an opening sentence, a precise step schema definition, and two critical behavioral notes in short sentences. No wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with no annotations, 2 parameters (one undocumented), and no output schema, the description provides most of what's needed to invoke it correctly: step format, duration cap, and release semantics. The omission of 'snapshot_after_ms' is a gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must explain parameters. It provides detailed step field semantics (buttons array, axis ranges -1..1, trigger ranges 0..1, duration_ms) and notes that no inputs means a pause. However, the second parameter 'snapshot_after_ms' is not mentioned at all.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states it runs timed states in order and releases everything, distinguishing it from single-input siblings like press or hold. It also defines the step structure explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for multi-step timed input sequences but does not explicitly state when to use this over chaining individual tools or how it compares to the 'hold' or 'play_recording' siblings. Some guidance implied by the step format.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
show_viewerA
Open the live video window for a human (keyboard in that window drives the console too).
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It usefully discloses a non-obvious trait: keyboard input in the viewer window also drives the console. However, it does not state whether the viewer blocks, requires a running app, how it interacts with other tools, or what happens if already open, leaving notable gaps.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence with no wasted words. It immediately states the action and then adds one important behavioral note in parentheses.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple 0-parameter tool with no output schema, the description is mostly complete. It covers the core action and a key interaction detail, though it could say more about lifecycle or side effects, such as whether the viewer must be closed or how it affects other sessions.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so there are no parameter semantics to document. Per the scoring rules, a 0-parameter tool receives a baseline of 4.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'Open the live video window for a human.' It clearly distinguishes this from sibling tools like snapshot or record_start, which capture or record rather than open a live viewer. It does not explicitly name a sibling alternative, so it falls short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the tool is used to view live video, but gives no explicit when-to-use guidance, no prerequisites, and no alternatives. An agent must infer that this is the right tool when a human needs to see the device screen.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
snapshotB
Current PS5 screen as a JPEG (scaled to max_width; the source is 1920x1080).
| Name | Required | Description | Default |
|---|---|---|---|
| max_width | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the disclosure burden, and it does add useful behavioral context: the output format (JPEG), the scaling behavior, and the source resolution (1920x1080). However, it says nothing about latency, whether the screen must be active, or what happens on failure, which are meaningful gaps for a no-annotation tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence that states the resource, format, and scaling in order of importance. No filler, though it is quite terse for a tool with no annotations.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple one-parameter screenshot tool this covers the essentials (format, dimensions, scaling), and no output schema is needed since the return is a JPEG image. It stops short of operational context like preconditions or failure behavior, leaving moderate gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the schema only shows 'max_width' as an integer with default 1280. The description compensates by explaining that the image is scaled to max_width and that the source is 1920x1080, which adds real meaning, but it omits whether aspect ratio is preserved or units.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific resource (the current PS5 screen) and output format (JPEG), which clearly distinguishes it from control-oriented siblings like press, hold, and stick. It lacks an explicit verb ('capture'/'take') but the noun phrase is unambiguous about what the tool returns.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to call this versus alternatives such as show_viewer, nor any mention of prerequisites like needing an active screen. Usage must be entirely inferred from the name and output description.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
statusC
Payload link, payload version and counters, capture state and frame age.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden, yet it says nothing about whether this is a read-only query, whether it blocks, or freshness guarantees. It only enumerates returned fields without behavioral context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single compact sentence, but it is a fragment with no front-loaded purpose. The structure prioritizes enumerating output fields over stating the tool's function.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
This is a no-parameter status query with no output schema and no annotations. The agent has nothing to confirm when to invoke it or what to expect behaviorally, only a partial list of returned data.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so per the rubric the baseline is 4. No parameter information is needed and none is missing.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description lists what the tool returns (payload link, version, counters, capture state, frame age) but never states what action it performs or that it retrieves current status. It reads as a field list rather than a clear verb+resource statement.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No indication of when to call this versus siblings like snapshot or wait_for_change. The sibling set contains many status-adjacent tools and the description offers no routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
stickC
Deflect the left or right stick. x, y in -1..1 (left/up negative), then return to centre.
| Name | Required | Description | Default |
|---|---|---|---|
| x | Yes | ||
| y | Yes | ||
| side | Yes | ||
| duration_ms | Yes | ||
| snapshot_after_ms | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries full behavioral burden. It discloses x/y range and the return-to-centre behavior, but omits critical details such as side values, duration_ms semantics, snapshot_after_ms, and whether the deflection is held or immediately released.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, front-loaded sentence with zero wasted words. It efficiently combines purpose, coordinate semantics, and release behavior.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with five parameters (four required), no annotations, no output schema, and 0% schema description coverage, the description is far too sparse. It covers x/y but leaves essential required parameters and overall call behavior unexplained.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for five parameters. The description adds useful meaning for x and y (range -1..1, left/up negative), but completely omits side, duration_ms, and snapshot_after_ms, leaving required parameters undocumented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Deflect') and resource ('stick'), and clarifies left/right and return-to-centre behavior. It is clear, but does not explicitly differentiate from sibling tools like press, hold, or touchpad, which prevents a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no when-to-use guidance, no conditions for choosing this tool over alternatives like touchpad or hold, and no prerequisites. Usage is only implied by the action described.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
touchpadC
Click the touchpad. Coordinates are sent but not yet delivered by padd (touch mapping is unknown).
| Name | Required | Description | Default |
|---|---|---|---|
| x | No | ||
| y | No | ||
| duration_ms | No | ||
| snapshot_after_ms | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It usefully discloses a real behavioral limitation — coordinates are sent but not delivered because touch mapping is unknown — which an agent must know. It does not disclose whether the click has destructive side effects, what the returned result is, or how duration_ms and snapshot_after_ms affect behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences with the action front-loaded and the important caveat second. No wasted words, though the second sentence is cryptic and could be phrased more clearly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No annotations, no output schema, and four undocumented parameters. The description covers the action and one caveat but leaves parameter meaning and behavioral consequences almost entirely to inference, which is insufficient for a mutation-style interaction tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across 4 parameters. The description only gestures at the x/y coordinates ('Coordinates are sent') without format, range, or meaning, and says nothing about duration_ms or snapshot_after_ms, leaving two parameters entirely unexplained beyond their defaults.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Click the touchpad'), so the basic action is unambiguous. However, it never distinguishes itself from sibling tools like press, hold, stick, or trigger, which likely operate on the same touch surface, leaving the agent to guess which to use.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when this tool should be used versus press/hold/trigger, nor any stated preconditions. The caveat that coordinates are 'not yet delivered' hints the tool may be unreliable, but this is not framed as an explicit usage rule.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
triggerC
Pull L2 or R2 (side: left/right) to value 0..1, then release.
| Name | Required | Description | Default |
|---|---|---|---|
| side | Yes | ||
| value | Yes | ||
| duration_ms | Yes | ||
| snapshot_after_ms | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, and it does disclose the key behavioral trait: the trigger is pulled and then automatically released, implying a momentary action rather than a latched one. However it says nothing about timing/units for duration_ms, whether the call blocks, whether a snapshot is taken, or any rate/permission constraints.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler; the action and its argument summary are stated immediately. It is efficient, though arguably under-specified rather than truly concise.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 4-parameter tool with no annotations and no output schema, the description omits usage guidance relative to siblings, units for duration_ms, the meaning of snapshot_after_ms, and any return behavior. Required parameters remain unelaborated, so an agent would be guessing on invocation details.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across 4 parameters. The description only hints at 'side' (left/right) and 'value' (0..1 range), leaving 'duration_ms' — a required parameter — and 'snapshot_after_ms' entirely unexplained in both the schema and the description. Partial compensation at best for a required, undocumented parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb+resource pair ('Pull L2 or R2') and states the full action lifecycle (pull to a value, then release), so an agent can tell it apart from siblings like press, hold, or release_all. It stops short of explicitly contrasting itself with those siblings, but the mapped button and analog-value framing makes the intent unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to choose this over press, hold, stick, or release_all, and no stated prerequisites. The auto-release behavior is implied but the agent must infer that release_all is not needed afterward. No exclusion conditions are stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
uninstall_appsA
DESTRUCTIVE: uninstall installed titles by id (e.g. ["PPSA01325"]); their data is deleted.
Only uninstall what the user asked for by name. For titles ShadowMountPlus manages, their source folder or image is deleted too (otherwise it installs them again). The console refuses while a game is running; close_app first. Removal finishes in the background after each request is accepted. Needs padd 1.3 or later.
| Name | Required | Description | Default |
|---|---|---|---|
| title_ids | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden and does so richly: it flags DESTRUCTIVE, states that title data is deleted, that ShadowMountPlus-managed source folders/images are also removed, that the console rejects calls while a game runs, that removal completes asynchronously in the background, and that padd 1.3+ is required. Nothing behaviorally important is left implicit.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the DESTRUCTIVE warning, then ordered by operational importance (scope, cascade deletion, runtime precondition, async behavior, version requirement). Dense but every sentence earns its place with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive, no-annotation tool with no output schema and 0% param coverage, the description supplies safety, cascade effects, prerequisite sequencing, async semantics, and a compatibility floor. An agent has everything needed to invoke it correctly and safely.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% for the single title_ids parameter, so the description must compensate. It does so by giving a literal example of the id array format (["PPSA01325"]), which is the key detail needed to call it, though it doesn't explain where ids come from or whether multiple ids are allowed.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (uninstall) and resource (installed titles) and gives a concrete id example ("PPSA01325"), so the agent knows exactly what the tool does. It does not explicitly distinguish itself from the inverse sibling 'install', so it falls just short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
"Only uninstall what the user asked for by name" sets scope, and "The console refuses while a game is running; close_app first" names the sibling to call first as a prerequisite. Clear usage context, though no explicit when-not or alternative-selection reasoning.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
wait_forB
Wait until a saved template appears on screen (or disappears, with gone=true).
Matching is normalised cross-correlation over the whole frame; score 1.0 is identical. Returns where it was found, the score and the frame.
| Name | Required | Description | Default |
|---|---|---|---|
| gone | No | ||
| template | Yes | ||
| min_score | No | ||
| timeout_ms | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It usefully discloses the matching method (normalised cross-correlation over the whole frame, score 1.0 = identical) and the return contents, which is genuine behavioural context. It omits blocking semantics, what happens on timeout, and permission requirements.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two compact sentences, front-loaded with the primary action and the mode switch, followed by the matching detail. No filler, though the parenthetical mode note could be folded more cleanly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description correctly covers return values (location, score, frame). But with 0% schema coverage and no annotations, the timeout and min_score parameters remain undocumented, leaving the agent short of what it needs to call the tool well.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It explains template (via 'saved template') and gone=true, and clarifies the score scale relevant to min_score, but leaves min_score's threshold meaning and timeout_ms entirely unexplained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: waits until a saved template appears on screen, with a clear inverse mode via gone=true. It does not differentiate itself from the sibling wait_for_change, leaving the agent to infer the distinction, so it falls short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use it (polling for a template's presence/absence) and notes the gone=true alternative mode. However, it never says when to prefer this over wait_for_change or snapshot, and gives no guidance on timeout behaviour or failure conditions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
wait_for_changeB
Wait until the screen changes (mean luma difference >= threshold on a 32x18 grid); returns the new frame.
| Name | Required | Description | Default |
|---|---|---|---|
| threshold | No | ||
| timeout_ms | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It helpfully discloses the change-detection algorithm and that it 'returns the new frame', but is silent on the most important behavioral question: what happens on timeout (default 3000ms) — does it raise, return, or return the stale frame?
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, front-loaded sentence with zero filler. It packs the purpose, mechanism, and return value efficiently, though the parenthetical grid detail could arguably be trimmed.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Covers purpose, detection mechanism and return value for a 2-param tool with no output schema. However, missing timeout semantics and relying on defaults that aren't explained leaves meaningful gaps for an agent invoking it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It partially explains 'threshold' (luma difference on a 32x18 grid), but 'timeout_ms' is never described or referenced, leaving half the parameters undocumented in both schema and description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Wait until') and resource ('the screen changes') and even exposes the detection mechanism (mean luma difference >= threshold on a 32x18 grid). This clearly separates it from the fixed-duration sibling 'wait_for', though it does not name that sibling directly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by the semantics of waiting for a screen change, but there is no explicit when-to-use vs when-to-use-alternative guidance and the closely related 'wait_for' sibling is never mentioned or differentiated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
24 tool updates
v0.1.0- First observed
close_app - First observed
hold - First observed
home - First observed
install - First observed
launch - First observed
list_apps - First observed
list_recordings - First observed
list_templates - First observed
play_recording - First observed
press - First observed
record_start - First observed
record_stop - First observed
release_all - First observed
save_template - First observed
sequence - First observed
show_viewer - First observed
snapshot - First observed
status - First observed
stick - First observed
touchpad - First observed
trigger - First observed
uninstall_apps - First observed
wait_for - First observed
wait_for_change
TDQS
Scored across 24 tools
Tools split cleanly into input emulation, recording, vision/template matching, and app management, with distinct purposes. Minor overlap between press/hold/sequence and between wait_for/wait_for_change, but the descriptions clearly differentiate each.
Almost entirely lowercase snake_case with predictable verb_noun or noun forms (record_start, list_recordings, save_template, wait_for_change, uninstall_apps). A few bare single-word names (snapshot, press, touchpad, home, status) deviate slightly but remain readable and consistent in style.
At 24 tools it sits at the heavy end, but the domain genuinely spans capture, input, recording, vision, and app lifecycle, so each tool earns its place. Slightly over-scoped but reasonable for the feature breadth.
Covers the full lifecycle: capture/view, input, record/replay, template-based waits, app install/launch/uninstall, and status. Minor gaps like no delete for templates or recordings, but agents can work around these.
Maintenance
Related MCP Connectors
Remote MCP server for AI.TV creators — delegate account operations to your AI agent over MCP.
A real Android phone in the cloud for MCP agents: observe the screen, act, reset to a saved state.
Drive real devices from your AI Coding tool. Embed a client SDK (Unity, Godot, Flutter, iOS/macOS, Android, React Native, Web) in your app, then capture screenshots, traverse the UI tree, inject taps and key events, and run automated test tasks on the physical device over a secure relay.
OCR, transcription, file extraction, and image generation for AI agents via MCP.
Related MCP Servers
- AlicenseNot gradedqualityCmaintenanceEnables AI coding agents to work inside a live Unreal Editor for Fortnite session: write and compile Verse, place and wire Creative devices, edit Scene Graph entities, manipulate actors and assets, take screenshots, run playtests, and read editor logs via MCP.MIT
- AlicenseNot gradedqualityAmaintenanceEnables AI agents to operate and view the Vega Virtual Device through MCP: press remote keys, take screenshots, record video with sound, wait for screen changes, and check the TV safe area.680 npmMIT
- AlicenseAqualityBmaintenanceEnables MCP clients to capture screenshots and control a Raspberry Pi's Wayland desktop over an existing SSH connection with mouse, drag, scroll, keyboard, and Unicode typing actions. Input tools return fresh screenshots, enforce single-client sessions with idle release, and reject stale or uncertain state until verified.11MIT
- AlicenseNot gradedqualityAmaintenanceLets AI agents build, run and live-test FiveM servers by sending server and client console commands over RCON and devcon, driving the game window with keyboard, mouse and screenshot automation, and invoking in-game natives, exports and NUI callbacks through a companion bridge resource.79 npm15MIT