Skip to main content
Glama
AstralVoidZ
by AstralVoidZ

ppsspp-dfx-mcp

English | 中文

PyPI CI Python License: MIT

一个把 PPSSPP 变成 AI 可调试目标的 MCP(Model Context Protocol)服务器。它把 PSP 模拟器的 WebSocket 调试器封装为面向 LLM agent 的工具面: 会话生命周期、内存读写、反汇编、断点、CPU 控制、输入自动化、截图、回放录制与诊断 脚本——并内建结构化契约、防御性错误分类法与任务级评估。

文档:docs/SCOPE.md(范围与协议面边界)· CHANGELOG.md(变更记录)

项目状态

项目处于 alpha 阶段并快速迭代,工具面与配置格式可能出现不兼容变更。 运行前请阅读 SECURITY.md。

Related MCP server: GDB MCP Server

功能特性

  • 37 个静态工具,全部带结构化 inputSchema / outputSchema——没有无约束的 返回值,每个参数都有类型和说明。

  • 动态脚本工具:项目专属的诊断脚本通过 scripts.manifest.yaml 暴露为 ppsspp_script_<name> 工具,输入类型由脚本自带的 Pydantic model 决定; ppsspp_run_script 调用未暴露的脚本,ppsspp_list_scripts 查看清单—— 详见配置。

  • 会话模型:支持多个并发 PPSSPP 会话、就绪探测(wait_ready)与楔死自愈 (resilient 启动)。

  • 面向 agent 的人体工学:组合工具(ppsspp_frame_snapshot、 ppsspp_breakpoint(action="wait"/"trace"/"stats")、ppsspp_batch_step)、 session_id 自动解析、 防御性错误码([CODE] message 格式、CPU 冻结与连接断开的区分),错误文本内嵌 恢复建议。

  • 后台自动化:批量任务跑在独立的服务端任务上,不受 MCP 客户端工具调用超时的 影响;支持状态轮询、取消与注册表盘点(ppsspp_batch_status(batch_id 省略))。

  • 内建评估体系(evals/):49 张场景卡 + 确定性门禁 + 对录制夹具的盲测 runner + 汇总报告——工具面按 agent 实际使用的方式被测试。

  • 诚实的协议面:能力只在其背后存在可用实现时才声明;刻意置 false 的开关 附有设计理由说明。

运行

环境要求:Python 3.13+(配合独立 venv,原因见从源码运行); 带 WebSocket 调试器的 PPSSPP——官方发行版即可,开启方法见 docs/ppsspp-build.md(服务器负责启动它,并连接 ws://<host>:<port>/debugger);一个 MCP 客户端(ZCode、Claude Desktop、 MCP Inspector 等)。

从 PyPI 安装运行

版本状态:PyPI 上是 alpha 阶段的发布快照,可能落后于仓库 main。以你实际 装到的 wheel 版本为准(pip show ppsspp-dfx-mcp),变更记录见 CHANGELOG.md。

使用独立 venv——本服务器的 MCP SDK v2 无法与其他 MCP 服务器锁定的 1.x mcp 包共存:

# Windows:
py -3.14 -m venv .venv
# POSIX:
python3.14 -m venv .venv
.venv\Scripts\python -m pip install ppsspp-dfx-mcp     # Windows
.venv/bin/python -m pip install ppsspp-dfx-mcp         # POSIX

把服务器注册到你的 MCP 客户端(入口由安装包提供,无需指向仓库内脚本):

{
  "mcpServers": {
    "ppsspp-dfx": {
      "command": "C:/absolute/path/to/.venv/Scripts/ppsspp-dfx-mcp.exe",
      "cwd": "C:/absolute/path/to/your-project"
    }
  }
}

command 指向安装 venv 内的入口可执行文件,POSIX 上为 .venv/bin/ppsspp-dfx-mcp;cwd 是服务器发现 .ppsspp-dfx/ 配置的目录 (见配置)。

cwd 可以省略(部分 harness 不支持该字段):把项目根同时写进 env.PPSSPP_DFX_PROJECT_ROOT 即可,两者等价。实测从 %TEMP% 启动、不设 cwd、仅给该环境变量,服务器仍能正确定位 project_root / config_dir / output_dir 并完成握手;反之若两者都缺,会退化为告警并丢失 scripts.manifest.yaml(ppsspp_script_* 工具静默消失)。

手动启动验证:

.venv/Scripts/ppsspp-dfx-mcp.exe   # Windows
.venv/bin/ppsspp-dfx-mcp           # POSIX

然后直接给 agent 派任务:"启动模拟器加载这个 ISO,告诉我当前 PC"——服务器 负责会话启动、就绪探测与状态读取。工具描述遵循 PURPOSE / USAGE / BEHAVIOR / RETURNS 约定,错误路径内嵌恢复指引,agent 无需示例即可自助。

不要在客户端与服务器之间插入包装脚本:Windows 上 os.execv 是 CreateProcess + 父进程等待(不是 POSIX 进程替换),多一层会让最内层服务器 立即读到 stdin EOF 并静默退出——表面现象只是 -32000: Connection closed。

从源码运行

仓库检出内自带引导脚本与可直接使用的 .mcp.json:

git clone https://github.com/AstralVoidZ/ppsspp-dfx-mcp.git
cd ppsspp-dfx-mcp

# 在本目录执行——创建 .venv/ppsspp-dfx-mcp 并安装(editable,含 dev 依赖):
python scripts/check_env.py --bootstrap

# 校验解释器 / SDK 版本 / 包导入 / 每份 .mcp.json:
python scripts/check_env.py --check

--bootstrap 只准备主 venv(.venv/ppsspp-dfx-mcp,跑测试与全量回归用)。 服务器启动不再需要它:提交的 .mcp.json 走通用运行器形式,环境由运行器首次运行时自建。

.mcp.json 的两种合法形式(--check 会逐份校验):

配置所在位置

形式

说明

包根(独立检出的仓库根)

["run", "ppsspp-dfx-mcp"]

自定位:运行器按工作目录找到项目

非包根(monorepo 的工作区根)

["run", "--directory", "<包目录>", "ppsspp-dfx-mcp"]

必须钉住包目录,否则会落到工作区自己的项目上

工作目录假设:MCP 客户端通常以 .mcp.json 所在目录为工作目录启动子进程; 相对定位依赖这一点。config.project_root() 的解析规则是 PPSSPP_DFX_PROJECT_ROOT 环境变量 > 工作目录,且不做向上搜索——工作目录一旦落到别处, 输出目录与脚本清单都会错位(表现为动态脚本工具静默消失)。

不要在提交的配置里写平台绑定的解释器路径(…/Scripts/python.exe 或 …/bin/python): 它在另一个操作系统上必然失效。

把服务器注册到你的 MCP 客户端——让客户端读取仓库里的 .mcp.json,或按同样的 结构内联(以包根那份为例):

{
  "mcpServers": {
    "ppsspp-dfx": {
      "command": "uv",
      "args": ["run", "ppsspp-dfx-mcp"]
    }
  }
}

若客户端把 command 解析到别的工作目录(不遵守上述假设),改用 python scripts/check_env.py --print-config 打印的片段——它给出同形的运行器形式 并用 --directory 钉住包目录的绝对路径,可直接粘贴。 另请确保通用运行器在 PATH 上(--check 会检查;缺失时 stdio 传输下客户端只看得见 子进程退出,看不到原因)。

python scripts/check_env.py --print-config

手动启动验证:

.venv/ppsspp-dfx-mcp/Scripts/python -m ppsspp_dfx_mcp   # Windows
.venv/ppsspp-dfx-mcp/bin/python -m ppsspp_dfx_mcp        # POSIX

配置

环境变量(全部可选):

变量

默认值

说明

PPSSPP_DFX_LOG_LEVEL

INFO

日志级别

PPSSPP_DFX_LOG_FORMAT

text

日志格式(text 或 json)

PPSSPP_DFX_RATE_LIMIT

60

单工具限流(次/分钟,0 为关闭)

PPSSPP_DFX_WS_HOST

127.0.0.1

PPSSPP WebSocket 主机

PPSSPP_DFX_WS_PORT

12345

PPSSPP WebSocket 端口

PPSSPP_DFX_EXE_PATH

(来自 yaml)

PPSSPP 可执行文件路径

PPSSPP_DFX_SESSIONS_PATH

~/.ppsspp-dfx/sessions.json

会话状态路径

PPSSPP_DFX_PROJECT_ROOT

cwd

项目根目录,覆盖 cwd 发现(MCP host 从临时目录启动时需设置)。路径不存在时报 CONFIG_INVALID;路径有效但无 .ppsspp-dfx/ 标记目录时告警不阻断

PPSSPP_DFX_CONFIG_DIR

<project_root>/.ppsspp-dfx/config

配置目录路径,覆盖默认发现逻辑

PPSSPP_DFX_IR

未设

设为 1 强制 CPUCore=2 解释器模式。内存断点必需——JIT fastmem 直写下断点不触发;代价是速度,只用于 trace/breakpoint 会话

PPSSPP_DFX_BOOT_HEAL_QUARANTINE

1

1 = 允许 boot 楔死自愈隔离 GPU 后端黑名单文件(仅重命名,从不删除);0 = 该步不执行

PPSSPP_DFX_MEMSTICK_DIR

自动探测

memstick 目录(日志/截图捕获用)

PPSSPP_DFX_WORKSPACE_ROOT

自动探测

scripts/_wire.py 的工作区根(取最近的含 .mcp.json 的祖先目录)。仅引导脚本使用

PPSSPP_DFX_ALLOW_REMOTE_DEBUGGER

未设

设为 1 时接受绑定在非回环地址上的无鉴权调试器(默认 fail-closed 拒绝)。详见 SECURITY.md

PPSSPP_DFX_ALLOW_ABS_SCRIPT

未设

设为 1 允许脚本 manifest 使用绝对路径(默认拒绝,仅相对路径)。详见 SECURITY.md

PPSSPP_DFX_ISO_ROOT

未设

ISO 路径白名单根——设置后 iso_path 仅接受该树内路径。详见 SECURITY.md

PPSSPP_DFX_WS_PORT 只在连接「已在运行的 PPSSPP」时生效:服务器自己启动 PPSSPP 时会随机选空闲端口并在启动后发现它(避开 12345 冲突),此时该变量被忽略。

限流只覆盖协议分发:PPSSPP_DFX_RATE_LIMIT 由中间件在 JSON-RPC tools/call 分发层执行,按 (session_id, tool_name) 分桶——不同会话互不占用额度;无法归属的 请求(参数结构异常、工具名未知/未注册)统一落入 __unknown__ 桶并被限流 (fail-closed)。进程内调用(如测试里的 mcp.call_tool(...)、工具直接调用另一工具) 不经过分发层,因此不受限流约束——它不是进程级全局限流器。

评测专用变量(跑 evals/ 时才需要,日常使用无需设置): PPSSPP_DFX_TEST_MODE、PPSSPP_DFX_FIXTURE_DIR、PPSSPP_DFX_TEST_EXE_PATH、 PPSSPP_DFX_TEST_ISO_PATH、PPSSPP_DFX_TEST_PPSSPP_LOG、PPSSPP_DFX_SKILL_DIR、 PPSSPP_DFX_EVALS_LLM_API_PATH、PPSSPP_DFX_EVAL_GAME_*。

完整可复制的配置模板(含全部变量注释、按用途分组)见源码检出中的 examples/mcp.json.template(PyPI 安装的用户可从 GitHub 仓库 获取)。

项目级 YAML 配置位于 .ppsspp-dfx/config/(相对工作目录):

  • project.yaml — ppsspp_exe 路径与项目元数据

  • addresses.yaml — 命名地址常量(同时为内存向导的 completions 能力提供候选)

  • scripts.manifest.yaml — 诊断脚本清单。每个条目带机器可读的 status (migrated = 可运行,skeleton = 方法体返回 not_implemented)。标记 exposed: true 的脚本在启动时注册为 ppsspp_script_<name> 工具——skeleton 除外,preflight 会拒绝它们。ppsspp_reload_scripts 将动态工具注册表与清单 重新同步(无需重启),并报告声明与注册的对账结果。

独立部署快速开始

三份配置文件(project.yaml / addresses.yaml / scripts.manifest.yaml)有 开箱模板——从这里开始,不要从零手写 YAML。模板的取法取决于安装方式:

从源码检出(模板就在仓库里):

mkdir -p .ppsspp-dfx/config
cp examples/project.yaml examples/addresses.yaml \
   examples/scripts.manifest.yaml .ppsspp-dfx/config/

从 PyPI 安装(wheel 只打包 src/ppsspp_dfx_mcp,不含 examples/, 请在 GitHub 上取同一份模板):

mkdir -p .ppsspp-dfx/config
for f in project.yaml addresses.yaml scripts.manifest.yaml; do
  curl -fsSL "https://raw.githubusercontent.com/AstralVoidZ/ppsspp-dfx-mcp/main/examples/$f" \
    -o ".ppsspp-dfx/config/$f"
done

也可在 examples/ 目录里逐个浏览/下载。取到模板后编辑 .ppsspp-dfx/config/project.yaml:把 ppsspp_exe 指向你的带 WebSocket 调试器的 PPSSPP 构建;把 addresses.yaml 里的 PLACEHOLDER 地址替换为你自己逆向得到的值。

首次会话前需要知道的两件事:

  • 没有 scripts.manifest.yaml 服务器仍能启动,但所有 ppsspp_script_* 工具会 静默消失——即使 scripts: 列表为空也请保留模板(check_env.py --check 报的正是这个警告)。

  • 配置为空且无占位值时,服务器侧一切功能可用;只有会话启动需要真实的 ppsspp_exe(或 PPSSPP_DFX_EXE_PATH),地址常量也只有在你提供自己游戏的 数值后才有意义。

协议面

在 initialize 握手时声明——且只声明实际注册的能力(SDK 从请求处理器 是否存在来推导各项能力,所以这里出现的每一项背后都有可用实现):

能力

声明

说明

tools

✅

37 个静态工具 + 动态 ppsspp_script_<name>

resources

✅

ppsspp://game-state、ppsspp://registers(快照)

prompts

✅

memory-breakpoint-wizard、memory-trace-wizard

completions

✅

两个内存向导的 address 参数,候选来自 addresses.yaml

logging

❌

协议修订 2026-07-28 移除了 logging/setLevel

tasks

❌

仅 SDK 2.2.0 的类型定义,无服务器端实现

tools.list_changed 与 resources.subscribe 刻意置 false。SDK 2.2.0 的 MCPServer 没有暴露握手期设置 notification_options 的入口,声明它们等于承诺 一个服务器发不出的通知。现有替代:

  • ppsspp_reload_scripts 会报告变化内容(exposed_added / exposed_removed),agent 无需通知通道即可响应。

  • 服务器 instructions 字符串告诉新 agent 工具面包含什么。

若未来 SDK 开放了该入口,翻转开关并补上 send_*_list_changed 调用即可——L2 契约测试(tests/unit/l2_mcp_contract/test_capabilities_contract.py)断言当前 的 false 状态并会失败,这是设计信号:该决策需要重新审视,而非回归。

返回形态

图像类工具(ppsspp_screenshot、ppsspp_dump)返回拆成两半的 CallToolResult:

  • content — 一个携带像素的 ImageContent 块。

  • structuredContent — 仅元数据(file_path / size_bytes / format,加上 mode、width、height、empty 等各工具自有字段)。图像的 base64 副本 不在这个通道里——那会膨胀 schema,且重复 content 已承载的内容。

每个工具都声明结构化 outputSchema——没有工具返回无约束对象或 items 为空的 数组。ppsspp_run_script 的 input 参数是唯一注册在案的例外:其形状由被调用的 脚本决定,因此只描述而不约束。

structuredContent 的序列化行为以 mcp SDK 2.2.0(mcp.server.mcpserver)实测 为准。多形态工具(同一工具不同 action 返回不同形状,如 ppsspp_breakpoint / ppsspp_diff_memory / ppsspp_scan / ppsspp_session / ppsspp_batch_step / ppsspp_frame_snapshot 等)的输出契约声明为 partial:完整负载始终经 content 文本通道以 JSON 返回,而 structuredContent 对部分形态的填充行为在不同 SDK 版本上可能不同——机器可读消费方请以文本通道 JSON 为兜底。

错误处理

当被模拟的 CPU 冻结(死循环 / HLE 阻塞 / GPU 管线停滞)时,服务器返回 CPU_FREEZE_SUSPECTED 而不是笼统的 WS_DISCONNECTED——区分"PPSSPP 进程还 活着但 CPU 冻结"与"进程已死 / WebSocket 断开"。

对 CPU_FREEZE_SUSPECTED 的建议处理:

  • 不要重启会话——PPSSPP 还在运行。

  • 用 ppsspp_screenshot 截取当前画面辅助诊断。

  • 尝试 step(action='resume')(对真正的死循环可能无效)。

  • 用 hle.thread.list 查看线程状态(可能暴露 HLE 阻塞)。

  • 在当前 PC 处用 ppsspp_disassemble 检查指令流。

相关错误码:WS_DISCONNECTED(PID 已死,真断开)、WS_TIMEOUT(带票据的 RPC 超时,保守默认)、CPU_STATE_ERROR(当前 CPU 状态不适合该操作)。错误 文本始终以 [CODE] 开头,agent 可编程分类;存在下一步的地方都内嵌了恢复建议。

故障排查速查表

症状

原因 / 修复

-32000: Connection closed(无任何信息)

MCP 客户端与服务器之间有包装脚本:Windows 上 os.execv 实为 CreateProcess + 父进程等待(非 POSIX 替换),内层 server 的 stdin 立即 EOF 静默退出。去掉中间层,直接以 venv 解释器为 command(见从源码运行)

客户端启动 server 报 -32000: Connection closed / command 路径不存在

该配置的相对解释器路径未被 provision(其目录旁没有对应 venv),或客户端把相对 command 解析到了另一个工作目录。运行 python scripts/check_env.py --bootstrap 在各配置目录旁建 venv,或改用 python scripts/check_env.py --print-config 输出的绝对路径片段(见从源码运行)

check_env 报「独立 venv 缺失」

.venv/ 被 gitignore 排除,新 clone 必然没有。运行 python scripts/check_env.py --bootstrap(见从源码运行)

mcp SDK 版本不满足 / 导入期崩溃

系统 Python 的 mcp 包常被其他 MCP server 钉在 1.x,与 SDK v2 不可调和。不要全局安装——用 check_env.py --bootstrap 建独立 venv,或按从 PyPI 安装运行安装到独立 venv

ppsspp_script_* 工具全部消失(服务器正常启动)

.ppsspp-dfx/config/scripts.manifest.yaml 缺失——缺失仅告警不阻断,动态工具静默清空。按独立部署快速开始取三份模板修复(check_env.py --check 会提示;PyPI 安装时模板不在 wheel 内,需从 GitHub 取)

[PPSSPP_NOT_FOUND]

PPSSPP 可执行文件未配置。设 PPSSPP_DFX_EXE_PATH,或 .ppsspp-dfx/config/project.yaml 的 ppsspp_exe(优先级 env > yaml)

找不到 .ppsspp-dfx/config

配置目录按 cwd 发现(无父级上溯)。从含 .ppsspp-dfx/ 的目录启动,或设 PPSSPP_DFX_CONFIG_DIR 指向它

cwd 无 .ppsspp-dfx/ 标记目录(告警)

MCP host 从临时目录启动服务器。设 PPSSPP_DFX_PROJECT_ROOT 显式 pin 项目根目录(告警不阻断,向后兼容)

[CONFIG_INVALID] PPSSPP_DFX_PROJECT_ROOT=... does not exist

环境变量指向的路径不存在。这是显式配置错误——修正路径或取消设置该环境变量以回退到 cwd

WebSocket 连接失败 / WS_DISCONNECTED

PPSSPP 未运行、端口不对,或未启用 WebSocket debugger。check_env.py --check 验证环境,ppsspp_session(action='get') 验证会话

工具调用挂起 / 超时(WS_TIMEOUT)

PPSSPP 主循环负责 dispatch WebSocket 请求:UI 卡死、模态对话框弹出或模拟暂停时请求不会被处理。先截图确认 UI 状态

boot 阶段 [BOOT_TIMEOUT]

启动楔死疑似。start(resilient=true) 会隔离 GPU 后端黑名单(仅重命名 FailedGraphicsBackends.txt,不删除)并自愈重启(≤2 次重试)

已知限制

诚实声明协议面的边界——以下各项均已在对应工具的描述中标注,此处汇总:

  • IR 编码无法在 MCP 侧可靠判别:PPSSPP 的 JIT-IR 代码段用 read_u32 读取不会报错, 但可能得到无意义的值(不是真实 MIPS 指令)——要读代码段请改用 ppsspp_disassemble (搜索指令用 ppsspp_search_disasm)。

  • 条件断点由 MCP 侧求值:该构建的 IR 模式忽略寄存器条件(上游缺陷,已实机建档), 因此 breakpoint 的 condition 不下发 PPSSPP,改由 action='wait' 在命中时用 cpu.evaluate 求值——求值器只在 wait 运行期间生效;假命中自动 resume 并计入 filtered_hits,同一地址 ≥10 次命中且间隔 <1s 触发风暴熔断(自动撤防 + storm_break=true)。CPU 若在布防前已处于暂停态,则无法归因(手动暂停与命中不可 区分):返回 hit=true 并附 note 说明注册的条件未被求值。

  • 无存档 API:PPSSPP 的 WebSocket debugger 不暴露 savestate.* 事件, 服务器无法提供存档保存/加载。用 PPSSPP 的 UI 快捷键(F1-F8 存档槽)。

  • 帧推进只有指令级:step 走 cpu.stepInto。整帧推进的替代:在 vblank 处理器设断点后 resume。

  • analog 摇杆是持久共享态:send_analog 写入后保持到下次写入,无自动复位。

  • VRAM 直读截图不可靠:直读 VRAM 与 GPU 渲染输出不同步,颜色可能失真; 默认走 render 通道。source='output' 在部分游戏上有崩溃风险,仅在 render 通道空帧回退时使用。

  • replay 时钟锚定:replay 时间线使用录制会话 boot 时刻的绝对游戏时钟, 只能在全新 boot 后按 boot 对齐序列注入(工具返回体带对齐序列说明)。

  • 保护地址段写入需显式 force=true:内核内存与 top.prx 代码段默认拒绝 写入/汇编码——这是防误写设计,不是限制性 bug。

  • 会话状态单写者:~/.ppsspp-dfx/sessions.json 跨进程共享会话登记, 并发多个 MCP 服务器实例指向同一路径时后写覆盖。

  • trace 只编排内存断点:ppsspp_breakpoint(action='trace') 布防的是 内存访问断点(默认读访问)。执行断点的一次性等待用 action='set' + action='wait' 组合。

性能参考(本机实测)

参考环境:Windows x64,PPSSPP v1.20.4-605,服务器与 PPSSPP 同机(localhost WS)。 数字随机器与游戏负载浮动,供超时预算估量,非性能承诺:

操作

实测

单次 WS 往返(game.status 级别的轻量调用)

p50 ≈ 0.21 ms,p95 ≈ 0.28 ms(n=60)

全频段 24 MB pattern 扫描(ppsspp_scan background=true,64 KiB 分块)

≈ 40 s(384 次分块读)

断点命中→可观测(热地址 set + wait,resume 后到 wait 确认)

p50 ≈ 11 ms(n=30)

社区与支持

贡献

见 CONTRIBUTING.md。

开发

从 docs/SCOPE.md(范围与协议面边界)与 evals/README.md(盲测评估体系:场景卡、确定性门禁、 runner、报告)入手。

# 前置:pytest 在 dev 依赖组中(默认安装不含)——二选一:
#   uv sync                                # 装入 dependency-groups(含 pytest)
#   pip install -e ".[dev]"                # 或装 dev extra
# 全量测试套件(单元 + 契约 + 集成;1700+ 个测试用例(不含参数化展开)):
.venv/ppsspp-dfx-mcp/Scripts/python -m pytest tests -q

# 工具签名/描述变更后重新生成工具面基线(与变更同笔提交):
.venv/ppsspp-dfx-mcp/Scripts/python scripts/dump_tool_surface.py

CI 门禁边界——勿把「CI 全绿」读作「真机已验证」:CI (.github/workflows/ci.yml)只执行 python -m pytest tests -q,不设置 PPSSPP_DFX_TEST_EXE_PATH / PPSSPP_DFX_TEST_ISO_PATH,因此依赖真实 PPSSPP 与游戏 ISO 的集成用例在 CI 中 一律 skip、不会执行。这些用例属本地真机门控:需在本机显式导出上述两个 环境变量后运行。python scripts/check_skips.py 仅审计「跳过理由是否已登记」, 不改变这一事实。

致谢

  • PPSSPP —— 被调试目标本身。本服务的 WebSocket 调试协议契约(debugger.ppsspp.org 子协议、事件语义与 HLE 内省字段)对照其源码逐项梳理并建档(见 docs/SCOPE.md)。

  • mcp-ppsspp、mcp-bizhawk、 mcp-mgba —— 同类模拟器-MCP 桥接方案;本服务的覆盖定位以它们为对照 (见 docs/SCOPE.md 的「与同类项目的覆盖对比」)。

  • 运行时依赖(MCP Python SDK、 pydantic、PyYAML、websockets)声明于 pyproject.toml。

引用

@misc{ppssppdfxmcp2026,
  title={ppsspp-dfx-mcp: a PPSSPP debug MCP server for PSP game localization},
  author={AstralVoidZ and contributors},
  year={2026},
  publisher={GitHub},
  howpublished={\url{https://github.com/AstralVoidZ/ppsspp-dfx-mcp}},
}

许可证

MIT

Available Tools

37 tools
ppsspp_analyze_logA
Read-onlyIdempotent

PURPOSE: Filter a PPSSPP log file for ERROR / WARNING / CRASH lines.

USAGE: log_path optional (defaults to the server-mirrored PPSSPP broadcast log at .ppsspp-dfx/output/ppsspp.log, written while a session runs); filter optional (keyword); session_id optional.

BEHAVIOR: READ-ONLY. Reads and filters a log file. Does not contact PPSSPP.

RETURNS: {log_path, matches: [{line_no, text}...], count, filter, filter_mode, total_matches, truncated}. truncated=true covers both an explicit limit cut and the internal 500-match cap; total_matches is a lower bound (>= the value) whenever truncated=true.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoCap the returned match list (0 = no cap beyond the hard internal cap of 500). When truncation happens the response carries total_matches and truncated=true; total_matches is a LOWER BOUND whenever truncated=true (the scan stops at 500 matches, so the real count can be higher). Use this on long logs instead of receiving 200KB+ of matches.
filterNoAdditional keyword to filter for. Case-sensitive substring match; combined with the severity keywords per filter_mode.
log_pathNoPath to the log file to analyze. Whitelist: must be a file anywhere inside the server-managed .ppsspp-dfx tree (the whole tree is allowed, wider than output/ or config/ — verified against the runtime whitelist); arbitrary filesystem paths are rejected. If None, reads the server-mirrored PPSSPP broadcast log (.ppsspp-dfx/output/ppsspp.log — the running game's own ERROR/WARNING lines, captured while a session runs).
session_idNoOptional session ID (reserved for future use; ignored).
filter_modeNo'any' (default, legacy): a line matches if it contains a severity keyword OR the filter. 'all': a line must contain a severity keyword AND the filter — use this to narrow (e.g. filter='GPU', filter_mode='all' → only GPU-related ERROR/WARNING/CRASH lines).any

Output Schema

ParametersJSON Schema
NameRequiredDescription
countYesNumber of matches (after limit truncation).
filterYesUser-supplied keyword filter.
matchesYesMatching log lines.
log_pathYesPath to the log file (or '(default)' if from launcher).
truncatedYesTrue when limit truncated the match list.
filter_modeYesFilter combination mode: 'any' (legacy OR) or 'all' (severity AND filter).
total_matchesYesMatch count before limit truncation.

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false, so the safety profile is covered. The description adds genuinely new context: it does not contact PPSSPP, the default source is the server-mirrored broadcast log, and it explains that the 500-match internal cap makes total_matches a lower bound — behavioral detail absent from annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded PURPOSE/USAGE/BEHAVIOR/RETURNS structure with no filler; every sentence carries load-bearing information and the most decision-relevant facts come first.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With annotations covering safety, an output schema present, and 100% schema description coverage, the description supplies the remaining gaps (default source, truncation lower-bound semantics) and omits return-value detail that the output schema already owns.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and each parameter already carries a detailed description, so the baseline is 3. The USAGE line restates optional/default status but adds no syntax or format detail beyond what the schema provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a precise verb+resource+scope: filter a PPSSPP log file for ERROR/WARNING/CRASH lines. No sibling tool covers this surface, and the purpose is unambiguous without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly documents that log_path, filter, and session_id are optional and what the defaults resolve to, and the schema's filter_mode documents the 'all' vs 'any' narrowing with a concrete example. It stops short of naming when a sibling tool would be preferable, but no obvious alternative exists.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ppsspp_assembleA
Destructive

PURPOSE: Assemble MIPS instruction(s) and write the resulting bytes to memory.

USAGE: session_id + address + code ('

' or ';' separated — PPSSPP assembles one line per call so the tool loops; armips-style ';' comments are NOT supported here).

BEHAVIOR: DESTRUCTIVE. Protected ranges (kernel, top.prx code) need force=true. A partial write is reported with an error directing you to disassemble and inspect.

RETURNS: {address, code, bytes_written, response, text}.
ParametersJSON Schema
NameRequiredDescriptionDefault
codeYesMIPS assembly source. May be a single instruction ('nop', 'addiu r5, r0, 0x10') or multiple instructions separated by '\n' or ';'. The assembler is PPSSPP's built-in MIPS encoder.
forceNoSet to True to write assembled bytes to protected code/data regions of the modules loaded in THIS session (kernel memory below 0x08800000, plus the top.prx code section as reported by the live module list); declared data addresses from addresses.yaml are exempt. Writing to those ranges without force=True raises ToolError to prevent accidental crashes.
addressYesTarget address where assembled bytes will be written, as a hex string (e.g. '0x08804000'). Must be a valid MIPS-aligned address for the ISA.
session_idYesActive session ID.

Output Schema

ParametersJSON Schema
NameRequiredDescription
codeYesAssembly source string passed in.
textYesUnified text representation: 'Assembled N bytes → 0x{ADDR:08X}'.
addressYesTarget address, hex string (e.g. '0x08804000').
responseYesRaw PPSSPP WebSocket response.
bytes_writtenYesNumber of bytes written. Best-effort: derived from response['bytes'] length when available, else 0.

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true and readOnlyHint=false, so safety is partly covered. The description goes further by naming the protected ranges (kernel, top.prx code), the force=true escape hatch, and the partial-write error directing the agent to disassemble and inspect — meaningful operational detail beyond the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Structured as PURPOSE/USAGE/BEHAVIOR/RETURNS with no filler; the critical constraint (one line per call) is front-loaded. The RETURNS line slightly duplicates the existing output schema, which is the only wasted sentence.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive mutation tool, the description covers the force requirement, protected-range behavior, comment syntax limits, invocation loop, and error path; return values are already documented by the output schema and annotations cover the safety profile. Nothing essential to correct invocation is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so baseline is 3. The description adds value the schema does not: the assembler loops one line per call, ';' comments are not supported, and the partial-write/force semantics are summarized as an operational consequence rather than just a field definition.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (assemble) plus resource (MIPS instructions) and the side effect (write resulting bytes to memory). An agent can immediately distinguish this from ppsspp_write_memory (raw bytes) and ppsspp_disassemble (reverse direction) without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The USAGE line names the required parameter triple and warns that ';' comments are unsupported and that PPSSPP assembles one line per call, so the tool loops. However, it never states when to prefer this over ppsspp_write_memory or ppsspp_run_script, and gives no exclusions or alternative routing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ppsspp_batch_cancelA

PURPOSE: Request cancellation of a queued or running background batch job.

USAGE: batch_id from ppsspp_batch_step(background=true); omit batch_id for survey/list mode (e.g. to recover a lost batch_id or audit background activity before touching the session).

BEHAVIOR: STATE-CHANGE. Cancels the detached task; the job's own finally block releases the session lock, so subsequent tool calls are free to use the session immediately. The abort happens at the current step boundary (a press finishes, a mid-wait cuts within ~1s). Cancelling an already-finished job is an error — check ppsspp_batch_status first if unsure.

RETURNS: {batch_id, status, note} — poll ppsspp_batch_status for the terminal state.

ParametersJSON Schema
NameRequiredDescriptionDefault
batch_idYesJob id returned by ppsspp_batch_step(background=true).

Output Schema

ParametersJSON Schema
NameRequiredDescription
noteYes
statusYes
batch_idYes

TDQS

A4.3/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare the mutation/safety profile (readOnlyHint=false, idempotentHint=false, destructiveHint=false); the description goes well beyond by disclosing that the job's finally block releases the session lock immediately, that the abort lands at the step boundary (~1s mid-wait), and that cancelling a finished job errors. This is exactly the operational context an agent needs and cannot get from structured fields.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Sectioned PURPOSE/USAGE/BEHAVIOR/RETURNS layout is front-loaded with the essential verb and every sentence carries information: id sourcing, lock-release timing, abort granularity, and the error case. No filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Output schema exists so return values need not be detailed, yet the description still summarizes the return shape and directs polling to ppsspp_batch_status. Annotations plus description cover safety and side effects well; the only real gap is the unresolved required-vs-optional contradiction for batch_id.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% (baseline 3), but the description actively contradicts the schema by instructing the agent to omit batch_id for 'survey/list mode' while the input schema marks batch_id required. A misleading claim about a required parameter can cause a failed call, which is worse than saying nothing.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (cancel) and resource (queued or running background batch job), and implicitly distinguishes itself from ppsspp_batch_step and ppsspp_batch_status by naming them as the source of the id and the pre-check tool. An agent can route to it without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit sourcing ('batch_id from ppsspp_batch_step(background=true)') and a clear guard ('check ppsspp_batch_status first if unsure'). The claimed 'omit batch_id for survey/list mode' is useful context but conflicts with the schema's required field, so guidance is not fully trustworthy.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ppsspp_batch_statusA
Read-onlyIdempotent

PURPOSE: Poll the state and progress of a background batch job without touching the session.

USAGE: batch_id from ppsspp_batch_step(background=true); omit batch_id for survey/list mode (e.g. to recover a lost batch_id or audit background activity before touching the session).

BEHAVIOR: Lock-free registry read — never opens the WS transport and never waits for the per-session lock, so it is safe to call while a background batch (or any other tool) owns the session. Executed-step count updates as steps complete; 'completed' carries the full foreground-shaped result. READ-ONLY.

RETURNS: with batch_id → {batch_id, session_id, status: queued|running|completed|failed|cancelled, executed, total, error, result, retention_jobs}; with batch_id omitted (survey) → {jobs: [{batch_id, session_id, status, executed, total, error, result_present}], retention_jobs} in submission order (finished jobs beyond retention are evicted and absent).

ParametersJSON Schema
NameRequiredDescriptionDefault
batch_idNoJob id returned by ppsspp_batch_step(background=true). Omit to survey ALL retained jobs instead (list mode; recovers a batch_id after the submit response was lost).

Output Schema

ParametersJSON Schema
NameRequiredDescription
jobsNoAll retained jobs, in submission order
errorNoError message if failed/cancelled (null when none; real runner emits null)
totalNoTotal steps in the batch
resultNoFinal ppsspp_batch_step-shaped response; present once the batch completed
statusNo'queued' / 'running' / 'completed' / 'failed' / 'cancelled' (protocol Tasks mapping: 'queued'→'working')
batch_idNoJob id
executedNoSteps executed so far
session_idNoSession the batch runs on
retention_jobsNoFinished-job retention window (in job count): finished jobs beyond the oldest this many are evicted and absent from jobs

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare the safety profile (readOnly/idempotent/non-destructive); the description goes further by explaining WHY it is safe — lock-free registry read, no WS transport, no per-session lock wait — which tells the agent it can be called concurrently with a batch owning the session. It also discloses eviction behavior (finished jobs beyond retention are absent).

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four labeled sections (PURPOSE/USAGE/BEHAVIOR/RETURNS) with the most decision-relevant fact — what it does and that it is session-safe — front-loaded. Dense but every sentence carries operational information; no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Even though an output schema exists, the description supplies the concurrency/safety model and retention eviction caveat that an agent needs to call this correctly alongside a running batch. Nothing required for correct invocation is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and the schema already documents the omit-to-survey semantics, so the baseline is 3. The description earns an extra point by tying each parameter mode to a distinct return shape (single job vs. jobs array in submission order), which is meaning the schema alone does not convey.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource — 'poll the state and progress of a background batch job' — and immediately scopes it with 'without touching the session'. An agent can distinguish it from ppsspp_batch_step and ppsspp_batch_cancel purely from this line.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly names where batch_id comes from (ppsspp_batch_step with background=true) and defines the no-argument survey/list mode with concrete use cases (recovering a lost batch_id, auditing background activity before touching the session). Both modes and their selection conditions are stated, not implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ppsspp_batch_stepA

PURPOSE: Execute an ordered automation sequence of press / wait / state_probe / screenshot steps in one call, optionally on a detached background task that outlives the client timeout.

USAGE: session_id + steps:[{type: press|wait|state_probe|screenshot, ...}]; on_failure='continue'|'abort' (default continue); background=false|true.

ROUTING: ordered multi-step automation -> here; single CPU-step -> ppsspp_step; background job survey -> ppsspp_batch_status(batch_id omitted). BEHAVIOR: STATE-CHANGE. Foreground (default) holds the session lock for the whole batch; frames are 60fps wall-clock equivalents; sequences estimated >25s are rejected up-front with BATCH_BUDGET_EXCEEDED (the MCP client aborts tool calls at ~30s, killing the remaining steps server-side). Per-step MCP progress is reported when the client requests it. background=true validates and submits instantly, returns {action:'submitted', batch_id,...}, keeps the session lock for the batch duration, and reports progress via ppsspp_batch_status. If any foreground step fails the whole call is isError BATCH_STEP_FAILED — inspect results[] per step. Screenshots are auto-skipped during replay recording.

RETURNS: foreground {total, executed, succeeded, failed, skipped, recording_mode, results[], aborted}; background {action:'submitted', batch_id, session_id, total, estimated_s, hint}.

ParametersJSON Schema
NameRequiredDescriptionDefault
stepsYesOrdered list of step dicts to execute. Each step must have a 'type' field. Supported types: - press: {type:'press', button:'cross', duration:30} - wait: {type:'wait', frames:60} - state_probe: {type:'state_probe', names:'game_mode', samples:1} - screenshot: {type:'screenshot', source:'render'}
backgroundNoRun on a detached server task that survives the MCP client's ~30s tool-call timeout (default false). Foreground calls exceeding the 25s budget are rejected with BATCH_BUDGET_EXCEEDED; background calls return a batch_id immediately — poll ppsspp_batch_status for progress and the final result, cancel via ppsspp_batch_cancel. Background jobs have no per-step MCP progress notifications; use the status poll.
on_failureNoWhat to do when a step fails (default 'continue'). 'continue' keeps running subsequent steps; 'abort' stops the batch immediately. Screenshot steps that are skipped due to recording mode are NOT failures.continue
session_idYesActive session ID.

Output Schema

ParametersJSON Schema
NameRequiredDescription
hintNoHow to poll / cancel this background batch
totalNoTotal steps in the batch
actionNoAlways 'submitted'
failedNoSteps with status='failure'
abortedNoWhether batch aborted early on failure
resultsNoPer-step results, in order
skippedNoSteps with status='skipped'
batch_idNoJob id for ppsspp_batch_status / _cancel
executedNoSteps actually executed (excludes skipped)
succeededNoSteps with status='success'
session_idNoSession the batch will execute on
estimated_sNoHeuristic wall-clock estimate in seconds
abort_reasonNoEmpty if not aborted
recording_modeNoWhether session was recording a replay when batch ran

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare the mutation profile (readOnlyHint=false, idempotentHint=false) but the description adds substantial context beyond them: the session lock is held for the whole batch, sequences estimated >25s are rejected with BATCH_BUDGET_EXCEEDED because the client aborts at ~30s, foreground failure makes the whole call isError BATCH_STEP_FAILED, background returns a batch_id immediately, and screenshots are auto-skipped during replay. This is a rich, non-obvious behavioral disclosure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with PURPOSE/USAGE/ROUTING/BEHAVIOR/RETURNS headers that make it scannable, and each section earns its place given the tool's complexity. It is dense and long, but the structure keeps it navigable rather than wasteful.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Despite an existing output schema, the definition covers everything an agent needs: supported step types, budget/timeout constraints, lock behavior, failure semantics, background submission flow, and progress reporting. Nothing critical for correct invocation is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the schema already documents session_id, steps, on_failure, and background with detailed per-field descriptions and a discriminator. The description largely restates the defaults already present in the schema, adding only marginal framing, so baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Execute'), the resource ('ordered automation sequence'), and the exact step types (press/wait/state_probe/screenshot), plus the background variant. The ROUTING line explicitly separates it from ppsspp_step and ppsspp_batch_status, so an agent can identify it without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

ROUTING gives explicit when-to-use routing: ordered multi-step automation goes here, single CPU-step goes to ppsspp_step, background job survey goes to ppsspp_batch_status. Alternatives and the condition selecting them are named outright.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ppsspp_breakpointA

PURPOSE: Manage breakpoints AND consume their hits — set/remove/update/list CPU execution breakpoints and memory watchpoints, strict-wait for a hit, or one-call arm-hit-capture-resume tracing.

USAGE: session_id optional when exactly one session is active; management actions as below; action='wait' blocks until any breakpoint is hit (set/mem_set first; lock-free; breakpoint stays armed); action='trace' arms a temporary MEMORY breakpoint at address, waits, captures pc/registers/backtrace, always removes it and resumes (defaults to read access; narrow with read/write/size). For EXECUTION breakpoints use action='set' + 'wait'.

ROUTING: persistent breakpoint management -> here; one-shot strict-wait -> action='wait'; armed hit-capture -> action='trace'. BEHAVIOR: MUTATING. trace arms/removes and set/mem_* manage state; Reliable hits need CPUCore=2 (IR Interpreter). mem_remove resolves the watchpoint's real size via mem_list first (address+size matching); mem_update merges existing read/write/change unconditionally (PPSSPP zero-omits omitted bools). CPU set/remove return no data — the tool follows with a list for verification. wait/trace are lock-free during the wait itself (concurrent reads keep working); do NOT submit step/pause/resume during a wait. CONDITION SEMANTICS: PPSSPP's IR mode ignores register conditions, so any condition is enforced MCP-side — the breakpoint is armed unconditionally and each hit's expression is evaluated with cpu.evaluate; a falsy hit is auto-resumed (not surfaced) and counted in filtered_hits.

RETURNS: stats → {mode:"stats", window_s, total_hits, by_pc: [{pc, count, first_seen, last_seen}]} (fixed ~30s sampling window — no shorter-window option, probe_changes?: [{probe, old, new, ts}], note}; management actions → {action, address, enabled, breakpoints[]}; wait → {hit, already_paused, timeout_s, pc, reason, related_address, ticks, condition, condition_filtered, filtered_hits, storm_break}; trace → {hit, already_paused, address, access, timeout_s, hits: [{pc, related_address, reason, ticks, mem_hits?, registers?, backtrace?}], bp_removed, resumed, note}.

ParametersJSON Schema
NameRequiredDescriptionDefault
logNoLog flag (update / mem_set / mem_update only; None = don't change). For mem_set, defaults to False when None.
readNoTrigger on read access (mem_set only; None defaults to True). For mem_update, passing read triggers a merge query — omit to leave read unchanged.
sizeNoMemory breakpoint watch size in bytes (mem_set / mem_remove / mem_update; default 4). Fixed-width watches use 1/2/4; larger sizes are passed through to PPSSPP as a range watch. NOTE: PPSSPP matches a memory watchpoint by the exact address+size pair (BreakpointSubscriber.cpp) -- the size is part of the match key, not just bookkeeping. mem_remove therefore resolves the real size via mem_list first, because removing with the caller's size alone silently fails when it differs (e.g. a 16-byte watch removed with the default 4).
writeNoTrigger on write access (mem_set only; None defaults to True). For mem_update, passing write triggers a merge query — omit to leave write unchanged.
actionYesBreakpoint operation. Valid values: Consumption actions (lifecycle orchestration): - 'wait': STRICT-WAIT — block until any breakpoint is hit (arm nothing; set/mem_set first). Lock-free: concurrent reads keep working. The breakpoint stays armed. - 'stats': HIT-FREQUENCY — count breakpoint hits by pc over a time window; optionally samples probe value changes via state_observer. Read-only. - 'trace': HIT-SNAPSHOT-RESUME — arm a temporary MEMORY breakpoint at `address`, wait for the hit, capture pc/registers/backtrace, ALWAYS remove it, then resume (defaults to read access; narrow with read/write/size). For EXECUTION breakpoints use action='set' + 'wait' instead. CPU breakpoint actions: - 'set': add a CPU execution breakpoint (requires address; enabled? defaults to True; condition? optional — enforced MCP-side (falsy hits auto-resumed, counted in filtered_hits), NOT sent to PPSSPP). - 'remove': delete a CPU breakpoint by address. - 'list': list all current CPU breakpoints. - 'update': update a CPU breakpoint's enabled/log/condition/log_format (requires address; all other params optional; condition='' clears it). Memory breakpoint actions: - 'mem_set': add a memory access breakpoint (requires address; size?/read?/write?/enabled?/log?/condition?/log_format?). - 'mem_remove': delete a memory breakpoint by address. Delete semantics are STRICT: removing a non-existent memcheck is an ERROR (unlike ppsspp_state_observer action=clear, which is idempotent-ok). - 'mem_list': list all current memory breakpoints. - 'mem_update': update a memory breakpoint's enabled/log/condition/log_format (requires address).
addressNoRequired for set / remove / update / mem_set / mem_remove / mem_update. Breakpoint address, as a hex string (e.g. '0x08804000'). Not used by list / mem_list. The schema default of '0x0' exists for legacy callers -- do NOT rely on it when the action is one of the above.0x0
enabledNoBreakpoint enable flag. For action='set' / 'mem_set', defaults to True when None. For action='update' / 'mem_update', None means 'don't change'. Ignored for remove / list actions.
conditionNoBreak condition expression (set / update / mem_set / mem_update; None = don't send). Enforced MCP-side: PPSSPP's IR mode silently ignores register conditions, so the breakpoint is armed UNCONDITIONALLY and falsy hits are auto-resumed and counted in filtered_hits.
timeout_sNoWait budget in seconds (wait / trace only; default 30, clamped to [0.5, 300]). On timeout: hit=false — NOT an error — so callers can poll.
log_formatNoLog format string (update / mem_set / mem_update only; None = don't change).
session_idNoActive session ID; omit to auto-resolve when exactly one session is active.
want_backtraceNoInclude the HLE backtrace in the hit (trace only; CPU is paused at the hit, so the trace is valid).
want_registersNoInclude the full CPU register dump in the hit (trace only).

Output Schema

ParametersJSON Schema
NameRequiredDescription
pcNo
hitNo
hitsNo
modeNo
noteNo
by_pcNo
ticksNo
accessNo
actionNo'set' / 'remove' / 'list' / 'update' / 'mem_set' / 'mem_remove' / 'mem_list' / 'mem_update'.
reasonNo
addressNo
enabledNoEnabled flag (set / mem_set / update / mem_update only).
resumedNo
mem_hitsNo
window_sNo
conditionNo
timeout_sNo
bp_removedNo
total_hitsNo
breakpointsNoBreakpoint list (list / mem_list only).
storm_breakNo
filtered_hitsNo
probe_changesNo
already_pausedNo
related_addressNo
condition_filteredNo

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the annotations, it discloses that the tool is mutating, that trace arms and always removes a temporary memory breakpoint and resumes, that reliable hits need CPUCore=2, that mem_remove resolves the real size via mem_list, that mem_update merges omitted booleans, that CPU set/remove return no data and are followed by a list, and that wait/trace remain lock-free. These are material behavioral traits not captured by the structured fields.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long, but it is front-loaded with PURPOSE and organized into USAGE, ROUTING, BEHAVIOR, and RETURNS sections. The length is defensible for an 11-action tool, though some material repeats the schema or output schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the 11 actions, 13 parameters, annotations, and output schema, the description is complete. It covers action semantics, mutation behavior, condition filtering, timeout behavior, return-shape expectations, and important edge cases such as strict mem_remove and PPSSPP IR condition handling.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all 13 parameters in detail. The description adds cross-action routing and some caveats, but most parameter-specific meaning is redundant with the schema, so the baseline of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb-and-resource statement: it manages CPU execution breakpoints and memory watchpoints and consumes their hits. It distinguishes management, wait, trace, and stats actions, and routes execution breakpoints to set + wait, so an agent can identify what this tool does without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly states when to use each action: persistent management, one-shot strict-wait via action='wait', and armed hit-capture via action='trace'. It also says when to use set + wait for execution breakpoints and warns against submitting step/pause/resume during a wait, giving clear alternatives and exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ppsspp_contextA
Read-onlyIdempotent

PURPOSE: One-call crash-triage pack — identity + disassembly window + optional backtrace for an address.

USAGE: pass a crash PC or call target; identity resolves via addresses.yaml known_functions (IDA offset applied); disasm covers window instructions before/after; include_backtrace=true adds the call stack (pauses the CPU briefly).

BEHAVIOR: READ-ONLY. Unknown addresses return identity=null and the raw window instead of failing; backtrace is skipped (not an error) when the CPU is running — the note field says why. SCOPE: identity/region come from a TOPX-specific address knowledge base (addresses.yaml). For any other ISO they come back empty and note says so explicitly -- an empty identity means 'not in the tables', NOT 'bad address'.

ROUTING: persistent breakpoints around this address -> ppsspp_breakpoint; one armed hit-capture -> ppsspp_breakpoint(action="trace"); recurring sampling -> ppsspp_state_observer.

RETURNS: {address, identity: {name, start, offset} | null, region, disasm: [{address, text}], backtrace: [...], backtrace_note}.

ParametersJSON Schema
NameRequiredDescriptionDefault
windowNoDisassembly instructions BEFORE the address (same count after, so 2*window+1 total; default 8, cap 32).
addressYesAddress to triage, hex ('0x088EF0F4') or decimal — typically a crash PC or a call target.
session_idNoActive session ID; auto-resolved when exactly one session is active.
include_backtraceNoInclude the call stack (default false). Requires a paused CPU — the tool pauses/resumes around it; concurrent readers block briefly.

Output Schema

ParametersJSON Schema
NameRequiredDescription
disasmYesInstructions around the address (window before / at / after).
regionYesMemory-map region containing the address (e.g. 'user'), '' when unmapped.
addressYesQueried address (hex).
identityYesNearest known function at/below the address; None when unknown.
backtraceYesCall stack (paused CPU only; empty when skipped).
backtrace_noteYesWhy the backtrace is empty (e.g. 'CPU running — pause first').

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well beyond the read-only/idempotent annotations: unknown addresses return identity=null rather than failing, backtrace is silently skipped (not an error) when the CPU is running with the reason in `note`, the CPU is briefly paused for backtrace and concurrent readers block, and identity comes from a TOPX-specific KB where an empty result means 'not in tables' not 'bad address'.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Header-labeled sections (PURPOSE/USAGE/BEHAVIOR/SCOPE/ROUTING/RETURNS) make it front-loaded and scannable, and each sentence carries signal. The RETURNS block is somewhat redundant given an output schema exists, costing a point.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values needn't be spelled out, yet the description still covers the non-obvious cases (null identity, skipped backtrace, ISO-specific empty results, brief CPU pause) that an agent must know to interpret responses correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description reinforces that address is a crash PC/call target and that window covers instructions before/after, but the schema already documents the window count, cap, and the backtrace pause/concurrency behavior, so little is added beyond structured fields.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource pack ('one-call crash-triage pack — identity + disassembly window + optional backtrace for an address') and explicitly names sibling tools in the ROUTING section, so an agent can distinguish it from ppsspp_breakpoint and ppsspp_disassemble without opening schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides explicit when-to-use ('pass a crash PC or call target') and explicit alternatives with their selecting conditions: persistent breakpoints -> ppsspp_breakpoint, one armed hit-capture -> ppsspp_breakpoint(action="trace"), recurring sampling -> ppsspp_state_observer.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ppsspp_diff_memoryA
Read-only

PURPOSE: Snapshot a memory range and diff it against current memory — the classic "what changed?" variable-locator.

USAGE: diff_memory(action="snapshot", start=..., end=...) → handle; act in game; diff_memory(action="compare", handle=...) → changed-byte list; action="drop"/"list" manage handles.

BEHAVIOR: READ-ONLY. Memory is never written — only the per-server snapshot registry mutates. Large ranges are read across multiple reads; per-snapshot cap 8 MiB; registry cap 4 with FIFO eviction. The registry is process-global: parallel sessions share one cap and FIFO order, so another session's snapshots can evict yours under load. compare requires the handle's exact range.

RETURNS: snapshot → {handle, start, size_bytes, checksum}; compare → {handle, start, size_bytes, changed_count, truncated, changes: [{address, old, new}]}; drop → {handle, dropped}; list → {handles: [...], count, max_snapshots}.

ParametersJSON Schema
NameRequiredDescriptionDefault
endNoRange end (exclusive), same format — required for snapshot unless 'address'+'size' are given.
sizeNoRange length in bytes. Provided for naming consistency with read_memory/write_memory/scan, which all take 'size'. Use either address+size or start+end.
startNoRange start, hex ('0x08804000') or decimal — required for snapshot unless 'address'+'size' are given.
actionYesDiff operation. Valid values: - 'snapshot': read [start, start+size) and store it under a new handle (registry cap 4, FIFO eviction). - 'compare': read the same range now and diff against the handle's snapshot; returns a changed-byte list (inline cap 256, truncated flag keeps the true count). - 'drop': release a handle. - 'list': live handles.
handleNoSnapshot handle — required for compare/drop.
addressNoRange start, same format as 'start'. Provided for naming consistency with read_memory/write_memory/scan, which all take 'address'. Use either address+size or start+end.
session_idNoActive session ID; auto-resolved when exactly one session is active (required for snapshot/compare, ignored for drop/list).

Output Schema

ParametersJSON Schema
NameRequiredDescription
countNoNumber of live snapshots.
startNoStart address (hex).
handleNoDropped handle.
changesNoFirst N changed bytes (address + old/new).
droppedNoTrue when the handle existed and was removed.
handlesNoLive snapshot handles (oldest first).
checksumNosha256 of the snapshot bytes (first 16 hex).
truncatedNoTrue when changes exceed the inline cap (see changes list).
size_bytesNoCompared range size in bytes.
changed_countNoTotal changed bytes.
max_snapshotsNoRegistry capacity (FIFO eviction).

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false, yet the description adds substantial behavior beyond that: the per-snapshot 8 MiB cap, registry cap of 4 with FIFO eviction, and critically the process-global registry shared across parallel sessions so another session can evict your snapshots. These are non-obvious operational traits that materially change how an agent should use the tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four labeled sections (PURPOSE/USAGE/BEHAVIOR/RETURNS) front-load the most important information and every sentence carries payload — caps, eviction semantics, return shapes. No filler or redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 7-parameter tool with annotations and an output schema, the description closes the remaining gaps: cross-action parameter requirements, registry capacity and eviction behavior, session-sharing hazards, and return payloads. An agent has everything needed to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is already 100%, so the baseline is 3, but the USAGE and RETURNS blocks add call-shape meaning beyond the schema: which params belong to which action (handle required for compare/drop, session_id ignored for drop/list) and what each action returns. The clarification that compare requires the handle's exact range is genuine added semantics.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The PURPOSE line gives a specific verb+resource with an evocative framing (“classic 'what changed?' variable-locator”) that makes the tool instantly identifiable. It is clearly distinguishable from read_memory, write_memory, and scan, which the description later references by name.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The USAGE section lays out the exact workflow (snapshot → act in game → compare) and enumerates how each of the four actions is invoked. It does not, however, state when to prefer this over overlapping siblings like scan or watch_value, so an agent gets strong invocation guidance but no exclusion criteria.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ppsspp_disassembleA
Read-onlyIdempotent

PURPOSE: Disassemble N MIPS instructions at a given address.

USAGE: address required; session_id optional when exactly one session is active; count optional (default 10).

BEHAVIOR: READ-ONLY. Calls memory.disasm via WebSocket. Does not modify memory or CPU state.

RETURNS: {address, count, instructions: [{address, text}...]}.

ParametersJSON Schema
NameRequiredDescriptionDefault
countNoNumber of instructions to disassemble. Default 10. count=0 is treated as 'use the default' and returns 10 instructions (the response carries a note saying so). Capped at 100 to prevent oversized responses; a larger value is clamped and the note reports the clamp.
addressYesStarting address for disassembly, as a hex string (e.g. '0x08804000').
session_idNoActive session ID; omit to auto-resolve when exactly one session is active.

Output Schema

ParametersJSON Schema
NameRequiredDescription
countYesNumber of instructions disassembled.
addressYesStarting address, hex string (e.g. '0x08804000').
instructionsYesList of disasm line dicts (text + address).

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint/idempotentHint=true and destructiveHint=false, so the safety profile is pre-covered; the description reinforces this and adds genuine context beyond structured fields by naming the underlying mechanism ('Calls memory.disasm via WebSocket'). It also confirms no memory or CPU state is modified, but adds no rate-limit or failure-mode detail.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four labeled sections (PURPOSE/USAGE/BEHAVIOR/RETURNS) with zero filler; the essential purpose leads and each line carries distinct information. Well front-loaded for fast agent scanning.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so the RETURNS block is a convenience rather than a necessity. Combined with full 100% schema coverage and annotations covering the safety profile, an agent has everything needed to invoke this correctly; the added WebSocket/read-only context closes the remaining gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%: the schema already documents count's default, count=0 handling, the 100 cap, the clamp note, address hex format, and session_id auto-resolution. The description only restates defaults and never adds syntax or edge-case meaning beyond the schema, so baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a precise verb+resource: 'Disassemble N MIPS instructions at a given address.' This clearly distinguishes it from siblings like ppsspp_assemble (the inverse) and ppsspp_search_disasm (which searches rather than dumps at a fixed address).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

USAGE states parameter requirements ('address required; session_id optional when exactly one session is active; count optional'), which is helpful, but it never says when to prefer this tool over alternatives such as ppsspp_search_disasm or ppsspp_list_addresses. Usage is implied rather than explicitly scoped.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ppsspp_dumpA
Read-onlyIdempotent

PURPOSE: Dump the currently-bound GPU texture OR CLUT palette as an image plus metadata.

USAGE: kind='texture' → the bound texture (level selects mipmap; PPSSPP captures the currently-bound texture — it does NOT support capture by VRAM address); kind='clut' → the bound palette (level must be 0).

BEHAVIOR: READ-ONLY. An empty capture raises CAPTURE_EMPTY — enter a scene that renders (texture) or uses the palette (clut) and retry.

RETURNS: structuredContent metadata (kind/level/file_path/size_bytes/format); the image itself arrives as an ImageContent block.

ParametersJSON Schema
NameRequiredDescriptionDefault
kindYesWhat to capture from the CURRENTLY bound GPU state (no VRAM-address targeting): 'texture' = the bound texture (use level for mipmap); 'clut' = the bound CLUT palette.
levelNoTexture mipmap level (default 0). kind=texture only — a non-zero level with kind=clut is rejected.
session_idYesActive session ID.

Output Schema

ParametersJSON Schema
NameRequiredDescription
levelYesTexture mipmap level dumped.
formatYesImage format ('png' or 'jpeg').
file_pathYesPath where image was saved.
size_bytesYesDecoded image size in bytes.

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover read-only/idempotent/non-destructive, and the description adds real behavior beyond that: a documented failure mode (CAPTURE_EMPTY) with a retry remedy, the constraint that capture is by currently-bound state only (no VRAM-address targeting), and the response shape. This is genuinely useful error/limitation disclosure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Labeled sections (PURPOSE/USAGE/BEHAVIOR/RETURNS) front-load the key information and every sentence carries weight. The only cost is some duplication of the schema's kind/level descriptions, which slightly dilutes conciseness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, the description needn't detail return values, and it still summarizes structuredContent fields and the ImageContent block. Between annotations, schema, and the error guidance, an agent has everything needed to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so all three parameters are already documented, including the enum and the clut/level rejection rule. The description largely restates the same semantics (no VRAM-address targeting is the only marginal addition), so baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource: dump the currently-bound GPU texture or CLUT palette as image plus metadata. It is clearly distinguishable from siblings like ppsspp_screenshot and ppsspp_gpu_record because it names GPU texture/CLUT capture as the scope.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit per-value guidance (kind='texture' uses level for mipmap; kind='clut' requires level=0) and tells the agent what to do when capture is empty. It does not, however, name an alternative tool for related needs (e.g., screenshot), so it stops short of full when/when-not routing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ppsspp_evaluateA
Read-onlyIdempotent

PURPOSE: Evaluate a debugger expression (register names, hex literals, simple arithmetic).

USAGE: session_id + expression. No '*addr' dereference syntax — read memory with read_u32 instead.

BEHAVIOR: READ-ONLY. Pauses/resumes the CPU internally.

RETURNS: {expression, value, response, text}.

ParametersJSON Schema
NameRequiredDescriptionDefault
expressionYesDebugger expression to evaluate. Examples: 'r5 + 0x10', 'pc', 'r5 + r6'. The supported syntax is whatever PPSSPP's expression evaluator accepts. Note: PPSSPP's evaluator does NOT support dereference syntax like '*0x08804000' — use read_u32 / read_bytes instead to read memory at an address.
session_idYesActive session ID.

Output Schema

ParametersJSON Schema
NameRequiredDescription
textYesUnified text representation: '{EXPR} = 0x{VAL:X}' when value is an int, or '{EXPR} = {value!r}' otherwise.
valueYesEvaluated value (int) if numeric, else None.
responseYesRaw PPSSPP WebSocket response.
expressionYesThe expression evaluated.

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and destructiveHint, so the description's 'READ-ONLY' label mostly repeats structured data. However, 'Pauses/resumes the CPU internally' adds a meaningful behavioral side effect beyond the annotations exceptional. This helps an agent understand why a seemingly safe read may still briefly affect execution state.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and well-organized into PURPOSE, USAGE, BEHAVIOR, and RETURNS. Every line carries meaningful operational information without redundancy or fluff. The most important constraints are front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the two parameters, high schema coverage, and rich annotations, the description provides everything needed to invoke the tool correctly: purpose, required inputs, a behavioral side effect, an alternative for an unsupported use case, and the return shape. Nothing essential is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and the expression parameter already documents supported syntax, examples, and the no-dereference caveat. The description's 'session_id + expression' adds no new semantic value beyond the schema. There is no coverage gap for the description to compensate for, so baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Evaluate a debugger expression', and enumerates the accepted operand classes ('register names, hex literals, simple arithmetic'). It also separates itself from memory-reading tools by explicitly stating dereference syntax is not supported and pointing to read_u32 instead.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The 'USAGE' line indicates that session_id and expression are required. More importantly, the exclusion 'No '*addr' dereference syntax — read memory with read_u32 instead' gives an explicit when-not-to-use rule and names the alternative. It does not cover all possible alternatives among the many sibling tools, but the key routing edge case is handled.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ppsspp_frame_snapshotA

PURPOSE: One-call paused scene snapshot — pause (unless already paused), capture pc + registers + optional named probes, then resume.

USAGE: session_id; probes = optional comma-separated state_observer registry names; want_registers default true. Prefer this over a manual pause + query(registers) + resume sequence.

ROUTING: pause+capture+resume in one call -> here; cheap PC-only check -> ppsspp_query(action='register', name='pc', safe=true); recurring sampled probes -> ppsspp_state_observer. BEHAVIOR: STATE-CHANGE. The session lock is held for the whole call (pause→capture→resume is short). A CPU we paused is resumed before returning; an already-paused CPU stays paused. A failing capture never leaves the game frozen.

RETURNS: {was_stepping, resumed, pc, trust_level, registers, probes} — registers/probes keys are ALWAYS present; they carry null when opted out (want_registers=false / probes omitted) — nullable-key contract, 2026-09-08.

ParametersJSON Schema
NameRequiredDescriptionDefault
probesNoOptional comma-separated state_observer registry probe names to capture alongside the CPU state (empty = none).
session_idYesActive session ID.
want_registersNoInclude the full CPU register dump (GPR/FPU/VFPU).

Output Schema

ParametersJSON Schema
NameRequiredDescription
pcNoProgram counter, hex string (high trust).
probesNostate_observer capture block (when requested).
resumedNoTrue when the tool resumed a CPU it had paused (an already-paused CPU is left paused).
registersNoFull GPR/FPU/VFPU register dump (when requested).
trust_levelNosafe_get_pc trust level.
was_steppingNoTrue when the CPU was already paused at entry.

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare readOnlyHint=false, idempotentHint=false, destructiveHint=false. The description goes well beyond them: it discloses the session-lock scope, that a CPU it paused is resumed before returning while an already-paused CPU stays paused, and that a failing capture never leaves the game frozen — exactly the safety contract an agent needs for a state-mutating call.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The PURPOSE/USAGE/ROUTING/BEHAVIOR/RETURNS layout is front-loaded and scannable, with the routing decision placed early. It is slightly verbose, and the RETURNS block partially duplicates the existing output schema, but each section still earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given an output schema already exists, the description needn't restate return fields, yet its nullable-key contract note and behavior guarantees are relevant extras. Routing, preconditions, failure behavior, and pause/resume semantics are all covered for a short state-mutating snapshot tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and every parameter already carries a description and default, so the baseline is 3. The description's USAGE line largely restates the schema (probes = comma-separated registry names, want_registers default true) rather than adding syntax or format detail beyond it.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The PURPOSE line gives a specific verb+resource ('paused scene snapshot') and spells out the full sequence (pause, capture pc + registers + optional probes, resume). Siblings like ppsspp_query and ppsspp_state_observer are explicitly named as alternatives, so the agent can distinguish this tool without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The ROUTING line states exactly when to use this tool ('pause+capture+resume in one call') and names two alternatives with their triggering conditions ('cheap PC-only check -> ppsspp_query', 'recurring sampled probes -> ppsspp_state_observer'). Nothing is left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ppsspp_gpu_recordA
Read-onlyIdempotent

PURPOSE: Capture the next rendered frame's GE command stream as a binary dump file.

USAGE: session_id. The CPU must be RUNNING — a paused GPU never flips a frame; the MCP pre-probe converts that into a clean CPU_STATE_ERROR.

BEHAVIOR: READ-ONLY. Captures to a binary file under output/gpu_dumps/ (not JSON).

RETURNS: {size_bytes, file_path, raw, text}.

ParametersJSON Schema
NameRequiredDescriptionDefault
session_idYesActive session ID.

Output Schema

ParametersJSON Schema
NameRequiredDescription
rawYesRaw `gpu.record.dump` response metadata from PPSSPP.
textYesUnified text representation: 'dumped {N} bytes → {file_path}'.
file_pathYesLocal file path where the dump was auto-saved (.ppsspp-dfx/output/gpu_dumps/<timestamp>.dump). Empty if the dump was empty or save failed.
size_bytesYesSize of the decoded GE command dump in bytes. 0 if PPSSPP returned no data (e.g., no game running, or timeout).

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the readOnlyHint annotation, the description reveals the output destination (output/gpu_dumps/), the binary (non-JSON) format, and the MCP pre-probe error conversion. It explicitly says READ-ONLY, which is consistent with annotations, and adds meaningful behavioral context not visible in the schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is organized into four compact labeled sections (PURPOSE, USAGE, BEHAVIOR, RETURNS) and front-loads the key facts. Every sentence carries operational value with no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter tool with a rich annotation set and an existing output schema, the description covers purpose, precondition, output behavior, and return shape. An agent has enough information to invoke it correctly and understand the result without looking up additional context.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and the session_id parameter is already documented as 'Active session ID.' The description only repeats 'session_id' in USAGE without adding extra syntax, validation, or format semantics, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb and resource: 'Capture the next rendered frame's GE command stream as a binary dump file.' This clearly identifies the tool's operation and distinguishes it from sibling GPU tools like ppsspp_gpu_stats or ppsspp_screenshot.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It states the only required input, session_id, and gives an explicit precondition: the CPU must be RUNNING, with a concrete explanation of why and the resulting CPU_STATE_ERROR. It does not name alternative sibling tools, but the when-to-use condition is unambiguous.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ppsspp_gpu_statsA
Read-onlyIdempotent

PURPOSE: Query GPU counters — fps, vblanks per second, timing info.

USAGE: session_id. The CPU must be RUNNING; paused, the MCP pre-probe returns CPU_STATE_ERROR instead of hanging — which doubles as the cheapest paused-CPU probe.

BEHAVIOR: READ-ONLY.

ON TIMEOUT: the error names WHY it timed out -- expected_stall (the CPU was stepping, so no frame is coming), pairing_broken (a broadcast arrived meanwhile, so ticket pairing failed), or no_producer (nothing was broadcast at all, so the emulator is not producing frames -- check for a modal dialog blocking it). A CPU_FREEZE_SUSPECTED is re-checked against the frame heartbeat first: if no frames are arriving it is reported as WS_TIMEOUT with a no-producer attribution, NOT as a CPU freeze.

RETURNS: {fps, vblanks_per_second, info, timing, raw, text}.

ParametersJSON Schema
NameRequiredDescriptionDefault
session_idYesActive session ID.

Output Schema

ParametersJSON Schema
NameRequiredDescription
fpsYesFrames per second. None if PPSSPP didn't return it (e.g., no game running, or response shape differs).
rawYesRaw `gpu.stats.get` response from PPSSPP.
infoYesGPU info dict (vendor / name / version, etc.).
textYesUnified text representation: 'fps={FPS} vblanks={VBLANKS} info_keys={N} timing_keys={N}'.
timingYesGPU timing dict (frame / block / vertex timing, etc.).
vblanks_per_secondYesVBlanks per second. None if PPSSPP didn't return it.

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the readOnly/idempotent annotations, it discloses a full timeout taxonomy (expected_stall, pairing_broken, no_producer) and the escalation rule that CPU_FREEZE_SUSPECTED is demoted to WS_TIMEOUT with no-producer attribution when the frame heartbeat is silent. That is behavioral context an agent cannot get from the annotations or schema, plus an operational remedy hint (check for a modal dialog).

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Sectioned PURPOSE/USAGE/BEHAVIOR/ON TIMEOUT/RETURNS with the core verb front-loaded, and each block carries distinct information. The RETURNS list slightly duplicates the existing output schema, which is the only mild redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-arg-complexity read tool with an output schema already present, the description covers everything an agent needs: what it returns, the required CPU state, and how to interpret every failure mode. Preconditions and error semantics are fully specified.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% for the single session_id parameter, so the schema already documents it fully. The description restates 'USAGE: session_id' and adds a state precondition, but no format, sourcing, or scoping detail beyond the schema, so baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource — 'Query GPU counters — fps, vblanks per second, timing info' — and enumerates exactly which counters. An agent can distinguish it from the sibling ppsspp_gpu_record (which captures/records rather than queries) without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives a hard precondition ('The CPU must be RUNNING') and even a non-obvious secondary use ('doubles as the cheapest paused-CPU probe'), which is real when-to-use guidance. It stops short of naming alternative tools for the same data, so it falls just below the top band.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ppsspp_healthA
Read-onlyIdempotent

PURPOSE: Probe MCP server liveness and readiness — plus an optional four-point session health battery.

USAGE: no args for the server-level probe (does NOT contact PPSSPP); pass session_id to also run the session battery (iso_loaded / cpu_running / ws_connected / game_mode_valid — absorbed from the former ppsspp_smoke_test tool).

BEHAVIOR: READ-ONLY. Server counters are read in-memory; the session battery (when requested) contacts PPSSPP over the session transport but never mutates state.

READING session_checks: each entry carries value_status besides passed -- 'ok' (the probe really read), 'stale_address_suspected' (the read succeeded and returned zero on several consecutive readings, so the probe address may have drifted -- a suspicion, not a verdict), 'failed' (the read raised or the data was absent; no value is reported), 'not_configured' (no probe address, so nothing was read). A probe that READ ZERO and one that COULD NOT READ both show passed=false while meaning opposite things: the first is a fact about the game, the second about the tooling. Do not read passed=false alone as a finding about the emulated game.

RETURNS: Dict with status ('ok'/'degraded'), version, python_version, pydantic_version, uptime_s, tool_count, session_count — plus session_checks: [{name, passed, detail, value_status, value?}] and overall_session_status when session_id is provided.

ParametersJSON Schema
NameRequiredDescriptionDefault
session_idNoOptional session ID — when provided, appends session_checks (the four-point battery: iso_loaded / cpu_running / ws_connected / game_mode_valid, absorbed from ppsspp_smoke_test) to the server-level report. Omit for the zero-contact server liveness probe.

Output Schema

ParametersJSON Schema
NameRequiredDescription
statusYesServer status: 'ok' when all subsystems healthy; 'degraded' when sessions.json is inaccessible/corrupted but server is otherwise functional.
versionYesMCP server version.
uptime_sYesServer uptime in seconds.
tool_countYesNumber of registered MCP tools.
session_countYesNumber of active sessions.
session_errorYesWhen status='degraded', describes the sessions.json issue (e.g. 'FileNotFoundError: ...' or 'JSONDecodeError: ...'). None when sessions.json is healthy.
python_versionYesPython interpreter version.
session_checksNo
pydantic_versionYesPydantic version.
overall_session_statusNo

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly/idempotent/non-destructive, so the safety profile is covered; the description nonetheless adds meaningful context: the server probe never contacts PPSSPP, the session battery contacts the transport but never mutates, and it deeply explains value_status semantics, crucially warning that passed=false can mean opposite things (read-zero vs could-not-read). This is above-and-beyond disclosure, though it stops short of e.g. counter/pagination caveats.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded sectioned structure (PURPOSE/USAGE/BEHAVIOR/READING/RETURNS) with every block earning its place; the READING section is long but carries high-value disambiguation for interpreting passed vs value_status. Slightly verbose but not wasteful.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, yet the description still summarizes the return shape (status, version, uptime, tool_count, session_checks, overall_session_status) and explains how to interpret the fields, which is exactly what an agent needs to act on the result. Nothing required to call or interpret the tool is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and the single parameter is already well documented, so baseline is 3; the description adds the enumeration of the four battery checks and clarifies that omitting session_id yields the zero-contact probe, giving more meaning than the schema alone.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource (probe server liveness/readiness plus an optional session health battery) and explicitly separates the server-level probe from the session battery, distinguishing it from siblings like ppsspp_state_observer and ppsspp_session. An agent can identify this as the health/diagnostic tool without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit routing: no args for the server-level probe (no PPSSPP contact), pass session_id to also run the four-point session battery. It also names the absorbed former tool (ppsspp_smoke_test), so an agent migrating from that tool knows where it went.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ppsspp_hold_buttonsA

PURPOSE: Hold a combination of PSP buttons until a subsequent call changes the state.

USAGE: session_id + buttons required (pipe-separated combination, e.g. 'cross|circle'). Valid names: cross / circle / triangle / square / up / down / left / right / start / select / ltrigger / rtrigger.

BEHAVIOR: STATE-CHANGE. Sets the button-held state in PPSSPP; it persists until the next hold_buttons call. To release, call hold_buttons with buttons='' (every button is then sent as released). Note: send_analog does NOT release buttons -- it drives the analog axes on a separate PPSSPP event.

RETURNS: {buttons}.

ParametersJSON Schema
NameRequiredDescriptionDefault
buttonsYesButton combination string. Multiple buttons separated by '|' (e.g. 'cross|circle'). Held until released (send an empty combination or use press_button to clear). Valid names: cross / circle / triangle / square / up / down / left / right / start / select / ltrigger / rtrigger.
session_idYesActive session ID.

Output Schema

ParametersJSON Schema
NameRequiredDescription
buttonsYesButton combination string (e.g. 'cross|circle').

TDQS

A4.1/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only mark it non-read-only and non-idempotent; the description adds the crucial state-machine semantics: the held state persists across calls until the next hold_buttons call, and must be explicitly cleared with an empty combination. It also warns that send_analog does not clear buttons, which prevents a realistic misuse. This is substantive behavior beyond the structured fields.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The PURPOSE/USAGE/BEHAVIOR/RETURNS labeling front-loads the essential fact (persistent hold) and makes scanning easy. It is slightly redundant with the 100%-covered schema, listing valid button names twice across description and schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 2-parameter state-mutation tool, the definition covers required inputs, the full button vocabulary, persistence semantics, and the release path, and return values are captured by the output schema plus the brief RETURNS line. Nothing needed to invoke this correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both parameters and the pipe-separated name list are already documented in the schema. The description repeats the same format guidance and valid name list rather than adding new semantics, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Hold a combination of PSP buttons') and scope ('until a subsequent call changes the state'), which is far more specific than a tautology. It distinguishes itself from send_analog ('send_analog does NOT release buttons'), but does not directly contrast itself with the closest sibling, ppsspp_press_button (momentary vs. persistent hold) — that contrast appears only in the schema, not the description.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear context for when to use it (persistent held state), how to release it (call with buttons=''), and an explicit exclusion (send_analog does not release buttons). It stops short of naming press_button as the alternative for one-shot presses in the description text, leaving part of the routing to the schema.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ppsspp_list_addressesA
Read-onlyIdempotent

PURPOSE: List the project's known address constants from addresses.yaml — the single source of truth; never guess hex addresses.

USAGE: optional section filter; an unknown section returns an error listing the valid ones.

CONVERSION: IDA <-> PPSSPP address conversion is plain arithmetic — ppsspp_addr = ida_addr + (top_base.ppsspp - top_base.ida) (defaults 0x08804000 - 0x00000000). ppsspp_convert_address was un-tooled in v0.1.6. BEHAVIOR: READ-ONLY. Int values ≥0x1000 are returned as hex strings that can be pasted straight into address parameters.

RETURNS: {sections, count, section_filter}.

ParametersJSON Schema
NameRequiredDescriptionDefault
sectionNoOptional section filter (e.g. 'known_functions', 'state_probes', 'top_base'). If omitted, returns all sections.

Output Schema

ParametersJSON Schema
NameRequiredDescription
countYes
sectionsYes
section_filterYes

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly/idempotent/non-destructive, so 'READ-ONLY' is largely redundant, but the description adds genuinely new behavior: ints ≥0x1000 are returned as hex strings safe to paste into address parameters, and the exact IDA↔PPSSPP conversion arithmetic (with the v0.1.6 note that convert_address no longer exists). Those are real operational facts an agent cannot get from the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Labeled PURPOSE/USAGE/CONVERSION/BEHAVIOR/RETURNS sections are front-loaded and each carries usable content; nothing is padding. It is denser than strictly needed for a one-parameter list tool — the conversion formula and version-history note are slightly tangential — but they are relevant and compact.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, and the description still summarizes the return shape ({sections, count, section_filter}) and the hex-string formatting of values. With safety covered by annotations, filtering, error behavior, conversion, and return content all addressed, an agent has everything needed to call and interpret this tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and the schema already documents the section parameter with examples, so the baseline is 3. The description adds error semantics the schema lacks: an unknown section fails with an error enumerating valid sections, which changes how an agent should recover from a bad filter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (list), a specific resource (address constants), and the authoritative source (addresses.yaml), plus the explicit anti-goal 'never guess hex addresses'. An agent can distinguish this from memory-map/scan/read-memory siblings without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear context for use (the source of truth for addresses, used before supplying address parameters) and documents the optional section filter and its error behavior. It does not explicitly name which sibling to use instead for related lookups (e.g. memory_map vs. this list), so it stops short of full routing guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ppsspp_list_scriptsA
Read-onlyIdempotent

PURPOSE: List diagnostic scripts declared in .ppsspp-dfx/config/scripts.manifest.yaml.

USAGE: category optional filter (eboot / state / p0ab / ndx / memory / misc / recipe).

BEHAVIOR: READ-ONLY. Reads the in-memory manifest registry (loaded at startup). Does not execute any script.

RETURNS: {scripts: [ScriptEntryView...], count, category}.

ParametersJSON Schema
NameRequiredDescriptionDefault
categoryNoOptional category filter. Valid values: eboot / state / p0ab / ndx / memory / misc / recipe. If omitted, all manifest entries are returned.

Output Schema

ParametersJSON Schema
NameRequiredDescription
countYesNumber of entries returned.
scriptsYesManifest entries (filtered by category if requested).
categoryYesCategory filter applied (None = no filter).

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already provide readOnlyHint, openWorldHint, idempotentHint, and destructiveHint. The description adds valuable behavioral context by stating that the tool reads the in-memory manifest registry loaded at startup and does not execute any script. This goes beyond the structured annotations and clarifies the side-effect-free nature.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact, front-loaded with purpose, and uses clear labeled sections (PURPOSE, USAGE, BEHAVIOR, RETURNS). Every sentence contributes useful information with no filler or redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a low-complexity read-only listing tool with one optional parameter and an output schema, the description is complete. It covers what the tool does, the filter options, behavioral guarantees, and the return envelope. Nothing needed for correct invocation is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the input schema fully documents the category parameter and its valid values. The description repeats this information and adds a return-shape hint, but it does not provide meaningful new semantics beyond what the schema already includes.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('List') and a precise resource (diagnostic scripts declared in .ppsspp-dfx/config/scripts.manifest.yaml). It is clearly distinct from sibling tools like ppsspp_run_script and ppsspp_reload_scripts, and the 'Does not execute any script' note reinforces what this tool is for.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The 'USAGE' section explains the optional category filter and its valid values, but it does not explicitly say when to choose this tool over siblings or when not to use it. The intended use is implied rather than stated, and no alternatives are mentioned.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ppsspp_memory_mapB
Read-onlyIdempotent

PURPOSE: Get the PPSSPP memory region map (user / kernel / VRAM ranges).

USAGE: session_id.

BEHAVIOR: READ-ONLY.

RETURNS: {ranges[], mapping, text}.

ParametersJSON Schema
NameRequiredDescriptionDefault
session_idYesActive session ID.

Output Schema

ParametersJSON Schema
NameRequiredDescription
textYesUnified text representation: one line per range, formatted as '0x{ADDR:08X}-0x{END:08X} {TYPE}/{subtype} {NAME}'.
rangesYesMemory ranges from `memory.mapping`. Each entry has 'type' (ram/vram/sram), 'subtype' (primary/mirror), 'name', 'address', and 'size'.
mappingYesRaw `memory.mapping` response, minus the volatile per-call protocol `ticket` (stripped so the business fields are byte-stable across identical calls).

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so 'BEHAVIOR: READ-ONLY' is a restatement that adds no new information. The RETURNS line hints at output shape, but an output schema already exists, so nothing about side effects, session prerequisites, or failure modes is added beyond structured data.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The PURPOSE/USAGE/BEHAVIOR/RETURNS labeling is front-loaded and scannable, and the whole description is four short lines with no prose padding. The USAGE and RETURNS lines are largely redundant with the schema and output schema, which is a small waste rather than a structural flaw.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter read tool, the safety profile (annotations), parameter (schema), and return shape (output schema) are all covered by structured fields, and the purpose line names the exact resource. The only real omission — routing guidance among the many memory-related siblings — is a usage-guideline gap rather than a completeness one.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the single parameter is documented as 'Active session ID.' in the schema; 'USAGE: session_id' in the description adds no syntax, format, or validity detail. Baseline 3 applies when the schema carries the parameter semantics.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource — 'Get the PPSSPP memory region map' — and enumerates the sub-regions (user/kernel/VRAM), so an agent knows it retrieves the address-space layout rather than bytes. It does not explicitly contrast itself with close siblings like ppsspp_read_memory or ppsspp_list_addresses, which is the only thing keeping it from a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'USAGE: session_id' merely restates the sole required parameter and gives no when-to-use, when-not-to-use, or alternative routing. An agent cannot tell from this description whether to prefer it over ppsspp_read_memory, ppsspp_list_addresses, or ppsspp_search_memory_info.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ppsspp_press_buttonA

PURPOSE: Simulate a single PSP button press for a duration.

USAGE: session_id + button required; duration optional (default 1 frame). Valid button names: cross / circle / triangle / square / up / down / left / right / start / select / ltrigger / rtrigger.

BEHAVIOR: STATE-CHANGE. Sends input events to PPSSPP. Button state returns to released after the duration elapses.

RETURNS: {button, duration}.

ParametersJSON Schema
NameRequiredDescriptionDefault
buttonYesButton name. Valid: cross / circle / triangle / square / up / down / left / right / start / select / ltrigger / rtrigger.
durationNoPress duration in frames (default 1; 60fps wall-clock, cap 18000 ≈ 300s — values above are rejected). duration=0 is accepted and is a no-op (the button is pressed and released in the same frame). The call blocks for the duration.
session_idYesActive session ID.

Output Schema

ParametersJSON Schema
NameRequiredDescription
buttonYesButton name pressed.
durationYesPress duration in frames.

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare it is a non-readonly, non-idempotent, non-destructive mutation. The description adds genuine behavioral context beyond that: it labels itself 'STATE-CHANGE', explains that the button returns to released after the duration, and the schema notes the call blocks and the 18000-frame cap. It stops short of noting anything about concurrent input or failure modes.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The PURPOSE/USAGE/BEHAVIOR/RETURNS structure front-loads the essentials and is easy to scan. The button enumeration is duplicated from the schema, which is mild redundancy, but overall there is little waste.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 3-param input tool with an output schema, annotations, and full schema coverage, the description covers purpose, required inputs, mutation semantics, and duration behavior. The extra RETURNS line is redundant against the output schema, but nothing critical is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all three parameters (including button names, duration default/cap, and blocking behavior) thoroughly. The description largely repeats the button list and default already in the schema, adding no format or syntax detail beyond it. Baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource: 'Simulate a single PSP button press for a duration.' An agent immediately knows what it does. However, it does not distinguish itself from the sibling ppsspp_hold_buttons or ppsspp_send_analog, which are plausibly related input-simulation tools, so it falls short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It clearly states the required parameters (session_id + button) and that duration is optional with a default, which is useful invocation guidance. But it gives no when-to-use guidance relative to alternatives like hold_buttons or send_analog, nor any prereq like requiring an active session. Usage is implied rather than routed.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ppsspp_queryA
Read-onlyIdempotent

PURPOSE: Aggregate game-state queries — game_state, registers (all or one), backtrace, threads, modules, and function-list management (funcs/func_scan/func_add/func_remove).

USAGE: action + session_id; 'register' needs name; func_scan/func_remove need address; top_n defaults to 100 (pass 0 for the full list — hle.func.list can reach 700+KB).

ROUTING: one-shot PC read -> query(action='register', name='pc') (safe=true pauses for consistency; safe=false for hot-path polling); pause+capture -> ppsspp_frame_snapshot; recurring named probes -> ppsspp_state_observer; game_state / backtrace / threads / modules / HLE func management also here. BEHAVIOR: READ-ONLY. Lookups only — func_add/func_remove mutate the debugger function list. Verified on a live game: threads / modules / funcs / func_scan respond while the CPU is RUNNING (no pause needed); running-state PC/isCurrent reads are LOW trust unless safe=true (which pauses briefly for a consistent, high-trust read).

RETURNS: {action, data, trust_level} — data shape depends on the action.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameNoRequired for action='register' (the register name to read) and for action='func_add'. Ignored by func_remove because PPSSPP's hle.func.remove protocol does not accept a name parameter.
safeNoaction=register/registers only: pause the CPU for a consistent read (trust_level='high', same as the retired ppsspp_get_pc) — or read without pausing (trust_level='low', zero cost, racy while running; for hot-path polling).
sizeNoFunction size in bytes, 'func_add' only. When omitted the server sends no size — on PPSSPP builds where the omit path underflows (v1.20.4-1845 and earlier) this produces an unusable zero-size function, so this tool defaults to sending 4. Pass an explicit size to override.
top_nNoLimit the number of entries returned for 'funcs' / 'func_scan' actions (default 100). 0 = no limit — hle.func.list can reach 700+KB, pass 0 only when the full list is genuinely needed.
actionYesQuery action. Valid values: - 'game_state': PPSSPP game status (paused / game title). - 'registers': all CPU registers (GPR + FPU + VFPU). - 'register': single register by name (MIPS ABI name like 'a0'/'v0'/'t9', or 'pc'/'hi'/'lo'). - 'backtrace': HLE call stack (thread optional). - 'threads': PSP thread list (safe: stepping → query → resume). - 'modules': list all loaded HLE modules. - 'funcs': list registered HLE function tracking entries. - 'func_scan': scan HLE functions in a 64KB range starting at address (requires address; CPU must be stepping). - 'func_add': add HLE function tracking (name? and/or address?). - 'func_remove': remove HLE function tracking (address required; PPSSPP protocol only accepts address, no name).
threadNoThread ID (backtrace action only; None = current).
addressNoRequired for func_remove and func_scan. Function address as a hex string (e.g. '0x08804000'). Not used by the other actions. The schema default of '0x0' exists for legacy callers -- do NOT rely on it when the action is one of the above.0x0
session_idYesActive session ID.

Output Schema

ParametersJSON Schema
NameRequiredDescription
dataYesRaw result payload.
textYesUnified text representation. Populated for action='registers' with grouped '── GPR ──' / '── FPU ──' / '── VFPU ──' headers and ' name = 0xVAL' lines, and for action='register' with a single 'name = 0xVAL' line (the requested register name echoed with its hex value). Empty for other actions (use the structured `data` field).
actionYes'game_state' / 'registers' / 'backtrace' / 'threads' / 'modules' / 'funcs' / 'func_scan' / 'func_add' / 'func_remove'.
trust_levelYesTrust annotation (threads + pc only). Lowercase enum value: 'high' (stepping-verified) / 'medium' / 'low'.

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well beyond the annotations: it discloses that threads/modules/funcs/func_scan respond while the CPU is RUNNING without a pause, that running-state PC reads are LOW trust unless safe=true, and that func_add/func_remove mutate the debugger-side function list (a nuance annotations alone would not convey). It also notes top_n=0 can produce a 700+KB payload, which is real behavioral context for an agent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with PURPOSE/USAGE/ROUTING/BEHAVIOR/RETURNS headers, so an agent can skim. It is dense but largely earns its length; a few clauses (e.g. the 'same as the retired ppsspp_get_pc' aside) restate what the annotation or schema already implies.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, the description need not document return fields, and it correctly just lists {action, data, trust_level}. Given 10 actions, 8 parameters, and an enum, the routing, prerequisites, trust-level semantics, and mutation caveats together make this complete enough to call correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is already 100%, so the baseline is 3, but the description adds genuine value: top_n's default and the rationale for pass-0 (700+KB hle.func.list), the safe flag's effect on trust_level, and which actions require name/address. It stops short of explaining the output data shape per action.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The PURPOSE line names a specific verb (aggregate queries) and enumerates the concrete resources covered — game_state, registers, backtrace, threads, modules, and HLE function-list management. Combined with the ROUTING section, an agent can distinguish this multi-action query tool from ppsspp_frame_snapshot and ppsspp_state_observer without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The ROUTING block gives explicit when-to-use-this vs when-to-use-a-sibling guidance: one-shot PC read -> query(action='register', name='pc'), pause+capture -> ppsspp_frame_snapshot, recurring probes -> ppsspp_state_observer. It also states the per-action parameter prerequisites (name for 'register', address for func_scan/func_remove) and the safe=true/false tradeoff for PC reads.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ppsspp_read_memoryA
Read-onlyIdempotent

PURPOSE: Read memory (read_bytes / read_u32 / read_string).

USAGE: action; session_id optional when exactly one session is active; address as '0x' hex string; read_bytes ≤65536 per call (split larger reads); Memory scanning has moved to ppsspp_scan.

BEHAVIOR: READ-ONLY. read_string is ASCII-only (use read_bytes + Shift-JIS decode for game text). Reading code segments: use ppsspp_disassemble — MCP provides no IR-encoding detection (a read_u32 over JIT-IR bytes just returns the raw value).

RETURNS: {action, address, value, size, text, file} — read_bytes has output=value (default; byte list + hex text) / hex (text only, value=null) / file (paths + 64-byte preview; payload saved under .ppsspp-dfx/output/memory_reads/).

ParametersJSON Schema
NameRequiredDescriptionDefault
sizeNoNumber of bytes to read (read_bytes only). Max 65536 per call (MAX_SINGLE_READ_BYTES); larger reads are rejected with ARGS_INVALID -- chunk them instead.
actionYesRead action. Valid values: - 'read_bytes': read raw bytes (requires address + size). - 'read_u32': read a 32-bit unsigned int (requires address). - 'read_string': read a string (requires address).
lengthNo(deprecated, ignored) PPSSPP memory.readString does not accept a length parameter. Kept for backward schema compatibility.
outputNoPayload channel for read_bytes (ignored by other actions). 'value' (default) returns the byte list inline plus a hex dump in `text`. 'hex' keeps only the hex dump in `text` (value=null) — roughly half the characters. 'file' saves raw bytes + hex dump under .ppsspp-dfx/output/memory_reads/ and returns the paths plus a 64-byte preview — use for reads near the 65536-byte cap.value
addressNoRequired for every action. Starting address for read_bytes/read_u32/read_string, as a hex string (e.g. '0x08804000'). The schema default of '0x0' exists for legacy callers -- do NOT rely on it.0x0
max_lenNoMaximum string length in bytes for read_string (0 = default cap 4096). Values are clamped to 65536.
session_idNoActive session ID; omit to auto-resolve when exactly one session is active.

Output Schema

ParametersJSON Schema
NameRequiredDescription
fileYesAbsolute path of the saved raw-byte file when read_bytes ran with output='file' (hex dump sits beside it as <file>.hex.txt); empty string otherwise.
sizeYesNumber of bytes read (read_bytes), or number of matches (scan). Unused for read_u32 / read_string.
textYesUnified text representation following spec conventions: '0xADDR: VAL (0xVAL_HEX)' for read_u32 (8-digit zero-padded hex), hex dump for read_bytes, repr for read_string, 'scan: N matches at 0xA1, 0xA2, ...' for scan.
valueYesRead value (int/str/list[int]/list[dict] depending on action).
actionYesRead action performed ('read_bytes'/'read_u32'/'read_string'/'scan').
addressYesStarting address, hex string (e.g. '0x08804000').

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly/idempotent/non-destructive, but the description adds real value beyond that: read_string is ASCII-only (with a workaround for game text), MCP does no IR-encoding detection so read_u32 on JIT-IR returns raw bytes, and the size cap behavior (rejected with ARGS_INVALID). This is the substantive behavioral detail the annotations cannot express.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with PURPOSE/USAGE/BEHAVIOR/RETURNS sections, each sentence carrying distinct load. The structured layout lets an agent extract the scoping, routing, and output-channel rules quickly without wasted prose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 7-parameter read tool with an output schema, the definition covers the action enum, the address/size constraints, the ASCII-only limitation, the disassembly routing, and the RETURN shape with the three output channels. An agent has everything needed to call it correctly in one pass.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the schema already documents every parameter in detail; baseline would be 3. The description still adds cross-parameter relationships not obvious from the schema alone (read_string is ASCII-only, read_bytes outputs to value/hex/file, chunking semantics for size). It doesn't add much beyond the schema, but the added interaction rules are genuinely useful.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Opens with an explicit verb+resource ('Read memory') and immediately enumerates the three concrete actions (read_bytes/read_u32/read_string). It also distinguishes itself from siblings by routing scanning to ppsspp_scan and disassembly to ppsspp_disassemble, so an agent can place it precisely among the 37 tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit when-to-use rules: session_id optional when exactly one session is active, address must be '0x' hex, reads over 65536 bytes must be chunked. It also names the alternative destination for related work ('Memory scanning has moved to ppsspp_scan', disassembly to ppsspp_disassemble), which is exactly the sibling routing a good definition provides.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ppsspp_reload_scriptsA
Idempotent

PURPOSE: Manually reload the script manifest YAML and clear the script module cache.

USAGE: No parameters.

BEHAVIOR: MUTATING. Re-reads the manifest file and invalidates cached script modules. Reversible: re-reading an unchanged file produces an equal registry. When the exposed tool set ACTUALLY changes (tools added or removed), the server notifies the client with a tool-list-changed notification so cached tools/list results are invalidated; a no-op reload sends nothing.

RETURNS: {reloaded_count, exposed_count, manifest_path, scripts: [ScriptEntryView...]}.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
scriptsYesAll entries after reload (post-reload snapshot).
exposed_addedYesScript names newly registered by this reload's sync.
exposed_countYesNumber of exposed scripts (exposed=true) after reload.
manifest_pathYesAbsolute path to the manifest YAML.
reloaded_countYesTotal script entries after reload (manifest-wide).
exposed_removedYesScript names unregistered by this reload's sync.
restart_requiredYesTrue when some registration changes could not be applied at runtime (SDK limitation) and a server restart is needed to fully reconcile exposed tools.
exposed_registeredYesNumber of exposed scripts ACTUALLY registered as dynamic tools after the reload sync (declared != registered is now surfaced instead of silently diverging).

TDQS

A4.2/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well past the annotations by disclosing that it re-reads the manifest, invalidates the module cache, is reversible because an unchanged file yields an equal registry, and that a tool-list-changed notification is emitted only when the exposed set actually changes. Annotations only say it is mutating, idempotent and non-destructive; this adds the observable side effects.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Labeled PURPOSE/USAGE/BEHAVIOR/RETURNS sections front-load the essential information, and the conditional notification detail is packed into one dense sentence without padding. Every sentence contributes.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists and the description still summarizes the return shape; annotations cover the safety profile and the description adds the cache/manifest semantics and notification behavior. Nothing an agent needs to call this mutation correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters (schema coverage 100%), and the description correctly states 'No parameters', so there is nothing further to document. Baseline 4 applies for a parameterless tool.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('reload the script manifest YAML and clear the script module cache'), which is far more than a restatement of the name. It does not, however, explicitly contrast itself with siblings like ppsspp_list_scripts or ppsspp_run_script, so an agent must infer the distinction.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'USAGE: No parameters' describes invocation mechanics, not when to reach for this tool. There is no statement of when a reload is warranted (e.g., after the manifest changes) or how it relates to list_scripts/run_script, leaving usage implied rather than guided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ppsspp_replayA

PURPOSE: Aggregate PPSSPP replay subsystem — record input sequences, execute them, and save/load .ppr recordings.

USAGE: session_id optional when exactly one session is active; actions: begin/abort/flush/execute/status/time_get/time_set/save/load/wait_complete; execute needs version + base64_input; time_set needs value; save/load take a bare file name (always under output/replays/).

BEHAVIOR: STATE-CHANGE. Recording requires the CPU RUNNING (real input timing); screenshots are rejected while recording. Replay timelines use ABSOLUTE game-clock timestamps anchored at the RECORDING session's boot — a replay only injects correctly when a fresh boot's clock is aligned to them: execute/load ONLY loads the event table and returns t0_s / estimated_end_s + the boot-aligned sequence (reset -> wait_ready -> wait boot+estimated_end_s -> abort); it does NOT play by itself. executing/saving NEVER clear on their own — only abort clears them — so wait_complete times out on any un-aborted replay; completion = the timeline estimate + explicit abort. execute/load auto-abort a live executing/saving state first. restore_rtc defaults to False: setting it rewinds the game-visible wall clock of the RUNNING session and pollutes every in-game timer; when needed, set it before the boot-aligned reset.

RETURNS: {action, executing, saving, version, size, base64, base_rtc, data} — execute/load data carries t0_s, estimated_end_s, event_count and boot_aligned_sequence; fields depend on the action.

ParametersJSON Schema
NameRequiredDescriptionDefault
valueNoRequired for action='time_set'. Base RTC value in seconds (uint32). Not used by the other actions. The schema default of 0 exists for legacy callers -- do NOT rely on it when the action is 'time_set'.
actionYesReplay operation. Valid values: - 'begin': begin/resume recording. - 'abort': abort any recording or execution. - 'flush': flush recorded data (returns version + base64). - 'execute': execute a replay (requires version + base64_input). ONLY loads the event table — follow the boot-aligned sequence in the response (reset + wait + abort) or input never injects. - 'status': query {executing, saving}. - 'time_get': get base RTC. - 'time_set': set base RTC (requires value). WARNING: rewinds the game-visible wall clock on the RUNNING session — pollutes every in-game timer. - 'save': flush + time_get + write .ppr file (requires file_path: bare file name under output/replays/). - 'load': read .ppr + execute (requires file_path; same containment). Returns t0_s / estimated_end_s and the boot-aligned sequence. - 'wait_complete': poll replay.status until executing=False — NOTE: executing never clears on its own (only abort clears it), so this always times out on an un-aborted replay; kept for recording-completion checks and backwards compatibility.
versionNoRequired for action='execute'. Replay format version (from a prior replay.flush). Not used by the other actions. The schema default of 0 exists for legacy callers -- do NOT rely on it when the action is 'execute'.
file_pathNoBare .ppr file NAME (no directory parts) for action='save' / action='load'. The file is always placed under the server-managed directory .ppsspp-dfx/output/replays/ — absolute paths and path separators are rejected. Required for save / load; ignored for all other actions.
session_idNoActive session ID; omit to auto-resolve when exactly one session is active.
timeout_msNoTotal timeout in milliseconds for action='wait_complete' (default 10000 = 10s, clamped 100..25000). The ceiling is 25s because the MCP client aborts a tool call at ~30s: a larger value could never be honoured. Ignored for all other actions.
interval_msNoPolling interval in milliseconds for action='wait_complete' (default 100ms, clamped 10..5000 — below 10 the poll degenerates to a busy loop on the WS). Ignored for all other actions.
restore_rtcNoWhether to restore base_rtc via replay.time_set before execute when action='load' (default False). true sets the game-visible wall clock back to the recording moment — pollutes EVERY timer of the running session (attract timeouts, clocks, cooldowns) because game time = rtcBaseTime + elapsed. Only use for deterministic replays, and prefer setting it BEFORE the reset of the boot-aligned sequence so the game boots on the shifted base. Ignored for all other actions.
base64_inputNoBase64-encoded replay data (from a prior replay.flush). Required for action='execute'; ignored for all other actions.
session_noteNoOptional human-readable note embedded in the .ppr file when action='save'. Ignored for all other actions.

Output Schema

ParametersJSON Schema
NameRequiredDescription
dataYesRaw PPSSPP response dict (echoed for diagnostic / future field extraction). Empty dict when no additional fields.
sizeYesRecording size in bytes from `replay.flush`. 0 when the action does not return a size.
actionYesReplay action executed: 'begin' / 'abort' / 'flush' / 'execute' / 'status' / 'time_get' / 'time_set' / 'save' / 'load' / 'wait_complete'.
base64YesBase64-encoded recording payload from `replay.flush`, or the input payload passed to `replay.execute`. Empty string when the action does not carry a payload.
savingYesTrue if a replay recording is in progress. After `begin` → True; after `flush` or `abort` → False.
versionYesRecording format version from `replay.flush` (currently 1). 0 when the action does not return a version.
base_rtcYesBase RTC timestamp (seconds) from `replay.time.get` / `replay.time.set`. 0 when the action does not return it.
executingYesTrue if a replay is currently executing. Drives `wait_complete`'s exit condition (polls until False).
wait_iterationsYesNumber of `replay.status` polls performed by `wait_complete` before exiting. 0 for non-wait actions.

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes far beyond the annotations (readOnly=false, destructive=false, idempotent=false) by disclosing that recording requires the CPU RUNNING, that screenshots are rejected while recording, that timelines are absolute game-clock anchored at the recording boot, that execute/load only load the event table and must be followed by the boot-aligned reset/wait/abort sequence, and that executing/saving never self-clear so wait_complete always times out on un-aborted replays.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded PURPOSE/USAGE/BEHAVIOR/RETURNS structure makes a dense description scannable, and nearly every clause carries actionable detail (timing anchor, auto-abort, restore_rtc pollution). It is long and occasionally repetitive with the schema, but not padded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a stateful 10-parameter replay tool with an output schema present, the description covers the full lifecycle semantics an agent needs: action prerequisites, cross-action state interactions, timing anchoring, and which return fields appear per action. Nothing material is left unstated.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is already 100%, so the baseline is 3, but the description adds cross-parameter context the schema treats piecewise: which fields pair with which action, that the schema defaults for value/version exist for legacy callers and must not be relied on, and the ordering constraint on restore_rtc relative to the reset.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb set and resource: record input sequences, execute them, save/load .ppr recordings in the PPSSPP replay subsystem. This is clearly distinct from siblings like ppsspp_press_button or ppsspp_wait_frames, which provide inputs/frames rather than aggregate replay lifecycle control.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly enumerates the ten actions and their per-action prerequisites (execute needs version + base64_input, time_set needs value, save/load take a bare file name under output/replays/, session_id optional when exactly one session is active). It also states when not to rely on defaults and warns when restore_rtc is needed and when it must be set relative to the boot-aligned reset.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ppsspp_run_scriptA

PURPOSE: Invoke a manifest-registered diagnostic script by name with validated input.

USAGE: name (see ppsspp_list_scripts; skeleton scripts return not_implemented); input dict validated against the script's Pydantic model; session_id required when the script declares requires_ppsspp (missing → SESSION_NOT_FOUND).

BEHAVIOR: STATE-CHANGE. Runs manifest-registered script code. Unknown names → SCRIPT_NOT_FOUND.

RETURNS: {name, output, output_model}.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYesScript name (must appear in manifest).
inputNoScript input as a JSON dict. Validated against the script's Pydantic Input model. Pass {} for scripts with no required fields.
session_idNoOptional session ID. Enforced by this tool for scripts with requires_ppsspp=true (fails with SESSION_NOT_FOUND when none resolves). Priority: this parameter > an optional session_id field on the script's Input model.

Output Schema

ParametersJSON Schema
NameRequiredDescription
nameYesScript name that was executed.
outputYesScript output (serialized Pydantic Output model).
output_modelYesPydantic Output model class name (for type introspection).

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description explicitly declares 'STATE-CHANGE', matching the annotations (readOnlyHint=false, idempotentHint=false). It adds behavioral context beyond annotations by naming error codes (SCRIPT_NOT_FOUND, SESSION_NOT_FOUND) and noting that skeleton scripts are non-functional. While it doesn't detail what state changes occur, the description covers the critical edge cases and the fact that it runs arbitrary registered code.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is exceptionally well-structured with clear PURPOSE, USAGE, BEHAVIOR, and RETURNS sections. It is concise, with every sentence providing necessary information and no filler. The front-loaded purpose statement immediately tells the agent what the tool does.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool that runs scripts with three parameters, an output schema, and a rich set of sibling tools, the description is complete. It covers how to obtain valid names, handles the session_id requirement, describes error behavior, and specifies the return shape. The agent has everything needed to call it correctly without additional lookups.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all three parameters (name, input, session_id) are already well-documented in the schema. The description adds marginal value by referencing ppsspp_list_scripts for name discovery and clarifying the session_id priority, but these are largely restatements of schema hints. The baseline of 3 is appropriate because the schema carries the load, and the description does not meaningfully enhance parameter meaning beyond that.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Invoke') and resource ('manifest-registered diagnostic script') with validated input, clearly distinguishing it from siblings like ppsspp_list_scripts (which lists scripts) and ppsspp_reload_scripts (which reloads). It even notes that skeleton scripts return not_implemented, further disambiguating expected behavior.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The USAGE section explicitly tells the agent to get the script name from ppsspp_list_scripts, mentions that skeleton scripts return not_implemented, and clarifies the session_id requirement with the error SESSION_NOT_FOUND when missing. It also implies when not to use the tool (for skeleton scripts) without being verbose.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ppsspp_scanA
Read-only

PURPOSE: Three-mode memory scanner — byte-pattern search, Cheat-Engine-style value scan with narrowing sessions, and charset-aware string harvesting.

USAGE: mode='pattern' + pattern + start_addr/end_addr (migrated from read_memory scan); mode='value' + phase='initial'/value/width → handle, then phase='narrow'/op/value to converge, 'list'/'drop' to manage; mode='strings' + charset + start_addr/end_addr → [{address, text}]. background=true submits a detached job instead (recommended for full-band scans) and returns {action:'submitted', batch_id, ...} — poll ppsspp_batch_status(batch_id=...).

BEHAVIOR: READ-ONLY. Large ranges are read across multiple reads; unreadable regions are skipped per-chunk (one WS round-trip each), but CONSECUTIVE read timeouts (10 s each, >5 in a row) abort the scan — a wedged PPSSPP fails the scan cleanly instead of pinning the session lock. pattern/strings ranges over 2 MiB are AUTO-BACKGROUNDED (returns {action:'submitted', batch_id, ...} even with background=false) — measured: 24 MB @ 4 KiB chunks takes 53-96 s depending on PPSSPP build, always past the ~30s client timeout, while @ 64 KiB chunks it is 3.4-40 s (build-dependent). Value sessions live in a bounded per-server registry (cap 4, FIFO), are bound to the creating session, and the initial-scan cap is 8 MiB foreground / 32 MiB background. Background scans carry a 600 s wall-clock budget; exceeding it fails the job and releases the session. The registry is process-global: parallel sessions share one cap and FIFO order, so another session's scans can evict your handle under load.

ROUTING: what-changed-between-two-points -> ppsspp_diff_memory (snapshots); who-accesses-this-address -> ppsspp_breakpoint(action='trace'); value candidates with known addresses -> read_memory directly.

RETURNS: pattern → {action, address, value: [matches], size}; value initial → {scan_handle, width, candidates, passes}; value narrow → {scan_handle, candidates, passes}; value list → {scan_handle, addresses: [...]}; value drop → {scan_handle, dropped}; strings → {charset, count, strings: [{address, text}]}; background submission (explicit background=true OR pattern/strings range > 2 MiB) → {action: 'submitted', batch_id, session_id, estimated_s}.

ParametersJSON Schema
NameRequiredDescriptionDefault
opNoComparison for value scans (initial + narrow; default eq).eq
modeYesScan mode: - 'pattern': byte-pattern search (hex/ascii) over a range — migrated from read_memory(action=scan). - 'value': Cheat-Engine-style value scan with narrowing sessions (phase: initial → narrow → list → drop; width u8/u16/u32, op eq/ne/lt/gt). - 'strings': charset-aware string harvesting (charset shift_jis/utf8/ascii, min_len, quality filter) — returns [{address, text}].
phaseNoValue-scan phase (value mode): 'initial' scans the range for `value`; 'narrow' re-reads candidates and filters by `op`+`value` (requires explicit session_id; auto-resolve not supported for this phase); 'list' returns current candidates; 'drop' releases the session.
valueNoValue to scan/narrow for (value mode).
widthNoValue width (value mode; default u16).u16
charsetNoString charset (strings mode; default shift_jis).shift_jis
min_lenNoMinimum string length (strings mode; default 6).
patternNoPattern to scan for (pattern mode). Interpreted per pattern_type: 'hex' (default, e.g. 'AABBCCDD') or 'ascii'.
qualityNoCJK-ratio quality floor for shift_jis (strings mode; 0..1, default 0.2; 0 disables). Random bytes can chance-decode to kana — the filter keeps signal. ascii/utf8 have no quality filter — expect noise in code regions.
end_addrNoRange end, exclusive, hex string (same format as `address`).
backgroundNoRun as a detached background job: returns a batch_id immediately; poll ppsspp_batch_status(batch_id=...), cancel via ppsspp_batch_cancel. Value initial cap lifts 8 MiB → 32 MiB in background mode. NOTE: pattern/strings ranges over 2 MiB are auto-backgrounded even when this is false — a foreground scan that outlives the ~30s client timeout is the classic 'frozen session' trap.
chunk_sizeNoBytes per read request during chunked scans (default 65536 — measured ~6x faster end-to-end than the old 4096 default; clamped to [64, 65536]).
session_idNoActive session ID; auto-resolved when exactly one session is active. Required for the pattern / value initial / strings phases, which start a new scan. The narrow / list / drop phases only re-read addresses already recorded by an earlier phase, so they do not need it passed -- but it must still resolve to the same session.
start_addrNoRange start, inclusive, hex string (same format as `address`).
max_resultsNoMaximum number of matches (pattern mode, default 100).
scan_handleNoValue-scan session handle (narrow/list/drop phases).
pattern_typeNoHow to interpret `pattern` (pattern mode). 'hex' (default) or 'ascii'.hex

Output Schema

ParametersJSON Schema
NameRequiredDescription
hintNo
modeNoScan mode: pattern / value / strings.
sizeNoMatch count (pattern mode).
countNoHit count (pattern & strings modes; matches the value/strings list in this response).
valueNoMatches (pattern mode).
widthNoValue width (u8/u16/u32).
actionNo
passesNoCompleted passes (value narrow).
addressNoScan start (pattern mode).
charsetNoCharset used (strings mode).
droppedNoTrue when the session was dropped.
stringsNoHarvested strings (strings mode).
batch_idNo
addressesNoCandidate addresses (value list).
truncatedNoTrue when the strings hit cap was reached and remaining matches were dropped (narrow the range or raise min_len).
candidatesNoCandidate count (value initial/narrow).
session_idNo
estimated_sNo
scan_handleNoValue-scan session handle.

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare readOnly/destructive/idempotent; the description adds substantive behavior the annotations cannot carry — consecutive-timeout abort rules, per-chunk skipping of unreadable regions, auto-backgrounding above 2 MiB, a FIFO registry cap of 4 shared process-globally (other sessions can evict your handle), an 8/32 MiB initial-scan cap, and a 600 s job budget. This is exactly the operational context an agent needs before committing to a long scan.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The DESCRIPTION is front-loaded and sectioned (PURPOSE/USAGE/BEHAVIOR/ROUTING/RETURNS), so an agent can stop reading early. It is long, and the auto-backgrounding rule is repeated in USAGE, BEHAVIOR, and the background parameter, which is mild redundancy for a 17-parameter tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a three-mode, 17-parameter tool with concurrency and timeout hazards, the description covers mode selection, session lifecycle, failure modes, and background handoff. An output schema exists, so the RETURNS section is a convenience rather than a necessity, and nothing an agent needs to invoke this correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3; the description goes further by specifying cross-parameter contracts (which params matter per mode, that narrow/list/drop must resolve to the same creating session, that the initial cap lifts under background=true). It doesn't restate hex-format syntax or enum values already in the schema, which is the right restraint.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The PURPOSE line names a specific verb+resource split into three concrete modes (byte-pattern search, value scan with narrowing, charset-aware string harvesting), which is far more than a restatement of the name. Combined with the ROUTING section naming siblings (ppsspp_diff_memory, ppsspp_breakpoint, read_memory), an agent can place this tool precisely in the family.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The USAGE block gives explicit parameter combinations per mode and the phase workflow (initial → narrow → list/drop), and the ROUTING block states when to use a sibling instead ('what-changed-between-two-points -> ppsspp_diff_memory', 'who-accesses-this-address -> ppsspp_breakpoint', 'known addresses -> read_memory directly'). When-not guidance is explicit, including the 'frozen session' trap.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ppsspp_screenshotA
Read-onlyIdempotent

PURPOSE: Capture the framebuffer as an image (ImageContent) plus metadata.

USAGE: session_id optional when exactly one session is active; source='render' (default; empty frames auto-fall back to VRAM — colors unreliable there) or 'output' (CRASH-RISK, do not use); mutually exclusive with the deprecated mode param.

BEHAVIOR: READ-ONLY. An empty capture returns empty=true instead of an error — advance to a rendered scene and retry.

RETURNS: structuredContent metadata (mode/source/size_bytes/width/height/file_path/format/empty); the image itself arrives as an ImageContent block. The auto-saved PNG/JPG path is in file_path.

ParametersJSON Schema
NameRequiredDescriptionDefault
modeNoDEPRECATED — use `source` instead. Legacy Win32 / VRAM fallback paths. 'auto' = three-tier fallback (wm_command → printwindow → vram). 'wm_command' / 'printwindow' / 'vram' = the specific strategy. Mutually exclusive with `source`.
sourceNoCapture source (new, preferred). 'render' = with_stepping + gpu.buffer.renderColor (default when neither source nor mode is given). 'output' = gpu.buffer.screenshot (CRASH-RISK on some games). Mutually exclusive with `mode`.
session_idNoActive session ID; omit to auto-resolve when exactly one session is active.

Output Schema

ParametersJSON Schema
NameRequiredDescription
modeYesCapture mode used. Echoes the caller's `mode` value ('auto'/'wm_command'/'printwindow'/'vram') on the deprecated path, or the `source` value ('render'/'output') on the new path.
emptyYesTrue when the capture produced no pixels (size_bytes=0 and no ImageContent). Agents can branch on this instead of parsing size_bytes heuristics — mirrors dump_texture's CAPTURE_EMPTY error.
widthYesImage width in pixels (0 if unknown).
formatYesImage format ('png' or 'jpeg').
heightYesImage height in pixels (0 if unknown).
sourceYesThe `source` parameter value when the new path was taken, or None when the deprecated `mode` path was used.
file_pathYesPath where image was saved.
size_bytesYesDecoded image size in bytes.

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the readOnly/idempotent/destructive annotations, it discloses the render-to-VRAM auto-fallback, unreliable colors on VRAM, crash risk of 'output', and the empty=true instead of error behavior. These are precisely the behavioral traits that affect invocation and interpretation of results.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Organized into PURPOSE, USAGE, BEHAVIOR, and RETURNS with each sentence carrying distinct information and no filler. The critical caveats are near the top and easy to scan.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With annotations covering safety and an output schema covering metadata, the description adds exactly what is needed: purpose, source selection rules, empty-capture semantics, and how the image is returned and where it is saved. No critical invocation detail is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the schema already documents the parameters, but the description adds operationally important semantics: source default plus fallback behavior, the 'output' risk warning, mode deprecation and exclusivity, and the single-active-session condition for omitting session_id.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb, resource, and output: captures the framebuffer as an ImageContent plus metadata. The purpose is clear, but it does not explicitly distinguish itself from close siblings such as ppsspp_frame_snapshot or ppsspp_dump_texture.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides actionable usage conditions: session_id may be omitted when exactly one session is active, source='render' is the safe default, and source='output' is explicitly marked crash-risk and not to be used. It explains mutual exclusivity with the deprecated mode param, though it does not name alternative tools for when this tool should be avoided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ppsspp_search_disasmA
Read-onlyIdempotent

PURPOSE: Loop-search disassembly for a substring, collecting matching instructions with context.

USAGE: session_id + match (a leading '$' is stripped); start address; end=0 wraps the search around the whole region; max_results default 100.

BEHAVIOR: READ-ONLY.

RETURNS: {address, match, end, results[{address, text, name, params}], text}.

ParametersJSON Schema
NameRequiredDescriptionDefault
endNoEnd address for the search range, as a hex string (e.g. '0x08810000'). If 0 or equal to `address`, PPSSPP performs a loop search (wraps around memory).0x0
matchYesCase-insensitive substring to search for in the disassembly text (e.g., 'jal', 'addiu', 'lw r5'). May include '$' register prefix (e.g. 'jr $ra') — automatically stripped before forwarding to PPSSPP (PPSSPP register names have no '$' prefix). Required by PPSSPP's memory.searchDisasm event.
addressYesStarting address for the disassembly search, as a hex string (e.g. '0x08804000').
session_idYesActive session ID.
max_resultsNoMaximum number of matches to collect (default 100). PPSSPP returns only the first match per call; the tool loops from each match address+4 until no more matches, this cap is reached, or a loop is detected.

Output Schema

ParametersJSON Schema
NameRequiredDescription
endYesEnd address, hex string. '0x00000000' or equal to `address` means loop search.
textYesUnified multi-line text representation. Each line is '0x{ADDR:08X}: {text}'.
matchYesCase-insensitive substring matched.
addressYesStarting address for the search, hex string (e.g. '0x08804000').
resultsYesList of disasm line dicts (address / text / name / params).

TDQS

A4.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, and the description repeats 'READ-ONLY.' It also discloses loop/wrap behavior and match collection, but these details are already in the schema's parameter descriptions, so the description adds little new behavioral context beyond the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description uses four labeled one-line sections (PURPOSE, USAGE, BEHAVIOR, RETURNS) and front-loads the purpose. Every line earns its place, and there is no fluff or redundant elaboration.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With the detailed input schema covering all parameters, the read-only/idempotent annotations, and an output schema present, the description gives enough operational context including defaults, wrap behavior, and return envelope. An agent has all needed information to select and invoke this tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and every parameter already has thorough semantic explanation (e.g., end=0 wrap behavior, '$' stripping, max_results loop mechanism). The description's USAGE line merely summarizes these details without adding further meaning, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Loop-search disassembly for a substring, collecting matching instructions with context.' This makes the tool's function immediately clear and inherently differentiates it from siblings like ppsspp_disassemble (one-shot disassembly) and ppsspp_memory_info_search (searching memory values rather than disassembly text).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The USAGE section provides concrete invocation guidance: required session_id and match, leading '$' stripping, start address, end=0 wrap-around, and max_results default. It does not explicitly name alternatives or state when not to use this tool, but the context is clear enough to proceed.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ppsspp_search_memory_infoA
Read-onlyIdempotent

PURPOSE: Search PPSSPP's memory-tracking metadata for allocation/texture tags matching a string.

USAGE: session_id + match (case-insensitive substring, required); optional address/end/type filters.

BEHAVIOR: READ-ONLY. Returns a single extent per matching tag.

RETURNS: {regions[], count, raw, text}.

ParametersJSON Schema
NameRequiredDescriptionDefault
endNoOptional end address for the search range, as a hex string (e.g. '0x08810000'). If omitted, PPSSPP searches to the end of the address space.
typeNoOptional type filter (e.g., 'texture', 'vertex'). If omitted, all region types are returned.
matchYesCase-insensitive substring to match against memory region tags (e.g., 'texture', 'vertex', 'framebuf').
addressNoOptional start address for the search range, as a hex string (e.g. '0x08804000'). If omitted, PPSSPP searches the entire address space.
session_idYesActive session ID.

Output Schema

ParametersJSON Schema
NameRequiredDescription
rawYesRaw `memory.info.search` response from PPSSPP.
textYesUnified text representation: one line per region, formatted as '0x{ADDR:08X}-0x{END:08X} {TYPE} {TAG}'.
countYesNumber of regions in the result.
regionsYesMatching memory regions. Each entry has type / address / size / ticks / pc / tag / allocated (field presence depends on PPSSPP version).

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The 'READ-ONLY' statement merely restates readOnlyHint/destructiveHint=false already in annotations, and the RETURNS shape is largely covered by the output schema. The one genuinely additive detail is 'Returns a single extent per matching tag', which clarifies result granularity, but permissions, pagination/limits, and failure modes are unaddressed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four labeled sections, front-loaded PURPOSE, zero filler sentences. The agent can parse purpose, inputs, safety, and return shape at a glance.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With annotations carrying the safety profile, a 100%-covered schema, and an output schema describing the {regions[], count, raw, text} return, the description covers what an agent needs to invoke correctly. Only cross-tool routing guidance is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all five parameters (including hex-string address/end semantics and the type filter examples) are already documented in the schema. The description only repeats the case-insensitive substring nature of 'match' and the optionality of the filters, adding nothing new.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: searching PPSSPP's memory-tracking metadata for allocation/texture tags. This is distinguishable from siblings like ppsspp_search_disasm (disassembly) and ppsspp_scan (memory scanning), though no sibling is named explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The USAGE line identifies the required inputs (session_id + match) and the optional filters, which implies when the tool applies. However, it gives no guidance on when to prefer this over ppsspp_search_disasm, ppsspp_scan, or ppsspp_read_memory, and no exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ppsspp_send_analogA

PURPOSE: Send an analog stick position (x, y in [0, 255], 128 = center).

USAGE: session_id + x + y required. 0 = full left / up, 255 = full right / down.

BEHAVIOR: STATE-CHANGE. Sets analog stick position; persists until next send_analog call.

RETURNS: {x, y}.

ParametersJSON Schema
NameRequiredDescriptionDefault
xYesAnalog X coordinate in [0, 255] (128 = center). 0 = full left, 255 = full right.
yYesAnalog Y coordinate in [0, 255] (128 = center). 0 = full up, 255 = full down.
session_idYesActive session ID.

Output Schema

ParametersJSON Schema
NameRequiredDescription
xYesX coordinate in [0, 255] (128 = center).
yYesY coordinate in [0, 255] (128 = center).

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The BEHAVIOR section explicitly labels this as a STATE-CHANGE operation, states that the position persists until the next send_analog call, and notes that it is not idempotent in effect. This adds meaningful context beyond the annotations (which only say readOnlyHint=false, idempotentHint=false) by explaining the persistence semantics.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and well-structured with clear PURPOSE, USAGE, BEHAVIOR, and RETURNS sections. Every sentence earns its place, and the most important information (what the tool does and required parameters) is front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers purpose, usage, behavior, and return value. With an output schema present and full parameter schema coverage, nothing critical is missing. It could slightly improve by noting whether the analog position resets on session end, but this is a minor gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already fully documents all three parameters. The description adds a concise summary of the coordinate system but does not add new information beyond what the schema provides. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Send') and resource ('analog stick position'), and precisely defines the coordinate semantics (x, y in [0, 255], 128 = center). It clearly distinguishes this from sibling input tools like ppsspp_press_button and ppsspp_hold_buttons by focusing on analog state rather than discrete button presses.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The USAGE section explicitly lists required parameters (session_id + x + y) and explains the value mapping (0 = full left/up, 255 = full right/down). It does not explicitly name alternative tools or when-not-to-use conditions, but the purpose is specific enough that an agent can infer when to use it versus siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ppsspp_sessionA

PURPOSE: Start / stop / inspect PPSSPP debug sessions — action=list / start / stop / get / wait_ready; wait_ready blocks until the emulated CPU is up.

USAGE: action='list' takes no other params and returns {sessions, count} (idle sessions >30min are auto-GC'd as a side effect; NOT a per-session health probe — use ppsspp_health(session_id=...) for that); action='start' needs iso_path (pass wait_ready=true to block until the CPU is up in the same call); stop/get/wait_ready need session_id. Call wait_ready AFTER start and BEFORE any memory tool — PPSSPP answers WebSocket before the CPU boots. start(resilient=true) self-heals boot wedges (blacklist quarantine + relaunch with the same session_id, ≤2 retries).

BEHAVIOR: STATE-CHANGE. start spawns a PPSSPP subprocess + WS debugger; stop terminates it (never taskkill the process yourself); wait_ready polls the probe lock-free and fails [BOOT_TIMEOUT] on wedge suspicion; list/get are read-only.

RETURNS: action=start/get → SessionResponse {session_id, iso_path, pid, ws_url, created_at, last_active_at, exec_count, ws_connected, recovered, restored (1 = session record was restored from sessions.json after a server restart, 0 = created in this process), ppsspp_version}; action=wait_ready → {action, ready, elapsed_s, probe_addr, probe_value, note} — elapsed_s is THIS call's wait duration, not the start→ready total; action=list → {sessions: [SessionResponse...], count}.

ParametersJSON Schema
NameRequiredDescriptionDefault
actionYesSession operation. Valid values: - 'list': list all active sessions (no other params). Idle sessions (>30 min) are auto-GC'd as a side effect; returns {sessions, count}. NOT a per-session health probe — use ppsspp_health(session_id=…) for that. - 'start': launch a new PPSSPP session (requires iso_path). Set wait_ready=true to block until the emulated CPU is up (same probe/budget semantics as 'wait_ready'). - 'stop': terminate an existing session (requires session_id). - 'get': query session health (requires session_id). - 'wait_ready': block until the emulated CPU has started (requires session_id). Call this after 'start' BEFORE any memory/disassembly tool — PPSSPP answers WebSocket before the ISO finishes booting, and early reads fail with 'CPU not started'.
iso_pathNoAbsolute path to the ISO file (required when action=start).
resilientNoaction=start only: self-healing boot — on wedge evidence (CPU-ready probe exhausted, handshake never accepted, process died) the launcher is torn down, the GPU-backend failure blacklist is quarantined (rename), and the session relaunches with the SAME session_id up to 2 retries; the response carries recovered=N (0 = first launch). Exhaustion raises [BOOT_TIMEOUT]. Ignored in fake mode.
timeout_sNoBoot budget in seconds (action=wait_ready, action=start with wait_ready=true, or the per-attempt CPU-ready budget when action=start with resilient=true; default 75, clamped to [1, 300]).
probe_addrNoHex address polled by the readiness probe (action=wait_ready, action=start with wait_ready=true, or the resilient-start gate; default '0x08804000', the project's top.prx load base).0x08804000
session_idNoSession ID (required when action=stop / get / wait_ready).
wait_readyNoaction=start only: block until the emulated CPU is ready before returning (same probe/budget as action=wait_ready; raises [BOOT_TIMEOUT] on wedge suspicion). Fake test mode is ready immediately. Default true: start() returns a session that is ready to use, so the first tool call does not fail with a version-handshake timeout. Pass false only when you want the raw launch without waiting.

Output Schema

ParametersJSON Schema
NameRequiredDescription
pidNoPPSSPP process PID (None if stopped).
noteNoOptional human context (e.g. fake-mode short-circuit).
countNoNumber of sessions.
readyNoTrue when the CPU-start probe succeeded.
actionNoLiteral 'wait_ready' (echoes the session action).
ws_urlNoWebSocket URL (ws://host:port/debugger).
iso_pathNoAbsolute path to the ISO file.
restoredNo1 when this session was restored from sessions.json (a previous server run left it behind) rather than started fresh in this process — its game state may be stale.
sessionsNoActive sessions.
elapsed_sNoWall-clock seconds spent polling.
recoveredNoresilient-start relaunch count (0 = the first launch succeeded; >0 means the game state was reset by a wedge heal — breakpoints need re-arming).
created_atNoISO 8601 timestamp of session creation.
exec_countNoNumber of tool calls made against this session.
probe_addrNoPolled address, hex string (default top.prx base).
session_idNoSession UUID-like identifier.
probe_valueNou32 read at probe_addr once ready, hex string (None in fake mode).
ws_connectedNoTrue if WebSocket is currently connected.
last_active_atNoISO 8601 timestamp of last tool call.
ppsspp_versionNoPPSSPP build fingerprint captured from the version handshake (None until the session transport binds).

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well beyond the annotations: declares START spawns a subprocess + WS debugger, STOP terminates it with an explicit 'never taskkill the process yourself' warning, wait_ready's lock-free polling and [BOOT_TIMEOUT] failure mode, the 30-min idle GC side effect, and resilient self-heal behavior. None of this is recoverable from the readOnlyHint=false/destructiveHint=false hints alone.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Labeled PURPOSE/USAGE/BEHAVIOR/RETURNS sections are front-loaded and scannable. Some content (the RETURNS breakdown) is duplicated by the schema and output schema, so a few sentences do not strictly earn their place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 7-parameter, multi-action tool with no idempotency and open-world behavior, the description covers triggers, ordering, failure modes, and state semantics. An output schema exists, yet the description pre-explains the return shapes, leaving no gap an agent would need to fill.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 100% schema coverage the schema already documents each parameter, so the baseline is 3, but the description adds cross-parameter semantics the schema does not: which action consumes which argument, and the resilient-start reinterpretation of timeout_s as a per-attempt budget. It restates some field meanings rather than extending all of them, keeping it off a 5.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The PURPOSE line names a specific resource (PPSSPP debug sessions) and enumerates the exact verb set (list/start/stop/get/wait_ready), so the agent knows the tool's scope instantly. It also explicitly differentiates from the sibling ppsspp_health for per-session health probing.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

USAGE spells out per-action parameter requirements and a hard ordering rule ('Call wait_ready AFTER start and BEFORE any memory tool') with the reason (WS answers before CPU boots). It names the alternative (ppsspp_health) and the fallback (pass wait_ready=true on start) rather than leaving routing to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ppsspp_state_observerA

PURPOSE: Named memory-probe registry plus running-state observation — register probes once, then sample them cheaply every loop.

USAGE: action + session_id for observe; register needs name + address (+size 1/2/4, description); observe takes comma-separated names and samples.

ROUTING: recurring sampled probes across loops -> here; one-shot paused snapshot -> ppsspp_frame_snapshot; single-address access watch -> ppsspp_breakpoint(action="trace"). BEHAVIOR: STATE-CHANGE. register/clear mutate the registry; observe is reliable while RUNNING. The registry is PER-SESSION (keyed by session_id), seeded from addresses.yaml state_probes; clear removes user-registered probes only, so the configured baseline survives. Delete semantics are IDEMPOTENT: clearing an unknown probe name succeeds (ok), unlike ppsspp_breakpoint mem_remove which rejects missing targets.

RETURNS: {registered|probes|observations, count, success_count, failure_count} — shape depends on the action.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameNoProbe name. Required for action='register'; optional for action='observe' (comma-separated names; omit to observe all registered probes). Ignored for list / clear.
sizeNoRead width in bytes (1 = u8, 2 = u16, 4 = u32). Default 4. Used by action='register'. Ignored for all other actions (probe's stored size is used at observe time).
namesNoComma-separated probe names for action='observe'. If empty, all registered probes are observed. Ignored for all other actions.
actionYesObserver operation. Valid values: - 'register': add a probe to the runtime registry (requires name + address; optional size default 4, optional description). - 'list': list all registered probes. - 'observe': read current value of named probe(s) (optional names — omit to observe all); optional samples (default 1) for multi-sample median. - 'clear': clear the runtime registry.
addressNoRequired for action='register'. Absolute runtime address to read, as a hex string (e.g. '0x08804000'). Not used by the other actions. The schema default of '0x0' exists for legacy callers -- do NOT rely on it when the action is 'register'.0x0
samplesNoNumber of samples to take per probe for action='observe' (default 1). If >1, samples are taken with a short yield between reads; the final value is the last read (caller can inspect stability by comparing samples externally).
session_idYesActive session ID.
descriptionNoOptional human-readable note for action='register'.

Output Schema

ParametersJSON Schema
NameRequiredDescription
dataYesRaw PPSSPP echo (reserved)
countYesProbe count (register/list/clear) or observation count (observe)
actionYesObserver action executed
probesYesAll probes (action=list only)
registeredYesProbe added (action=register only)
observationsYesPer-probe readings (action=observe only)
failure_countYesFailed observations (action=observe only)
success_countYesSuccessful observations (action=observe only)

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the annotations, the description discloses that register/clear mutate state, the registry is PER-SESSION keyed by session_id, it is seeded from addresses.yaml, and clear removes only user-registered probes so the configured baseline survives. It even contrasts delete semantics with ppsspp_breakpoint mem_remove. The one nuance is that the description calls clear idempotent while idempotentHint=false applies to the whole multi-action tool — a refinement rather than a hard contradiction, since register is not idempotent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Labeled sections (PURPOSE/USAGE/ROUTING/BEHAVIOR/RETURNS) make it front-loaded and scannable, and every section carries non-redundant information. It is somewhat long for a single tool, and the RETURNS line partly duplicates the existing output schema, keeping it just below a 5.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a high-complexity tool (8 params, 4 actions) the description covers the action model, per-session registry lifetime, mutation semantics, idempotent delete behavior, and the return shape — with an output schema and annotations also present. Nothing an agent needs to invoke it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents name, size, names, samples, address and the action enum in detail. The description restates the action/parameter pairing ('register needs name + address (+size 1/2/4, description); observe takes comma-separated names') but adds no syntax or format meaning beyond what the schema provides. Baseline 3 is correct.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb+resource pairing — 'Named memory-probe registry plus running-state observation' — and immediately distinguishes the two modes (register once, sample cheaply). The ROUTING section names the exact siblings it is not (ppsspp_frame_snapshot, ppsspp_breakpoint trace), so an agent can separate this from adjacent memory tools without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

ROUTING gives explicit when-to-use-this-vs-alternatives: recurring sampled probes across loops go here, one-shot paused snapshots go to ppsspp_frame_snapshot, single-address access watches go to ppsspp_breakpoint(action='trace'). The USAGE line also spells out which parameters each action requires.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ppsspp_stepA

PURPOSE: Aggregate CPU run-state control (pause / resume / reset / run_until / next_hle).

USAGE: action='pause' / 'resume' / 'reset' / 'next_hle' take only session_id (optional when exactly one session is active); 'run_until' requires address.

NOTE (v0.1.6): single-stepping (into/over/out) moved to ppsspp_batch_step's cpu_step step type — this tool no longer accepts those actions.

ROUTING: single run-state operations -> here (run_until for run-to-address); multi-step press/wait/probe sequences and cpu_step -> ppsspp_batch_step. BEHAVIOR: STATE-CHANGE. Advances or changes CPU run state. 'reset' reboots the game (lost in-memory state). 'run_until' sets a temp breakpoint and resumes.

RETURNS: {action, address, pc, ticks, reason, related_address}. pc is stepping-verified (HIGH trust) for 'pause'; for 'resume' pc/ticks are a best-effort LOW-trust cpu.status snapshot of the running CPU (0/0.0 if that read failed); 'reset' reports 0.

ParametersJSON Schema
NameRequiredDescriptionDefault
actionYesCPU step / run-state operation. Valid values: - 'pause': pause CPU (enter stepping mode). - 'resume': resume CPU (exit stepping mode); the response pc/ticks are a LOW-trust snapshot of the running CPU (inaccurate unless stepping), not a precise resume location. - 'reset': reset the game (reboot). - 'run_until': run until the specified address is reached (requires address). - 'next_hle': step to next HLE callback. NOTE: single-stepping (into/over/out) lives in ppsspp_batch_step as the 'cpu_step' step type (mode='into'|'over'|'out', count 1..1000) — it requires the CPU to enter stepping mode, which the executor handles automatically.
addressNoRequired for action='run_until'. Target address, as a hex string (e.g. '0x08804000'). Not used by the other actions. The schema default of '0x0' exists for legacy callers -- do NOT rely on it when the action is 'run_until'. run_until is fire-and-forget: it returns immediately with no hit confirmation — poll PC or set a breakpoint to observe arrival.0x0
session_idNoActive session ID; omit to auto-resolve when exactly one session is active.

Output Schema

ParametersJSON Schema
NameRequiredDescription
pcYesProgram counter after step, hex string. For into/over/out/run_until/next_hle this comes from the cpu.stepping broadcast. For pause, extracted via safe_get_pc after CPU enters stepping. For resume, a best-effort LOW-trust cpu.status snapshot of the running CPU (inaccurate unless stepping — CPUCoreSubscriber.cpp:105). '0x00000000' for reset.
ticksYesCPU ticks at step completion (from cpu.stepping broadcast). For resume, a best-effort LOW-trust cpu.status snapshot. 0.0 for pause/reset.
actionYes'into' / 'over' / 'out' / 'pause' / 'resume' / 'reset' / 'run_until' / 'next_hle'.
reasonYesStep reason from cpu.stepping broadcast (e.g. 'cpu.stepInto'). Empty for pause/resume/reset.
addressYesTarget address for run_until, hex string (e.g. '0x08804000'); '0x00000000' for other actions.
related_addressYesRelated address for temporary breakpoints, hex string. '0x00000000' when absent.

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Declares the tool as STATE-CHANGE, discloses that 'reset' reboots the game and loses in-memory state, that 'run_until' sets a temp breakpoint and is fire-and-forget with no hit confirmation, and that resume pc/ticks are LOW-trust. This goes well beyond the readOnly/destructive/idempotent annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with PURPOSE/USAGE/ROUTING labels and tightly organized, but it is dense and partially duplicates the enum descriptions in the schema, which costs a little economy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Covers purpose, routing, per-action behavior, and even return-field trust levels. With an output schema already present, nothing an agent needs to invoke this correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is already 100%, so baseline is 3, but the description adds genuine nuance: the schema default '0x0' should not be relied on for run_until, run_until returns immediately with no confirmation, and resume pc/ticks are a best-effort snapshot. This meaningfully supplements the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Aggregate CPU run-state control') and enumerates the exact actions. It also contrasts itself with the sibling ppsspp_batch_step, so an agent can distinguish the two without opening schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit routing rules: single run-state operations go here, multi-step press/wait/probe sequences and cpu_step go to ppsspp_batch_step. It also states per-action requirements ('run_until' requires address) and the session_id omission rule.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ppsspp_wait_framesB
Idempotent

PURPOSE: Wait N frames (wall-clock sleep at 60 FPS by default) to let the emulator advance.

USAGE: session_id + frames required; interval optional (default 1/60 s).

BEHAVIOR: STATE-CHANGE. Sleeps the caller; emulator advances N frames. Session must be alive (validated before sleep).

RETURNS: {frames, elapsed_s}.

ParametersJSON Schema
NameRequiredDescriptionDefault
framesYesNumber of frames to wait (at 60 FPS, frames/60 seconds). MUST be >= 1 — frames=0 is rejected (it would be a silent no-op returning elapsed_s=0 without advancing the emulator).
intervalNoPer-frame interval in seconds (default 1/60). Total wait = frames * interval.
session_idYesActive session ID.

Output Schema

ParametersJSON Schema
NameRequiredDescription
framesYesNumber of frames waited.
elapsed_sYesWall-clock seconds elapsed.

TDQS

B3.2/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare idempotentHint=true, but the description states 'STATE-CHANGE. Sleeps the caller; emulator advances N frames.' This implies each call advances the emulator, so repeated calls with the same arguments would have additional effect, directly contradicting idempotentHint=true. Per the rules, this is an annotation contradiction, so score 1.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is organized into PURPOSE, USAGE, BEHAVIOR, and RETURNS sections. It is front-loaded with the purpose and every sentence is concise and informative, with no wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With annotations, a full output schema, and 100% schema coverage, the description supplies the core operational context (state-change, session validation, return shape). It is adequate for invocation, though it lacks sibling differentiation, which is more properly a purpose/usage concern.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and the schema already documents frames, interval, and session_id in detail. The description only restates that session_id and frames are required and interval is optional, adding no extra semantics beyond the schema. Baseline 3 applies when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'Wait N frames ... to let the emulator advance.' It is clear what the tool does, but it does not name or differentiate from siblings like ppsspp_step or ppsspp_batch_step, so it falls short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description lists required parameters (session_id + frames) and optional interval, but gives no explicit guidance on when to prefer this tool over alternatives such as ppsspp_step or ppsspp_batch_step. Usage context is implied by the purpose, but there are no when/when-not conditions or alternatives named.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ppsspp_watch_valueA
Read-only

PURPOSE: Value-change watch on an address — zero-pause alternative to a read watchpoint.

USAGE: session_id + address + mode + interval. Polls the value every interval_frames and records changes (old/new/frame/time). Pure reads: the CPU is never paused, so hot addresses are safe (storm-free).

BEHAVIOR: READ-ONLY. Polling loop; never mutates state; blocks ~duration_frames/60 seconds (cap 18000 frames).

RETURNS: {address, size, mode, samples, first_value, last_value, changes: [{frame, t_s, old, new}], change_count}.

ParametersJSON Schema
NameRequiredDescriptionDefault
modeNoRead width for change detection: 'u8', 'u16' or 'u32' (default u32). Must match how the watched value is meant to be read.u32
addressYesAddress to watch, hex string (e.g. '0x08A0D000').
session_idYesActive session ID.
duration_framesNoTotal watch window in frames (cap 18000 ~ 5 min at 60fps). A-19: holds the per-session lock for the WHOLE window — concurrent tool calls on this session get SESSION_BUSY after ~5s; use short windows or poll repeatedly instead.
interval_framesNoPoll cadence in frames (60 = 1s wall clock).

Output Schema

ParametersJSON Schema
NameRequiredDescription
modeYes
sizeYes
addressYes
changesYes
samplesYes
last_valueYes
first_valueYes
change_countYes

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, destructiveHint=false, so safety is partly covered, but the description adds genuinely new context: it is a polling loop that never mutates state, it blocks ~duration_frames/60 seconds (cap 18000 frames), and it is 'storm-free' on hot addresses. The concurrency/lock behavior (SESSION_BUSY) is only in the schema, not here, keeping it short of a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Labeled PURPOSE/USAGE/BEHAVIOR/RETURNS sections front-load the key facts with no filler. The RETURNS block duplicates information already carried by the output schema, costing a little economy, but overall it is tight and skimmable.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With full schema coverage, rich annotations and an output schema, the description supplies the extra behavioral context (blocking duration, cap, lock implications implied via the schema) that an agent needs to call it correctly. It is complete for this tool's complexity, though it could have surfaced the session-lock behavior in prose.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all five parameters including mode, interval_frames and the lock caveat on duration_frames are already documented in the schema. The description adds only a light gloss ('Polls the value every interval_frames and records changes'), which is marginal over what structured fields provide, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The opening sentence names a specific verb+resource ('Value-change watch on an address') and immediately positions it against a sibling concept ('zero-pause alternative to a read watchpoint'). An agent can distinguish it from ppsspp_read_memory and ppsspp_breakpoint without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The USAGE line specifies the required inputs and the description names one alternative (read watchpoint) with the condition that selects this tool (zero-pause, storm-free on hot addresses). It lacks explicit when-not guidance against other siblings like ppsspp_breakpoint or ppsspp_state_observer, so it is clear context but not full routing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ppsspp_write_memoryA
Destructive

PURPOSE: Write u8/u16/u32 or raw bytes to memory.

USAGE: session_id + address ('0x' hex) + data + format ('u8'|'u16'|'u32'|'bytes'; bytes accepts hex or base64).

BEHAVIOR: DESTRUCTIVE. Protected ranges (kernel, top.prx code) need force=true (PROTECTED_ADDRESS).

RETURNS: {address, format, bytes_written, value, text}.

ParametersJSON Schema
NameRequiredDescriptionDefault
dataNoValue to write, as a string. For format='u8'/'u16'/'u32', a hex string (e.g. '0x00000001') or decimal string (e.g. '1'). For format='bytes', a hex string (e.g. 'AABBCCDD') or base64 string. Optional when 'value' is supplied (alias); omitting both is an error.
forceNoSet to True to write to protected code/data regions of the modules loaded in THIS session (kernel memory below 0x08800000, plus the top.prx code section as reported by the live module list); declared data addresses from addresses.yaml are exempt. Writing to those ranges without force=True raises ToolError to prevent accidental crashes.
valueNoAlias for 'data'. Accepted because sibling tools take a 'value'; 'data' remains the canonical name and wins when both are supplied.
formatNoWrite format. 'u32' (default) writes a 32-bit int. 'u8'/'u16' write byte/halfword granules (byte patches). 'bytes' writes raw bytes (data is hex-decoded).u32
addressYesTarget address, as a hex string (e.g. '0x08804000').
session_idYesActive session ID.

Output Schema

ParametersJSON Schema
NameRequiredDescription
textYesUnified text representation: 'Wrote 0xVAL → 0xADDR' for u32, 'Wrote N bytes → 0xADDR' for bytes.
valueYesFor format='u32', the value written as hex string (e.g. '0x00000001'); None for 'bytes'.
formatYes'u32' or 'bytes'.
addressYesTarget address, hex string (e.g. '0x08804000').
bytes_writtenYesNumber of bytes written.

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true and idempotentHint=false, but the description adds real value beyond them: it defines which ranges are protected (kernel below 0x08800000 and top.prx code), states the exempt case (addresses.yaml), and notes that a ToolError (PROTECTED_ADDRESS) is raised without force. This is meaningful operational detail, not repetition of the annotation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The labeled PURPOSE/USAGE/BEHAVIOR/RETURNS structure is front-loaded and scannable, with no filler sentences. It is appropriately sized for a 6-parameter mutation tool, though the terse phrasings occasionally assume reader familiarity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With annotations covering the safety profile, a full output schema documenting the return shape, and 100% schema coverage on all six parameters, the description supplies exactly the remaining context an agent needs: format options, destructive nature, force requirement, and error behavior.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents data, value, force, format, address and session_id in depth. The description echoes the format enum and the data/value roles without adding syntax or edge cases beyond the schema, matching the baseline 3 when structured fields do the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb+resource ('Write u8/u16/u32 or raw bytes to memory') and names the accepted formats, so the tool's function is unambiguous. It does not explicitly contrast itself with siblings like ppsspp_read_memory or ppsspp_write_register, leaving differentiation to the tool name.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The USAGE block lists the required inputs (session_id + address + data + format) and the BEHAVIOR block explains when force=true is needed, which implies usage context. However, it never states when to choose this tool over ppsspp_write_register or other mutation siblings, so alternatives are left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ppsspp_write_registerA
Destructive

PURPOSE: Set a CPU register (GPR/FPU/VFPU names, plus pc/hi/lo).

USAGE: session_id + name (MIPS ABI names preferred; numeric aliases like 'r5' normalize to the GPR of that index, i.e. r5 -> a1) + value (hex).

BEHAVIOR: DESTRUCTIVE. Pauses and resumes the CPU automatically (REQUIRED_STEPPING handled internally) — no manual pause needed.

RETURNS: {name, value, response, text}.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYesCPU register name (MIPS standard names). GPRs: 'v0'/'v1'/'a0'-'a3'/'t0'-'t9'/'s0'-'s7'/'gp'/'sp'/'fp'/'ra'/'hi'/'lo'/'pc'. FPU: 'f0'-'f31'. VFPU: 'v0'-'v127'. Numeric aliases like 'r5' are accepted and normalized (r5 -> a1, i.e. GPR index 5) — prefer the MIPS standard name. Case-sensitive (lowercase by convention).
valueYesValue to write, as a hex string (e.g. '0x00000001'). Treated as an unsigned 32-bit int; values outside [0, 0xFFFFFFFF] are REJECTED (out-of-32-bit-range error), not wrapped.
session_idYesActive session ID.

Output Schema

ParametersJSON Schema
NameRequiredDescription
nameYesRegister name (e.g., 'v0', 'a0', 'pc', 'hi', 'lo').
textYesUnified text representation: 'Wrote 0x{VAL:X} → {REG}'.
valueYesValue written, hex string (e.g. '0x00000001').
responseYesRaw PPSSPP WebSocket response (may be empty).

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true and readOnlyHint=false, so the safety profile is covered. The description adds genuine value beyond them by disclosing that the CPU is paused and resumed internally and that no manual pause is required.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The PURPOSE/USAGE/BEHAVIOR/RETURNS structure is front-loaded and easy to scan. The USAGE line slightly duplicates schema information, but no sentence is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, the RETURNS line is a bonus rather than a necessity, and the pause/resume disclosure closes the main behavioral gap. Complete enough for a destructive mutation tool whose annotations already carry the safety signals.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and the name/value descriptions already document alias normalization, hex format, and out-of-range rejection. The description largely echoes the schema, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a precise verb (Set) plus resource (CPU register) and enumerates scope: GPR/FPU/VFPU names plus pc/hi/lo. This is enough to distinguish it from the sibling ppsspp_write_memory without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The USAGE line only restates the required parameter trio already declared in the schema and adds no when-to-use or when-not guidance. It never says when to prefer this over ppsspp_write_memory or other mutation tools, so usage is only implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 19 tool updatesv0.1.7
    • Changedppsspp_analyze_log1 field changed
      • changedInput schema / properties / log_path / description
        Previous value: -"Path to the log file to analyze. S3 whitelist: must be a file anywhere inside the server-managed .ppsspp-dfx tree (the whole tree is allowed, wider than output/ or config/ — verified against the runtime whitelist); arbitrary filesystem paths are rejected. If None, reads the server-mirrored PPSSPP broadcast log (.ppsspp-dfx/output/ppsspp.log — the running game's own ERROR/WARNING lines, captured while a session runs)."New value: +"Path to the log file to analyze. Whitelist: must be a file anywhere inside the server-managed .ppsspp-dfx tree (the whole tree is allowed, wider than output/ or config/ — verified against the runtime whitelist); arbitrary filesystem paths are rejected. If None, reads the server-mirrored PPSSPP broadcast log (.ppsspp-dfx/output/ppsspp.log — the running game's own ERROR/WARNING lines, captured while a session runs)."
    • Changedppsspp_assemble1 field changed
      • changedInput schema / properties / force / description
        Previous value: -"Set to True to write assembled bytes to protected code-section addresses (kernel memory < 0x08800000 or top.prx code section 0x08804000-0x08D34000). Writing to these ranges without force=True raises ToolError to prevent accidental crashes."New value: +"Set to True to write assembled bytes to protected code/data regions of the modules loaded in THIS session (kernel memory below 0x08800000, plus the top.prx code section as reported by the live module list); declared data addresses from addresses.yaml are exempt. Writing to those ranges without force=True raises ToolError to prevent accidental crashes."
    • Changedppsspp_breakpoint8 fields changed
      • changedInput schema / properties / action / description
        Previous value: -"Breakpoint operation. Valid values:\nConsumption actions (lifecycle orchestration):\n- 'wait': STRICT-WAIT — block until any breakpoint is hit \n(arm nothing; set/mem_set first). Lock-free: concurrent \nreads keep working. The breakpoint stays armed.\n- 'stats': HIT-FREQUENCY — count breakpoint hits by pc \nover a time window; optionally samples probe value \nchanges via state_observer. Read-only.\n- 'trace': HIT-SNAPSHOT-RESUME — arm a temporary \nMEMORY breakpoint at `address`, wait for the hit, \ncapture pc/registers/backtrace, ALWAYS remove it, then \nresume (defaults to read access; narrow with \nread/write/size). For EXECUTION breakpoints use \naction='set' + 'wait' instead.\nCPU breakpoint actions:\n- 'set': add a CPU execution breakpoint (requires address; enabled? defaults to True; condition? optional — enforced MCP-side (falsy hits auto-resumed, counted in filtered_hits), NOT sent to PPSSPP).\n- 'remove': delete a CPU breakpoint by address.\n- 'list': list all current CPU breakpoints.\n- 'update': update a CPU breakpoint's enabled/log/condition/log_format (requires address; all other params optional; condition='' clears it).\nMemory breakpoint actions:\n- 'mem_set': add a memory access breakpoint (requires address; size?/read?/write?/enabled?/log?/condition?/log_format?).\n- 'mem_remove': delete a memory breakpoint by address. Delete semantics are STRICT: removing a non-existent memcheck is an ERROR (unlike ppsspp_state_observer action=clear, which is idempotent-ok — F-5 contract, 2026-09-08).\n- 'mem_list': list all current memory breakpoints.\n- 'mem_update': update a memory breakpoint's enabled/log/condition/log_format (requires address)."New value: +"Breakpoint operation. Valid values:\nConsumption actions (lifecycle orchestration):\n- 'wait': STRICT-WAIT — block until any breakpoint is hit \n(arm nothing; set/mem_set first). Lock-free: concurrent \nreads keep working. The breakpoint stays armed.\n- 'stats': HIT-FREQUENCY — count breakpoint hits by pc \nover a time window; optionally samples probe value \nchanges via state_observer. Read-only.\n- 'trace': HIT-SNAPSHOT-RESUME — arm a temporary \nMEMORY breakpoint at `address`, wait for the hit, \ncapture pc/registers/backtrace, ALWAYS remove it, then \nresume (defaults to read access; narrow with \nread/write/size). For EXECUTION breakpoints use \naction='set' + 'wait' instead.\nCPU breakpoint actions:\n- 'set': add a CPU execution breakpoint (requires address; enabled? defaults to True; condition? optional — enforced MCP-side (falsy hits auto-resumed, counted in filtered_hits), NOT sent to PPSSPP).\n- 'remove': delete a CPU breakpoint by address.\n- 'list': list all current CPU breakpoints.\n- 'update': update a CPU breakpoint's enabled/log/condition/log_format (requires address; all other params optional; condition='' clears it).\nMemory breakpoint actions:\n- 'mem_set': add a memory access breakpoint (requires address; size?/read?/write?/enabled?/log?/condition?/log_format?).\n- 'mem_remove': delete a memory breakpoint by address. Delete semantics are STRICT: removing a non-existent memcheck is an ERROR (unlike ppsspp_state_observer action=clear, which is idempotent-ok).\n- 'mem_list': list all current memory breakpoints.\n- 'mem_update': update a memory breakpoint's enabled/log/condition/log_format (requires address)."
      • changedInput schema / properties / address / description
        Previous value: -"Breakpoint address, as a hex string (e.g. '0x08804000'). Required for set / remove / update / mem_set / mem_remove / mem_update; ignored for list / mem_list."New value: +"Required for set / remove / update / mem_set / mem_remove / mem_update. Breakpoint address, as a hex string (e.g. '0x08804000'). Not used by list / mem_list. The schema default of '0x0' exists for legacy callers -- do NOT rely on it when the action is one of the above."
      • addedInput schema / properties / session_id / anyOf
        Added value: +[
        +  {
        +    "type": "string"
        +  },
        +  {
        +    "type": "null"
        +  }
        +]
      • addedInput schema / properties / session_id / default
        Added value: +null
      • changedInput schema / properties / session_id / description
        Previous value: -"Active session ID."New value: +"Active session ID; omit to auto-resolve when exactly one session is active."
      • removedInput schema / properties / session_id / type
        Removed value: -"string"
      • changedInput schema / properties / size / description
        Previous value: -"Memory breakpoint watch size in bytes (mem_set / mem_remove / mem_update; default 4). Fixed-width watches use 1/2/4; larger sizes are passed through to PPSSPP as a range watch. NOTE (verified): mem_remove matches by ADDRESS only — a wrong or omitted size still removes the breakpoint at that address, so keep the size you set for bookkeeping, not for matching."New value: +"Memory breakpoint watch size in bytes (mem_set / mem_remove / mem_update; default 4). Fixed-width watches use 1/2/4; larger sizes are passed through to PPSSPP as a range watch. NOTE: PPSSPP matches a memory watchpoint by the exact address+size pair (BreakpointSubscriber.cpp) -- the size is part of the match key, not just bookkeeping. mem_remove therefore resolves the real size via mem_list first, because removing with the caller's size alone silently fails when it differs (e.g. a 16-byte watch removed with the default 4)."
      • changedInput schema / required
        Previous value: -[
        -  "session_id",
        -  "action"
        -]New value: +[
        +  "action"
        +]
    • Changedppsspp_diff_memory5 fields changed
      • changedInput schema / properties / action / description
        Previous value: -"Diff operation. Valid values:\n- 'snapshot': read [start, start+size) and store it under a new handle (registry cap 4, FIFO eviction).\n- 'compare': read the same range now and diff against the handle's snapshot — returns changed-byte list (inline cap 256, truncated flag keeps the true count).\n- 'drop': release a handle.\n- 'list': live handles."New value: +"Diff operation. Valid values:\n- 'snapshot': read [start, start+size) and store it under a new handle (registry cap 4, FIFO eviction).\n- 'compare': read the same range now and diff against the handle's snapshot; returns a changed-byte list (inline cap 256, truncated flag keeps the true count).\n- 'drop': release a handle.\n- 'list': live handles."
      • addedInput schema / properties / address
        Added value: +{
        +  "anyOf": [
        +    {
        +      "type": "string"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null,
        +  "description": "Range start, same format as 'start'. Provided for naming consistency with read_memory/write_memory/scan, which all take 'address'. Use either address+size or start+end.",
        +  "title": "Address"
        +}
      • changedInput schema / properties / end / description
        Previous value: -"Range end (exclusive), same format — required for snapshot."New value: +"Range end (exclusive), same format — required for snapshot unless 'address'+'size' are given."
      • addedInput schema / properties / size
        Added value: +{
        +  "anyOf": [
        +    {
        +      "type": "integer"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null,
        +  "description": "Range length in bytes. Provided for naming consistency with read_memory/write_memory/scan, which all take 'size'. Use either address+size or start+end.",
        +  "title": "Size"
        +}
      • changedInput schema / properties / start / description
        Previous value: -"Range start, hex ('0x08804000') or decimal — required for snapshot."New value: +"Range start, hex ('0x08804000') or decimal — required for snapshot unless 'address'+'size' are given."
    • Changedppsspp_disassemble1 field changed
      • changedInput schema / properties / count / description
        Previous value: -"Number of instructions to disassemble. Capped at 100 to prevent oversized responses."New value: +"Number of instructions to disassemble. Default 10. count=0 is treated as 'use the default' and returns 10 instructions (the response carries a note saying so). Capped at 100 to prevent oversized responses; a larger value is clamped and the note reports the clamp."
    • Changedppsspp_health3 fields changed
      • addedOutput schema / $defs / _HealthSessionCheck / properties / value
        Added value: +{
        +  "title": "Value",
        +  "type": "integer"
        +}
      • addedOutput schema / $defs / _HealthSessionCheck / properties / value_status
        Added value: +{
        +  "title": "Value Status",
        +  "type": "string"
        +}
      • changedOutput schema / description
        Previous value: -"HealthOutput + the per-session battery keys added when `session_id`\nis given (M9: they previously existed only in the text channel — the\noutput contract dropped them from structuredContent)."New value: +"HealthOutput + the per-session battery keys added when `session_id`\nis given. They previously existed only in the text channel — the\noutput contract dropped them from structuredContent."
    • Changedppsspp_memory_map1 field changed
      • changedOutput schema / properties / mapping / description
        Previous value: -"Raw `memory.mapping` response."New value: +"Raw `memory.mapping` response, minus the volatile per-call protocol `ticket` (stripped so the business fields are byte-stable across identical calls)."
    • Changedppsspp_press_button1 field changed
      • changedInput schema / properties / duration / description
        Previous value: -"Press duration in frames (default 1; 60fps wall-clock, cap 18000 ≈ 300s — values above are rejected). The call blocks for the duration."New value: +"Press duration in frames (default 1; 60fps wall-clock, cap 18000 ≈ 300s — values above are rejected). duration=0 is accepted and is a no-op (the button is pressed and released in the same frame). The call blocks for the duration."
    • Changedppsspp_query3 fields changed
      • changedInput schema / properties / address / description
        Previous value: -"Function address as a hex string (e.g. '0x08804000'). Required for func_remove and func_scan."New value: +"Required for func_remove and func_scan. Function address as a hex string (e.g. '0x08804000'). Not used by the other actions. The schema default of '0x0' exists for legacy callers -- do NOT rely on it when the action is one of the above."
      • changedInput schema / properties / name / description
        Previous value: -"Function name (func_add only; ignored by func_remove because PPSSPP's hle.func.remove protocol does not accept a name parameter)."New value: +"Required for action='register' (the register name to read) and for action='func_add'. Ignored by func_remove because PPSSPP's hle.func.remove protocol does not accept a name parameter."
      • changedOutput schema / properties / text / description
        Previous value: -"Unified text representation. Populated for action='registers' with grouped '── GPR ──' / '── FPU ──' / '── VFPU ──' headers and '  name = 0xVAL' lines. Empty for other actions (use the structured `data` field)."New value: +"Unified text representation. Populated for action='registers' with grouped '── GPR ──' / '── FPU ──' / '── VFPU ──' headers and '  name = 0xVAL' lines, and for action='register' with a single 'name = 0xVAL' line (the requested register name echoed with its hex value). Empty for other actions (use the structured `data` field)."
    • Changedppsspp_read_memory4 fields changed
      • changedInput schema / properties / action / description
        Previous value: -"Read action. Valid values:\n- 'read_bytes': read raw bytes (requires address + size).\n- 'read_u32': read a 32-bit unsigned int (requires address).\n- 'read_string': read a string (requires address).\npattern + start_addr + end_addr). Optional max_results (default 100)."New value: +"Read action. Valid values:\n- 'read_bytes': read raw bytes (requires address + size).\n- 'read_u32': read a 32-bit unsigned int (requires address).\n- 'read_string': read a string (requires address).\n"
      • changedInput schema / properties / address / description
        Previous value: -"Starting address for read_bytes/read_u32/read_string, as a hex string (e.g. '0x08804000')."New value: +"Required for every action. Starting address for read_bytes/read_u32/read_string, as a hex string (e.g. '0x08804000'). The schema default of '0x0' exists for legacy callers -- do NOT rely on it."
      • changedInput schema / properties / size / description
        Previous value: -"Number of bytes to read (read_bytes only)."New value: +"Number of bytes to read (read_bytes only). Max 65536 per call (MAX_SINGLE_READ_BYTES); larger reads are rejected with ARGS_INVALID -- chunk them instead."
      • changedOutput schema / properties / text / description
        Previous value: -"Unified text representation following spec conventions: '0xADDR: VAL (0xVAL_HEX)' for read_u32, hex dump for read_bytes, repr for read_string, 'scan: N matches at 0xA1, 0xA2, ...' for scan."New value: +"Unified text representation following spec conventions: '0xADDR: VAL (0xVAL_HEX)' for read_u32 (8-digit zero-padded hex), hex dump for read_bytes, repr for read_string, 'scan: N matches at 0xA1, 0xA2, ...' for scan."
    • Changedppsspp_reload_scripts1 field changed
      • changedOutput schema / properties / exposed_registered / description
        Previous value: -"Number of exposed scripts ACTUALLY registered as dynamic tools after the reload sync (F4/F5: declared != registered is now surfaced instead of silently diverging)."New value: +"Number of exposed scripts ACTUALLY registered as dynamic tools after the reload sync (declared != registered is now surfaced instead of silently diverging)."
    • Changedppsspp_replay11 fields changed
      • changedInput schema / properties / action / description
        Previous value: -"Replay operation. Valid values:\n- 'begin': begin/resume recording.\n- 'abort': abort any recording or execution.\n- 'flush': flush recorded data (returns version + base64).\n- 'execute': execute a replay (requires version + base64_input). ONLY loads the event table — follow the boot-aligned sequence in the response (reset + wait + abort) or input never injects (U7 root cause).\n- 'status': query {executing, saving}.\n- 'time_get': get base RTC.\n- 'time_set': set base RTC (requires value). WARNING: rewinds the game-visible wall clock on the RUNNING session — pollutes every in-game timer (R2).\n- 'save': flush + time_get + write .ppr file (requires file_path: bare file name under output/replays/).\n- 'load': read .ppr + execute (requires file_path; same containment). Returns t0_s / estimated_end_s and the boot-aligned sequence.\n- 'wait_complete': poll replay.status until executing=False — NOTE: executing never clears on its own (only abort clears it), so this always times out on an un-aborted replay; kept for recording-completion checks and backwards compatibility."New value: +"Replay operation. Valid values:\n- 'begin': begin/resume recording.\n- 'abort': abort any recording or execution.\n- 'flush': flush recorded data (returns version + base64).\n- 'execute': execute a replay (requires version + base64_input). ONLY loads the event table — follow the boot-aligned sequence in the response (reset + wait + abort) or input never injects.\n- 'status': query {executing, saving}.\n- 'time_get': get base RTC.\n- 'time_set': set base RTC (requires value). WARNING: rewinds the game-visible wall clock on the RUNNING session — pollutes every in-game timer.\n- 'save': flush + time_get + write .ppr file (requires file_path: bare file name under output/replays/).\n- 'load': read .ppr + execute (requires file_path; same containment). Returns t0_s / estimated_end_s and the boot-aligned sequence.\n- 'wait_complete': poll replay.status until executing=False — NOTE: executing never clears on its own (only abort clears it), so this always times out on an un-aborted replay; kept for recording-completion checks and backwards compatibility."
      • changedInput schema / properties / restore_rtc / description
        Previous value: -"Whether to restore base_rtc via replay.time_set before execute when action='load' (default False; R2). true sets the game-visible wall clock back to the recording moment — pollutes EVERY timer of the running session (attract timeouts, clocks, cooldowns) because game time = rtcBaseTime + elapsed. Only use for deterministic replays, and prefer setting it BEFORE the reset of the boot-aligned sequence so the game boots on the shifted base. Ignored for all other actions."New value: +"Whether to restore base_rtc via replay.time_set before execute when action='load' (default False). true sets the game-visible wall clock back to the recording moment — pollutes EVERY timer of the running session (attract timeouts, clocks, cooldowns) because game time = rtcBaseTime + elapsed. Only use for deterministic replays, and prefer setting it BEFORE the reset of the boot-aligned sequence so the game boots on the shifted base. Ignored for all other actions."
      • addedInput schema / properties / session_id / anyOf
        Added value: +[
        +  {
        +    "type": "string"
        +  },
        +  {
        +    "type": "null"
        +  }
        +]
      • addedInput schema / properties / session_id / default
        Added value: +null
      • changedInput schema / properties / session_id / description
        Previous value: -"Active session ID."New value: +"Active session ID; omit to auto-resolve when exactly one session is active."
      • removedInput schema / properties / session_id / type
        Removed value: -"string"
      • changedInput schema / properties / timeout_ms / description
        Previous value: -"Total timeout in milliseconds for action='wait_complete' (default 10000 = 10s, clamped 100..60000). Ignored for all other actions."New value: +"Total timeout in milliseconds for action='wait_complete' (default 10000 = 10s, clamped 100..25000). The ceiling is 25s because the MCP client aborts a tool call at ~30s: a larger value could never be honoured. Ignored for all other actions."
      • changedInput schema / properties / timeout_ms / maximum
        Previous value: -60000New value: +25000
      • changedInput schema / properties / value / description
        Previous value: -"Base RTC value in seconds (uint32). Required for action='time_set'; ignored for all other actions."New value: +"Required for action='time_set'. Base RTC value in seconds (uint32). Not used by the other actions. The schema default of 0 exists for legacy callers -- do NOT rely on it when the action is 'time_set'."
      • changedInput schema / properties / version / description
        Previous value: -"Replay format version (from a prior replay.flush). Required for action='execute'; ignored for all other actions."New value: +"Required for action='execute'. Replay format version (from a prior replay.flush). Not used by the other actions. The schema default of 0 exists for legacy callers -- do NOT rely on it when the action is 'execute'."
      • changedInput schema / required
        Previous value: -[
        -  "session_id",
        -  "action"
        -]New value: +[
        +  "action"
        +]
    • Changedppsspp_scan1 field changed
      • changedInput schema / properties / session_id / description
        Previous value: -"Active session ID; auto-resolved when exactly one session is active (required for pattern/value initial/strings; ignored for narrow/list/drop phases which only re-read candidate addresses — narrow still needs it)."New value: +"Active session ID; auto-resolved when exactly one session is active. Required for the pattern / value initial / strings phases, which start a new scan. The narrow / list / drop phases only re-read addresses already recorded by an earlier phase, so they do not need it passed -- but it must still resolve to the same session."
    • Changedppsspp_session6 fields changed
      • changedInput schema / properties / wait_ready / default
        Previous value: -falseNew value: +true
      • changedInput schema / properties / wait_ready / description
        Previous value: -"action=start only: block until the emulated CPU is ready before returning (same probe/budget as action=wait_ready; raises [BOOT_TIMEOUT] on wedge suspicion). Fake test mode is ready immediately. Default false keeps the historical two-call flow."New value: +"action=start only: block until the emulated CPU is ready before returning (same probe/budget as action=wait_ready; raises [BOOT_TIMEOUT] on wedge suspicion). Fake test mode is ready immediately. Default true: start() returns a session that is ready to use, so the first tool call does not fail with a version-handshake timeout. Pass false only when you want the raw launch without waiting."
      • changedOutput schema / $defs / SessionResponse / properties / recovered / description
        Previous value: -"H2: resilient-start relaunch count (0 = the first launch succeeded; >0 means the game state was reset by a wedge heal — breakpoints need re-arming)."New value: +"resilient-start relaunch count (0 = the first launch succeeded; >0 means the game state was reset by a wedge heal — breakpoints need re-arming)."
      • changedOutput schema / $defs / SessionResponse / properties / restored / description
        Previous value: -"F-6(a): 1 when this session was restored from sessions.json (a previous server run left it behind) rather than started fresh in this process — its game state may be stale."New value: +"1 when this session was restored from sessions.json (a previous server run left it behind) rather than started fresh in this process — its game state may be stale."
      • changedOutput schema / properties / recovered / description
        Previous value: -"H2: resilient-start relaunch count (0 = the first launch succeeded; >0 means the game state was reset by a wedge heal — breakpoints need re-arming)."New value: +"resilient-start relaunch count (0 = the first launch succeeded; >0 means the game state was reset by a wedge heal — breakpoints need re-arming)."
      • changedOutput schema / properties / restored / description
        Previous value: -"F-6(a): 1 when this session was restored from sessions.json (a previous server run left it behind) rather than started fresh in this process — its game state may be stale."New value: +"1 when this session was restored from sessions.json (a previous server run left it behind) rather than started fresh in this process — its game state may be stale."
    • Changedppsspp_state_observer3 fields changed
      • changedInput schema / properties / address / description
        Previous value: -"Absolute runtime address to read, as a hex string (e.g. '0x08804000'). Required for action='register'; ignored for all other actions."New value: +"Required for action='register'. Absolute runtime address to read, as a hex string (e.g. '0x08804000'). Not used by the other actions. The schema default of '0x0' exists for legacy callers -- do NOT rely on it when the action is 'register'."
      • addedOutput schema / $defs / ProbeObservationView / properties / note
        Added value: +{
        +  "default": "",
        +  "description": "Empty unless value_status is suspicious",
        +  "title": "Note",
        +  "type": "string"
        +}
      • addedOutput schema / $defs / ProbeObservationView / properties / value_status
        Added value: +{
        +  "default": "ok",
        +  "description": "'ok' or 'stale_address_suspected'. The latter means this address has read zero on several consecutive readings, so the probe address may have drifted -- a suspicion, not a verdict",
        +  "title": "Value Status",
        +  "type": "string"
        +}
    • Changedppsspp_step4 fields changed
      • changedInput schema / properties / action / description
        Previous value: -"CPU step / run-state operation. Valid values:\n- 'pause': pause CPU (enter stepping mode).\n- 'resume': resume CPU (exit stepping mode).\n- 'reset': reset the game (reboot).\n- 'run_until': run until the specified address is reached (requires address).\n- 'next_hle': step to next HLE callback.\nNOTE: single-stepping (into/over/out) lives in ppsspp_batch_step as the 'cpu_step' step type (mode='into'|'over'|'out', count 1..1000) — it requires the CPU to enter stepping mode, which the executor handles automatically."New value: +"CPU step / run-state operation. Valid values:\n- 'pause': pause CPU (enter stepping mode).\n- 'resume': resume CPU (exit stepping mode); the response pc/ticks are a LOW-trust snapshot of the running CPU (inaccurate unless stepping), not a precise resume location.\n- 'reset': reset the game (reboot).\n- 'run_until': run until the specified address is reached (requires address).\n- 'next_hle': step to next HLE callback.\nNOTE: single-stepping (into/over/out) lives in ppsspp_batch_step as the 'cpu_step' step type (mode='into'|'over'|'out', count 1..1000) — it requires the CPU to enter stepping mode, which the executor handles automatically."
      • changedInput schema / properties / address / description
        Previous value: -"Target address, as a hex string (e.g. '0x08804000'). Required for action='run_until'; ignored for all other actions. run_until is fire-and-forget: it returns immediately with no hit confirmation — poll PC or set a breakpoint to observe arrival."New value: +"Required for action='run_until'. Target address, as a hex string (e.g. '0x08804000'). Not used by the other actions. The schema default of '0x0' exists for legacy callers -- do NOT rely on it when the action is 'run_until'. run_until is fire-and-forget: it returns immediately with no hit confirmation — poll PC or set a breakpoint to observe arrival."
      • changedOutput schema / properties / pc / description
        Previous value: -"Program counter after step, hex string. For into/over/out/run_until/next_hle this comes from the cpu.stepping broadcast. For pause, extracted via safe_get_pc after CPU enters stepping. '0x00000000' for resume/reset (no trustworthy PC available)."New value: +"Program counter after step, hex string. For into/over/out/run_until/next_hle this comes from the cpu.stepping broadcast. For pause, extracted via safe_get_pc after CPU enters stepping. For resume, a best-effort LOW-trust cpu.status snapshot of the running CPU (inaccurate unless stepping — CPUCoreSubscriber.cpp:105). '0x00000000' for reset."
      • changedOutput schema / properties / ticks / description
        Previous value: -"CPU ticks at step completion (0.0 for pause/resume/reset)."New value: +"CPU ticks at step completion (from cpu.stepping broadcast). For resume, a best-effort LOW-trust cpu.status snapshot. 0.0 for pause/reset."
    • Changedppsspp_wait_frames1 field changed
      • changedInput schema / properties / frames / description
        Previous value: -"Number of frames to wait (at 60 FPS, frames/60 seconds)."New value: +"Number of frames to wait (at 60 FPS, frames/60 seconds). MUST be >= 1 — frames=0 is rejected (it would be a silent no-op returning elapsed_s=0 without advancing the emulator)."
    • Changedppsspp_watch_value2 fields changed
      • changedInput schema / properties / duration_frames / description
        Previous value: -"Total watch window in frames (cap 18000 ~ 5 min at 60fps)."New value: +"Total watch window in frames (cap 18000 ~ 5 min at 60fps). A-19: holds the per-session lock for the WHOLE window — concurrent tool calls on this session get SESSION_BUSY after ~5s; use short windows or poll repeatedly instead."
      • changedInput schema / properties / mode / description
        Previous value: -"Value interpretation for change detection."New value: +"Read width for change detection: 'u8', 'u16' or 'u32' (default u32). Must match how the watched value is meant to be read."
    • Changedppsspp_write_memory7 fields changed
      • addedInput schema / properties / data / anyOf
        Added value: +[
        +  {
        +    "type": "string"
        +  },
        +  {
        +    "type": "null"
        +  }
        +]
      • addedInput schema / properties / data / default
        Added value: +null
      • changedInput schema / properties / data / description
        Previous value: -"Value to write, as a string. For format='u8'/'u16'/'u32', a hex string (e.g. '0x00000001') or decimal string (e.g. '1'). For format='bytes', a hex string (e.g. 'AABBCCDD') or base64 string."New value: +"Value to write, as a string. For format='u8'/'u16'/'u32', a hex string (e.g. '0x00000001') or decimal string (e.g. '1'). For format='bytes', a hex string (e.g. 'AABBCCDD') or base64 string. Optional when 'value' is supplied (alias); omitting both is an error."
      • removedInput schema / properties / data / type
        Removed value: -"string"
      • changedInput schema / properties / force / description
        Previous value: -"Set to True to write to protected code-section addresses (kernel memory < 0x08800000 or top.prx code section 0x08804000-0x08D34000). Writing to these ranges without force=True raises ToolError to prevent accidental crashes."New value: +"Set to True to write to protected code/data regions of the modules loaded in THIS session (kernel memory below 0x08800000, plus the top.prx code section as reported by the live module list); declared data addresses from addresses.yaml are exempt. Writing to those ranges without force=True raises ToolError to prevent accidental crashes."
      • addedInput schema / properties / value
        Added value: +{
        +  "anyOf": [
        +    {
        +      "type": "string"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null,
        +  "description": "Alias for 'data'. Accepted because sibling tools take a 'value'; 'data' remains the canonical name and wins when both are supplied.",
        +  "title": "Value"
        +}
      • changedInput schema / required
        Previous value: -[
        -  "session_id",
        -  "address",
        -  "data"
        -]New value: +[
        +  "session_id",
        +  "address"
        +]
  2. 28 tool updatesv0.1.6
    • Changedppsspp_analyze_log9 fields changed
      • changedInput schema / properties / filter / description
        Previous value: -"Additional keyword to filter for (in addition to ERROR/WARNING/CRASH). Case-sensitive substring match."New value: +"Additional keyword to filter for. Case-sensitive substring match; combined with the severity keywords per filter_mode."
      • addedInput schema / properties / filter_mode
        Added value: +{
        +  "default": "any",
        +  "description": "'any' (default, legacy): a line matches if it contains a severity keyword OR the filter. 'all': a line must contain a severity keyword AND the filter — use this to narrow (e.g. filter='GPU', filter_mode='all' → only GPU-related ERROR/WARNING/CRASH lines).",
        +  "enum": [
        +    "any",
        +    "all"
        +  ],
        +  "title": "Filter Mode",
        +  "type": "string"
        +}
      • addedInput schema / properties / limit
        Added value: +{
        +  "default": 0,
        +  "description": "Cap the returned match list (0 = no cap beyond the hard internal cap of 500). When truncation happens the response carries total_matches and truncated=true; total_matches is a LOWER BOUND whenever truncated=true (the scan stops at 500 matches, so the real count can be higher). Use this on long logs instead of receiving 200KB+ of matches.",
        +  "title": "Limit",
        +  "type": "integer"
        +}
      • changedInput schema / properties / log_path / description
        Previous value: -"Path to the log file to analyze. S3 whitelist: must be a file under the server-managed .ppsspp-dfx tree (.ppsspp-dfx/output/ or .ppsspp-dfx/config/); arbitrary filesystem paths are rejected. If None, reads the server-mirrored PPSSPP broadcast log (.ppsspp-dfx/output/ppsspp.log — the running game's own ERROR/WARNING lines, captured while a session runs)."New value: +"Path to the log file to analyze. S3 whitelist: must be a file anywhere inside the server-managed .ppsspp-dfx tree (the whole tree is allowed, wider than output/ or config/ — verified against the runtime whitelist); arbitrary filesystem paths are rejected. If None, reads the server-mirrored PPSSPP broadcast log (.ppsspp-dfx/output/ppsspp.log — the running game's own ERROR/WARNING lines, captured while a session runs)."
      • changedOutput schema / properties / count / description
        Previous value: -"Number of matches."New value: +"Number of matches (after limit truncation)."
      • addedOutput schema / properties / filter_mode
        Added value: +{
        +  "description": "Filter combination mode: 'any' (legacy OR) or 'all' (severity AND filter).",
        +  "title": "Filter Mode",
        +  "type": "string"
        +}
      • addedOutput schema / properties / total_matches
        Added value: +{
        +  "description": "Match count before limit truncation.",
        +  "title": "Total Matches",
        +  "type": "integer"
        +}
      • addedOutput schema / properties / truncated
        Added value: +{
        +  "description": "True when limit truncated the match list.",
        +  "title": "Truncated",
        +  "type": "boolean"
        +}
      • changedOutput schema / required
        Previous value: -[
        -  "log_path",
        -  "matches",
        -  "count",
        -  "filter"
        -]New value: +[
        +  "log_path",
        +  "matches",
        +  "count",
        +  "filter",
        +  "filter_mode",
        +  "total_matches",
        +  "truncated"
        +]
    • Removedppsspp_batch_list
    • Changedppsspp_batch_status12 fields changed
      • addedInput schema / properties / batch_id / anyOf
        Added value: +[
        +  {
        +    "type": "string"
        +  },
        +  {
        +    "type": "null"
        +  }
        +]
      • addedInput schema / properties / batch_id / default
        Added value: +null
      • changedInput schema / properties / batch_id / description
        Previous value: -"Job id returned by ppsspp_batch_step(background=true)."New value: +"Job id returned by ppsspp_batch_step(background=true). Omit to survey ALL retained jobs instead (list mode; recovers a batch_id after the submit response was lost)."
      • removedInput schema / properties / batch_id / type
        Removed value: -"string"
      • removedInput schema / required
        Removed value: -[
        -  "batch_id"
        -]
      • addedOutput schema / $defs
        Added value: +{
        +  "BatchJobSummaryView": {
        +    "additionalProperties": false,
        +    "description": "One retained background batch job in the ppsspp_batch_list survey.",
        +    "properties": {
        +      "batch_id": {
        +        "description": "Job id (for ppsspp_batch_status / _cancel)",
        +        "title": "Batch Id",
        +        "type": "string"
        +      },
        +      "error": {
        +        "default": "",
        +        "description": "Error message if failed/cancelled",
        +        "title": "Error",
        +        "type": "string"
        +      },
        +      "executed": {
        +        "description": "Steps executed so far",
        +        "title": "Executed",
        +        "type": "integer"
        +      },
        +      "result_present": {
        +        "description": "Whether the final result is retained (true once completed; fetch it via ppsspp_batch_status)",
        +        "title": "Result Present",
        +        "type": "boolean"
        +      },
        +      "session_id": {
        +        "description": "Session the batch runs / ran on",
        +        "title": "Session Id",
        +        "type": "string"
        +      },
        +      "status": {
        +        "description": "'queued' / 'running' / 'completed' / 'failed' / 'cancelled' (protocol Tasks mapping: 'queued'→'working')",
        +        "title": "Status",
        +        "type": "string"
        +      },
        +      "total": {
        +        "description": "Total steps in the batch",
        +        "title": "Total",
        +        "type": "integer"
        +      }
        +    },
        +    "required": [
        +      "batch_id",
        +      "session_id",
        +      "status",
        +      "executed",
        +      "total",
        +      "result_present"
        +    ],
        +    "title": "BatchJobSummaryView",
        +    "type": "object"
        +  }
        +}
      • addedOutput schema / properties / error / anyOf
        Added value: +[
        +  {
        +    "type": "string"
        +  },
        +  {
        +    "type": "null"
        +  }
        +]
      • changedOutput schema / properties / error / description
        Previous value: -"Error message if failed/cancelled"New value: +"Error message if failed/cancelled (null when none; real runner emits null)"
      • removedOutput schema / properties / error / type
        Removed value: -"string"
      • addedOutput schema / properties / jobs
        Added value: +{
        +  "description": "All retained jobs, in submission order",
        +  "items": {
        +    "$ref": "#/$defs/BatchJobSummaryView"
        +  },
        +  "title": "Jobs",
        +  "type": "array"
        +}
      • changedOutput schema / properties / retention_jobs / description
        Previous value: -"Finished-job retention window of the registry (in job count, not seconds): completed/failed/cancelled jobs beyond the oldest this many are evicted. Precursor of the protocol-level Task ttl."New value: +"Finished-job retention window (in job count): finished jobs beyond the oldest this many are evicted and absent from jobs"
      • removedOutput schema / required
        Removed value: -[
        -  "batch_id",
        -  "session_id",
        -  "status",
        -  "executed",
        -  "total",
        -  "error",
        -  "result",
        -  "retention_jobs"
        -]
    • Changedppsspp_batch_step8 fields changed
      • addedInput schema / $defs / CpuStepStep
        Added value: +{
        +  "description": "{type:'cpu_step', mode, count} — N CPU instructions (into/over/out).",
        +  "properties": {
        +    "count": {
        +      "description": "Number of CPU steps to execute (1..1000; default 1).",
        +      "title": "Count",
        +      "type": "integer"
        +    },
        +    "mode": {
        +      "description": "CPU stepping mode: 'into', 'over', or 'out'.",
        +      "enum": [
        +        "into",
        +        "over",
        +        "out"
        +      ],
        +      "title": "Mode",
        +      "type": "string"
        +    },
        +    "type": {
        +      "const": "cpu_step",
        +      "title": "Type",
        +      "type": "string"
        +    }
        +  },
        +  "required": [
        +    "type",
        +    "mode",
        +    "count"
        +  ],
        +  "title": "CpuStepStep",
        +  "type": "object"
        +}
      • addedInput schema / properties / steps / items / discriminator / mapping / cpu_step
        Added value: +"#/$defs/CpuStepStep"
      • changedInput schema / properties / steps / items / oneOf
        Previous value: -[
        -  {
        -    "$ref": "#/$defs/PressStep"
        -  },
        -  {
        -    "$ref": "#/$defs/WaitStep"
        -  },
        -  {
        -    "$ref": "#/$defs/StateProbeStep"
        -  },
        -  {
        -    "$ref": "#/$defs/ScreenshotStep"
        -  }
        -]New value: +[
        +  {
        +    "$ref": "#/$defs/PressStep"
        +  },
        +  {
        +    "$ref": "#/$defs/WaitStep"
        +  },
        +  {
        +    "$ref": "#/$defs/StateProbeStep"
        +  },
        +  {
        +    "$ref": "#/$defs/ScreenshotStep"
        +  },
        +  {
        +    "$ref": "#/$defs/CpuStepStep"
        +  }
        +]
      • changedOutput schema / properties / action / description
        Previous value: -"Always 'run'"New value: +"Always 'submitted'"
      • addedOutput schema / properties / batch_id
        Added value: +{
        +  "description": "Job id for ppsspp_batch_status / _cancel",
        +  "title": "Batch Id",
        +  "type": "string"
        +}
      • addedOutput schema / properties / estimated_s
        Added value: +{
        +  "description": "Heuristic wall-clock estimate in seconds",
        +  "title": "Estimated S",
        +  "type": "number"
        +}
      • addedOutput schema / properties / hint
        Added value: +{
        +  "description": "How to poll / cancel this background batch",
        +  "title": "Hint",
        +  "type": "string"
        +}
      • addedOutput schema / properties / session_id
        Added value: +{
        +  "description": "Session the batch will execute on",
        +  "title": "Session Id",
        +  "type": "string"
        +}
    • Changedppsspp_breakpoint31 fields changed
      • changedInput schema / properties / action / description
        Previous value: -"Breakpoint operation. Valid values:\nCPU breakpoint actions:\n- 'set': add a CPU execution breakpoint (requires address; enabled? defaults to True; condition? optional).\n- 'remove': delete a CPU breakpoint by address.\n- 'list': list all current CPU breakpoints.\n- 'update': update a CPU breakpoint's enabled/log/condition/log_format (requires address; all other params optional).\nMemory breakpoint actions:\n- 'mem_set': add a memory access breakpoint (requires address; size?/read?/write?/enabled?/log?/condition?/log_format?).\n- 'mem_remove': delete a memory breakpoint by address. Delete semantics are STRICT: removing a non-existent memcheck is an ERROR (unlike ppsspp_state_observer action=clear, which is idempotent-ok — F-5 contract, 2026-09-08).\n- 'mem_list': list all current memory breakpoints.\n- 'mem_update': update a memory breakpoint's enabled/log/condition/log_format (requires address)."New value: +"Breakpoint operation. Valid values:\nConsumption actions (lifecycle orchestration):\n- 'wait': STRICT-WAIT — block until any breakpoint is hit \n(arm nothing; set/mem_set first). Lock-free: concurrent \nreads keep working. The breakpoint stays armed.\n- 'stats': HIT-FREQUENCY — count breakpoint hits by pc \nover a time window; optionally samples probe value \nchanges via state_observer. Read-only.\n- 'trace': HIT-SNAPSHOT-RESUME — arm a temporary \nMEMORY breakpoint at `address`, wait for the hit, \ncapture pc/registers/backtrace, ALWAYS remove it, then \nresume (defaults to read access; narrow with \nread/write/size). For EXECUTION breakpoints use \naction='set' + 'wait' instead.\nCPU breakpoint actions:\n- 'set': add a CPU execution breakpoint (requires address; enabled? defaults to True; condition? optional — enforced MCP-side (falsy hits auto-resumed, counted in filtered_hits), NOT sent to PPSSPP).\n- 'remove': delete a CPU breakpoint by address.\n- 'list': list all current CPU breakpoints.\n- 'update': update a CPU breakpoint's enabled/log/condition/log_format (requires address; all other params optional; condition='' clears it).\nMemory breakpoint actions:\n- 'mem_set': add a memory access breakpoint (requires address; size?/read?/write?/enabled?/log?/condition?/log_format?).\n- 'mem_remove': delete a memory breakpoint by address. Delete semantics are STRICT: removing a non-existent memcheck is an ERROR (unlike ppsspp_state_observer action=clear, which is idempotent-ok — F-5 contract, 2026-09-08).\n- 'mem_list': list all current memory breakpoints.\n- 'mem_update': update a memory breakpoint's enabled/log/condition/log_format (requires address)."
      • changedInput schema / properties / action / enum
        Previous value: -[
        -  "set",
        -  "remove",
        -  "list",
        -  "update",
        -  "mem_set",
        -  "mem_remove",
        -  "mem_list",
        -  "mem_update"
        -]New value: +[
        +  "wait",
        +  "trace",
        +  "stats",
        +  "set",
        +  "remove",
        +  "list",
        +  "update",
        +  "mem_set",
        +  "mem_remove",
        +  "mem_list",
        +  "mem_update"
        +]
      • changedInput schema / properties / condition / description
        Previous value: -"Break condition expression (set / update / mem_set / mem_update; None = don't send)."New value: +"Break condition expression (set / update / mem_set / mem_update; None = don't send). Enforced MCP-side: PPSSPP's IR mode silently ignores register conditions, so the breakpoint is armed UNCONDITIONALLY and falsy hits are auto-resumed and counted in filtered_hits."
      • changedInput schema / properties / size / description
        Previous value: -"Memory breakpoint watch size in bytes (mem_set / mem_remove / mem_update; default 4). Fixed-width watches use 1/2/4; larger sizes are passed through to PPSSPP as a range watch. PPSSPP matches memory breakpoints by address+size pair, so remove/update must pass the exact size recorded at set time."New value: +"Memory breakpoint watch size in bytes (mem_set / mem_remove / mem_update; default 4). Fixed-width watches use 1/2/4; larger sizes are passed through to PPSSPP as a range watch. NOTE (verified): mem_remove matches by ADDRESS only — a wrong or omitted size still removes the breakpoint at that address, so keep the size you set for bookkeeping, not for matching."
      • addedInput schema / properties / timeout_s
        Added value: +{
        +  "default": 30,
        +  "description": "Wait budget in seconds (wait / trace only; default 30, clamped to [0.5, 300]). On timeout: hit=false — NOT an error — so callers can poll.",
        +  "title": "Timeout S",
        +  "type": "number"
        +}
      • addedInput schema / properties / want_backtrace
        Added value: +{
        +  "default": false,
        +  "description": "Include the HLE backtrace in the hit (trace only; CPU is paused at the hit, so the trace is valid).",
        +  "title": "Want Backtrace",
        +  "type": "boolean"
        +}
      • addedInput schema / properties / want_registers
        Added value: +{
        +  "default": false,
        +  "description": "Include the full CPU register dump in the hit (trace only).",
        +  "title": "Want Registers",
        +  "type": "boolean"
        +}
      • addedOutput schema / properties / access
        Added value: +{
        +  "title": "Access",
        +  "type": "string"
        +}
      • removedOutput schema / properties / address / description
        Removed value: -"Breakpoint address, hex string (e.g. '0x08804000'); '0x00000000' for list / mem_list."
      • addedOutput schema / properties / already_paused
        Added value: +{
        +  "title": "Already Paused",
        +  "type": "boolean"
        +}
      • addedOutput schema / properties / bp_removed
        Added value: +{
        +  "title": "Bp Removed",
        +  "type": "boolean"
        +}
      • addedOutput schema / properties / by_pc
        Added value: +{
        +  "items": {
        +    "additionalProperties": true,
        +    "type": "object"
        +  },
        +  "title": "By Pc",
        +  "type": "array"
        +}
      • addedOutput schema / properties / condition
        Added value: +{
        +  "anyOf": [
        +    {
        +      "type": "string"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "title": "Condition"
        +}
      • addedOutput schema / properties / condition_filtered
        Added value: +{
        +  "title": "Condition Filtered",
        +  "type": "integer"
        +}
      • addedOutput schema / properties / filtered_hits
        Added value: +{
        +  "title": "Filtered Hits",
        +  "type": "integer"
        +}
      • addedOutput schema / properties / hit
        Added value: +{
        +  "title": "Hit",
        +  "type": "boolean"
        +}
      • addedOutput schema / properties / hits
        Added value: +{
        +  "items": {
        +    "additionalProperties": true,
        +    "type": "object"
        +  },
        +  "title": "Hits",
        +  "type": "array"
        +}
      • addedOutput schema / properties / mem_hits
        Added value: +{
        +  "items": {
        +    "additionalProperties": true,
        +    "type": "object"
        +  },
        +  "title": "Mem Hits",
        +  "type": "array"
        +}
      • addedOutput schema / properties / mode
        Added value: +{
        +  "title": "Mode",
        +  "type": "string"
        +}
      • addedOutput schema / properties / note
        Added value: +{
        +  "anyOf": [
        +    {
        +      "type": "string"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "title": "Note"
        +}
      • addedOutput schema / properties / pc
        Added value: +{
        +  "anyOf": [
        +    {
        +      "type": "string"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "title": "Pc"
        +}
      • addedOutput schema / properties / probe_changes
        Added value: +{
        +  "items": {
        +    "additionalProperties": true,
        +    "type": "object"
        +  },
        +  "title": "Probe Changes",
        +  "type": "array"
        +}
      • addedOutput schema / properties / reason
        Added value: +{
        +  "anyOf": [
        +    {
        +      "type": "string"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "title": "Reason"
        +}
      • addedOutput schema / properties / related_address
        Added value: +{
        +  "anyOf": [
        +    {
        +      "type": "string"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "title": "Related Address"
        +}
      • addedOutput schema / properties / resumed
        Added value: +{
        +  "title": "Resumed",
        +  "type": "boolean"
        +}
      • addedOutput schema / properties / storm_break
        Added value: +{
        +  "title": "Storm Break",
        +  "type": "boolean"
        +}
      • addedOutput schema / properties / ticks
        Added value: +{
        +  "anyOf": [
        +    {
        +      "type": "integer"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "title": "Ticks"
        +}
      • addedOutput schema / properties / timeout_s
        Added value: +{
        +  "title": "Timeout S",
        +  "type": "number"
        +}
      • addedOutput schema / properties / total_hits
        Added value: +{
        +  "title": "Total Hits",
        +  "type": "integer"
        +}
      • addedOutput schema / properties / window_s
        Added value: +{
        +  "title": "Window S",
        +  "type": "number"
        +}
      • removedOutput schema / required
        Removed value: -[
        -  "action",
        -  "address",
        -  "enabled",
        -  "breakpoints"
        -]
    • Addedppsspp_context
    • Removedppsspp_convert_address
    • Addedppsspp_diff_memory
    • Addedppsspp_dump
    • Removedppsspp_dump_clut
    • Removedppsspp_dump_texture
    • Removedppsspp_get_pc
    • Changedppsspp_health6 fields changed
      • addedInput schema / properties / session_id
        Added value: +{
        +  "anyOf": [
        +    {
        +      "type": "string"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null,
        +  "description": "Optional session ID — when provided, appends session_checks (the four-point battery: iso_loaded / cpu_running / ws_connected / game_mode_valid, absorbed from ppsspp_smoke_test) to the server-level report. Omit for the zero-contact server liveness probe.",
        +  "title": "Session Id"
        +}
      • addedOutput schema / $defs
        Added value: +{
        +  "_HealthSessionCheck": {
        +    "properties": {
        +      "detail": {
        +        "title": "Detail",
        +        "type": "string"
        +      },
        +      "name": {
        +        "title": "Name",
        +        "type": "string"
        +      },
        +      "passed": {
        +        "title": "Passed",
        +        "type": "boolean"
        +      }
        +    },
        +    "title": "_HealthSessionCheck",
        +    "type": "object"
        +  }
        +}
      • addedOutput schema / description
        Added value: +"HealthOutput + the per-session battery keys added when `session_id`\nis given (M9: they previously existed only in the text channel — the\noutput contract dropped them from structuredContent)."
      • addedOutput schema / properties / overall_session_status
        Added value: +{
        +  "title": "Overall Session Status",
        +  "type": "string"
        +}
      • addedOutput schema / properties / session_checks
        Added value: +{
        +  "items": {
        +    "$ref": "#/$defs/_HealthSessionCheck"
        +  },
        +  "title": "Session Checks",
        +  "type": "array"
        +}
      • changedOutput schema / title
        Previous value: -"HealthOutput"New value: +"HealthWithSessionChecks"
    • Removedppsspp_memory_info_search
    • Changedppsspp_press_button1 field changed
      • changedInput schema / properties / duration / description
        Previous value: -"Press duration in frames (default 1)."New value: +"Press duration in frames (default 1; 60fps wall-clock, cap 18000 ≈ 300s — values above are rejected). The call blocks for the duration."
    • Changedppsspp_query2 fields changed
      • addedInput schema / properties / safe
        Added value: +{
        +  "default": true,
        +  "description": "action=register/registers only: pause the CPU for a consistent read (trust_level='high', same as the retired ppsspp_get_pc) — or read without pausing (trust_level='low', zero cost, racy while running; for hot-path polling).",
        +  "title": "Safe",
        +  "type": "boolean"
        +}
      • addedInput schema / properties / size
        Added value: +{
        +  "anyOf": [
        +    {
        +      "type": "integer"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null,
        +  "description": "Function size in bytes, 'func_add' only. When omitted the server sends no size — on PPSSPP builds where the omit path underflows (v1.20.4-1845 and earlier) this produces an unusable zero-size function, so this tool defaults to sending 4. Pass an explicit size to override.",
        +  "title": "Size"
        +}
    • Changedppsspp_read_memory10 fields changed
      • changedInput schema / properties / action / description
        Previous value: -"Read action. Valid values:\n- 'read_bytes': read raw bytes (requires address + size).\n- 'read_u32': read a 32-bit unsigned int (requires address).\n- 'read_string': read a string (requires address).\n- 'scan': scan memory for a pattern (requires pattern + start_addr + end_addr). Optional max_results (default 100)."New value: +"Read action. Valid values:\n- 'read_bytes': read raw bytes (requires address + size).\n- 'read_u32': read a 32-bit unsigned int (requires address).\n- 'read_string': read a string (requires address).\npattern + start_addr + end_addr). Optional max_results (default 100)."
      • changedInput schema / properties / action / enum
        Previous value: -[
        -  "read_bytes",
        -  "read_u32",
        -  "read_string",
        -  "scan"
        -]New value: +[
        +  "read_bytes",
        +  "read_u32",
        +  "read_string"
        +]
      • changedInput schema / properties / address / description
        Previous value: -"Starting address for read_bytes/read_u32/read_string, as a hex string (e.g. '0x08804000'). Ignored for scan (use start_addr)."New value: +"Starting address for read_bytes/read_u32/read_string, as a hex string (e.g. '0x08804000')."
      • removedInput schema / properties / chunk_size
        Removed value: -{
        -  "default": 4096,
        -  "description": "Bytes per read request during scan (scan only, default 4096). Larger values reduce round-trips but increase per-read latency.",
        -  "title": "Chunk Size",
        -  "type": "integer"
        -}
      • removedInput schema / properties / end_addr
        Removed value: -{
        -  "default": "0x0",
        -  "description": "Scan end address, exclusive (scan only), hex string (same format as `address`).",
        -  "title": "End Addr",
        -  "type": "string"
        -}
      • changedInput schema / properties / max_len / description
        Previous value: -"Maximum string length in bytes for read_string (0 = default cap 4096). Always uses read_bytes + local NUL scan — PPSSPP memory.readString is never called (its strnlen scans to memory end and a giant response can kill the WebSocket). Values are clamped to 65536. Ignored for other actions."New value: +"Maximum string length in bytes for read_string (0 = default cap 4096). Values are clamped to 65536."
      • removedInput schema / properties / max_results
        Removed value: -{
        -  "default": 100,
        -  "description": "Maximum number of matches to return (scan only, default 100).",
        -  "title": "Max Results",
        -  "type": "integer"
        -}
      • removedInput schema / properties / pattern
        Removed value: -{
        -  "anyOf": [
        -    {
        -      "type": "string"
        -    },
        -    {
        -      "type": "null"
        -    }
        -  ],
        -  "default": null,
        -  "description": "Pattern to scan for (scan only). Interpreted according to `pattern_type`: 'hex' (default) expects even-length hex digits like 'AABBCCDD'; 'ascii' treats the string as literal ASCII bytes like 'hello'.",
        -  "title": "Pattern"
        -}
      • removedInput schema / properties / pattern_type
        Removed value: -{
        -  "default": "hex",
        -  "description": "How to interpret `pattern` (scan only). 'hex' (default) decodes as hex string; 'ascii' encodes the pattern string as literal ASCII bytes.",
        -  "enum": [
        -    "hex",
        -    "ascii"
        -  ],
        -  "title": "Pattern Type",
        -  "type": "string"
        -}
      • removedInput schema / properties / start_addr
        Removed value: -{
        -  "default": "0x0",
        -  "description": "Scan start address, inclusive (scan only), hex string (same format as `address`).",
        -  "title": "Start Addr",
        -  "type": "string"
        -}
    • Changedppsspp_replay6 fields changed
      • changedInput schema / properties / interval_ms / description
        Previous value: -"Polling interval in milliseconds for action='wait_complete' (default 100ms). Ignored for all other actions."New value: +"Polling interval in milliseconds for action='wait_complete' (default 100ms, clamped 10..5000 — below 10 the poll degenerates to a busy loop on the WS). Ignored for all other actions."
      • addedInput schema / properties / interval_ms / maximum
        Added value: +5000
      • addedInput schema / properties / interval_ms / minimum
        Added value: +10
      • changedInput schema / properties / timeout_ms / description
        Previous value: -"Total timeout in milliseconds for action='wait_complete' (default 10000 = 10s). Ignored for all other actions."New value: +"Total timeout in milliseconds for action='wait_complete' (default 10000 = 10s, clamped 100..60000). Ignored for all other actions."
      • addedInput schema / properties / timeout_ms / maximum
        Added value: +60000
      • addedInput schema / properties / timeout_ms / minimum
        Added value: +100
    • Addedppsspp_scan
    • Addedppsspp_search_memory_info
    • Changedppsspp_session11 fields changed
      • changedInput schema / properties / action / description
        Previous value: -"Session operation. Valid values:\n- 'start': launch a new PPSSPP session (requires iso_path). Set wait_ready=true to block until the emulated CPU is up (same probe/budget semantics as 'wait_ready').\n- 'stop': terminate an existing session (requires session_id).\n- 'get': query session health (requires session_id).\n- 'wait_ready': block until the emulated CPU has started (requires session_id). Call this after 'start' BEFORE any memory/disassembly tool — PPSSPP answers WebSocket before the ISO finishes booting, and early reads fail with 'CPU not started'."New value: +"Session operation. Valid values:\n- 'list': list all active sessions (no other params). Idle sessions (>30 min) are auto-GC'd as a side effect; returns {sessions, count}. NOT a per-session health probe — use ppsspp_health(session_id=…) for that.\n- 'start': launch a new PPSSPP session (requires iso_path). Set wait_ready=true to block until the emulated CPU is up (same probe/budget semantics as 'wait_ready').\n- 'stop': terminate an existing session (requires session_id).\n- 'get': query session health (requires session_id).\n- 'wait_ready': block until the emulated CPU has started (requires session_id). Call this after 'start' BEFORE any memory/disassembly tool — PPSSPP answers WebSocket before the ISO finishes booting, and early reads fail with 'CPU not started'."
      • changedInput schema / properties / action / enum
        Previous value: -[
        -  "start",
        -  "stop",
        -  "get",
        -  "wait_ready"
        -]New value: +[
        +  "list",
        +  "start",
        +  "stop",
        +  "get",
        +  "wait_ready"
        +]
      • addedOutput schema / $defs
        Added value: +{
        +  "SessionResponse": {
        +    "additionalProperties": false,
        +    "description": "Response view for a single session (start/stop/get).",
        +    "properties": {
        +      "created_at": {
        +        "description": "ISO 8601 timestamp of session creation.",
        +        "title": "Created At",
        +        "type": "string"
        +      },
        +      "exec_count": {
        +        "default": 0,
        +        "description": "Number of tool calls made against this session.",
        +        "title": "Exec Count",
        +        "type": "integer"
        +      },
        +      "iso_path": {
        +        "description": "Absolute path to the ISO file.",
        +        "title": "Iso Path",
        +        "type": "string"
        +      },
        +      "last_active_at": {
        +        "description": "ISO 8601 timestamp of last tool call.",
        +        "title": "Last Active At",
        +        "type": "string"
        +      },
        +      "pid": {
        +        "anyOf": [
        +          {
        +            "type": "integer"
        +          },
        +          {
        +            "type": "null"
        +          }
        +        ],
        +        "default": null,
        +        "description": "PPSSPP process PID (None if stopped).",
        +        "title": "Pid"
        +      },
        +      "ppsspp_version": {
        +        "anyOf": [
        +          {
        +            "additionalProperties": true,
        +            "type": "object"
        +          },
        +          {
        +            "type": "null"
        +          }
        +        ],
        +        "default": null,
        +        "description": "PPSSPP build fingerprint captured from the version handshake (None until the session transport binds).",
        +        "title": "Ppsspp Version"
        +      },
        +      "recovered": {
        +        "default": 0,
        +        "description": "H2: resilient-start relaunch count (0 = the first launch succeeded; >0 means the game state was reset by a wedge heal — breakpoints need re-arming).",
        +        "title": "Recovered",
        +        "type": "integer"
        +      },
        +      "restored": {
        +        "default": 0,
        +        "description": "F-6(a): 1 when this session was restored from sessions.json (a previous server run left it behind) rather than started fresh in this process — its game state may be stale.",
        +        "title": "Restored",
        +        "type": "integer"
        +      },
        +      "session_id": {
        +        "description": "Session UUID-like identifier.",
        +        "title": "Session Id",
        +        "type": "string"
        +      },
        +      "ws_connected": {
        +        "default": false,
        +        "description": "True if WebSocket is currently connected.",
        +        "title": "Ws Connected",
        +        "type": "boolean"
        +      },
        +      "ws_url": {
        +        "description": "WebSocket URL (ws://host:port/debugger).",
        +        "title": "Ws Url",
        +        "type": "string"
        +      }
        +    },
        +    "required": [
        +      "session_id",
        +      "iso_path",
        +      "ws_url",
        +      "created_at",
        +      "last_active_at"
        +    ],
        +    "title": "SessionResponse",
        +    "type": "object"
        +  }
        +}
      • addedOutput schema / properties / action
        Added value: +{
        +  "description": "Literal 'wait_ready' (echoes the session action).",
        +  "title": "Action",
        +  "type": "string"
        +}
      • addedOutput schema / properties / count
        Added value: +{
        +  "description": "Number of sessions.",
        +  "title": "Count",
        +  "type": "integer"
        +}
      • addedOutput schema / properties / elapsed_s
        Added value: +{
        +  "description": "Wall-clock seconds spent polling.",
        +  "title": "Elapsed S",
        +  "type": "number"
        +}
      • addedOutput schema / properties / note
        Added value: +{
        +  "anyOf": [
        +    {
        +      "type": "string"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "description": "Optional human context (e.g. fake-mode short-circuit).",
        +  "title": "Note"
        +}
      • addedOutput schema / properties / probe_addr
        Added value: +{
        +  "description": "Polled address, hex string (default top.prx base).",
        +  "title": "Probe Addr",
        +  "type": "string"
        +}
      • addedOutput schema / properties / probe_value
        Added value: +{
        +  "anyOf": [
        +    {
        +      "type": "string"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "description": "u32 read at probe_addr once ready, hex string (None in fake mode).",
        +  "title": "Probe Value"
        +}
      • addedOutput schema / properties / ready
        Added value: +{
        +  "description": "True when the CPU-start probe succeeded.",
        +  "title": "Ready",
        +  "type": "boolean"
        +}
      • addedOutput schema / properties / sessions
        Added value: +{
        +  "description": "Active sessions.",
        +  "items": {
        +    "$ref": "#/$defs/SessionResponse"
        +  },
        +  "title": "Sessions",
        +  "type": "array"
        +}
    • Removedppsspp_session_list
    • Removedppsspp_smoke_test
    • Changedppsspp_step3 fields changed
      • changedInput schema / properties / action / description
        Previous value: -"CPU step / run-state operation. Valid values:\n- 'into': step into (including delay slot).\n- 'over': step over (skip function calls).\n- 'out': step out of current function.\n- 'pause': pause CPU (enter stepping mode).\n- 'resume': resume CPU (exit stepping mode).\n- 'reset': reset the game (reboot).\n- 'run_until': run until the specified address is reached (requires address).\n- 'next_hle': step to next HLE callback."New value: +"CPU step / run-state operation. Valid values:\n- 'pause': pause CPU (enter stepping mode).\n- 'resume': resume CPU (exit stepping mode).\n- 'reset': reset the game (reboot).\n- 'run_until': run until the specified address is reached (requires address).\n- 'next_hle': step to next HLE callback.\nNOTE: single-stepping (into/over/out) lives in ppsspp_batch_step as the 'cpu_step' step type (mode='into'|'over'|'out', count 1..1000) — it requires the CPU to enter stepping mode, which the executor handles automatically."
      • changedInput schema / properties / action / enum
        Previous value: -[
        -  "into",
        -  "over",
        -  "out",
        -  "pause",
        -  "resume",
        -  "reset",
        -  "run_until",
        -  "next_hle"
        -]New value: +[
        +  "pause",
        +  "resume",
        +  "reset",
        +  "run_until",
        +  "next_hle"
        +]
      • changedInput schema / properties / address / description
        Previous value: -"Target address, as a hex string (e.g. '0x08804000'). Required for action='run_until'; ignored for all other actions."New value: +"Target address, as a hex string (e.g. '0x08804000'). Required for action='run_until'; ignored for all other actions. run_until is fire-and-forget: it returns immediately with no hit confirmation — poll PC or set a breakpoint to observe arrival."
    • Removedppsspp_trace_memory_access
    • Removedppsspp_wait_breakpoint
    • Addedppsspp_watch_value
    • Changedppsspp_write_register2 fields changed
      • changedInput schema / properties / name / description
        Previous value: -"CPU register name (MIPS standard names). GPRs: 'v0'/'v1'/'a0'-'a3'/'t0'-'t9'/'s0'-'s7'/'gp'/'sp'/'fp'/'ra'/'hi'/'lo'/'pc'. FPU: 'f0'-'f31'. VFPU: 'v0'-'v127'. Numeric aliases like 'r5' are NOT accepted by PPSSPP — use the MIPS standard name (e.g., 'a1' instead of 'r5'). Case-sensitive (lowercase by convention)."New value: +"CPU register name (MIPS standard names). GPRs: 'v0'/'v1'/'a0'-'a3'/'t0'-'t9'/'s0'-'s7'/'gp'/'sp'/'fp'/'ra'/'hi'/'lo'/'pc'. FPU: 'f0'-'f31'. VFPU: 'v0'-'v127'. Numeric aliases like 'r5' are accepted and normalized (r5 -> a1, i.e. GPR index 5) — prefer the MIPS standard name. Case-sensitive (lowercase by convention)."
      • changedInput schema / properties / value / description
        Previous value: -"Value to write, as a hex string (e.g. '0x00000001'). Treated as an unsigned 32-bit int; values outside [0, 0xFFFFFFFF] are wrapped by PPSSPP."New value: +"Value to write, as a hex string (e.g. '0x00000001'). Treated as an unsigned 32-bit int; values outside [0, 0xFFFFFFFF] are REJECTED (out-of-32-bit-range error), not wrapped."
  3. 1 tool updatev0.1.4
    • Changedppsspp_batch_step1 field changed
      • addedInput schema / $defs / PressStep / properties / button / enum
        Added value: +[
        +  "cross",
        +  "circle",
        +  "triangle",
        +  "square",
        +  "up",
        +  "down",
        +  "left",
        +  "right",
        +  "start",
        +  "select",
        +  "home",
        +  "screen",
        +  "note",
        +  "ltrigger",
        +  "rtrigger",
        +  "hold",
        +  "wlan",
        +  "remote_hold",
        +  "vol_up",
        +  "vol_down",
        +  "disc",
        +  "memstick",
        +  "forward",
        +  "back",
        +  "playpause"
        +]
  4. 41 tool updatesv0.1.0
    • First observedppsspp_analyze_log
    • First observedppsspp_assemble
    • First observedppsspp_batch_cancel
    • First observedppsspp_batch_list
    • First observedppsspp_batch_status
    • First observedppsspp_batch_step
    • First observedppsspp_breakpoint
    • First observedppsspp_convert_address
    • First observedppsspp_disassemble
    • First observedppsspp_dump_clut
    • First observedppsspp_dump_texture
    • First observedppsspp_evaluate
    • First observedppsspp_frame_snapshot
    • First observedppsspp_get_pc
    • First observedppsspp_gpu_record
    • First observedppsspp_gpu_stats
    • First observedppsspp_health
    • First observedppsspp_hold_buttons
    • First observedppsspp_list_addresses
    • First observedppsspp_list_scripts
    • First observedppsspp_memory_info_search
    • First observedppsspp_memory_map
    • First observedppsspp_press_button
    • First observedppsspp_query
    • First observedppsspp_read_memory
    • First observedppsspp_reload_scripts
    • First observedppsspp_replay
    • First observedppsspp_run_script
    • First observedppsspp_screenshot
    • First observedppsspp_search_disasm
    • First observedppsspp_send_analog
    • First observedppsspp_session
    • First observedppsspp_session_list
    • First observedppsspp_smoke_test
    • First observedppsspp_state_observer
    • First observedppsspp_step
    • First observedppsspp_trace_memory_access
    • First observedppsspp_wait_breakpoint
    • First observedppsspp_wait_frames
    • First observedppsspp_write_memory
    • First observedppsspp_write_register

TDQS

A3.9/5.0

Scored across 37 tools

Disambiguation4/5

The tool set has several clusters that a naive agent could confuse (read_memory / query / state_observer / frame_snapshot / watch_value for state reads; step / batch_step for stepping; watch_value / breakpoint(trace) / state_observer for observing), but nearly every description carries an explicit ROUTING block that names the intended tool for each case, which sharply reduces misselection. Boundaries remain slightly blurred by design given the mega-tools (breakpoint, query, scan, replay) that fold many actions into one.

Naming Consistency5/5

Every tool uses the identical ppsspp_ snake_case prefix and mostly a verb_noun shape (read_memory, write_register, press_button, hold_buttons, batch_step), with a few noun-named aggregators (context, health, session, replay). The convention is uniform and predictable throughout.

Tool Count3/5

37 tools is on the heavy side for a single MCP server and several obvious groupings (batch_step/batch_status/batch_cancel, list_scripts/run_script/reload_scripts) could plausibly be folded together. The domain genuinely spans sessions, memory, breakpoints, GPU, input, replay, and scripting, so the breadth is partly earned, but it sits at the upper edge of comfortable.

Completeness5/5

The surface covers the full debugging lifecycle end to end: session start/stop/wait, memory read/write/scan/diff, disassembly and search, breakpoints and value watches, register/PC evaluation, GPU stats/dump/record, screenshots, input, replay, batching, and script management. There are no obvious dead ends for the stated PPSSPP-debugging purpose.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    A
    quality
    B
    maintenance
    An MCP server that exposes PPSSPP — the PlayStation Portable emulator — to any MCP-compatible client (Claude Desktop, Claude Code, etc.) via PPSSPP's built-in WebSocket debugger interface. Read and write PSP memory, drive games with button input, capture screenshots, set CPU breakpoints, inspect MIPS Allegrex registers — all through a clean tool interface. No bridge plugin needed; PPSSPP's debugg
    23
    10 npm
    9
    MIT
  • A
    license
    Not graded
    quality
    C
    maintenance
    Enables AI agents to interact with GDB for debugging via the MCP protocol. Supports setting breakpoints, stepping through code, inspecting memory and registers, and more.
    85
    MIT
  • A
    license
    B
    quality
    D
    maintenance
    Enables AI agents to programmatically control Cheat Engine for memory scanning, editing, cheat table management, and reverse engineering tasks.
    42
    MIT