VoltageInputMcp
VoltageInputMcp
一个 MCP 服务器,让前沿模型以输入速度而非工具调用速度驱动计算机。
问题
计算机使用工具每次操作都要与远程模型往返一次。截图上传、决策下发、一次点击。这对填写表单没问题,但对任何需要快速连续输入序列的场景毫无用处——玩游戏、操作模态对话框、驱动时间轴、任何第三个输入依赖于前两个已经落地的 UI。瓶颈不是模型的智能。而是智能在 800 毫秒之外,而输入需要间隔 8 毫秒。
Related MCP server: live-mcp
答案的形态
将决策与执行分离,把执行放在与键盘同一台机器上。
┌─────────────────────────────────────────────────────────────────┐
│ Layer 1 — the orchestrator (Claude, or any MCP client) │
│ Writes a Playbook: states, what to look for, what is allowed, │
│ when to move on. Thinks once, up front. Watches and corrects. │
└───────────────────────────┬─────────────────────────────────────┘
│ MCP
┌───────────────────────────▼─────────────────────────────────────┐
│ Layer 2 — two small local models, on your GPU │
│ │
│ vision (Qwen2.5-VL-3B) "of these specific things, │
│ which are on screen, and where?" │
│ actuator (Qwen3-1.7B) "given that, which inputs?" │
│ │
│ Neither plans. Both answer one closed question per cycle. │
└───────────────────────────┬─────────────────────────────────────┘
│
┌───────────────────────────▼─────────────────────────────────────┐
│ safety governor → /dev/uinput → the actual desktop │
└─────────────────────────────────────────────────────────────────┘编排器是大脑。小模型是手臂。手臂不聪明,也从不被要求聪明。
速度实际上来自哪里
不是来自小模型速度快——一个 3B VLM 仍然需要约 300 毫秒。它来自四件事,按影响从大到小排列:
突发(Bursts)。 执行器不发出单个输入。它发出一个突发:一个由专用执行器运行、无模型参与的定时输入程序。
g:0;c:l;w:150;t:"README.md";k:enter;w:80;k:ctrl+s那是一次决策和七个输入,跨越约 400 毫秒,精确到毫秒级调度。一个 40 动作的突发仍然只花一次决策。输入速率由突发决定,而非模型。
反射(Reflexes)。 在决策之间,以微秒级触发廉价屏幕探针——一个像素、一个区域平均值——完全不需要模型。
{"id": "heal", "when": "probe('health') < 0.25", "do": "k:q;w:60", "cooldown_ms": 800}跳过感知。 大多数循环看到的屏幕没有变化。一个 40 微秒的帧差决定是花 300 毫秒在视觉模型上,还是复用上一次的观察结果。在普通桌面工作中,这会在大多数循环中跳过 VLM。
提示缓存局部性。 提示按静态优先排序,使 llama.cpp 复用 KV 缓存,只对变化的尾部重新预填充。
为什么小模型虽小却可靠
因为它们不被要求可靠——它们被约束。
在 llama.cpp 下,两个模型都根据GBNF 语法生成,该语法每个循环根据当前状态重新生成。语法不是建议。它屏蔽 logits,使得只有延续有效解析的 token 可达。具体来说,执行器不能:
发出格式错误的突发
命名策略拒绝的键——该键不在语法中
引用未被观察到的元素——索引范围是根据本循环的元素数量构建的
提出 Playbook 未声明的状态转换
而视觉模型不能凭空发明 UI 元素名称:其标签词汇表是你编写的 watch 列表,外加一小套通用集合。因此 sees("address bar") 守卫比较的是封闭词汇表,而不是 3B 模型随意产出的任何名词。
没有重试循环,也没有防御性 JSON 解析,因为格式错误的输出不是不太可能——而是不可表示。
Playbook
你不给小模型一个目标。你给它们一个状态机。转换是守卫表达式,由运行时求值,而非模型。
{
"name": "open_downloads",
"goal": "Open the file manager at ~/Downloads. Delete nothing, confirm nothing.",
"initial": "launch",
"policy": {
"dry_run": true,
"allow_verbs": ["g", "c", "k", "t", "w"],
"deny_labels": ["delete", "trash", "confirm", "empty trash"]
},
"budget": { "max_cycles": 60, "max_seconds": 90 },
"states": {
"launch": {
"brief": "Open the application launcher and start the file manager.",
"watch": ["application launcher", "search field", "file manager icon"],
"on_enter": "k:meta;w:400",
"transitions": [
{ "when": "sees('search field')", "to": "type_name" },
{ "when": "cycles() > 6", "to": "@failure", "note": "launcher never opened" }
]
},
"navigate": {
"brief": "Focus the location bar with ctrl+l, type the path, press Enter.",
"watch": ["location bar", "file list", "error message"],
"on_enter": "k:ctrl+l;w:200",
"transitions": [
{ "when": "text('Downloads')", "to": "@success" },
{ "when": "sees('error message')", "to": "@failure" }
]
}
},
"success_when": "text('Downloads') and not flag('loading')"
}voltage_reference 返回完整的 DSL、JSON schema 和守卫函数表,因此编排器无需阅读本仓库即可编写 Playbook。
性能调优
以下所有数字都是在参考机器(RTX 3050 6 GB 笔记本,llama.cpp 下的 Qwen2.5-VL-3B + Qwen3-1.7B)上实测的,而非推导。
两个模型都受解码限制。输出 token 是唯一重要的杠杆。
这出乎意料——设计最初假设视觉受预填充限制,但事实并非如此。预填充实测约 28 毫秒且平坦,从 448×252 到 896×504 不变。解码运行在约 22 毫秒/token。因此:
项目 | 成本 |
一个输出 token | 约 22 毫秒 |
一个报告的元素 | 约 21 个 token ≈ 500 毫秒 |
视觉,2 个元素 | 约 1.0 秒 |
视觉,4 个元素 | 约 2.2 秒 |
执行器,缓存前缀 | 140–400 毫秒,取决于 note 长度 |
三个后果,每一个都改变了一个默认值:
max_elements是视觉成本的主导因素。 默认值为 3。提高到 6 每个感知循环增加约 1.5 秒。将其设置为你守卫实际测试的数量。缩小
downscale_to没有帮助,通常反而有害。 448×252 实测比 896×504 慢 2.5 倍——图像越模糊,模型越不确定,因此输出更多 token。使用能容纳的最大尺寸。执行器的
note字段占其延迟的 55%。 它纯粹是诊断性的,48 字符时实测 412 毫秒/循环,而 12 字符时为 184 毫秒,0 字符时为 140 毫秒。现在默认为 12。
元素编码为 [label_index, x1, y1, x2, y2] 而非 {"l":"address bar","b":[...],"c":0.9},原因相同——实测减少 27–29% 的 token,降低 32–41% 的延迟。索引到封闭的 watch 词汇表也更安全:模型根本无法拼写标签,更不用说拼错了。
GBNF 求值在 CPU 上每个采样 token 运行一次,因此执行器比视觉模型获得更多 CPU 线程,尽管完全 GPU 卸载——而限制 allow_keys 不仅是安全优化,也是延迟优化。
两个设置如果错误会静默失败:
构建时设置
GGML_CUDA_FA_ALL_QUANTS=ON。 我们使用q8_0KV 缓存和闪存注意力提供服务。没有这个标志,llama.cpp 不会为该 KV 组合编译 FA 内核,并回退到慢速路径——没有错误,只是神秘地糟糕的数字。scripts/build-llama.sh会设置它。运行时设置
GGML_CUDA_ENABLE_UNIFIED_MEMORY=0。 如果为1,VRAM 溢出会静默地溢出到 PCIe 而不是失败。一切正常但慢约 10 倍。serve.sh将其固定为关闭。
测量而非猜测:
.venv/bin/voltage bench它使用循环使用的精确提示形状驱动两个后端,并报告冷启动与提示缓存延迟、三种输入尺寸下每视觉 token 的毫秒数,以及这些所隐含的循环时间。提示缓存加速低于约 1.5 倍意味着有动态内容泄漏到提示前缀中。
比较模型
显而易见的实验——"哪个模型写突发更好"——衡量的是错误的东西。语法已经保证每个突发都是有效的,所以更大的模型不能在语法上获胜。真正决定配置是否可用的是:
接地准确性。 一个快 200 毫秒但偏了 40 像素的模型毫无用处——点击会落空。以屏幕像素的中心距离衡量,而非 IoU,因为点击落在中心。
约束下的决策质量。 给定相同的观察,它是否选择正确的合法动作,以及它是否将整个序列链接到一个突发中,而不是每个循环发出一个胆小的动作?
延迟,只有在 1 和 2 可接受之后才重要。
.venv/bin/voltage fixture desktop # capture a real screen
.venv/bin/voltage compare # score whatever is running now地面真值来自由编排模型标注的真实截图——这与系统在运行时使用的参考相同。合成 UI 是一个陷阱:画出来的矩形在训练于真实界面的模型眼中不像按钮,因此针对它评分衡量的是错误的技能。
结果跨运行累积,因此工作流是:服务配置文件 A → compare → 服务配置文件 B → compare → 读取表格。voltage compare --list 无需重新运行即可打印。
夹具是你自己的,不提交。如果你的截图包含任何私密内容,请将 fixtures/ 添加到 .gitignore。
学习循环
针对陌生目标的第一个 Playbook 几乎从不是正确的。重要的是失败是具体的,并且下一次尝试从上一次学到的开始。
voltage_reference(section="loop") the loop itself, and what each failure means
voltage_reference(section="bursts") the burst cookbook: chaining, timing, game patterns
voltage_capture / voltage_observe look before writing — check your labels exist
voltage_validate_playbook dead guards, unreachable states, caught statically
voltage_run(dry_run=true) real models, real screen, nothing injected
voltage_diagnose(run_id) ← what to change, not raw data
voltage_learn(target=..., note=...) record it; persists across sessions
voltage_lessons(target=...) recall it before the next playbookvoltage_diagnose 是使这成为循环的关键部分。 它计算日志暗示但未陈述的内容,并为每项命名对应的编辑。在一次卡住的 Minecraft 运行中:
[BLOCKER] label_never_seen never reported: ['crosshair', 'health bar']
[BLOCKER] input_not_landing 14 bursts executed, but the screen never changed
[BLOCKER] state_never_left 'mine' ran 14 cycles and never transitioned
[PROBLEM] timid_bursts bursts averaged 1.0 actions
[HINT] vision_every_cycle vision ran on 100% of cycles它存在的意义在于区分:从未运行的突发和运行了但没起作用的突发在摘要中看起来相同,但原因无关。 前者是策略或语法问题。后者是窗口焦点、指针模式或忽略合成输入的应用程序。Diagnose 通过检查执行后帧是否实际变化来区分它们。
应用最高严重性的发现,重新运行,再次诊断。一次只改一个——同时改多个会使下一次诊断无法解读。
经验跨会话持久化,按目标为键,因此游戏的第二个 Playbook 从第一个发现的探针坐标和可用标签名称开始:
voltage_learn(target="minecraft", kind="label",
note="vision reports 'hotbar' reliably but never 'crosshair'")
voltage_learn(target="minecraft", kind="timing",
note="block placement needs w:100 after right click or it does not register")安全
生成输入的是一个 1.7B 模型。治理层不是建议性的:每个突发都经过它,包括反射突发和你自己编写的突发。
dry_run是默认值。 新的 Playbook 解析、检查并记录每个突发,同时不触碰任何东西。整个突发拒绝。 半执行一个预期的序列比不执行更糟。
deny_labels拒绝点击任何名为 Delete / Confirm / Purchase / Allow 的东西,无论它出现在哪里——这能捕获在意外位置弹出的对话框。区域围栏、键允许列表、被拒绝的和弦(
ctrl+alt+delete、alt+f4)、被拒绝的文本模式(rm -rf、sudo)、突发大小和每秒输入数上限。四个独立的停止机制:
voltage stop(写入文件——可通过 SSH 工作)、在循环卡住时在自己的线程上触发的安全定时器、物理输入竞争(触摸真实鼠标即停止),以及 Playbook 预算。按住的键始终释放——中止时、崩溃时、超时时。在
d:shift和u:shift之间中断的运行绝不能留下 Shift 卡住。
安装
从零到可用,两条命令。
Linux / macOS
git clone https://github.com/casualkre/voltage-input-mcp && cd voltage-input-mcp && ./install.shWindows(PowerShell)
git clone https://github.com/casualkre/voltage-input-mcp; cd voltage-input-mcp; powershell -ExecutionPolicy Bypass -File .\install.ps1然后,在任一系统上:
voltage setupinstall.sh 处理 Python、系统包、venv 和你的 PATH,并打印需要 root 权限的确切 sudo 命令,而不是请求它。然后 voltage setup 检测你已有的内容,只下载缺失的部分,启动模型服务器,并向你的 AI 客户端注册——运行每一步,而不是描述它。十到二十五分钟,几乎全部是下载时间。可安全重新运行;它会从上次中断的地方继续。
然后只需运行:
voltageSetup 检测你已有的内容并从那里继续。 它不假设起点:它探测你的操作系统、GPU、是否安装了 llama.cpp 或 Ollama、已拉取了哪些模型、输入和捕获是否工作、MCP 服务器是否已注册——然后只规划实际剩余步骤,并说明哪些需要你决定、哪些它可以直接做。如果你已有 Ollama,它会使用它。如果两个后端都没有,它用两行解释权衡并让你选择。
无参数运行会打开交互式控制台:实时状态、按依赖顺序修复未就绪内容的引导式设置、模型切换器、配置编辑器、一键注册 Claude Code,以及诊断。下面的每个子命令仍然可以非交互式工作,因此脚本和 CI 不受影响。
██╗ ██╗ ██████╗ ██╗ ████████╗ █████╗ ██████╗ ███████╗
██║ ██║██╔═══██╗██║ ╚══██╔══╝██╔══██╗██╔════╝ ██╔════╝
██║ ██║██║ ██║██║ ██║ ███████║██║ ███╗█████╗
╚██╗ ██╔╝██║ ██║██║ ██║ ██╔══██║██║ ██║██╔══╝
╚████╔╝ ╚██████╔╝███████╗██║ ██║ ██║╚██████╔╝███████╗
╚═══╝ ╚═════╝ ╚══════╝╚═╝ ╚═╝ ╚═╝ ╚═════╝ ╚══════╝
── status ──────────────────────────────────────────────
ok input device /dev/uinput
ok vision model http://127.0.0.1:8080
ok actuator model http://127.0.0.1:8081
ok mcp registered claude mcp list
ok voltage on PATH ~/.local/bin/voltage实验性配置文件
在 voltage → models 中单独列出,每个都在你必须接受的警告之后。它们存在是因为测量使权衡可预测:解码在约 22 毫秒/token 占主导地位,并随激活参数扩展,因此缩小模型确实会提高循环速率。代价是接地准确性。
profile | models | VRAM | trade |
| SmolVLM-500M + Qwen3-0.6B | ~2.2 GB | 3–4× 循环速率,接地(grounding)几乎不可用 |
| Qwen2.5-VL-3B + Qwen3-0.6B | ~3.8 GB | 决策更快,接地不变 |
| Qwen2.5-VL-32B + Qwen3-14B | ~34 GB | 最佳接地,1–2.5 秒/循环 |
| Qwen2.5-VL-32B + Qwen3-30B-A3B | ~43 GB | 30B 容量,~3B 解码速度 |
| 3B + 0.6B 在 CPU 上 | 无 | 无需 GPU 即可工作,每循环数秒 |
有两个值得单独说明:
hyper 是危险的那个。 SmolVLM-500M 不是接地模型。它会返回
边界框,而且经常是错的——而错误的边界框就是点到错误的位置,而不是
优雅降级。只在 watch 为空(由探针和反射完成实际工作)时使用它,
或者当每次点击都被 click_allow_regions 和
require_target_element 围栏保护时使用。
beefy_moe 是有趣的那个。 Qwen3-30B-A3B 是一个混合专家模型,有 ~3B
激活参数,因此它以大约 3B 的速度解码,同时以 30B 的容量进行推理——
而解码恰恰是这个循环的瓶颈。在相似的延迟下,它比稠密的 14B 是更好的执行器。
问题在于内存:只有激活的专家是快的,权重不是,所以全部 30B 仍然必须驻留在内存中。
recommend() 永远不会返回实验性 profile,并且有测试强制执行这一点。
自定义模型 profile
内置 profile 覆盖的是开发此项目所用的机器,而不是你的机器。从
voltage → profiles 添加你自己的,或者编辑配置文件旁边的 profiles.toml:
[my_rig]
description = "RTX 4090"
[my_rig.vision]
hf_repo = "ggml-org/Qwen2.5-VL-7B-Instruct-GGUF"
hf_file = "Qwen2.5-VL-7B-Instruct-Q4_K_M.gguf"
mmproj_file = "mmproj-Qwen2.5-VL-7B-Instruct-Q8_0.gguf"
params_b = 7.0
weights_mb = 4700
n_ctx = 4096
port = 8080
[my_rig.actuator]
hf_repo = "unsloth/Qwen3-4B-Instruct-2507-GGUF"
hf_file = "Qwen3-4B-Instruct-2507-Q4_K_M.gguf"
params_b = 4.0
weights_mb = 2500
port = 8081自定义 profile 按名称覆盖内置 profile,因此将一个命名为 lean 会重新调整
内置 profile 的参数,而无需分叉整个包。对于 Ollama 后端,使用 ollama_tag
而不是 hf_repo/hf_file。
有一个槽位很挑剔,另一个则不然。视觉模型必须能够在请求时发出接地边界框—— Qwen2.5-VL、Qwen3-VL、InternVL、MiniCPM-V 和 UI-TARS 都可以;而普通的 图像描述模型会漂亮地描述你的屏幕,却把边界框放在错误的位置。 执行器则很宽容:在 GBNF 语法下,它只是在少数几个合法的 续写选项中选择,所以几乎任何称职的 1B+ 指令模型都能工作。
Shell 命令 vs MCP 工具
两个不同的表面,混淆它们是最常见的第一个绊脚石:
调用方式 | 样子 | |
shell 命令 | 在终端中输入,带空格 |
|
MCP 工具 | 向 Claude 请求,带下划线 |
|
voltage_doctor 是 Claude 命名空间中的工具名,而不是磁盘上的程序。在终端中输入它
总会显示"未知命令"。请让 Claude 来运行它。
这会检查 /dev/uinput 访问权限、安装系统依赖、创建 venv,并
打印缺失的内容。然后:
./scripts/fetch-models.sh lean && ./scripts/serve.sh lean.venv/bin/voltage doctor将其连接到客户端
voltage connect显示已设置的内容、实时 URL、模型是否已启动,以及服务器是否已 注册——然后为每个客户端提供复制粘贴步骤,你的真实路径和环境 已经填好:
voltage connect --client claude-desktop
voltage connect --client cursor
voltage connect --json # just the mcpServers entry涵盖:Claude Code、Claude Desktop、claude.ai 自定义连接器、Cursor、Windsurf、Zed,
以及用于其他任何东西的通用 mcpServers 块。同样的内容也是 voltage
控制台中的第 4 屏,它还可以为你写入 Claude Desktop 配置(先备份
现有文件,如果它不是有效的 JSON 则拒绝修改)。
每个生成的配置都显式携带会话环境,因为这才是出错的地方:从没有
DBUS_SESSION_BUS_ADDRESS 的 shell 注册的服务器可以成功连接,但会静默失明——
输入正常,屏幕捕获不行。voltage connect 会检测到这种情况并明确说明。
将其添加为自定义连接器
通过 URL 添加 MCP 服务器的客户端需要 HTTP 而不是 stdio:
voltage serve --http然后添加 http://127.0.0.1:8765/mcp 作为自定义连接器。
绑定仅限于回环地址,并且需要 --allow-remote 才能更改。
这不是多余的样板:这个服务器的存在是为了移动鼠标、按键和读取
屏幕,而 MCP 本身没有认证机制。非回环绑定会发布
未经认证的桌面远程控制。如果你确实需要,请在前面放置一个带认证的
反向代理,并明白任何能访问该端口的人都拥有这台机器。
从 MCP 客户端启动
MCP 客户端以净化过的环境启动服务器——PATH、HOME 和
几乎没有其他东西。这是一个合理的默认值,但它会破坏屏幕捕获,因为访问
合成器需要 DBUS_SESSION_BUS_ADDRESS 和 WAYLAND_DISPLAY。输入注入
在没有它们的情况下仍然可以工作(uinput 是设备文件,不是会话服务),所以故障
看起来令人困惑地不完整:突发执行了,截图却没有。
显式传递它们:
claude mcp add voltage-input \
-e WAYLAND_DISPLAY="$WAYLAND_DISPLAY" \
-e DISPLAY="$DISPLAY" \
-e DBUS_SESSION_BUS_ADDRESS="$DBUS_SESSION_BUS_ADDRESS" \
-e XDG_RUNTIME_DIR="$XDG_RUNTIME_DIR" \
-- /absolute/path/to/voltage-input-mcp/.venv/bin/voltage-input-mcpvoltage_doctor 会准确报告哪些缺失,所以如果捕获失败,
这是第一个要查看的地方。
平台
输入 | 捕获 | 文本 | |
Linux |
| portal→PipeWire、KWin DBus、grim、X11 | 扫描码,非 ASCII 字符的剪贴板回退 |
Windows |
| GDI |
|
输入汇之上的所有内容——突发调度、时序、按住键跟踪、安全
调节器、整个运行时——都是共享的。每个平台实现五个方法
(key、button、move_abs、move_rel、scroll);参见 inputs/sink.py。
有两个值得了解的对称性差异:
在 Windows 上打字更正确。
KEYEVENTF_UNICODE传递 UTF-16 代码单元, 不涉及键盘布局。Linux uinput 发送扫描码,所以非美式 布局上的标点会出错——而且是静默的——这就是为什么剪贴板回退 在那里存在而在 Windows 上不需要。在 Linux 上捕获能力更强。 GDI
BitBlt无法看到某些硬件覆盖层 视频和全屏独占游戏;这些会捕获为黑色。请以 无边框窗口模式运行此类游戏。
在 Windows 上,SendInput 无法驱动属于提升进程的窗口(UIPI)——这
会静默失败,所以 voltage doctor 会报告你的提升状态。DPI 感知
在导入时声明;没有它,在缩放显示器上每个坐标都是错的。
要求
Linux(任何显示服务器)或 Windows 10/11
Python 3.11+
一个 GPU,
leanprofile 需要 ~5 GB 空闲;voltage profiles显示适合你的配置快速路径需要 llama.cpp,或者较慢的零构建路径需要 Ollama
已在 KDE Plasma 6 / Wayland / CUDA / Python 3.14 上端到端验证。Windows 路径 已实现并通过类型检查,但尚未在 Windows 机器上运行——请将其视为 未经测试,并报告任何问题。
编排器被告知它驱动的是哪个构建
同一个 Playbook 在一个配置上是正确的,在另一个配置上就是错误的,而远程模型 无法看到是哪个。因此服务器的 MCP 指令是在启动时根据实时配置 构建的,只携带那些会改变 Playbook 编写方式的条目:
ACTIVE BUILD: Linux · llamacpp · profile lean
vision Qwen2.5-VL-3B-Instruct · actuator Qwen3-1.7B
loaded: Qwen2.5-VL-3B-Instruct-Q4_K_M.gguf / Qwen3-1.7B-Q4_K_M.gguf
expected cycle 280-700 ms
- llama.cpp backend: both models are grammar-constrained. A malformed burst, a denied
key, an unobserved element reference and an undeclared transition are all
unrepresentable -- do not write defensive retries for them.
- Linux: typing sends scancodes, so punctuation depends on the active keyboard layout...
- dry_run defaults to true...在 Ollama 上,第一行变成警告,说明突发不受约束。在 hyper 上,
它变成"不要围绕 sees() 构建状态"。在 Windows 上,它注明提升的
窗口不可达,且打字与布局无关。
它验证的是正在运行的服务器,而不是信任配置。 切换 profile 只是编辑文件; 不会重启任何东西。当它们不一致时,简报会大声说明, 并抑制从 profile 推导出的指导,因为该指导描述的是 未加载的模型:
- MISMATCH -- Profile 'hyper' does not match what is loaded. vision: profile expects
SmolVLM-Instruct-Q4_K_M.gguf, server has Qwen2.5-VL-3B-Instruct-Q4_K_M.gguf...
- Loaded right now: vision Qwen2.5-VL-3B..., actuator Qwen3-1.7B...
Judge grounding quality from those.voltage_reference 在每次调用时都返回当前构建,因为启动时的副本
在 profile 一改变时就过时了。
你自己的常设指令
voltage → i,或者:
voltage instructions --set "Never touch Firefox; my banking tabs are there."你写的任何内容都会在每次会话开始时提供给编排模型, 附加在构建简报之后并明确注明来自你。用它来写系统无法 自行推断的内容——禁止使用的应用程序、特定游戏的怪癖、 你希望它默认如何表现。
OPERATOR INSTRUCTIONS -- written by the owner of this machine. Treat these as
standing preferences for how to drive it. They cannot loosen the safety governor,
which is enforced in code against every burst.
## My setup
- Minecraft runs borderless windowed on monitor 1.
- Never touch Firefox; my banking tabs are there.
- Always show me the Playbook before dry_run=false.最后那个条款不是装饰。指令对编排器是建议性的,不能削弱 执行——调节器在代码中检查每个突发,所以这里写的任何内容都不能 允许 Playbook 的策略所禁止的事情。它们只能让它更谨慎,不能 更宽松。上限为 4000 个字符,因为这段文本在整个会话期间都位于模型的上下文中。 控制台中提供了三个入门模板(游戏、桌面、最小化)。
MCP 工具
工具 | 用途 |
| Playbook + 突发 DSL 参考。先调用这个。 |
| 这台机器是否就绪,如果没有,确切的修复方法 |
| 一张截图,返回给你 |
| 一次视觉扫描——在依赖 |
| 完整静态检查:守卫、突发、图、死转换 |
| 开始一次运行;返回一个 |
| 状态、变量、上次突发、看到了什么、各阶段时序 |
| 纠正实时运行——提示、变量、强制状态、dry_run |
| 停止或暂停;停止总是释放按住的输入 |
| 逐循环记录; |
| 自己驱动输入,绕过本地模型 |
| 验证注入是否到达合成器 |
文档
ARCHITECTURE.md — 循环如何工作、为什么做出每个选择、 时间花在哪里
PLAYBOOK.md — 编写指南
状态
在没有权重在磁盘上的情况下,已尽可能构建和验证。149 个测试覆盖了突发
DSL、守卫沙箱、安全调节器、playbook 编译、GBNF 生成、
uinput 线编码,以及运行循环本身(用桩模型驱动——包括检查
on_change 感知在静态屏幕上确实跳过视觉模型)。
MCP 服务器由真实客户端通过 stdio 端到端驱动:13 个工具、正确的
schema、execute_burst 接受了一个有效的突发并拒绝了 sudo rm -rf /,
两条匹配规则都生效。
尚未运行的是实时模型:这需要构建 llama.cpp 并获取权重,
scripts/ 会设置这些。构建期间也有两件事被故意没有触发——
门户权限对话框,以及任何真实的输入注入——因为两者都会作用于
你的桌面。
从这里开始的执行顺序:
./scripts/setup.sh # reports what needs sudo, doesn't run it
./scripts/build-llama.sh # ~15 min with CUDA
./scripts/fetch-models.sh lean
./scripts/serve.sh lean
.venv/bin/voltage doctor # should now say READY然后在 MCP 客户端中:voltage_calibrate(观察光标实际移动)、voltage_observe(检查视觉模型能找到你的标签),然后运行一个 dry_run Playbook,并在设置 dry_run=false 之前阅读 voltage_journal。
作者署名
由 Claude Opus 5(Anthropic)在单次会话中端到端编写——架构、实现、测试和文档。人类指定了想法、设定了约束(KDE Wayland、6 GB VRAM、"比 computer-use 更快")并审阅了结果,但没有编写代码。
本仓库中沉淀的平台发现来自构建过程中对机器的探测,而非凭空假设——KWin 拒绝向非白名单可执行文件提供 ScreenShot2、grim 在 KWin 下无法工作、MCP 客户端会清除会话总线。每一条都在代码中迫使做出决策的位置有相应文档记录。
LICENSE 未将任何个人列为版权持有人,相关理由已在该文件中完整写明。
许可证
MIT。参见 LICENSE。
Available Tools
16 toolsvoltage_calibrateADestructive
Verify that input injection actually reaches the compositor.
Creates the virtual devices, moves the pointer to three known points, and captures after each to confirm the cursor moved. Reports whether absolute positioning works or whether the relative fallback is needed -- which cannot be known without trying, since it depends on how libinput classified the virtual device.
Run this once per machine before trusting a real (non-dry-run) Playbook.
| Name | Required | Description | Default |
|---|---|---|---|
| dry_run | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotations already declare destructiveHint=true and openWorldHint=true. The description adds valuable detail: it creates virtual devices, moves the pointer, and captures output—concrete side effects beyond the annotation. It also explains why these behaviors are unpredictable ('depends on how libinput classified the virtual device'), which aligns with openWorldHint. This goes beyond what annotations alone convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise and well-structured. It opens with the core purpose, immediately explains what the tool does, then provides the rationale and usage timing. Every sentence earns its place—no fluff or repetition. It is front-loaded with the most critical information and stays focused.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers the tool's purpose, mechanism, side effects, and when to run it. It also acknowledges an output summary (absolute vs relative fallback). However, it omits any explanation of the dry_run parameter, which is a key input that affects behavior. Given the presence of an output schema, return format doesn't need detailing, but the parameter gap leaves the description incomplete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The only parameter, dry_run, is entirely undocumented in the description. The schema gives its name, type, default, and requirement status but provides no semantic meaning. The description mentions 'non-dry-run' indirectly but never explains what dry_run does, when to set it to true, or what the difference is. With 0% schema description coverage, this is a significant gap—the description fails to compensate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's specific purpose: to verify that input injection reaches the compositor. It describes the concrete steps (creating virtual devices, moving pointer to three points, capturing) and the distinctive outcome (determining whether absolute positioning or relative fallback is needed). This distinguishes it from the many sibling tools, even without naming them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives a clear usage context: 'Run this once per machine before trusting a real (non-dry-run) Playbook.' It explains that this is a one-time calibration step and that the result cannot be known without trying, which implies this is the tool to use for that purpose. It doesn't explicitly mention alternatives or when not to use it, but the guidance is strong.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
voltage_captureARead-only
Take a screenshot and return it to you directly.
Use this to see the screen yourself -- before writing a Playbook, to pick coordinates for probes and click regions, or to work out why a run went wrong. This does not involve the local vision model.
region is [x, y, width, height] in desktop pixels; omit for the whole desktop.
| Name | Required | Description | Default |
|---|---|---|---|
| region | No | ||
| max_width | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=false, so the description's job is to add behavioral context. It does so by specifying 'return it to you directly' and explicitly noting it does not use the local vision model—information not in annotations. No contradictions exist.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is tightly written: a clear one-sentence purpose, a brief usage block, and a one-line parameter definition. No fluff or redundancy. The key purpose is front-loaded, and every sentence contributes value.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple screenshot tool, the description covers usage, region format, and direct return. However, the `max_width` parameter remains unexplained; an agent might not know how to set it or its effect on output. Given that the tool has no output schema, a bit more detail on max_width would make it fully complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description explains the `region` parameter fully: it is [x, y, width, height] in desktop pixels and can be omitted for the whole desktop. However, `max_width` is not described at all; the schema only shows it is an integer with default 1280. Since schema description coverage is 0%, the description should compensate for both parameters, but it only covers one.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening sentence 'Take a screenshot and return it to you directly' uses a specific verb and resource, and clearly states the result. It also distinguishes itself from the vision-model-based sibling by saying 'This does not involve the local vision model,' which makes its purpose unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It provides concrete use cases: 'before writing a Playbook, to pick coordinates for probes and click regions, or to work out why a run went wrong.' This tells the agent exactly when to invoke it. It does not explicitly mention alternatives or when not to use it, but the context is clear enough to guide selection.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
voltage_diagnoseARead-only
Explain why a run behaved as it did, and what to change.
Call this instead of reading the journal by hand. It computes what the journal
implies but does not state -- watch labels the vision model never once reported,
guards that never evaluated true, whether bursts actually moved the screen, whether
the actuator is chaining or emitting one action at a time -- and returns each with
the specific edit that fixes it, ordered blocker-first.
The distinction it exists for: a burst that never ran and a burst that ran and did nothing look identical in a summary and have unrelated causes. The first is policy or grammar; the second is window focus, pointer mode, or an application that ignores synthetic input.
Apply the highest-severity finding, re-run, diagnose again. Changing several things at once makes the next diagnosis uninterpretable.
| Name | Required | Description | Default |
|---|---|---|---|
| run_id | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=false, so the description doesn't need to restate safety. It adds valuable behavioral detail: it computes implicit journal information, returns specific edits ordered blocker-first, and distinguishes between a burst that never ran vs. ran but did nothing. This goes well beyond the annotation, providing non-obvious nuances.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with a crisp summary, then explains the key distinction and ends with an actionable workflow. Every sentence earns its place; there is no fluff or redundancy. Structure is clear and scannable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given that an output schema exists (the description doesn't need to detail return structure) and annotations cover safety, the description covers the essential context: the diagnostic purpose, the key distinction between two root causes, and the iterative workflow. The only minor gap is the run_id parameter semantics, which slightly detracts from completeness for an otherwise simple tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description carries the full burden for the single parameter run_id. It never mentions run_id, its format, how to obtain it, or whether it's required (though the schema marks it optional). The name 'run_id' is self-explanatory by convention, but the description provides no explicit guidance, and with only one parameter to cover, this is a noticeable gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Explain why a run behaved as it did, and what to change.' It then contrasts itself with reading the journal, making its purpose distinct from voltage_journal. No ambiguity about what the tool does.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly instructs to call this instead of reading the journal by hand, giving clear when-to-use context. It also provides a workflow (apply highest-severity finding, re-run, diagnose again). However, it doesn't name alternative siblings like voltage_doctor or voltage_observe, or describe conditions where those might be more appropriate, so it stops short of complete guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
voltage_doctorARead-only
Check that everything needed for a run is present and working.
Reports the session type, input-device permissions, which capture backends work, detected screen geometry, GPU memory versus the selected model profile, and whether both model backends respond. When something is missing it returns the exact command to fix it. Call this before the first run on a machine.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotations already declare readOnlyHint=true and openWorldHint=false, establishing a safe, closed-world read operation. The description adds valuable context beyond safety: it lists the specific components checked (session type, input-device permissions, capture backends, screen geometry, GPU memory vs model profile, both model backends) and states that it returns fix commands. This informs the agent about the scope of the check and the nature of the response, which is more than annotations provide. No contradiction found.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three sentences, tightly packed with information. The first sentence gives the core purpose, the second enumerates the checks and the fix-command behavior, and the third provides usage timing. Every sentence earns its place, and the most important information (purpose and when to use) is front-loaded. There is no fluff or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a no-argument diagnostic tool, the description is comprehensive. It clearly states what is checked, the response characteristic (returns fix commands), and when to call it. An output schema exists (though not shown), so the description doesn't need to detail return formats. Given the complexity of the checks and the existence of a schema, nothing essential is missing for the agent to use the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters and the schema shows none. The description implicitly confirms this by stating 'Call this before the first run on a machine' with no mention of inputs. Since there are no parameters to explain, the description effectively communicates that it requires no configuration. This is a perfect fit for the no-parameter case, and the baseline of 4 is exceeded because the description makes the absence of parameters obvious.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: 'Check that everything needed for a run is present and working.' It specifies a concrete action (check) and a distinct resource (run prerequisites). It differentiates from siblings like voltage_status and voltage_diagnose by enumerating the exact checklist items (session type, permissions, capture backends, geometry, GPU memory, model backends). This makes it unambiguous which tool to select for pre-flight validation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit timing guidance: 'Call this before the first run on a machine.' While it doesn't mention alternatives or when not to use it, the instruction is clear and actionable. It implies this is a single-use setup check, not a repeated monitoring tool. The guidance is sufficient for the agent to decide when to invoke it, though lacking explicit exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
voltage_execute_burstADestructive
Execute one input burst yourself, bypassing the local models entirely.
For moments that need your judgement rather than the actuator's: opening the right application, clicking a specific confirmed target, typing something exact. Also the fastest way to sanity-check that input injection works at all.
Syntax: m:640,360;c:l;w:120;t:"hello";k:enter. Call voltage_reference for the full
list. The safety policy still applies. Defaults to dry_run, so pass dry_run=false to
actually inject.
| Name | Required | Description | Default |
|---|---|---|---|
| burst | Yes | ||
| label | No | manual | |
| dry_run | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint: true, readOnlyHint: false, and openWorldHint: true. The description adds critical behavioral context: it defaults to dry_run ('pass dry_run=false to actually inject') and notes the safety policy. It also explains that this is a manual override path. These details go beyond the annotations and inform the agent about side effects and prerequisites.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-structured and efficient: it leads with the core action, then gives usage scenarios, then provides a syntax example and necessary caveats. Every sentence earns its place, and the dry_run warning is front-loaded within the critical context. No fluff or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (custom syntax, safety policy, dry_run default) and that an output schema exists, the description covers the essential aspects: purpose, when to use, how to construct the burst (via example and reference), and the dry_run behavior. The only gap is a full in-place explanation of the syntax and label, but the reference to voltage_reference and the presence of an output schema mitigate this. Overall, it is nearly complete for an agent to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It provides a concrete syntax example (`m:640,360;c:l;w:120;t:"hello";k:enter`) and explains the dry_run parameter clearly. However, burst syntax is not fully documented (only a pointer to voltage_reference) and the label parameter is not explained beyond its default. This is partial compensation—helpful but not exhaustive.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Execute one input burst yourself') and the resource (burst), and immediately differentiates from siblings by emphasizing 'bypassing the local models entirely' and 'moments that need your judgement rather than the actuator's'. It also names the exact use case (opening applications, clicking confirmed targets, typing exact text) and points to voltage_reference for full syntax, making the purpose unmistakable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly states when to use this tool ('For moments that need your judgement rather than the actuator's', 'the fastest way to sanity-check that input injection works at all'), implies alternatives by referencing voltage_reference for syntax, and reminds that 'the safety policy still applies'. This gives an agent clear decision-making guidance without ambiguity.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
voltage_journalARead-only
Read a run's cycle-by-cycle record: what was seen, decided, refused, executed.
only_refused=true filters to cycles the governor blocked, which is the fastest way
to see where a Playbook's policy and the actuator's intentions disagree.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| run_id | No | ||
| only_refused | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already mark readOnlyHint=true, and the description aligns with that by saying 'Read'. It adds value by explaining the behavioral semantics of the journal contents and the meaning of 'only_refused', which goes beyond the raw annotation. No contradiction exists.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two focused sentences with the core purpose front-loaded and the filter tip as a concise, well-formatted follow-up. No filler or repetition, and the code-styled parameter reference is efficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return format is covered. However, the description fails to explain the run_id parameter, which is central to selecting a run, and gives no mention of limit. The tool is simple with all optional params, but the missing parameter descriptions leave a gap in usability.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It explains only_refused in detail, but completely omits run_id and limit. run_id is critical for identifying which run to read, and limit is a common but still undocumented control. The description is inadequate for a zero-coverage schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Read') and resource ('a run's cycle-by-cycle record'), and lists the exact contents: what was seen, decided, refused, executed. This clearly distinguishes it from siblings like voltage_observe or voltage_diagnose by framing it as a chronological journal rather than a live observation or diagnostic.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives explicit context for using the 'only_refused' filter and explains the fastest way to see policy/actuator disagreement. While it doesn't mention sibling tools for comparison, the usage hint is concrete and actionable, and the description clearly implies this tool is for inspecting historical decisions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
voltage_learnADestructive
Record something worth carrying to the next run against this target.
Write these as concrete, reusable facts, not narration:
good "the health bar is at x=120..300, y=1010; region_mean on red channel works" good "vision reports 'hotbar' reliably but never 'crosshair' -- do not watch it" good "block placement needs w:100 after the right click or it does not register" bad "the run failed" bad "tried again and it worked better"
kind groups them: label (what the vision model does and does not recognise),
timing (waits that a specific application needs), policy (what the governor blocked
and whether that was right), burst (a sequence that works), observation (anything
else).
| Name | Required | Description | Default |
|---|---|---|---|
| kind | No | observation | |
| note | Yes | ||
| target | Yes | ||
| playbook | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already indicate a mutating, potentially destructive action (readOnlyHint=false, destructiveHint=true); the description does not contradict these and adds that notes are stored against a target. It does not describe side effects or permissions, but given annotation coverage it provides acceptable additional context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is efficiently structured: it opens with the core purpose, gives clear good/bad examples, and ends with a concise classification of kind values. Every sentence adds value, and the format is well-balanced for the tool's complexity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool that records notes, the description covers the purpose, content quality, and kind taxonomy, which is sufficient for basic use. Gaps remain around `playbook` and exact behavior (e.g., confirmation, persistence), but the presence of an output schema and annotations mitigates these. Overall it is fairly complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description compensates by explaining the meaning of `kind` (label, timing, policy, burst, observation) and prescribing the format for `note` via good/bad examples. It leaves `target` and `playbook` undefined, but `target` is self-evident and `playbook` remains ambiguous, so coverage is partial but effective.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool records reusable facts against a target, with concrete good/bad examples that make the purpose unmistakable. It does not explicitly differentiate from sibling tools like voltage_lessons, but the 'carrying to the next run' phrasing is specific enough to convey its unique role.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides strong guidance on what to record (concrete facts, not narration) and explains the kind grouping, but it never mentions alternative tools or conditions under which to avoid this tool. Usage context is implied rather than explicit, and no exclusions or comparisons are given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
voltage_lessonsARead-only
Recall what previous runs learned about driving something.
Call this before writing a Playbook for a target you have driven before. Lessons persist across sessions and are keyed by target ("minecraft", "roblox", "dolphin"), so a new Playbook can start from what the last one discovered -- which labels the vision model actually recognises, where the HUD probes are, what timing the game needs -- rather than rediscovering it.
Omit target to see everything recorded so far.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| target | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=true and the description aligns with that (no mutation implied). The description adds valuable behavioral context: lessons persist across sessions, are keyed by target, and include specific types of information (labels, HUD probes, timing). This goes beyond the annotation by describing persistence and content, which is useful for setting expectations about what the tool returns.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise and front-loaded with the purpose. It uses bold for emphasis ('before writing a Playbook') and keeps each sentence purposeful. There is no filler or redundant explanation. The structure guides the reader from what the tool does, to when to use it, to how to filter results.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has an output schema (as indicated by the context), so return values are documented elsewhere. The description provides sufficient context for an agent to decide when to call it: it explains the purpose, when it is appropriate (before writing a Playbook for a previously driven target), and how to control scope with the target parameter. No critical information is missing, given the read-only annotation and output schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 0% description coverage, so the description must compensate. It clearly explains the `target` parameter (keyed by target, omit to see everything) and gives examples of valid values. However, it does not mention the `limit` parameter at all, leaving its semantics to inference from the default value of 30. This is a partial compensation but not complete for both parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Recall what previous runs learned about driving something.' It then gives concrete examples of lesson content (labels, HUD probes, timing), which makes the tool's purpose unambiguous and distinct from any other sibling. The behavior is clearly scoped to recalling learned lessons, not a general-purpose query.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly instructs when to use the tool: 'Call this **before writing a Playbook** for a target you have driven before.' It also explains the benefit (start from previous discoveries rather than rediscovering) and provides parameter guidance: 'Omit `target` to see everything recorded so far.' This gives an agent clear, actionable context for choosing this tool over alternatives like voltage_learn.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
voltage_observeARead-only
Run one vision pass and return grounded elements in screen coordinates.
watch is the closed vocabulary the vision model may use -- it can only report
labels from this list, so name the things your Playbook's guards will test for.
Use this to check that the vision model can actually find what a state depends on
before committing to it in a Playbook. If an element does not come back here, a
sees(...) guard on it will never fire.
| Name | Required | Description | Default |
|---|---|---|---|
| watch | Yes | ||
| region | No | ||
| read_text | No | ||
| max_elements | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already convey read-only and closed-world hints. The description adds valuable behavioral context: it clarifies that 'watch' is a closed vocabulary, that the tool runs a single pass, and that missing elements imply guards never fire. This goes beyond the annotations and provides actionable insight into the tool's behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise, with two short paragraphs that are front-loaded with the core purpose. Every sentence adds distinct value—stating the action, vocabulary constraint, and practical implication. There is no fluff or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description captures the tool's primary purpose and a key behavioral consequence, and an output schema exists so return values are already documented. However, it does not explain non-required parameters (region, read_text, max_elements), which are likely needed for correct invocation. This gap reduces completeness, though the core use case is well covered.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It explains 'watch' as the closed vocabulary, which is essential, but it omits any explanation for 'region', 'read_text', and 'max_elements'. With only one parameter addressed, the description fails to adequately clarify the remaining parameters, leaving the agent with insufficient guidance.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states a specific verb and resource: 'Run one vision pass and return grounded elements in screen coordinates.' It also explains a distinct use case—checking if the vision model can find elements before committing to a Playbook. While it doesn't explicitly contrast with sibling tools, the purpose is specific and unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides explicit context for when to use the tool: 'Use this to check that the vision model can actually find what a state depends on before committing to it in a Playbook.' This is a clear directive without naming alternatives, but it effectively guides the agent on ideal usage scenarios.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
voltage_pauseBDestructive
Pause or resume a run. Held input is not released, so a paused run can continue.
| Name | Required | Description | Default |
|---|---|---|---|
| resume | No | ||
| run_id | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already flag destructiveHint=true, so the mutation nature is disclosed. The description adds the specific behavior that held input is retained, which goes beyond the annotations and gives the agent useful context about the pause/resume semantics. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with zero filler. The core action is front-loaded ('Pause or resume a run') and the clarifying detail about held input follows immediately. Every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the existence of an output schema and simple optional parameters, the description is far from complete. It lacks usage guidance, parameter semantics, and any mention of prerequisites or side effects beyond the held-input note. The agent would need to guess how to set 'resume' or when to pass 'run_id'.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% — neither 'resume' nor 'run_id' is explained in the schema. The description does not mention any parameters at all, so the agent has no idea that 'resume' likely indicates whether to resume or pause, or how 'run_id' selects the run. With two parameters and zero coverage, the description must compensate but fails completely.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a clear action (pause or resume), a specific resource (a run), and adds a key nuance (held input is not released). It distinguishes implicitly from voltage_stop but does not name sibling alternatives, so it falls short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No explicit guidance on when to use this tool versus alternatives like voltage_stop or voltage_run. The note about held input hints at a use case but does not state conditions or exclusions, leaving the agent to infer when pause is appropriate.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
voltage_referenceARead-only
Return everything needed to author and iterate on a run.
Call this before your first Playbook. Sections:
loop the learning loop -- how to go from a failed run to a working one, and what each failure mode actually means. Read this second. bursts the burst cookbook: how to chain inputs well, timing rules, ready-made patterns for desktop and for games, and the antipatterns that waste cycles. Read this if bursts are coming out one action at a time. burst the raw burst syntax playbook the state-machine JSON schema guards expression functions for transitions and reflexes example a complete working Playbook
| Name | Required | Description | Default |
|---|---|---|---|
| section | No | all |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, so the description need not restate safety. It adds value by explaining the content structure and the purpose of each section, which helps the agent understand what the tool actually returns. However, it does not disclose any potential caveats (e.g., response size, format specifics), though those may be covered by the output schema. The added context justifies a score slightly above baseline.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is efficiently organized: a one-line purpose, then a bulleted list of sections with clear labels and explanations. It front-loads the main instruction and uses formatting to allow fast scanning. No sentence is redundant; each adds useful detail about content or usage.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a reference tool, the description covers all essential information: what it returns, when to call it, what each section contains, and even contextual reading order. The read-only behavior is covered by annotations, and the output format is presumably defined by the output schema (present signal). Nothing necessary for an agent to select and invoke this tool is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must explain the 'section' parameter. It does so comprehensively by listing each enum value and its meaning, and even offers reading-order guidance (e.g., 'Read this second', 'Read this if...'). This fully compensates for the schema gap, making the parameter self-documenting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Return') and a resource ('everything needed to author and iterate on a run'), then enumerates the sections returned. It clearly distinguishes itself from sibling tools (e.g., voltage_execute_burst, voltage_validate_playbook) by being a reference/documentation tool, not an execution or validation tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly says 'Call this before your first Playbook,' giving a clear when-to-use directive. It also provides conditional reading order (e.g., 'Read this if bursts are coming out one action at a time') and labels like 'the learning loop,' which help an agent decide which section to request. This is strong, situation-specific guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
voltage_runADestructive
Start a Playbook. Returns immediately with a run_id; poll voltage_status.
dry_run overrides the Playbook's policy. Leave it unset for the Playbook's own
setting, which defaults to true. A dry run does everything except inject input, so
it is the correct way to check that your states, guards and transitions behave before
letting it touch the machine.
target_period_s is the loop period. 0.5 is a good default; lower it for games,
raise it for slow UI.
Stop a run with voltage_stop, adjust it live with voltage_steer. The run also stops on its own budget, on any physical keyboard or mouse input from the user, and on the panic file.
| Name | Required | Description | Default |
|---|---|---|---|
| dry_run | No | ||
| playbook | Yes | ||
| keep_frames | No | ||
| target_period_s | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already mark destructiveHint and openWorldHint, and the description complements these by explaining concrete behaviors: immediate return with run_id, polling requirement, dry_run overriding policy, and the specific conditions that terminate a run. It adds value beyond annotations without contradicting them.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is efficiently structured: the core action and return contract are front-loaded, followed by parameter guidance and termination behavior. Every sentence adds functional value, and no redundant or filler content is present.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers the essential lifecycle: starting, monitoring, adjusting, and stopping. It explains dry-run semantics and stopping triggers. However, it does not describe the structure of the `playbook` object or the meaning of `keep_frames`, which may be important for correct invocation. The presence of an output schema and related tools (voltage_validate_playbook) partially mitigates this.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must explain parameters. It does explain dry_run (including override semantics and default behavior) and target_period_s (with recommended values), but it does not explain `playbook` (the required parameter) or `keep_frames`. Since playbook is central and the schema offers no description, this leaves a gap for an agent constructing a valid call.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with 'Start a Playbook,' a specific verb-resource pairing that clearly states the tool's core function. It immediately distinguishes itself from siblings by mentioning polling with voltage_status, stopping with voltage_stop, and live adjustment with voltage_steer, so the agent can tell it apart without opening other schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides explicit context for when to use dry_run ('the correct way to check that your states, guards and transitions behave before letting it touch the machine'), recommends values for target_period_s, and explains how to stop or adjust a run using sibling tools. It also details automatic stopping conditions (budget, keyboard/mouse input, panic file), giving clear operational guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
voltage_statusARead-only
Poll a run: current state, variables, last burst, what the vision model sees.
Includes recent cycles, governor refusals, and per-stage timings so you can tell whether a slow loop is capture, vision, decision, or execution.
| Name | Required | Description | Default |
|---|---|---|---|
| run_id | No | ||
| journal_tail | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=false, so the read-only nature is covered. The description adds useful context beyond annotations: the specific data included (recent cycles, governor refusals, per-stage timings) and its diagnostic intent. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with no waste. The action is front-loaded ('Poll a run'), followed by a list of what it returns and the diagnostic purpose. Every phrase earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a monitoring tool with an output schema present, the description conveys enough about the returned data to be useful. However, the lack of parameter documentation is a notable gap that makes it slightly incomplete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description does not explain either parameter (run_id, journal_tail) at all. While run_id is somewhat inferable from its name, journal_tail is completely unexplained. The description fails to compensate for the schema's lack of documentation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb ('Poll') and resource ('a run'), then enumerates the returned data (state, variables, last burst, vision model view, cycles, refusals, timings). This clearly differentiates it from sibling tools like voltage_capture or voltage_execute_burst, which imply different actions.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It implies usage during a run to monitor state and diagnose slow loops ('so you can tell whether a slow loop is capture, vision, decision, or execution'). However, it doesn't explicitly state when not to use it or point to alternatives such as voltage_doctor or voltage_diagnose, leaving some ambiguity.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
voltage_steerADestructive
Correct a live run without restarting it.
hint is injected into the actuator's prompt as a supervisor note and persists until
changed -- use it when the actuator is doing something legal but wrong.
force_state jumps the machine on the next cycle. variables updates run variables.
dry_run can be flipped either way mid-run.
| Name | Required | Description | Default |
|---|---|---|---|
| hint | No | ||
| run_id | No | ||
| dry_run | No | ||
| variables | No | ||
| force_state | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already flag destructiveHint=true, so the description adds some context: hint persists, force_state jumps the machine, variables updates, dry_run can flip. However, it does not disclose potential side effects, irreversibility, or prerequisites despite the destructive nature. It does not contradict the annotations, but the coverage is not thorough.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise and well-structured: a one-sentence overview followed by per-parameter explanations. It is front-loaded with the main purpose, uses backticks for param names to aid scanning, and has no filler or redundant statements.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with 5 parameters, zero schema descriptions, a destructive annotation, and an output schema, the description covers the core actions but misses run_id semantics, any warning about destructive consequences, and what the output schema contains. It is usable but not fully complete for safe and correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description carries the full burden for parameter meaning. It explains hint, force_state, variables, and dry_run, but omits run_id entirely, leaving its role merely implied by the phrase 'a live run.' This is a partial but incomplete compensation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb-resource pair: 'Correct a live run without restarting it,' which clearly distinguishes this tool from siblings like voltage_stop, voltage_pause, or voltage_run. It also enumerates the effects of each parameter, leaving no ambiguity about what the tool does.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides a concrete usage scenario for `hint` ('when the actuator is doing something legal but wrong') and explains the function of each parameter (e.g., force_state jumps the machine, dry_run flips). It implies this tool is for mid-run corrections vs. restarting, but does not explicitly name alternatives or exclusion conditions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
voltage_stopADestructive
Stop a run and release every held key and button.
Safe to call at any time, including while a burst is mid-flight -- the burst is interrupted and anything held is released.
| Name | Required | Description | Default |
|---|---|---|---|
| reason | No | stopped by orchestrator | |
| run_id | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, and the description adds concrete behavior: releases every held key/button and interrupts bursts. This goes beyond the annotation's generic destroy flag without contradicting it, giving the agent a more precise model of consequences.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two concise sentences convey the purpose, safety, and edge-case behavior with zero filler. Information is front-loaded and every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple stop tool, the description covers the main behavior and safety profile. However, the lack of any parameter explanation means an agent might guess wrong about 'run_id' or 'reason' (e.g., whether run_id is required to target a specific run). Optional parameters with defaults mitigate, but the gap prevents full completeness.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the description does not explain either 'reason' or 'run_id.' The agent has no guidance on what these parameters control or when to provide them, though they are optional. With no parameter documentation anywhere, this is a notable gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Stop') and resource ('a run') while adding unique scope: 'release every held key and button.' This clearly distinguishes it from siblings like voltage_pause and voltage_run, and the mention of interrupting mid-flight bursts further clarifies its specific role.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context: 'Safe to call at any time' and explicitly covers the edge case of a mid-flight burst. However, it does not explicitly contrast with alternatives like voltage_pause or voltage_steer, leaving some ambiguity about when to choose this over a pause or a graceful stop.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
voltage_validate_playbookARead-only
Fully check a Playbook without running it.
Validates the schema, compiles every guard expression, parses every burst, checks that transition targets and probe references exist, and reports unreachable states and dead transitions. Errors come back as a complete list, not one at a time.
Always call this before voltage_run. Warnings are worth reading: "tests for X but X
is not in watch" means a transition that can never fire.
| Name | Required | Description | Default |
|---|---|---|---|
| playbook | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only mark readOnlyHint:true. The description adds substantial behavioral detail: it returns a complete list of errors rather than one at a time, reports unreachable states and dead transitions, and explains how to interpret warnings. This fully complements the annotation and does not contradict it.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is tightly written, starting with the primary purpose, then detailing checks, then error behavior, then usage guidance and a warning interpretation. Every sentence serves a purpose—no filler. It front-loads the action and clearly organizes information in short block format.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter validation tool with an output schema declared (though not shown explicitly), the description covers what it does, how it behaves, when to call it, and how to interpret results. With annotations covering read-only safety and the output schema expected to define return values, nothing essential is missing for an agent to decide and invoke correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema only defines a generic 'playbook' object with no description (0% coverage). The description compensates by making clear that the parameter is the Playbook being validated, and it describes what validation entails (schema, guards, bursts, references). This gives the agent enough context to pass the correct object, even without knowing its internal structure.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific, unambiguous statement: 'Fully check a Playbook without running it.' It enumerates the exact validations performed (schema, guards, bursts, transition targets, probe references) and reports unreachable states/dead transitions, distinguishing this validation tool from siblings like voltage_run and voltage_execute_burst. The verb 'validate' matches the tool name and clears its role.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly guides usage with 'Always call this before voltage_run,' which states when to use this tool relative to its primary sibling. It also adds a practical hint about interpreting warnings (e.g., 'tests for X but X is not in watch'). It does not list explicit exclusions, but the directive is clear and directly actionable.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
16 tool updates
v0.1.0- First observed
voltage_calibrate - First observed
voltage_capture - First observed
voltage_diagnose - First observed
voltage_doctor - First observed
voltage_execute_burst - First observed
voltage_journal - First observed
voltage_learn - First observed
voltage_lessons - First observed
voltage_observe - First observed
voltage_pause - First observed
voltage_reference - First observed
voltage_run - First observed
voltage_status - First observed
voltage_steer - First observed
voltage_stop - First observed
voltage_validate_playbook
TDQS
Scored across 16 tools
Each tool has a clearly distinct purpose: pre-flight checks, documentation, perception, input execution, validation, running, monitoring, control, and learning. Even similar tools like voltage_journal (raw data) and voltage_diagnose (analyzed explanation) are cleanly separated by their roles.
All tools follow a consistent voltage_ prefix with a verb or verb_noun pattern (capture, execute_burst, validate_playbook, etc.). No mixed conventions or ambiguous verbs; naming is predictable and intuitive.
16 tools is well-scoped for a comprehensive automation server covering setup, execution, monitoring, debugging, and learning. Each tool earns its place; the count supports the full workflow without bloat.
The tool surface covers the entire lifecycle: environment checks (doctor, calibrate), documentation (reference), perception (capture, observe), manual action (execute_burst), validation and execution (validate_playbook, run), live control (steer, stop, pause), monitoring (status, journal, diagnose), and cross-session learning (lessons, learn). No obvious gaps for the stated purpose.
Maintenance
Related MCP Connectors
Operate Linux, macOS and Windows from your LLM. Every action runs through an auditable allowlist.
Human-in-the-loop approval for agent actions, with verifiable action-bound receipts.
Control real Android and iOS devices with LLM agents — tap, swipe, type, automate flows.
Lets AI agents use a real human as a tool: visual checks, taste, phone calls, unblocking, approvals
Related MCP Servers
- AlicenseNot gradedqualityAmaintenanceA local autonomous AI agent that watches your screen, understands the visual layout, and executes native OS commands (clicking, typing) without cloud APIs.16MIT
- FlicenseNot gradedqualityBmaintenanceEnables AI agents to declaratively control web pages using real mouse and keyboard events via Chrome DevTools Protocol, without executing page JavaScript.10 npm1-
- AlicenseAqualityBmaintenanceEnables low-cost agent models to control Windows applications through a compact, state-safe proxy over Open Computer Use, reducing model-visible context by up to 99.8% with support for record/replay and reusable UI component memory.5MIT
- AlicenseNot gradedqualityAmaintenanceEnables AI agents to automate real desktop applications across Windows, Linux, and macOS using incremental screen perception, accessibility trees, OCR, and window management, dramatically reducing token usage compared to screenshot-per-step approaches.39 PyPI3MIT