VoltageInputMcp
VoltageInputMcp
フロンティアモデルがツール呼び出し速度ではなく入力速度でコンピュータを操作できるようにするMCPサーバー。
問題
コンピュータ操作ツールは、すべてのアクションでリモートモデルへの往復が発生します。スクリーンショット取得、判断、クリック1回。フォーム入力には十分ですが、連続した入力を素早く届ける必要があるものには役に立ちません——ゲームプレイ、モーダルダイアログの操作、タイムラインの操作、3番目の入力が最初の2つがすでに届いていることに依存するようなUI。ボトルネックはモデルの知能ではありません。知能が800ms先にあり、入力が8ms間隔で必要なことです。
Related MCP server: live-mcp
答えの形
判断と実行を分離し、実行をキーボードと同じマシンに置く。
┌─────────────────────────────────────────────────────────────────┐
│ Layer 1 — the orchestrator (Claude, or any MCP client) │
│ Writes a Playbook: states, what to look for, what is allowed, │
│ when to move on. Thinks once, up front. Watches and corrects. │
└───────────────────────────┬─────────────────────────────────────┘
│ MCP
┌───────────────────────────▼─────────────────────────────────────┐
│ Layer 2 — two small local models, on your GPU │
│ │
│ vision (Qwen2.5-VL-3B) "of these specific things, │
│ which are on screen, and where?" │
│ actuator (Qwen3-1.7B) "given that, which inputs?" │
│ │
│ Neither plans. Both answer one closed question per cycle. │
└───────────────────────────┬─────────────────────────────────────┘
│
┌───────────────────────────▼─────────────────────────────────────┐
│ safety governor → /dev/uinput → the actual desktop │
└─────────────────────────────────────────────────────────────────┘オーケストレーターが頭脳です。小規模モデルは腕です。腕は賢くなく、賢さを求められることもありません。
速度が実際にどこから来るか
小規模モデルが速いからではありません——3B VLMでも依然として約300msかかります。速度は4つの要素から来ており、影響の大きい順に並べると:
バースト。 アクチュエータは入力を1つ発行するのではありません。バーストを発行します:モデルを介さない専用エグゼキュータが実行する、タイミング指定された入力プログラムです。
g:0;c:l;w:150;t:"README.md";k:enter;w:80;k:ctrl+sこれは1回の判断と、約400msにわたる7つの入力で、ミリ秒単位でスケジュールされます。40アクションのバーストでも判断は1回です。入力レートはモデルではなくバーストによって決まります。
反射。 判断の合間に、モデルを一切使わず、マイクロ秒単位で安価な画面プローブ(1ピクセル、1領域の平均)を発火するルール。
{"id": "heal", "when": "probe('health') < 0.25", "do": "k:q;w:60", "cooldown_ms": 800}知覚のスキップ。 ほとんどのサイクルは、変化していない画面を見ています。40µsのフレーム差分で、ビジョンモデルに300msを費やすか、最後の観測結果を再利用するかを判断します。通常のデスクトップ作業では、ほとんどのサイクルでVLMをスキップします。
プロンプトキャッシュの局所性。 プロンプトは静的ファーストで順序付けされ、llama.cppがKVキャッシュを再利用し、変更された末尾のみを再プリフィルします。
小規模モデルが小さいのに信頼できる理由
信頼できることを求められていないからです——制約されているのです。
llama.cppでは、両モデルとも現在の状態から毎サイクル再生成されるGBNF文法に対して生成を行います。文法はアドバイスではありません。有効なパースを継続するトークンのみが到達可能になるようにロジットをマスクします。具体的には、アクチュエータはできない:
不正なバーストを出力する
ポリシーが拒否するキーを指定する——そのキーは文法に存在しない
観測されていない要素を参照する——インデックス範囲はこのサイクルの要素数から構築される
Playbookが宣言していない状態遷移を提案する
また、ビジョンモデルはUI要素名を発明できない:そのラベル語彙は、あなたが書いたwatchリストと小さな汎用セットだけです。したがって、sees("address bar")ガードは、3Bモデルがたまたま生成した名詞ではなく、閉じた語彙に対して比較します。
リトライループも防御的なJSONパースもありません。不正な出力が起こりにくいからではなく、表現不可能だからです。
Playbook
小規模モデルに目標を与えるのではありません。状態機械を与えます。遷移はランタイムによって評価されるガード式であり、モデルによって評価されるのではありません。
{
"name": "open_downloads",
"goal": "Open the file manager at ~/Downloads. Delete nothing, confirm nothing.",
"initial": "launch",
"policy": {
"dry_run": true,
"allow_verbs": ["g", "c", "k", "t", "w"],
"deny_labels": ["delete", "trash", "confirm", "empty trash"]
},
"budget": { "max_cycles": 60, "max_seconds": 90 },
"states": {
"launch": {
"brief": "Open the application launcher and start the file manager.",
"watch": ["application launcher", "search field", "file manager icon"],
"on_enter": "k:meta;w:400",
"transitions": [
{ "when": "sees('search field')", "to": "type_name" },
{ "when": "cycles() > 6", "to": "@failure", "note": "launcher never opened" }
]
},
"navigate": {
"brief": "Focus the location bar with ctrl+l, type the path, press Enter.",
"watch": ["location bar", "file list", "error message"],
"on_enter": "k:ctrl+l;w:200",
"transitions": [
{ "when": "text('Downloads')", "to": "@success" },
{ "when": "sees('error message')", "to": "@failure" }
]
}
},
"success_when": "text('Downloads') and not flag('loading')"
}voltage_referenceは完全なDSL、JSONスキーマ、ガード関数テーブルを返すため、オーケストレーターはこのリポジトリを読まずにPlaybookを作成できます。
パフォーマンスチューニング
以下の数値はすべて、リファレンスマシン(RTX 3050 6GBラップトップ、llama.cpp上のQwen2.5-VL-3B + Qwen3-1.7B)で測定されたものであり、推定ではありません。
両モデルともデコード律速です。出力トークンだけが意味のあるレバーです。
これは驚きでした——設計当初はビジョンがプリフィル律速だと思っていましたが、そうではありません。プリフィルは448×252から896×504まで約28msでフラットと測定されました。デコードは約22ms/トークンで動作します。つまり:
項目 | コスト |
出力トークン1個 | 約22 ms |
報告される要素1個 | 約21トークン ≈ 500 ms |
ビジョン、要素2個 | 約1.0 s |
ビジョン、要素4個 | 約2.2 s |
アクチュエータ、キャッシュ済みプレフィックス | ノート長に応じて140〜400 ms |
それぞれがデフォルトを変更した3つの結果:
max_elementsはビジョンの支配的なコストです。 デフォルトは3。6に上げると、知覚サイクルごとに約1.5秒追加されます。ガードが実際にテストする数に設定してください。downscale_toを縮小しても効果はなく、通常は悪化します。 448×252は896×504より2.5倍遅いと測定されました——画像がぼやけるとモデルの確信度が下がり、より多くのトークンを出力するためです。収まる最大サイズを使用してください。アクチュエータの
noteフィールドはレイテンシの55%を占めました。 純粋に診断用であり、48文字で412ms/サイクル、12文字で184ms、0文字で140msと測定されました。デフォルトは現在12です。
要素は{"l":"address bar","b":[...],"c":0.9}ではなく[label_index, x1, y1, x2, y2]としてエンコードされます——同じ理由で、トークンが27〜29%少なく、レイテンシが32〜41%低いと測定されました。閉じたwatch語彙へのインデックス付けはより安全でもあります:モデルはラベルを綴ることすらできず、綴り間違いはなおさらです。
GBNF評価はサンプリングされたトークンごとにCPUで1回実行されるため、アクチュエータは完全にGPUオフロードされているにもかかわらず、ビジョンモデルより多くのCPUスレッドを取得します——また、allow_keysの制限は安全性の最適化だけでなく、レイテンシの最適化でもあります。
間違っていると静かに失敗する2つの設定:
ビルド時の
GGML_CUDA_FA_ALL_QUANTS=ON。 私たちはq8_0KVキャッシュとフラッシュアテンションで提供しています。このフラグがないと、llama.cppはそのKV組み合わせ用のFAカーネルをコンパイルせず、低速パスにフォールバックします——エラーはなく、ただ不可解に悪い数値が出るだけです。scripts/build-llama.shがこれを設定します。実行時の
GGML_CUDA_ENABLE_UNIFIED_MEMORY=0。1の場合、VRAMオーバーフローは失敗せずに静かにPCIe経由でスピルオーバーします。すべて動作し、約10倍遅くなります。serve.shがこれをオフに固定します。
推測ではなく測定してください:
.venv/bin/voltage benchループが使用する正確なプロンプト形状で両バックエンドを駆動し、コールドvsプロンプトキャッシュ済みレイテンシ、3つの入力サイズでのms/ビジュアルトークン、およびそれらが示すサイクル時間を報告します。プロンプトキャッシュの高速化が約1.5倍未満の場合、動的な何かがプロンプトプレフィックスに漏れ込んでいることを意味します。
モデルの比較
明白な実験——「どのモデルがより良いバーストを書くか」——は間違ったものを測定しています。文法はすでにすべてのバーストが有効であることを保証しているため、より大きなモデルが構文で勝つことはできません。構成が使用可能かどうかを実際に決定するもの:
グラウンディング精度。 200ms速くて40pxずれているモデルは役に立ちません——クリックが外れます。クリックは中心に着地するため、IoUではなく画面ピクセル単位の中心距離として測定されます。
制約下での判断品質。 同じ観測結果が与えられたとき、正しい合法アクションを選択し、サイクルごとに1つの臆病なアクションを出力するのではなく、シーケンス全体を1つのバーストに連鎖させますか?
レイテンシ。 これは1と2が許容可能になって初めて重要になります。
.venv/bin/voltage fixture desktop # capture a real screen
.venv/bin/voltage compare # score whatever is running nowグラウンドトゥルースは、オーケストレーティングモデルによってラベル付けされた実際のスクリーンショットから来ます——これはこのシステムが実行時に使用するのと同じリファレンスです。合成UIは罠です:描かれた長方形は、実際のインターフェースで訓練されたモデルにはボタンとして読めないため、それに対するスコアリングは間違ったスキルを測定します。
結果は実行をまたいで蓄積されるため、ワークフローは:プロファイルAを提供 → compare → プロファイルBを提供 → compare → テーブルを読む。voltage compare --listは再実行せずにそれを表示します。
フィクスチャはあなたのものであり、コミットされません。スクリーンショットにプライベートなものが含まれる場合は、fixtures/を.gitignoreに追加してください。
学習ループ
馴染みのないターゲットに対する最初のPlaybookはほぼ間違いなく正しくありません。重要なのは、失敗が具体的であり、次の試行が前回の学習から始まることです。
voltage_reference(section="loop") the loop itself, and what each failure means
voltage_reference(section="bursts") the burst cookbook: chaining, timing, game patterns
voltage_capture / voltage_observe look before writing — check your labels exist
voltage_validate_playbook dead guards, unreachable states, caught statically
voltage_run(dry_run=true) real models, real screen, nothing injected
voltage_diagnose(run_id) ← what to change, not raw data
voltage_learn(target=..., note=...) record it; persists across sessions
voltage_lessons(target=...) recall it before the next playbookvoltage_diagnoseはこれをループにする部品です。 ジャーナルが示唆するが明示しないことを計算し、それぞれに対する編集を命名します。行き詰まったMinecraft実行では:
[BLOCKER] label_never_seen never reported: ['crosshair', 'health bar']
[BLOCKER] input_not_landing 14 bursts executed, but the screen never changed
[BLOCKER] state_never_left 'mine' ran 14 cycles and never transitioned
[PROBLEM] timid_bursts bursts averaged 1.0 actions
[HINT] vision_every_cycle vision ran on 100% of cyclesこれが存在するための区別:実行されなかったバーストと、実行されたが何もしなかったバーストは、サマリーでは同一に見え、無関係な原因を持ちます。 前者はポリシーまたは文法です。後者はウィンドウフォーカス、ポインターモード、または合成入力を無視するアプリです。Diagnoseは、実行後にフレームが実際に変化したかをチェックすることでそれらを分離します。
最も深刻度の高い発見を適用し、再実行し、再度診断します。一度に1つの変更——複数同時に行うと、次の診断が解釈不能になります。
レッスンはセッションをまたいで永続化され、ターゲットごとにキー付けされるため、ゲームの2番目のPlaybookは、最初のPlaybookが発見したプローブ座標と動作するラベル名から始まります:
voltage_learn(target="minecraft", kind="label",
note="vision reports 'hotbar' reliably but never 'crosshair'")
voltage_learn(target="minecraft", kind="timing",
note="block placement needs w:100 after right click or it does not register")安全性
入力を生成するのは1.7Bモデルです。ガバナーは助言ではない層です:すべてのバーストがそれを通過します——反射バーストや自分で書いたものも含みます。
dry_runがデフォルトです。 新しいPlaybookは、何にも触れずにすべてのバーストをパース、チェック、ジャーナルします。バースト全体の拒否。 意図したシーケンスの半分を実行することは、実行しないことより悪いです。
deny_labelsは、Delete / Confirm / Purchase / Allowと呼ばれるものへのクリックを、どこに現れても拒否します——これが予期しない場所にポップアップするダイアログを捕捉します。リージョンフェンシング、キー許可リスト、拒否されたコード(
ctrl+alt+delete、alt+f4)、拒否されたテキストパターン(rm -rf、sudo)、バーストサイズと1秒あたりの入力数上限。4つの独立した停止手段:
voltage stop(ファイルを書き込む——SSH経由で動作)、ループが固着した場合に独自のスレッドで発火するデッドマンタイマー、物理入力の競合(実際のマウスに触れると停止)、およびPlaybook予算。押されたキーは常に解放されます——アボート時、クラッシュ時、タイムアウト時。
d:shiftとu:shiftの間で中断された実行は、Shiftが押されたままになってはいけません。
インストール
ゼロから動作まで、2つのコマンド。
Linux / macOS
git clone https://github.com/casualkre/voltage-input-mcp && cd voltage-input-mcp && ./install.shWindows(PowerShell)
git clone https://github.com/casualkre/voltage-input-mcp; cd voltage-input-mcp; powershell -ExecutionPolicy Bypass -File .\install.ps1その後、どちらでも:
voltage setupinstall.shはPython、システムパッケージ、venv、PATHを処理し、rootが必要なものについては要求するのではなく正確なsudo行を表示します。その後voltage setupが既に持っているものを検出し、不足しているものだけをダウンロードし、モデルサーバーを起動し、AIクライアントに登録します——各ステップを説明するのではなく実行します。10〜25分で、ほぼすべてダウンロード時間です。再実行しても安全です。中断したところから再開します。
その後、ただ実行するだけです:
voltageセットアップは既に持っているものを検出し、そこから続行します。 開始点を想定しません:OS、GPU、llama.cppまたはOllamaがインストールされているか、どのモデルがすでにプルされているか、入力とキャプチャが機能するか、MCPサーバーが登録されているかをプローブし、実際に残っているステップだけを計画し、どれがあなたの判断を必要とし、どれを自動で実行できるかを示します。すでにOllamaを持っている場合はそれを使用します。どちらのバックエンドも持っていない場合は、トレードオフを2行で説明し、選択させます。
引数なしで対話型コンソールが開きます:ライブステータス、依存関係の順序で準備ができていないものを修正するガイド付きセットアップ、モデルスイッチャー、設定エディター、Claude Codeへのワンキー登録、診断。以下のすべてのサブコマンドは非対話型でも動作するため、スクリプトとCIは影響を受けません。
██╗ ██╗ ██████╗ ██╗ ████████╗ █████╗ ██████╗ ███████╗
██║ ██║██╔═══██╗██║ ╚══██╔══╝██╔══██╗██╔════╝ ██╔════╝
██║ ██║██║ ██║██║ ██║ ███████║██║ ███╗█████╗
╚██╗ ██╔╝██║ ██║██║ ██║ ██╔══██║██║ ██║██╔══╝
╚████╔╝ ╚██████╔╝███████╗██║ ██║ ██║╚██████╔╝███████╗
╚═══╝ ╚═════╝ ╚══════╝╚═╝ ╚═╝ ╚═╝ ╚═════╝ ╚══════╝
── status ──────────────────────────────────────────────
ok input device /dev/uinput
ok vision model http://127.0.0.1:8080
ok actuator model http://127.0.0.1:8081
ok mcp registered claude mcp list
ok voltage on PATH ~/.local/bin/voltage実験的プロファイル
voltage → modelsに別途リストされ、それぞれ受け入れなければならない警告の背後にあります。測定結果がトレードオフを予測可能にするため存在します:デコードは約22ms/トークンで支配的であり、アクティブパラメータに比例するため、モデルを縮小すると実際にループレートが上がります。その代償はグラウンディングです。
profile | models | VRAM | trade |
| SmolVLM-500M + Qwen3-0.6B | ~2.2 GB | ループ速度は3〜4倍、groundingはほぼ機能しない |
| Qwen2.5-VL-3B + Qwen3-0.6B | ~3.8 GB | 判断が速くなる、groundingは変わらない |
| Qwen2.5-VL-32B + Qwen3-14B | ~34 GB | 最高のgrounding、1〜2.5秒/サイクル |
| Qwen2.5-VL-32B + Qwen3-30B-A3B | ~43 GB | 30Bの能力を~3Bのデコード速度で |
| 3B + 0.6B on CPU | none | GPUなしで動作、1サイクル数秒 |
特に注目すべき2つ:
hyperは危険なプロファイルです。 SmolVLM-500Mはgroundingモデルではありません。ボックスを返すことはありますが、しばしば間違っています——そして間違ったボックスは、優雅な劣化ではなく、間違った場所へのクリックを意味します。watchが空の場合(probeとreflexが実際の作業を行う場合)、またはすべてのクリックがclick_allow_regionsとrequire_target_elementで囲まれている場合にのみ使用してください。
beefy_moeは興味深いプロファイルです。 Qwen3-30B-A3Bは~3Bのアクティブパラメータを持つmixture of expertsであり、30Bの能力で推論しながら~3B相当の速度でデコードします——そしてこのループのボトルネックはまさにデコードです。同程度のレイテンシでdense 14Bよりもはるかに優れたアクチュエータです。欠点はメモリです: 高速なのはアクティブなexpertだけで、重み自体ではないため、30Bすべてをメモリに常駐させる必要があります。
recommend()は実験的なプロファイルを返すことはなく、テストでそれを強制しています。
カスタムモデルプロファイル
組み込みプロファイルは、このシステムが開発されたマシンを対象としており、あなたのマシンではありません。voltage → profilesから、または設定ファイルの隣にあるprofiles.tomlを編集して、独自のプロファイルを追加してください:
[my_rig]
description = "RTX 4090"
[my_rig.vision]
hf_repo = "ggml-org/Qwen2.5-VL-7B-Instruct-GGUF"
hf_file = "Qwen2.5-VL-7B-Instruct-Q4_K_M.gguf"
mmproj_file = "mmproj-Qwen2.5-VL-7B-Instruct-Q8_0.gguf"
params_b = 7.0
weights_mb = 4700
n_ctx = 4096
port = 8080
[my_rig.actuator]
hf_repo = "unsloth/Qwen3-4B-Instruct-2507-GGUF"
hf_file = "Qwen3-4B-Instruct-2507-Q4_K_M.gguf"
params_b = 4.0
weights_mb = 2500
port = 8081カスタムプロファイルは名前で組み込みプロファイルの上にマージされるため、leanという名前を付けると、パッケージをフォークせずに組み込みプロファイルを調整できます。Ollamaバックエンドではhf_repo/hf_fileの代わりにollama_tagを使用してください。
一方は厳密で、もう一方は寛容です。ビジョンは要求に応じてgrounded bounding boxを出力できなければなりません——Qwen2.5-VL、Qwen3-VL、InternVL、MiniCPM-V、UI-TARSはすべて可能です。一般的なキャプションモデルは画面を美しく説明しますが、ボックスは間違った場所に置きます。アクチュエータは寛容です: GBNF文法の下では、わずかな合法な継続候補から選択するだけなので、ほぼすべての有能な1B以上のinstructモデルで動作します。
シェルコマンドとMCPツール
2つの異なるインターフェースであり、混同するのが最初のつまずきの典型です:
呼び出し方 | 見た目 | |
シェルコマンド | ターミナルでスペースを入れて入力 |
|
MCPツール | アンダースコアを付けてClaudeに依頼 |
|
voltage_doctorはClaudeの名前空間におけるツール名であり、ディスク上のプログラムではありません。ターミナルで入力すると常に「unknown command」と表示されます。代わりにClaudeに実行を依頼してください。
これにより/dev/uinputへのアクセスがチェックされ、システム依存関係がインストールされ、venvが作成され、不足しているものが表示されます。次に:
./scripts/fetch-models.sh lean && ./scripts/serve.sh lean.venv/bin/voltage doctorクライアントへの接続
voltage connectセットアップ済みの内容、ライブURL、モデルが起動しているか、サーバーが登録されているかを表示し、実際のパスと環境変数がすでに入力されたクライアント別のコピーペースト手順を提供します:
voltage connect --client claude-desktop
voltage connect --client cursor
voltage connect --json # just the mcpServers entry対応: Claude Code、Claude Desktop、claude.aiカスタムコネクタ、Cursor、Windsurf、Zed、およびその他のもの用の汎用mcpServersブロック。同じ内容はvoltageコンソールの画面4にも表示され、Claude Desktopの設定を書き込むこともできます(既存ファイルをバックアップし、有効なJSONでない場合は触ることを拒否します)。
生成されるすべての設定はセッション環境を明示的に含みます。なぜなら、そこで問題が発生するからです: DBUS_SESSION_BUS_ADDRESSなしのシェルから登録されたサーバーは接続に成功しますが、静かに見えなくなります——入力は機能し、画面キャプチャは機能しません。voltage connectはそのケースを検出して警告します。
カスタムコネクタとして追加
MCPサーバーをURLで追加するクライアントは、stdioではなくHTTPを必要とします:
voltage serve --http次に、http://127.0.0.1:8765/mcpをカスタムコネクタとして追加します。
バインドはループバックに制限されており、変更には--allow-remoteが必要です。これは定型文ではありません: このサーバーはマウスを動かし、キーを押し、画面を読むために存在し、MCPには独自の認証がありません。ループバック以外へのバインドは、認証なしのデスクトップのリモートコントロールを公開することになります。本当に必要な場合は、認証付きリバースプロキシを前面に置き、ポートに到達できる者は誰でもマシンを掌握できることを理解してください。
MCPクライアントからの起動
MCPクライアントはサニタイズされた環境でサーバーを起動します——PATH、HOME、その程度です。これは賢明なデフォルトですが、画面キャプチャを壊します。コンポジタに到達するにはDBUS_SESSION_BUS_ADDRESSとWAYLAND_DISPLAYが必要だからです。入力インジェクションはそれらなしでも機能します(uinputはデバイスファイルであり、セッションサービスではないため)、そのため障害は紛らわしいほど部分的なものに見えます: バーストは実行されるが、スクリーンショットは取得されません。
それらを明示的に渡してください:
claude mcp add voltage-input \
-e WAYLAND_DISPLAY="$WAYLAND_DISPLAY" \
-e DISPLAY="$DISPLAY" \
-e DBUS_SESSION_BUS_ADDRESS="$DBUS_SESSION_BUS_ADDRESS" \
-e XDG_RUNTIME_DIR="$XDG_RUNTIME_DIR" \
-- /absolute/path/to/voltage-input-mcp/.venv/bin/voltage-input-mcpvoltage_doctorはこれらのうちどれが欠けているかを正確に報告するため、キャプチャが失敗している場合は最初に確認すべき場所です。
プラットフォーム
入力 | キャプチャ | テキスト | |
Linux |
| portal→PipeWire、KWin DBus、grim、X11 | スキャンコード、非ASCII用のクリップボードフォールバック |
Windows |
| GDI |
|
入力シンクより上のすべて——バーストスケジューリング、タイミング、押下キー追跡、セーフティガバナー、ランタイム全体——は共有されています。各プラットフォームは5つのメソッド(key、button、move_abs、move_rel、scroll)を実装します。inputs/sink.pyを参照してください。
知っておくべき2つの非対称性:
Windowsの方がタイピングが正確です。
KEYEVENTF_UNICODEはキーボードレイアウトを介さずにUTF-16コードユニットを配信します。Linuxのuinputはスキャンコードを送信するため、非USレイアウトでは句読点が——静かに——間違って出力されます。これが、クリップボードフォールバックがLinuxに存在し、Windowsでは不要な理由です。Linuxの方がキャプチャの能力が高いです。 GDI
BitBltは一部のハードウェアオーバーレイ動画や全画面専用ゲームを認識できず、それらは黒くキャプチャされます。そのようなゲームはボーダーレスウィンドウモードで実行してください。
Windowsでは、SendInputは昇格したプロセスが所有するウィンドウを操作できません(UIPI)——これは静かに失敗するため、voltage doctorが昇格状態を報告します。DPI認識はインポート時に宣言されます。これがないと、スケーリングされたディスプレイではすべての座標が間違ってしまいます。
要件
Linux(任意のディスプレイサーバー)またはWindows 10/11
Python 3.11+
leanプロファイル用に~5 GBの空きがあるGPU。voltage profilesで自分の環境に合うものが表示されます高速パス用のllama.cpp、またはビルド不要の低速パス用のOllama
KDE Plasma 6 / Wayland / CUDA / Python 3.14でエンドツーエンドで検証済み。Windowsパスは実装され型チェックされていますが、Windowsマシンでは実行されていません——未テストとして扱い、壊れたものを報告してください。
オーケストレータは駆動しているビルドを通知される
同じPlaybookがある構成では正しく、別の構成では間違っている可能性があり、リモートモデルはどちらかを確認できません。そのため、サーバーのMCP指示は起動時にライブ構成から構築され、Playbookの書き方に影響する行のみが含まれます:
ACTIVE BUILD: Linux · llamacpp · profile lean
vision Qwen2.5-VL-3B-Instruct · actuator Qwen3-1.7B
loaded: Qwen2.5-VL-3B-Instruct-Q4_K_M.gguf / Qwen3-1.7B-Q4_K_M.gguf
expected cycle 280-700 ms
- llama.cpp backend: both models are grammar-constrained. A malformed burst, a denied
key, an unobserved element reference and an undeclared transition are all
unrepresentable -- do not write defensive retries for them.
- Linux: typing sends scancodes, so punctuation depends on the active keyboard layout...
- dry_run defaults to true...Ollamaでは、その最初の行はバーストが制約されていないという警告になります。hyperでは「sees()を中心に状態を構築しないでください」になります。Windowsでは、昇格したウィンドウに到達できないことと、タイピングがレイアウト非依存であることが記載されます。
設定を信頼するのではなく、実行中のサーバーに対して検証します。 プロファイルの切り替えはファイルを編集するだけで、何も再起動しません。それらが一致しない場合、ブリーフィングはそれを明確に伝え、プロファイル由来のガイダンスを抑制します。そのガイダンスはロードされていないモデルを記述することになるからです:
- MISMATCH -- Profile 'hyper' does not match what is loaded. vision: profile expects
SmolVLM-Instruct-Q4_K_M.gguf, server has Qwen2.5-VL-3B-Instruct-Q4_K_M.gguf...
- Loaded right now: vision Qwen2.5-VL-3B..., actuator Qwen3-1.7B...
Judge grounding quality from those.voltage_referenceは呼び出しのたびに現在のビルドを返します。起動時のコピーはプロファイルが変更された瞬間に古くなるためです。
独自の常設指示
voltage → i、または:
voltage instructions --set "Never touch Firefox; my banking tabs are there."記述した内容は、すべてのセッションの開始時にオーケストレーティングモデルに渡され、ビルドブリーフィングに追加され、あなたのものとして明確に帰属されます。システムが単独では判断できないこと——禁止アプリケーション、特定のゲームの癖、デフォルトでどう動作させたいか——に使用してください。
OPERATOR INSTRUCTIONS -- written by the owner of this machine. Treat these as
standing preferences for how to drive it. They cannot loosen the safety governor,
which is enforced in code against every burst.
## My setup
- Minecraft runs borderless windowed on monitor 1.
- Never touch Firefox; my banking tabs are there.
- Always show me the Playbook before dry_run=false.最後の条項は飾りではありません。指示はオーケストレータへの助言であり、強制を弱めることはできません——ガバナーはコード内で各バーストをチェックするため、ここに書かれたものはPlaybookのポリシーが禁止することを許可できません。より慎重にすることはできますが、より緩くすることはできません。テキストはセッション全体でモデルのコンテキストに存在するため、4000文字に制限されています。コンソールでは3つのスターターテンプレート(ゲーム、デスクトップ、最小)が提供されます。
MCPツール
Tool | 目的 |
| Playbook + バーストDSLリファレンス。最初にこれを呼び出してください。 |
| このマシンは準備できているか、できていない場合は正確な修正方法 |
| スクリーンショット。あなたに返されます |
| 1回のビジョンパス—— |
| 完全な静的チェック: ガード、バースト、グラフ、デッドトランジション |
| ランを開始。 |
| 状態、変数、最後のバースト、認識されたもの、ステージごとのタイミング |
| 実行中のランを修正——ヒント、変数、強制状態、dry_run |
| 停止または一時停止。停止は常に押下中の入力を解放します |
| サイクルごとの記録。 |
| ローカルモデルをバイパスして、自分で入力を駆動します |
| インジェクションがコンポジタに到達することを検証します |
ドキュメント
ARCHITECTURE.md — ループの仕組み、各選択の理由、時間の使われ方
PLAYBOOK.md — オーサリングガイド
ステータス
ディスク上に重みがない状態で可能な限り構築・検証済み。149のテストが、バーストDSL、ガードサンドボックス、セーフティガバナー、Playbookコンパイル、GBNF生成、uinputワイヤエンコーディング、およびランループ自体(スタブモデルで駆動——静的画面でon_change知覚が本当にビジョンモデルをスキップするかのチェックを含む)をカバーしています。
MCPサーバーは、実際のクライアントによってstdio経由でエンドツーエンドで駆動されました: 13ツール、正しいスキーマ、execute_burstは有効なバーストを受け入れ、sudo rm -rf /を両方の一致ルールで拒否しました。
実行されていないのはライブモデルです。それにはllama.cppのビルドと重みの取得が必要で、scripts/がセットアップします。ビルド中に意図的にトリガーしなかった2つのこともあります——ポータル権限ダイアログと、実際の入力インジェクション——どちらもデスクトップに作用するためです。
ここからの操作順序:
./scripts/setup.sh # reports what needs sudo, doesn't run it
./scripts/build-llama.sh # ~15 min with CUDA
./scripts/fetch-models.sh lean
./scripts/serve.sh lean
.venv/bin/voltage doctor # should now say READY次に、MCPクライアント内で voltage_calibrate(カーソルが実際に動くのを確認)、voltage_observe(ビジョンモデルがラベルを見つけられるか確認)、その後 dry_run のPlaybookを実行し、dry_run=false を設定する前に voltage_journal を読む。
著作者
Claude Opus 5(Anthropic)が単一セッションで最初から最後まで執筆 — アーキテクチャ、実装、テスト、ドキュメントのすべて。人間がアイデアを指定し、制約(KDE Wayland、6 GB VRAM、「コンピューター操作より高速」)を設定し、結果をレビューしたが、コードは書いていない。
このリポジトリに組み込まれたプラットフォームに関する発見は、推測ではなくビルド中のマシンの調査から得られたものである — KWin が許可リストにない実行ファイルへの ScreenShot2 を拒否すること、grim が KWin の下では動作しないこと、MCPクライアントがセッションバスを除去してしまうこと。それぞれが、コード上で判断を迫られた箇所に文書化されている。
LICENSE は個人を著作権者として指名しておらず、その理由はそこに明記されている。
ライセンス
MIT。 LICENSE を参照。
Available Tools
16 toolsvoltage_calibrateADestructive
Verify that input injection actually reaches the compositor.
Creates the virtual devices, moves the pointer to three known points, and captures after each to confirm the cursor moved. Reports whether absolute positioning works or whether the relative fallback is needed -- which cannot be known without trying, since it depends on how libinput classified the virtual device.
Run this once per machine before trusting a real (non-dry-run) Playbook.
| Name | Required | Description | Default |
|---|---|---|---|
| dry_run | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotations already declare destructiveHint=true and openWorldHint=true. The description adds valuable detail: it creates virtual devices, moves the pointer, and captures output—concrete side effects beyond the annotation. It also explains why these behaviors are unpredictable ('depends on how libinput classified the virtual device'), which aligns with openWorldHint. This goes beyond what annotations alone convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise and well-structured. It opens with the core purpose, immediately explains what the tool does, then provides the rationale and usage timing. Every sentence earns its place—no fluff or repetition. It is front-loaded with the most critical information and stays focused.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers the tool's purpose, mechanism, side effects, and when to run it. It also acknowledges an output summary (absolute vs relative fallback). However, it omits any explanation of the dry_run parameter, which is a key input that affects behavior. Given the presence of an output schema, return format doesn't need detailing, but the parameter gap leaves the description incomplete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The only parameter, dry_run, is entirely undocumented in the description. The schema gives its name, type, default, and requirement status but provides no semantic meaning. The description mentions 'non-dry-run' indirectly but never explains what dry_run does, when to set it to true, or what the difference is. With 0% schema description coverage, this is a significant gap—the description fails to compensate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's specific purpose: to verify that input injection reaches the compositor. It describes the concrete steps (creating virtual devices, moving pointer to three points, capturing) and the distinctive outcome (determining whether absolute positioning or relative fallback is needed). This distinguishes it from the many sibling tools, even without naming them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives a clear usage context: 'Run this once per machine before trusting a real (non-dry-run) Playbook.' It explains that this is a one-time calibration step and that the result cannot be known without trying, which implies this is the tool to use for that purpose. It doesn't explicitly mention alternatives or when not to use it, but the guidance is strong.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
voltage_captureARead-only
Take a screenshot and return it to you directly.
Use this to see the screen yourself -- before writing a Playbook, to pick coordinates for probes and click regions, or to work out why a run went wrong. This does not involve the local vision model.
region is [x, y, width, height] in desktop pixels; omit for the whole desktop.
| Name | Required | Description | Default |
|---|---|---|---|
| region | No | ||
| max_width | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=false, so the description's job is to add behavioral context. It does so by specifying 'return it to you directly' and explicitly noting it does not use the local vision model—information not in annotations. No contradictions exist.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is tightly written: a clear one-sentence purpose, a brief usage block, and a one-line parameter definition. No fluff or redundancy. The key purpose is front-loaded, and every sentence contributes value.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple screenshot tool, the description covers usage, region format, and direct return. However, the `max_width` parameter remains unexplained; an agent might not know how to set it or its effect on output. Given that the tool has no output schema, a bit more detail on max_width would make it fully complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description explains the `region` parameter fully: it is [x, y, width, height] in desktop pixels and can be omitted for the whole desktop. However, `max_width` is not described at all; the schema only shows it is an integer with default 1280. Since schema description coverage is 0%, the description should compensate for both parameters, but it only covers one.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening sentence 'Take a screenshot and return it to you directly' uses a specific verb and resource, and clearly states the result. It also distinguishes itself from the vision-model-based sibling by saying 'This does not involve the local vision model,' which makes its purpose unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It provides concrete use cases: 'before writing a Playbook, to pick coordinates for probes and click regions, or to work out why a run went wrong.' This tells the agent exactly when to invoke it. It does not explicitly mention alternatives or when not to use it, but the context is clear enough to guide selection.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
voltage_diagnoseARead-only
Explain why a run behaved as it did, and what to change.
Call this instead of reading the journal by hand. It computes what the journal
implies but does not state -- watch labels the vision model never once reported,
guards that never evaluated true, whether bursts actually moved the screen, whether
the actuator is chaining or emitting one action at a time -- and returns each with
the specific edit that fixes it, ordered blocker-first.
The distinction it exists for: a burst that never ran and a burst that ran and did nothing look identical in a summary and have unrelated causes. The first is policy or grammar; the second is window focus, pointer mode, or an application that ignores synthetic input.
Apply the highest-severity finding, re-run, diagnose again. Changing several things at once makes the next diagnosis uninterpretable.
| Name | Required | Description | Default |
|---|---|---|---|
| run_id | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=false, so the description doesn't need to restate safety. It adds valuable behavioral detail: it computes implicit journal information, returns specific edits ordered blocker-first, and distinguishes between a burst that never ran vs. ran but did nothing. This goes well beyond the annotation, providing non-obvious nuances.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with a crisp summary, then explains the key distinction and ends with an actionable workflow. Every sentence earns its place; there is no fluff or redundancy. Structure is clear and scannable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given that an output schema exists (the description doesn't need to detail return structure) and annotations cover safety, the description covers the essential context: the diagnostic purpose, the key distinction between two root causes, and the iterative workflow. The only minor gap is the run_id parameter semantics, which slightly detracts from completeness for an otherwise simple tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description carries the full burden for the single parameter run_id. It never mentions run_id, its format, how to obtain it, or whether it's required (though the schema marks it optional). The name 'run_id' is self-explanatory by convention, but the description provides no explicit guidance, and with only one parameter to cover, this is a noticeable gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Explain why a run behaved as it did, and what to change.' It then contrasts itself with reading the journal, making its purpose distinct from voltage_journal. No ambiguity about what the tool does.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly instructs to call this instead of reading the journal by hand, giving clear when-to-use context. It also provides a workflow (apply highest-severity finding, re-run, diagnose again). However, it doesn't name alternative siblings like voltage_doctor or voltage_observe, or describe conditions where those might be more appropriate, so it stops short of complete guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
voltage_doctorARead-only
Check that everything needed for a run is present and working.
Reports the session type, input-device permissions, which capture backends work, detected screen geometry, GPU memory versus the selected model profile, and whether both model backends respond. When something is missing it returns the exact command to fix it. Call this before the first run on a machine.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotations already declare readOnlyHint=true and openWorldHint=false, establishing a safe, closed-world read operation. The description adds valuable context beyond safety: it lists the specific components checked (session type, input-device permissions, capture backends, screen geometry, GPU memory vs model profile, both model backends) and states that it returns fix commands. This informs the agent about the scope of the check and the nature of the response, which is more than annotations provide. No contradiction found.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three sentences, tightly packed with information. The first sentence gives the core purpose, the second enumerates the checks and the fix-command behavior, and the third provides usage timing. Every sentence earns its place, and the most important information (purpose and when to use) is front-loaded. There is no fluff or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a no-argument diagnostic tool, the description is comprehensive. It clearly states what is checked, the response characteristic (returns fix commands), and when to call it. An output schema exists (though not shown), so the description doesn't need to detail return formats. Given the complexity of the checks and the existence of a schema, nothing essential is missing for the agent to use the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters and the schema shows none. The description implicitly confirms this by stating 'Call this before the first run on a machine' with no mention of inputs. Since there are no parameters to explain, the description effectively communicates that it requires no configuration. This is a perfect fit for the no-parameter case, and the baseline of 4 is exceeded because the description makes the absence of parameters obvious.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: 'Check that everything needed for a run is present and working.' It specifies a concrete action (check) and a distinct resource (run prerequisites). It differentiates from siblings like voltage_status and voltage_diagnose by enumerating the exact checklist items (session type, permissions, capture backends, geometry, GPU memory, model backends). This makes it unambiguous which tool to select for pre-flight validation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit timing guidance: 'Call this before the first run on a machine.' While it doesn't mention alternatives or when not to use it, the instruction is clear and actionable. It implies this is a single-use setup check, not a repeated monitoring tool. The guidance is sufficient for the agent to decide when to invoke it, though lacking explicit exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
voltage_execute_burstADestructive
Execute one input burst yourself, bypassing the local models entirely.
For moments that need your judgement rather than the actuator's: opening the right application, clicking a specific confirmed target, typing something exact. Also the fastest way to sanity-check that input injection works at all.
Syntax: m:640,360;c:l;w:120;t:"hello";k:enter. Call voltage_reference for the full
list. The safety policy still applies. Defaults to dry_run, so pass dry_run=false to
actually inject.
| Name | Required | Description | Default |
|---|---|---|---|
| burst | Yes | ||
| label | No | manual | |
| dry_run | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint: true, readOnlyHint: false, and openWorldHint: true. The description adds critical behavioral context: it defaults to dry_run ('pass dry_run=false to actually inject') and notes the safety policy. It also explains that this is a manual override path. These details go beyond the annotations and inform the agent about side effects and prerequisites.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-structured and efficient: it leads with the core action, then gives usage scenarios, then provides a syntax example and necessary caveats. Every sentence earns its place, and the dry_run warning is front-loaded within the critical context. No fluff or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (custom syntax, safety policy, dry_run default) and that an output schema exists, the description covers the essential aspects: purpose, when to use, how to construct the burst (via example and reference), and the dry_run behavior. The only gap is a full in-place explanation of the syntax and label, but the reference to voltage_reference and the presence of an output schema mitigate this. Overall, it is nearly complete for an agent to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It provides a concrete syntax example (`m:640,360;c:l;w:120;t:"hello";k:enter`) and explains the dry_run parameter clearly. However, burst syntax is not fully documented (only a pointer to voltage_reference) and the label parameter is not explained beyond its default. This is partial compensation—helpful but not exhaustive.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Execute one input burst yourself') and the resource (burst), and immediately differentiates from siblings by emphasizing 'bypassing the local models entirely' and 'moments that need your judgement rather than the actuator's'. It also names the exact use case (opening applications, clicking confirmed targets, typing exact text) and points to voltage_reference for full syntax, making the purpose unmistakable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly states when to use this tool ('For moments that need your judgement rather than the actuator's', 'the fastest way to sanity-check that input injection works at all'), implies alternatives by referencing voltage_reference for syntax, and reminds that 'the safety policy still applies'. This gives an agent clear decision-making guidance without ambiguity.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
voltage_journalARead-only
Read a run's cycle-by-cycle record: what was seen, decided, refused, executed.
only_refused=true filters to cycles the governor blocked, which is the fastest way
to see where a Playbook's policy and the actuator's intentions disagree.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| run_id | No | ||
| only_refused | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already mark readOnlyHint=true, and the description aligns with that by saying 'Read'. It adds value by explaining the behavioral semantics of the journal contents and the meaning of 'only_refused', which goes beyond the raw annotation. No contradiction exists.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two focused sentences with the core purpose front-loaded and the filter tip as a concise, well-formatted follow-up. No filler or repetition, and the code-styled parameter reference is efficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return format is covered. However, the description fails to explain the run_id parameter, which is central to selecting a run, and gives no mention of limit. The tool is simple with all optional params, but the missing parameter descriptions leave a gap in usability.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It explains only_refused in detail, but completely omits run_id and limit. run_id is critical for identifying which run to read, and limit is a common but still undocumented control. The description is inadequate for a zero-coverage schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Read') and resource ('a run's cycle-by-cycle record'), and lists the exact contents: what was seen, decided, refused, executed. This clearly distinguishes it from siblings like voltage_observe or voltage_diagnose by framing it as a chronological journal rather than a live observation or diagnostic.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives explicit context for using the 'only_refused' filter and explains the fastest way to see policy/actuator disagreement. While it doesn't mention sibling tools for comparison, the usage hint is concrete and actionable, and the description clearly implies this tool is for inspecting historical decisions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
voltage_learnADestructive
Record something worth carrying to the next run against this target.
Write these as concrete, reusable facts, not narration:
good "the health bar is at x=120..300, y=1010; region_mean on red channel works" good "vision reports 'hotbar' reliably but never 'crosshair' -- do not watch it" good "block placement needs w:100 after the right click or it does not register" bad "the run failed" bad "tried again and it worked better"
kind groups them: label (what the vision model does and does not recognise),
timing (waits that a specific application needs), policy (what the governor blocked
and whether that was right), burst (a sequence that works), observation (anything
else).
| Name | Required | Description | Default |
|---|---|---|---|
| kind | No | observation | |
| note | Yes | ||
| target | Yes | ||
| playbook | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already indicate a mutating, potentially destructive action (readOnlyHint=false, destructiveHint=true); the description does not contradict these and adds that notes are stored against a target. It does not describe side effects or permissions, but given annotation coverage it provides acceptable additional context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is efficiently structured: it opens with the core purpose, gives clear good/bad examples, and ends with a concise classification of kind values. Every sentence adds value, and the format is well-balanced for the tool's complexity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool that records notes, the description covers the purpose, content quality, and kind taxonomy, which is sufficient for basic use. Gaps remain around `playbook` and exact behavior (e.g., confirmation, persistence), but the presence of an output schema and annotations mitigates these. Overall it is fairly complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description compensates by explaining the meaning of `kind` (label, timing, policy, burst, observation) and prescribing the format for `note` via good/bad examples. It leaves `target` and `playbook` undefined, but `target` is self-evident and `playbook` remains ambiguous, so coverage is partial but effective.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool records reusable facts against a target, with concrete good/bad examples that make the purpose unmistakable. It does not explicitly differentiate from sibling tools like voltage_lessons, but the 'carrying to the next run' phrasing is specific enough to convey its unique role.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides strong guidance on what to record (concrete facts, not narration) and explains the kind grouping, but it never mentions alternative tools or conditions under which to avoid this tool. Usage context is implied rather than explicit, and no exclusions or comparisons are given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
voltage_lessonsARead-only
Recall what previous runs learned about driving something.
Call this before writing a Playbook for a target you have driven before. Lessons persist across sessions and are keyed by target ("minecraft", "roblox", "dolphin"), so a new Playbook can start from what the last one discovered -- which labels the vision model actually recognises, where the HUD probes are, what timing the game needs -- rather than rediscovering it.
Omit target to see everything recorded so far.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| target | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=true and the description aligns with that (no mutation implied). The description adds valuable behavioral context: lessons persist across sessions, are keyed by target, and include specific types of information (labels, HUD probes, timing). This goes beyond the annotation by describing persistence and content, which is useful for setting expectations about what the tool returns.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise and front-loaded with the purpose. It uses bold for emphasis ('before writing a Playbook') and keeps each sentence purposeful. There is no filler or redundant explanation. The structure guides the reader from what the tool does, to when to use it, to how to filter results.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has an output schema (as indicated by the context), so return values are documented elsewhere. The description provides sufficient context for an agent to decide when to call it: it explains the purpose, when it is appropriate (before writing a Playbook for a previously driven target), and how to control scope with the target parameter. No critical information is missing, given the read-only annotation and output schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 0% description coverage, so the description must compensate. It clearly explains the `target` parameter (keyed by target, omit to see everything) and gives examples of valid values. However, it does not mention the `limit` parameter at all, leaving its semantics to inference from the default value of 30. This is a partial compensation but not complete for both parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Recall what previous runs learned about driving something.' It then gives concrete examples of lesson content (labels, HUD probes, timing), which makes the tool's purpose unambiguous and distinct from any other sibling. The behavior is clearly scoped to recalling learned lessons, not a general-purpose query.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly instructs when to use the tool: 'Call this **before writing a Playbook** for a target you have driven before.' It also explains the benefit (start from previous discoveries rather than rediscovering) and provides parameter guidance: 'Omit `target` to see everything recorded so far.' This gives an agent clear, actionable context for choosing this tool over alternatives like voltage_learn.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
voltage_observeARead-only
Run one vision pass and return grounded elements in screen coordinates.
watch is the closed vocabulary the vision model may use -- it can only report
labels from this list, so name the things your Playbook's guards will test for.
Use this to check that the vision model can actually find what a state depends on
before committing to it in a Playbook. If an element does not come back here, a
sees(...) guard on it will never fire.
| Name | Required | Description | Default |
|---|---|---|---|
| watch | Yes | ||
| region | No | ||
| read_text | No | ||
| max_elements | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already convey read-only and closed-world hints. The description adds valuable behavioral context: it clarifies that 'watch' is a closed vocabulary, that the tool runs a single pass, and that missing elements imply guards never fire. This goes beyond the annotations and provides actionable insight into the tool's behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise, with two short paragraphs that are front-loaded with the core purpose. Every sentence adds distinct value—stating the action, vocabulary constraint, and practical implication. There is no fluff or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description captures the tool's primary purpose and a key behavioral consequence, and an output schema exists so return values are already documented. However, it does not explain non-required parameters (region, read_text, max_elements), which are likely needed for correct invocation. This gap reduces completeness, though the core use case is well covered.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It explains 'watch' as the closed vocabulary, which is essential, but it omits any explanation for 'region', 'read_text', and 'max_elements'. With only one parameter addressed, the description fails to adequately clarify the remaining parameters, leaving the agent with insufficient guidance.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states a specific verb and resource: 'Run one vision pass and return grounded elements in screen coordinates.' It also explains a distinct use case—checking if the vision model can find elements before committing to a Playbook. While it doesn't explicitly contrast with sibling tools, the purpose is specific and unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides explicit context for when to use the tool: 'Use this to check that the vision model can actually find what a state depends on before committing to it in a Playbook.' This is a clear directive without naming alternatives, but it effectively guides the agent on ideal usage scenarios.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
voltage_pauseBDestructive
Pause or resume a run. Held input is not released, so a paused run can continue.
| Name | Required | Description | Default |
|---|---|---|---|
| resume | No | ||
| run_id | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already flag destructiveHint=true, so the mutation nature is disclosed. The description adds the specific behavior that held input is retained, which goes beyond the annotations and gives the agent useful context about the pause/resume semantics. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with zero filler. The core action is front-loaded ('Pause or resume a run') and the clarifying detail about held input follows immediately. Every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the existence of an output schema and simple optional parameters, the description is far from complete. It lacks usage guidance, parameter semantics, and any mention of prerequisites or side effects beyond the held-input note. The agent would need to guess how to set 'resume' or when to pass 'run_id'.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% — neither 'resume' nor 'run_id' is explained in the schema. The description does not mention any parameters at all, so the agent has no idea that 'resume' likely indicates whether to resume or pause, or how 'run_id' selects the run. With two parameters and zero coverage, the description must compensate but fails completely.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a clear action (pause or resume), a specific resource (a run), and adds a key nuance (held input is not released). It distinguishes implicitly from voltage_stop but does not name sibling alternatives, so it falls short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No explicit guidance on when to use this tool versus alternatives like voltage_stop or voltage_run. The note about held input hints at a use case but does not state conditions or exclusions, leaving the agent to infer when pause is appropriate.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
voltage_referenceARead-only
Return everything needed to author and iterate on a run.
Call this before your first Playbook. Sections:
loop the learning loop -- how to go from a failed run to a working one, and what each failure mode actually means. Read this second. bursts the burst cookbook: how to chain inputs well, timing rules, ready-made patterns for desktop and for games, and the antipatterns that waste cycles. Read this if bursts are coming out one action at a time. burst the raw burst syntax playbook the state-machine JSON schema guards expression functions for transitions and reflexes example a complete working Playbook
| Name | Required | Description | Default |
|---|---|---|---|
| section | No | all |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, so the description need not restate safety. It adds value by explaining the content structure and the purpose of each section, which helps the agent understand what the tool actually returns. However, it does not disclose any potential caveats (e.g., response size, format specifics), though those may be covered by the output schema. The added context justifies a score slightly above baseline.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is efficiently organized: a one-line purpose, then a bulleted list of sections with clear labels and explanations. It front-loads the main instruction and uses formatting to allow fast scanning. No sentence is redundant; each adds useful detail about content or usage.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a reference tool, the description covers all essential information: what it returns, when to call it, what each section contains, and even contextual reading order. The read-only behavior is covered by annotations, and the output format is presumably defined by the output schema (present signal). Nothing necessary for an agent to select and invoke this tool is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must explain the 'section' parameter. It does so comprehensively by listing each enum value and its meaning, and even offers reading-order guidance (e.g., 'Read this second', 'Read this if...'). This fully compensates for the schema gap, making the parameter self-documenting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Return') and a resource ('everything needed to author and iterate on a run'), then enumerates the sections returned. It clearly distinguishes itself from sibling tools (e.g., voltage_execute_burst, voltage_validate_playbook) by being a reference/documentation tool, not an execution or validation tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly says 'Call this before your first Playbook,' giving a clear when-to-use directive. It also provides conditional reading order (e.g., 'Read this if bursts are coming out one action at a time') and labels like 'the learning loop,' which help an agent decide which section to request. This is strong, situation-specific guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
voltage_runADestructive
Start a Playbook. Returns immediately with a run_id; poll voltage_status.
dry_run overrides the Playbook's policy. Leave it unset for the Playbook's own
setting, which defaults to true. A dry run does everything except inject input, so
it is the correct way to check that your states, guards and transitions behave before
letting it touch the machine.
target_period_s is the loop period. 0.5 is a good default; lower it for games,
raise it for slow UI.
Stop a run with voltage_stop, adjust it live with voltage_steer. The run also stops on its own budget, on any physical keyboard or mouse input from the user, and on the panic file.
| Name | Required | Description | Default |
|---|---|---|---|
| dry_run | No | ||
| playbook | Yes | ||
| keep_frames | No | ||
| target_period_s | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already mark destructiveHint and openWorldHint, and the description complements these by explaining concrete behaviors: immediate return with run_id, polling requirement, dry_run overriding policy, and the specific conditions that terminate a run. It adds value beyond annotations without contradicting them.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is efficiently structured: the core action and return contract are front-loaded, followed by parameter guidance and termination behavior. Every sentence adds functional value, and no redundant or filler content is present.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers the essential lifecycle: starting, monitoring, adjusting, and stopping. It explains dry-run semantics and stopping triggers. However, it does not describe the structure of the `playbook` object or the meaning of `keep_frames`, which may be important for correct invocation. The presence of an output schema and related tools (voltage_validate_playbook) partially mitigates this.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must explain parameters. It does explain dry_run (including override semantics and default behavior) and target_period_s (with recommended values), but it does not explain `playbook` (the required parameter) or `keep_frames`. Since playbook is central and the schema offers no description, this leaves a gap for an agent constructing a valid call.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with 'Start a Playbook,' a specific verb-resource pairing that clearly states the tool's core function. It immediately distinguishes itself from siblings by mentioning polling with voltage_status, stopping with voltage_stop, and live adjustment with voltage_steer, so the agent can tell it apart without opening other schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides explicit context for when to use dry_run ('the correct way to check that your states, guards and transitions behave before letting it touch the machine'), recommends values for target_period_s, and explains how to stop or adjust a run using sibling tools. It also details automatic stopping conditions (budget, keyboard/mouse input, panic file), giving clear operational guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
voltage_statusARead-only
Poll a run: current state, variables, last burst, what the vision model sees.
Includes recent cycles, governor refusals, and per-stage timings so you can tell whether a slow loop is capture, vision, decision, or execution.
| Name | Required | Description | Default |
|---|---|---|---|
| run_id | No | ||
| journal_tail | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=false, so the read-only nature is covered. The description adds useful context beyond annotations: the specific data included (recent cycles, governor refusals, per-stage timings) and its diagnostic intent. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with no waste. The action is front-loaded ('Poll a run'), followed by a list of what it returns and the diagnostic purpose. Every phrase earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a monitoring tool with an output schema present, the description conveys enough about the returned data to be useful. However, the lack of parameter documentation is a notable gap that makes it slightly incomplete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description does not explain either parameter (run_id, journal_tail) at all. While run_id is somewhat inferable from its name, journal_tail is completely unexplained. The description fails to compensate for the schema's lack of documentation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb ('Poll') and resource ('a run'), then enumerates the returned data (state, variables, last burst, vision model view, cycles, refusals, timings). This clearly differentiates it from sibling tools like voltage_capture or voltage_execute_burst, which imply different actions.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It implies usage during a run to monitor state and diagnose slow loops ('so you can tell whether a slow loop is capture, vision, decision, or execution'). However, it doesn't explicitly state when not to use it or point to alternatives such as voltage_doctor or voltage_diagnose, leaving some ambiguity.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
voltage_steerADestructive
Correct a live run without restarting it.
hint is injected into the actuator's prompt as a supervisor note and persists until
changed -- use it when the actuator is doing something legal but wrong.
force_state jumps the machine on the next cycle. variables updates run variables.
dry_run can be flipped either way mid-run.
| Name | Required | Description | Default |
|---|---|---|---|
| hint | No | ||
| run_id | No | ||
| dry_run | No | ||
| variables | No | ||
| force_state | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already flag destructiveHint=true, so the description adds some context: hint persists, force_state jumps the machine, variables updates, dry_run can flip. However, it does not disclose potential side effects, irreversibility, or prerequisites despite the destructive nature. It does not contradict the annotations, but the coverage is not thorough.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise and well-structured: a one-sentence overview followed by per-parameter explanations. It is front-loaded with the main purpose, uses backticks for param names to aid scanning, and has no filler or redundant statements.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with 5 parameters, zero schema descriptions, a destructive annotation, and an output schema, the description covers the core actions but misses run_id semantics, any warning about destructive consequences, and what the output schema contains. It is usable but not fully complete for safe and correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description carries the full burden for parameter meaning. It explains hint, force_state, variables, and dry_run, but omits run_id entirely, leaving its role merely implied by the phrase 'a live run.' This is a partial but incomplete compensation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb-resource pair: 'Correct a live run without restarting it,' which clearly distinguishes this tool from siblings like voltage_stop, voltage_pause, or voltage_run. It also enumerates the effects of each parameter, leaving no ambiguity about what the tool does.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides a concrete usage scenario for `hint` ('when the actuator is doing something legal but wrong') and explains the function of each parameter (e.g., force_state jumps the machine, dry_run flips). It implies this tool is for mid-run corrections vs. restarting, but does not explicitly name alternatives or exclusion conditions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
voltage_stopADestructive
Stop a run and release every held key and button.
Safe to call at any time, including while a burst is mid-flight -- the burst is interrupted and anything held is released.
| Name | Required | Description | Default |
|---|---|---|---|
| reason | No | stopped by orchestrator | |
| run_id | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, and the description adds concrete behavior: releases every held key/button and interrupts bursts. This goes beyond the annotation's generic destroy flag without contradicting it, giving the agent a more precise model of consequences.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two concise sentences convey the purpose, safety, and edge-case behavior with zero filler. Information is front-loaded and every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple stop tool, the description covers the main behavior and safety profile. However, the lack of any parameter explanation means an agent might guess wrong about 'run_id' or 'reason' (e.g., whether run_id is required to target a specific run). Optional parameters with defaults mitigate, but the gap prevents full completeness.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the description does not explain either 'reason' or 'run_id.' The agent has no guidance on what these parameters control or when to provide them, though they are optional. With no parameter documentation anywhere, this is a notable gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Stop') and resource ('a run') while adding unique scope: 'release every held key and button.' This clearly distinguishes it from siblings like voltage_pause and voltage_run, and the mention of interrupting mid-flight bursts further clarifies its specific role.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context: 'Safe to call at any time' and explicitly covers the edge case of a mid-flight burst. However, it does not explicitly contrast with alternatives like voltage_pause or voltage_steer, leaving some ambiguity about when to choose this over a pause or a graceful stop.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
voltage_validate_playbookARead-only
Fully check a Playbook without running it.
Validates the schema, compiles every guard expression, parses every burst, checks that transition targets and probe references exist, and reports unreachable states and dead transitions. Errors come back as a complete list, not one at a time.
Always call this before voltage_run. Warnings are worth reading: "tests for X but X
is not in watch" means a transition that can never fire.
| Name | Required | Description | Default |
|---|---|---|---|
| playbook | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only mark readOnlyHint:true. The description adds substantial behavioral detail: it returns a complete list of errors rather than one at a time, reports unreachable states and dead transitions, and explains how to interpret warnings. This fully complements the annotation and does not contradict it.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is tightly written, starting with the primary purpose, then detailing checks, then error behavior, then usage guidance and a warning interpretation. Every sentence serves a purpose—no filler. It front-loads the action and clearly organizes information in short block format.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter validation tool with an output schema declared (though not shown explicitly), the description covers what it does, how it behaves, when to call it, and how to interpret results. With annotations covering read-only safety and the output schema expected to define return values, nothing essential is missing for an agent to decide and invoke correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema only defines a generic 'playbook' object with no description (0% coverage). The description compensates by making clear that the parameter is the Playbook being validated, and it describes what validation entails (schema, guards, bursts, references). This gives the agent enough context to pass the correct object, even without knowing its internal structure.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific, unambiguous statement: 'Fully check a Playbook without running it.' It enumerates the exact validations performed (schema, guards, bursts, transition targets, probe references) and reports unreachable states/dead transitions, distinguishing this validation tool from siblings like voltage_run and voltage_execute_burst. The verb 'validate' matches the tool name and clears its role.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly guides usage with 'Always call this before voltage_run,' which states when to use this tool relative to its primary sibling. It also adds a practical hint about interpreting warnings (e.g., 'tests for X but X is not in watch'). It does not list explicit exclusions, but the directive is clear and directly actionable.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
16 tool updates
v0.1.0- First observed
voltage_calibrate - First observed
voltage_capture - First observed
voltage_diagnose - First observed
voltage_doctor - First observed
voltage_execute_burst - First observed
voltage_journal - First observed
voltage_learn - First observed
voltage_lessons - First observed
voltage_observe - First observed
voltage_pause - First observed
voltage_reference - First observed
voltage_run - First observed
voltage_status - First observed
voltage_steer - First observed
voltage_stop - First observed
voltage_validate_playbook
TDQS
Scored across 16 tools
Each tool has a clearly distinct purpose: pre-flight checks, documentation, perception, input execution, validation, running, monitoring, control, and learning. Even similar tools like voltage_journal (raw data) and voltage_diagnose (analyzed explanation) are cleanly separated by their roles.
All tools follow a consistent voltage_ prefix with a verb or verb_noun pattern (capture, execute_burst, validate_playbook, etc.). No mixed conventions or ambiguous verbs; naming is predictable and intuitive.
16 tools is well-scoped for a comprehensive automation server covering setup, execution, monitoring, debugging, and learning. Each tool earns its place; the count supports the full workflow without bloat.
The tool surface covers the entire lifecycle: environment checks (doctor, calibrate), documentation (reference), perception (capture, observe), manual action (execute_burst), validation and execution (validate_playbook, run), live control (steer, stop, pause), monitoring (status, journal, diagnose), and cross-session learning (lessons, learn). No obvious gaps for the stated purpose.
Maintenance
Related MCP Connectors
Operate Linux, macOS and Windows from your LLM. Every action runs through an auditable allowlist.
Human-in-the-loop approval for agent actions, with verifiable action-bound receipts.
Control real Android and iOS devices with LLM agents — tap, swipe, type, automate flows.
Lets AI agents use a real human as a tool: visual checks, taste, phone calls, unblocking, approvals
Related MCP Servers
- AlicenseNot gradedqualityAmaintenanceA local autonomous AI agent that watches your screen, understands the visual layout, and executes native OS commands (clicking, typing) without cloud APIs.16MIT
- FlicenseNot gradedqualityBmaintenanceEnables AI agents to declaratively control web pages using real mouse and keyboard events via Chrome DevTools Protocol, without executing page JavaScript.10 npm1-
- AlicenseAqualityBmaintenanceEnables low-cost agent models to control Windows applications through a compact, state-safe proxy over Open Computer Use, reducing model-visible context by up to 99.8% with support for record/replay and reusable UI component memory.5MIT
- AlicenseNot gradedqualityAmaintenanceEnables AI agents to automate real desktop applications across Windows, Linux, and macOS using incremental screen perception, accessibility trees, OCR, and window management, dramatically reducing token usage compared to screenshot-per-step approaches.39 PyPI3MIT