VoltageInputMcp
VoltageInputMcp
프론티어 모델이 도구 호출 속도가 아닌 입력 속도로 컴퓨터를 조작할 수 있게 해주는 MCP 서버입니다.
문제
컴퓨터 사용 도구는 모든 동작마다 원격 모델에 왕복 요청을 보냅니다. 스크린샷 올리고, 결정 내리고, 클릭 한 번. 양식 작성에는 충분하지만 일련의 입력을 빠르게 전달해야 하는 모든 작업에는 쓸모가 없습니다 — 게임 플레이, 모달 대화상자 조작, 타임라인 구동, 세 번째 입력이 첫 두 입력이 이미 도착했을 때에만 의미가 있는 모든 UI가 그렇습니다. 병목은 모델의 지능이 아닙니다. 지능이 800ms 떨어져 있는데 입력은 8ms 간격으로 전달되어야 한다는 것이 문제입니다.
Related MCP server: live-mcp
해답의 형태
결정과 실행을 분리하고, 실행을 키보드와 같은 머신에 둡니다.
┌─────────────────────────────────────────────────────────────────┐
│ Layer 1 — the orchestrator (Claude, or any MCP client) │
│ Writes a Playbook: states, what to look for, what is allowed, │
│ when to move on. Thinks once, up front. Watches and corrects. │
└───────────────────────────┬─────────────────────────────────────┘
│ MCP
┌───────────────────────────▼─────────────────────────────────────┐
│ Layer 2 — two small local models, on your GPU │
│ │
│ vision (Qwen2.5-VL-3B) "of these specific things, │
│ which are on screen, and where?" │
│ actuator (Qwen3-1.7B) "given that, which inputs?" │
│ │
│ Neither plans. Both answer one closed question per cycle. │
└───────────────────────────┬─────────────────────────────────────┘
│
┌───────────────────────────▼─────────────────────────────────────┐
│ safety governor → /dev/uinput → the actual desktop │
└─────────────────────────────────────────────────────────────────┘오케스트레이터는 두뇌입니다. 작은 모델들은 팔입니다. 팔은 똑똑하지 않으며 똑똑할 것을 요구받지도 않습니다.
속도의 실제 원천
작은 모델이 빨라서가 아닙니다 — 3B VLM도 여전히 ~300ms가 걸립니다. 속도는 영향력이 큰 순서대로 네 가지에서 나옵니다:
버스트(Bursts). 액추에이터는 입력 하나를 내보내지 않습니다. 버스트를 내보냅니다: 모델이 개입하지 않는 전용 실행기가 실행하는 타이밍이 지정된 입력 프로그램입니다.
g:0;c:l;w:150;t:"README.md";k:enter;w:80;k:ctrl+s그것은 하나의 결정과 ~400ms에 걸친 일곱 개의 입력이며, 밀리초 단위로 스케줄링됩니다. 40개 동작의 버스트도 여전히 결정 하나만 필요합니다. 입력 속도는 모델이 아니라 버스트가 결정합니다.
반사(Reflexes). 결정 사이에 모델 없이 마이크로초 단위로 실행되는 저비용 화면 프로브(픽셀 하나, 영역 평균 하나)에서 발화하는 규칙입니다.
{"id": "heal", "when": "probe('health') < 0.25", "do": "k:q;w:60", "cooldown_ms": 800}지각(Perception) 건너뛰기. 대부분의 사이클에서 화면은 변하지 않았습니다. 40µs 프레임 차이 검사가 비전 모델에 300ms를 쓸지, 아니면 마지막 관측값을 재사용할지 결정합니다. 일반적인 데스크톱 작업에서는 대부분의 사이클에서 VLM을 건너뜁니다.
프롬프트 캐시 지역성(Prompt-cache locality). 프롬프트는 정적 우선 순서로 구성되어 llama.cpp가 KV 캐시를 재사용하고 변경된 꼬리 부분만 다시 프리필합니다.
작은 모델이 작음에도 불구하고 신뢰할 수 있는 이유
신뢰할 수 있도록 요구받지 않기 때문입니다 — 제약을 받습니다.
llama.cpp에서 두 모델 모두 현재 상태에서 매 사이클마다 재생성되는 GBNF 문법에 따라 생성합니다. 문법은 조언이 아닙니다. 유효한 파싱을 계속하는 토큰만 도달 가능하도록 로짓을 마스킹합니다. 구체적으로, 액추에이터는 다음을 할 수 없습니다:
잘못된 형식의 버스트를 내보내는 것
정책이 거부하는 키를 지정하는 것 — 키가 문법에 없습니다
관측되지 않은 요소를 참조하는 것 — 인덱스 범위는 이번 사이클의 요소 수로 구성됩니다
Playbook이 선언하지 않은 상태 전이를 제안하는 것
그리고 비전 모델은 UI 요소 이름을 발명할 수 없습니다: 라벨 어휘는 사용자가 작성한 watch 목록에 작은 일반 집합을 더한 것입니다. 따라서 sees("address bar") 가드는 3B 모델이 만들어낸 임의의 명사가 아니라 폐쇄된 어휘를 비교합니다.
재시도 루프도 방어적 JSON 파싱도 없습니다. 잘못된 형식의 출력이 드물어서가 아니라 — 표현 자체가 불가능하기 때문입니다.
Playbook
작은 모델에 목표를 주지 않습니다. 상태 머신을 줍니다. 전이는 런타임이 평가하는 가드 표현식이며, 모델이 아닙니다.
{
"name": "open_downloads",
"goal": "Open the file manager at ~/Downloads. Delete nothing, confirm nothing.",
"initial": "launch",
"policy": {
"dry_run": true,
"allow_verbs": ["g", "c", "k", "t", "w"],
"deny_labels": ["delete", "trash", "confirm", "empty trash"]
},
"budget": { "max_cycles": 60, "max_seconds": 90 },
"states": {
"launch": {
"brief": "Open the application launcher and start the file manager.",
"watch": ["application launcher", "search field", "file manager icon"],
"on_enter": "k:meta;w:400",
"transitions": [
{ "when": "sees('search field')", "to": "type_name" },
{ "when": "cycles() > 6", "to": "@failure", "note": "launcher never opened" }
]
},
"navigate": {
"brief": "Focus the location bar with ctrl+l, type the path, press Enter.",
"watch": ["location bar", "file list", "error message"],
"on_enter": "k:ctrl+l;w:200",
"transitions": [
{ "when": "text('Downloads')", "to": "@success" },
{ "when": "sees('error message')", "to": "@failure" }
]
}
},
"success_when": "text('Downloads') and not flag('loading')"
}voltage_reference는 전체 DSL, JSON 스키마, 가드 함수 테이블을 반환하므로 오케스트레이터가 이 저장소를 읽지 않고도 Playbook을 작성할 수 있습니다.
성능 튜닝
아래 모든 수치는 참조 머신(RTX 3050 6GB 노트북, llama.cpp에서 Qwen2.5-VL-3B + Qwen3-1.7B)에서 측정된 값이지 유도된 값이 아닙니다.
두 모델 모두 디코드 바운드입니다. 출력 토큰만이 유일하게 중요한 레버입니다.
이것은 놀라움이었습니다 — 설계는 원래 비전이 프리필 바운드라고 가정했지만, 그렇지 않습니다. 프리필은 448×252에서 896×504까지 ~28ms로 일정하게 측정되었습니다. 디코드는 ~22ms/토큰으로 실행됩니다. 따라서:
항목 | 비용 |
출력 토큰 하나 | ~22 ms |
보고된 요소 하나 | ~21 토큰 ≈ 500 ms |
비전, 요소 2개 | ~1.0 s |
비전, 요소 4개 | ~2.2 s |
액추에이터, 캐시된 프리픽스 | 노트 길이에 따라 140–400 ms |
각각 기본값을 바꾼 세 가지 결과:
max_elements는 비전 비용을 지배합니다. 기본값은 3입니다. 6으로 올리면 인지 사이클당 ~1.5s가 추가됩니다. 가드가 실제로 검사하는 수로 설정하세요.downscale_to를 줄이는 것은 도움이 되지 않으며 대개 해롭습니다. 448×252는 896×504보다 2.5배 느리게 측정되었습니다 — 더 흐릿한 이미지는 모델을 덜 확신하게 만들어 더 많은 토큰을 생성합니다. 들어맞는 가장 큰 크기를 사용하세요.액추에이터의
note필드는 지연 시간의 55%를 차지했습니다. 순수 진단용이며, 48자에서 사이클당 412ms로 측정된 반면 12자에서는 184ms, 0자에서는 140ms였습니다. 기본값은 이제 12입니다.
요소는 같은 이유로 {"l":"address bar","b":[...],"c":0.9} 대신 [label_index, x1, y1, x2, y2]로 인코딩됩니다 — 토큰 27–29% 감소, 지연 시간 32–41% 감소로 측정되었습니다. 폐쇄된 watch 어휘로 인덱싱하는 것도 더 안전합니다: 모델이 라벨을 철자할 수조차 없으며, 오타는 말할 것도 없습니다.
GBNF 평가는 샘플링된 토큰마다 CPU에서 한 번 실행되므로, 액추에이터는 완전히 GPU 오프로드됨에도 불구하고 비전 모델보다 더 많은 CPU 스레드를 받습니다 — 그리고 allow_keys를 제한하는 것은 안전성 최적화일 뿐만 아니라 지연 시간 최적화이기도 합니다.
잘못 설정하면 조용히 실패하는 두 가지 설정:
빌드 시
GGML_CUDA_FA_ALL_QUANTS=ON. 우리는q8_0KV 캐시 및 플래시 어텐션으로 서빙합니다. 이 플래그가 없으면 llama.cpp는 해당 KV 조합에 대한 FA 커널을 컴파일하지 않고 느린 경로로 폴백합니다 — 오류 없이, 그저 수상할 정도로 나쁜 수치만 나옵니다.scripts/build-llama.sh가 이 값을 설정합니다.런타임 시
GGML_CUDA_ENABLE_UNIFIED_MEMORY=0.1이면 VRAM 오버플로가 실패하는 대신 조용히 PCIe로 넘쳐 흐릅니다. 모든 것이 작동하고 ~10배 느립니다.serve.sh가 이 값을 고정합니다.
추측하지 말고 측정하세요:
.venv/bin/voltage bench루프가 사용하는 정확한 프롬프트 형태로 두 백엔드를 모두 구동하고 콜드 vs 프롬프트 캐시 지연 시간, 세 가지 입력 크기에서 ms/비주얼 토큰, 그리고 그로부터 암시되는 사이클 시간을 보고합니다. 프롬프트 캐시 속도 향상이 ~1.5× 미만이면 무언가 동적인 것이 프롬프트 프리픽스로 새어 들어갔다는 뜻입니다.
모델 비교
가장 obvious한 실험 — "어느 모델이 더 나은 버스트를 작성하는가" — 은 잘못된 것을 측정합니다. 문법이 이미 모든 버스트가 유효함을 보장하므로, 더 큰 모델이 구문에서 이길 수 없습니다. 구성이 사용 가능한지 여부를 실제로 결정하는 것은:
접지 정확도(Grounding accuracy). 200ms 더 빠르고 40px 벗어난 모델은 쓸모없습니다 — 클릭이 빗나갑니다. 클릭은 중심에 착지하므로 IoU가 아닌 화면 픽셀 단위 중심 거리로 측정합니다.
제약 하의 결정 품질. 동일한 관측이 주어졌을 때 올바른 합법적 동작을 선택하는가, 그리고 사이클당 소심한 동작 하나를 내보내는 대신 전체 시퀀스를 하나의 버스트로 연결하는가.
지연 시간. 1과 2가 수용 가능해진 후에만 중요합니다.
.venv/bin/voltage fixture desktop # capture a real screen
.venv/bin/voltage compare # score whatever is running nowGround truth는 오케스트레이팅 모델이 라벨링한 실제 스크린샷에서 나옵니다 — 이는 이 시스템이 런타임에 사용하는 것과 동일한 참조입니다. 합성 UI는 함정입니다: 그려진 사각형은 실제 인터페이스로 훈련된 모델에게 버튼으로 읽히지 않으므로, 그것에 대해 점수를 매기는 것은 잘못된 기술을 측정하는 것입니다.
결과는 실행 간에 누적되므로 워크플로는: 프로필 A 서빙 → compare → 프로필 B 서빙 → compare → 테이블 읽기입니다. voltage compare --list는 재실행 없이 출력합니다.
픽스처는 사용자의 것이며 커밋되지 않습니다. 스크린샷에 개인적인 것이 포함되어 있으면 fixtures/를 .gitignore에 추가하세요.
학습 루프
익숙하지 않은 대상을 위한 첫 번째 Playbook은 거의 항상 옳지 않습니다. 중요한 것은 실패가 구체적이고, 다음 시도가 마지막 시도가 배운 것에서 시작한다는 것입니다.
voltage_reference(section="loop") the loop itself, and what each failure means
voltage_reference(section="bursts") the burst cookbook: chaining, timing, game patterns
voltage_capture / voltage_observe look before writing — check your labels exist
voltage_validate_playbook dead guards, unreachable states, caught statically
voltage_run(dry_run=true) real models, real screen, nothing injected
voltage_diagnose(run_id) ← what to change, not raw data
voltage_learn(target=..., note=...) record it; persists across sessions
voltage_lessons(target=...) recall it before the next playbookvoltage_diagnose는 이것을 루프로 만드는 조각입니다. 저널이 암시하지만 명시하지 않는 것을 계산하고 각각에 대한 수정을 지정합니다. 막힌 Minecraft 실행에서:
[BLOCKER] label_never_seen never reported: ['crosshair', 'health bar']
[BLOCKER] input_not_landing 14 bursts executed, but the screen never changed
[BLOCKER] state_never_left 'mine' ran 14 cycles and never transitioned
[PROBLEM] timid_bursts bursts averaged 1.0 actions
[HINT] vision_every_cycle vision ran on 100% of cycles그것이 존재하는 이유의 구분: 실행된 적 없는 버스트와 실행되었지만 아무것도 하지 않은 버스트는 요약에서 동일하게 보이지만 원인은 무관합니다. 첫 번째는 정책 또는 문법 문제입니다. 두 번째는 창 포커스, 포인터 모드, 또는 합성 입력을 무시하는 앱입니다. Diagnose는 실행 후 프레임이 실제로 변경되었는지 확인하여 둘을 구분합니다.
가장 높은 심각도의 발견을 적용하고, 다시 실행하고, 다시 진단하세요. 한 번에 하나의 변경만 — 여러 개를 동시에 하면 다음 진단을 해석할 수 없게 됩니다.
교훈은 세션 간에 지속되며, 대상별로 키가 지정되므로 게임의 두 번째 Playbook은 첫 번째가 발견한 프로브 좌표와 작동하는 라벨 이름에서 시작합니다:
voltage_learn(target="minecraft", kind="label",
note="vision reports 'hotbar' reliably but never 'crosshair'")
voltage_learn(target="minecraft", kind="timing",
note="block placement needs w:100 after right click or it does not register")안전
입력을 생성하는 것은 1.7B 모델입니다. 거버너는 조언이 아닌 계층입니다: 반사 버스트와 직접 작성한 버스트를 포함한 모든 버스트가 그것을 통과합니다.
dry_run이 기본값입니다. 새 Playbook은 아무것도 건드리지 않으면서 모든 버스트를 파싱, 검사, 저널링합니다.전체 버스트 거부. 의도된 시퀀스를 절반만 실행하는 것은 실행하지 않는 것보다 나쁩니다.
deny_labels는 Delete / Confirm / Purchase / Allow라고 불리는 모든 것에 대한 클릭을 거부합니다 — 예상치 못한 곳에 나타나는 대화상자를 잡아내는 것이 바로 이것입니다.영역 펜싱, 키 허용 목록, 거부된 코드(chord)(
ctrl+alt+delete,alt+f4), 거부된 텍스트 패턴(rm -rf,sudo), 버스트 크기 및 초당 입력 수 상한.네 가지 독립 정지 장치:
voltage stop(파일을 작성 — SSH를 통해 작동), 루프가 멈추면 자체 스레드에서 발화하는 데드맨 타이머, 물리적 입력 경합(실제 마우스를 건드리면 정지), Playbook 예산.누른 키는 항상 해제됩니다 — 중단 시, 충돌 시, 시간 초과 시.
d:shift와u:shift사이에 중단된 실행은 Shift가 눌린 채로 남아서는 안 됩니다.
설치
아무것도 없는 상태에서 작동까지, 두 명령.
Linux / macOS
git clone https://github.com/casualkre/voltage-input-mcp && cd voltage-input-mcp && ./install.shWindows (PowerShell)
git clone https://github.com/casualkre/voltage-input-mcp; cd voltage-input-mcp; powershell -ExecutionPolicy Bypass -File .\install.ps1그런 다음, 둘 중 하나에서:
voltage setupinstall.sh는 Python, 시스템 패키지, venv 및 PATH를 처리하고 루트가 필요한 항목에 대해 요청하는 대신 정확한 sudo 줄을 출력합니다. 그런 다음 voltage setup은 이미 가지고 있는 것을 감지하고, 누락된 것만 다운로드하며, 모델 서버를 시작하고, AI 클라이언트에 등록합니다 — 각 단계를 설명하는 것이 아니라 실행합니다. 10~25분, 거의 모두 다운로드 시간입니다. 다시 실행해도 안전합니다; 중단된 지점에서 이어갑니다.
그런 다음 그냥 실행하세요:
voltage설정은 이미 가지고 있는 것을 감지하고 거기서부터 계속합니다. 시작점을 가정하지 않습니다: OS, GPU, llama.cpp 또는 Ollama 설치 여부, 이미 가져온 모델, 입력 및 캡처 작동 여부, MCP 서버 등록 여부를 프로브한 다음 — 실제로 남은 단계만 계획하고, 어떤 단계가 사용자의 결정이 필요한지, 어떤 단계를 그냥 수행할 수 있는지 말합니다. 이미 Ollama가 있으면 그것을 사용합니다. 백엔드가 둘 다 없으면 두 줄로 트레이드오프를 설명하고 선택하게 합니다.
인수 없이 실행하면 대화형 콘솔이 열립니다: 실시간 상태, 의존성 순서로 준비되지 않은 것을 고치는 안내 설정, 모델 전환기, 구성 편집기, Claude Code 원키 등록, 진단. 아래의 모든 하위 명령은 여전히 비대화형으로 작동하므로 스크립트와 CI는 영향을 받지 않습니다.
██╗ ██╗ ██████╗ ██╗ ████████╗ █████╗ ██████╗ ███████╗
██║ ██║██╔═══██╗██║ ╚══██╔══╝██╔══██╗██╔════╝ ██╔════╝
██║ ██║██║ ██║██║ ██║ ███████║██║ ███╗█████╗
╚██╗ ██╔╝██║ ██║██║ ██║ ██╔══██║██║ ██║██╔══╝
╚████╔╝ ╚██████╔╝███████╗██║ ██║ ██║╚██████╔╝███████╗
╚═══╝ ╚═════╝ ╚══════╝╚═╝ ╚═╝ ╚═╝ ╚═════╝ ╚══════╝
── status ──────────────────────────────────────────────
ok input device /dev/uinput
ok vision model http://127.0.0.1:8080
ok actuator model http://127.0.0.1:8081
ok mcp registered claude mcp list
ok voltage on PATH ~/.local/bin/voltage실험적 프로필
voltage → 모델에 별도로 나열되며, 각각 수락해야 하는 경고 뒤에 있습니다. 측정 결과가 트레이드오프를 예측 가능하게 만들기 때문에 존재합니다: 디코드가 ~22ms/토큰에서 지배적이고 활성 파라미터에 따라 확장되므로, 모델을 줄이면 실제로 루프 속도가 올라갑니다. 대가를 치르는 것은 접지 정확도입니다.
profile | models | VRAM | trade |
| SmolVLM-500M + Qwen3-0.6B | ~2.2 GB | 루프 속도 3–4배, 접지(grounding)가 거의 작동하지 않음 |
| Qwen2.5-VL-3B + Qwen3-0.6B | ~3.8 GB | 더 빠른 결정, 접지 성능은 동일 |
| Qwen2.5-VL-32B + Qwen3-14B | ~34 GB | 최상의 접지, 1–2.5초/사이클 |
| Qwen2.5-VL-32B + Qwen3-30B-A3B | ~43 GB | ~3B 디코딩 속도로 30B 용량 |
| 3B + 0.6B on CPU | 없음 | GPU 없이 작동, 사이클당 수 초 |
특별히 언급할 가치가 있는 두 가지:
hyper는 위험한 프로필입니다. SmolVLM-500M은 접지 모델이 아닙니다. 박스를 반환하긴
하겠지만 종종 틀릴 것입니다 — 그리고 틀린 박스는 우아한 성능 저하가 아니라 잘못된
위치에 대한 클릭입니다. watch가 비어 있는 경우(프로브와 리플렉스가 실제 작업을
수행하는 경우) 또는 모든 클릭이 click_allow_regions와 require_target_element로
보호되는 경우에만 사용하세요.
beefy_moe는 흥미로운 프로필입니다. Qwen3-30B-A3B는 ~3B의 활성 파라미터를 가진
mixture of experts 모델로, 30B 용량으로 추론하면서 대략 3B 속도로 디코딩합니다 — 그리고
디코딩이 바로 이 루프의 병목 지점입니다. 비슷한 지연 시간에서 밀집 14B 모델보다 훨씬
더 나은 액추에이터입니다. 단점은 메모리입니다: 빠른 것은 활성 전문가뿐이고 가중치는
아니므로, 30B 전체가 여전히 메모리에 상주해야 합니다.
recommend()는 실험적 프로필을 절대 반환하지 않으며, 테스트가 이를 강제합니다.
사용자 정의 모델 프로필
내장 프로필은 이 프로젝트가 개발된 머신을 기준으로 한 것이지, 여러분의 머신을 위한
것이 아닙니다. voltage → profiles에서 직접 추가하거나, 설정 파일 옆의
profiles.toml을 편집하여 추가하세요:
[my_rig]
description = "RTX 4090"
[my_rig.vision]
hf_repo = "ggml-org/Qwen2.5-VL-7B-Instruct-GGUF"
hf_file = "Qwen2.5-VL-7B-Instruct-Q4_K_M.gguf"
mmproj_file = "mmproj-Qwen2.5-VL-7B-Instruct-Q8_0.gguf"
params_b = 7.0
weights_mb = 4700
n_ctx = 4096
port = 8080
[my_rig.actuator]
hf_repo = "unsloth/Qwen3-4B-Instruct-2507-GGUF"
hf_file = "Qwen3-4B-Instruct-2507-Q4_K_M.gguf"
params_b = 4.0
weights_mb = 2500
port = 8081사용자 정의 프로필은 이름 기준으로 내장 프로필 위에 병합되므로, lean이라는 이름의
프로필을 만들면 패키지를 포크하지 않고도 내장 프로필을 재조정할 수 있습니다. Ollama
백엔드에서는 hf_repo/hf_file 대신 ollama_tag를 사용하세요.
한 슬롯은 까다롭고 한 슬롯은 그렇지 않습니다. 비전 모델은 요청 시 접지된 경계 박스를 출력할 수 있어야 합니다 — Qwen2.5-VL, Qwen3-VL, InternVL, MiniCPM-V, UI-TARS는 모두 가능합니다. 일반 캡셔너 모델은 화면을 아름답게 설명하겠지만 박스는 엉뚱한 곳에 놓을 것입니다. 액추에이터는 관대합니다: GBNF 문법 아래에서는 소수의 합법적인 연속 토큰 중에서 선택하는 것이므로, 거의 모든 유능한 1B+ instruct 모델이 작동합니다.
셸 명령 vs MCP 도구
두 가지 다른 표면이며, 이 둘을 혼동하는 것이 가장 흔한 첫 실수입니다:
호출 방식 | 형태 | |
셸 명령 | 터미널에 공백으로 입력 |
|
MCP 도구 | 밑줄로 Claude에게 요청 |
|
voltage_doctor는 디스크의 프로그램이 아니라 Claude의 네임스페이스에 있는 도구
이름입니다. 터미널에 입력하면 항상 "unknown command"라고 표시됩니다. 대신 Claude에게
실행을 요청하세요.
이것은 /dev/uinput 접근을 확인하고, 시스템 종속성을 설치하고, venv를 생성하고,
누락된 항목을 출력합니다. 그런 다음:
./scripts/fetch-models.sh lean && ./scripts/serve.sh lean.venv/bin/voltage doctor클라이언트에 연결하기
voltage connect설정된 내용, 라이브 URL, 모델이 실행 중인지, 서버가 등록되었는지 표시한 다음 — 실제 경로와 환경 변수가 이미 채워진 클라이언트별 복사-붙여넣기 단계를 제공합니다:
voltage connect --client claude-desktop
voltage connect --client cursor
voltage connect --json # just the mcpServers entry지원 대상: Claude Code, Claude Desktop, claude.ai custom connector, Cursor, Windsurf,
Zed, 그리고 그 외의 모든 것을 위한 일반 mcpServers 블록. 동일한 내용이 voltage
콘솔의 4번 화면이며, Claude Desktop 설정을 대신 작성해 줄 수도 있습니다(기존
파일을 먼저 백업하고, 유효한 JSON이 아니면 건드리지 않습니다).
모든 생성된 설정은 세션 환경을 명시적으로 포함합니다. 왜냐하면 그것이 바로 문제가
되는 부분이기 때문입니다: DBUS_SESSION_BUS_ADDRESS 없이 셸에서 등록된 서버는
연결에는 성공하지만 조용히 눈이 먼 상태가 됩니다 — 입력은 작동하지만 화면 캡처는
되지 않습니다. voltage connect는 이 경우를 감지하여 알려줍니다.
사용자 정의 커넥터로 추가하기
URL로 MCP 서버를 추가하는 클라이언트는 stdio 대신 HTTP가 필요합니다:
voltage serve --http그런 다음 http://127.0.0.1:8765/mcp를 사용자 정의 커넥터로 추가하세요.
바인딩은 루프백으로 제한되며, 이를 변경하려면 --allow-remote가 필요합니다.
이것은 단순한 보일러플레이트가 아닙니다: 이 서버는 마우스를 움직이고, 키를 누르고,
화면을 읽기 위해 존재하며, MCP에는 자체 인증이 없습니다. 루프백이 아닌 바인딩은
인증되지 않은 데스크톱 원격 제어를 공개하는 것입니다. 정말로 필요하다면 인증
리버스 프록시를 앞에 두고, 해당 포트에 도달하는 사람은 누구나 머신을 소유하게 된다는
것을 이해하세요.
MCP 클라이언트에서 실행하기
MCP 클라이언트는 정화된 환경으로 서버를 시작합니다 — PATH, HOME 및 그 외
거의 없음. 이것은 합리적인 기본값이지만 화면 캡처를 깨뜨립니다. 컴포지터에 도달하려면
DBUS_SESSION_BUS_ADDRESS와 WAYLAND_DISPLAY가 필요하기 때문입니다. 입력 주입은
이들 없이도 작동하므로(uinput은 세션 서비스가 아닌 디바이스 파일), 실패는 혼란스럽게
부분적으로 보입니다: 버스트는 실행되지만 스크린샷은 찍히지 않습니다.
이들을 명시적으로 전달하세요:
claude mcp add voltage-input \
-e WAYLAND_DISPLAY="$WAYLAND_DISPLAY" \
-e DISPLAY="$DISPLAY" \
-e DBUS_SESSION_BUS_ADDRESS="$DBUS_SESSION_BUS_ADDRESS" \
-e XDG_RUNTIME_DIR="$XDG_RUNTIME_DIR" \
-- /absolute/path/to/voltage-input-mcp/.venv/bin/voltage-input-mcpvoltage_doctor는 이 중 정확히 어떤 것이 누락되었는지 보고하므로, 캡처가 실패한다면
가장 먼저 확인할 곳입니다.
플랫폼
입력 | 캡처 | 텍스트 | |
Linux |
| portal→PipeWire, KWin DBus, grim, X11 | 스캔코드, 비-ASCII용 클립보드 폴백 |
Windows |
| GDI |
|
입력 싱크 위의 모든 것 — 버스트 스케줄링, 타이밍, 키 누름 추적, 안전 거버너, 전체
런타임 — 은 공유됩니다. 각 플랫폼은 5개의 메서드(key, button, move_abs,
move_rel, scroll)를 구현합니다. inputs/sink.py를 참조하세요.
알아둘 만한 두 가지 비대칭성:
Windows에서 타이핑이 더 정확합니다.
KEYEVENTF_UNICODE는 키보드 레이아웃이 개입되지 않은 UTF-16 코드 유닛을 전달합니다. Linux uinput은 스캔코드를 보내므로, 비-US 레이아웃에서 구두점이 — 조용히 — 잘못 입력됩니다. 이것이 클립보드 폴백이 Linux에 존재하고 Windows에는 필요 없는 이유입니다.Linux에서 캡처가 더 강력합니다. GDI
BitBlt는 일부 하드웨어 오버레이 비디오와 전체 화면 독점 게임을 볼 수 없습니다. 그런 경우 검은 화면이 캡처됩니다. 이러한 게임은 테두리 없는 창 모드로 실행하세요.
Windows에서 SendInput은 상승된 프로세스(UIPI)가 소유한 창을 구동할 수 없습니다 —
이것은 조용히 실패하므로, voltage doctor가 상승 상태를 보고합니다. DPI 인식은
import 시 선언됩니다. 이것 없이는 스케일된 디스플레이에서 모든 좌표가 잘못됩니다.
요구 사항
Linux(모든 디스플레이 서버) 또는 Windows 10/11
Python 3.11+
lean프로필에 ~5 GB 여유가 있는 GPU;voltage profiles가 여러분의 GPU에 맞는 것을 보여줍니다빠른 경로용 llama.cpp, 또는 더 느린 제로-빌드 경로용 Ollama
KDE Plasma 6 / Wayland / CUDA / Python 3.14에서 종단 간 검증되었습니다. Windows 경로는 구현되고 타입 검사되었지만 Windows 머신에서 실행된 적은 없습니다 — 테스트되지 않은 것으로 간주하고 문제가 발생하면 보고해 주세요.
오케스트레이터는 어떤 빌드를 구동 중인지 알 수 있습니다
동일한 Playbook이 한 구성에서는 타당하고 다른 구성에서는 틀릴 수 있으며, 원격 모델은 어느 쪽인지 볼 수 없습니다. 따라서 서버의 MCP 지침은 시작 시 라이브 구성에서 작성되며, Playbook 작성 방식을 바꾸는 줄만 포함합니다:
ACTIVE BUILD: Linux · llamacpp · profile lean
vision Qwen2.5-VL-3B-Instruct · actuator Qwen3-1.7B
loaded: Qwen2.5-VL-3B-Instruct-Q4_K_M.gguf / Qwen3-1.7B-Q4_K_M.gguf
expected cycle 280-700 ms
- llama.cpp backend: both models are grammar-constrained. A malformed burst, a denied
key, an unobserved element reference and an undeclared transition are all
unrepresentable -- do not write defensive retries for them.
- Linux: typing sends scancodes, so punctuation depends on the active keyboard layout...
- dry_run defaults to true...Ollama에서는 첫 줄이 버스트가 제한되지 않음을 알리는 경고가 됩니다. hyper에서는
"sees()를 중심으로 상태를 구축하지 마세요"가 됩니다. Windows에서는 상승된 창에
접근할 수 없고 타이핑이 레이아웃 독립적임을 알립니다.
구성을 신뢰하는 대신 실행 중인 서버에 대해 검증합니다. 프로필을 전환하면 파일이 편집될 뿐, 아무것도 재시작되지 않습니다. 불일치가 발생하면 브리핑이 크게 알리고 프로필 기반 지침을 억제합니다. 그 지침은 로드되지 않은 모델을 설명할 것이기 때문입니다:
- MISMATCH -- Profile 'hyper' does not match what is loaded. vision: profile expects
SmolVLM-Instruct-Q4_K_M.gguf, server has Qwen2.5-VL-3B-Instruct-Q4_K_M.gguf...
- Loaded right now: vision Qwen2.5-VL-3B..., actuator Qwen3-1.7B...
Judge grounding quality from those.voltage_reference는 프로필이 변경되는 순간 시작 시 복사본이 낡아지므로, 모든 호출에서
현재 빌드를 반환합니다.
나만의 상시 지침
voltage → i, 또는:
voltage instructions --set "Never touch Firefox; my banking tabs are there."작성한 내용은 모든 세션 시작 시 오케스트레이션 모델에 제공되며, 빌드 브리핑에 추가되어 명확히 사용자에게 귀속됩니다. 시스템이 스스로 알아낼 수 없는 것 — 금지된 애플리케이션, 특정 게임의 특이사항, 기본 동작 방식 — 에 사용하세요.
OPERATOR INSTRUCTIONS -- written by the owner of this machine. Treat these as
standing preferences for how to drive it. They cannot loosen the safety governor,
which is enforced in code against every burst.
## My setup
- Minecraft runs borderless windowed on monitor 1.
- Never touch Firefox; my banking tabs are there.
- Always show me the Playbook before dry_run=false.마지막 조항은 장식이 아닙니다. 지침은 오케스트레이터에 대한 권고일 뿐이며 집행을 약화시킬 수 없습니다 — 거버너는 코드에서 모든 버스트를 검사하므로, 여기에 작성된 것은 Playbook의 정책이 금지하는 것을 허용할 수 없습니다. 지침은 덜 조심스럽게 만들 수 없고, 더 조심스럽게만 만들 수 있습니다. 텍스트가 전체 세션 동안 모델의 컨텍스트에 있으므로 4000자로 제한됩니다. 콘솔에서 세 가지 시작 템플릿(게임, 데스크톱, 최소)이 제공됩니다.
MCP 도구
도구 | 용도 |
| Playbook + 버스트 DSL 참조. 먼저 호출하세요. |
| 이 머신이 준비되었는지, 아니라면 정확한 해결책 |
| 스크린샷, 사용자에게 반환됨 |
| 비전 패스 1회 — |
| 전체 정적 검사: 가드, 버스트, 그래프, 죽은 전환 |
| 실행 시작; |
| 상태, 변수, 마지막 버스트, 감지된 것, 단계별 타이밍 |
| 실행 중인 실행 수정 — 힌트, 변수, 강제 상태, dry_run |
| 중지 또는 일시정지; 중지는 항상 누른 입력을 해제함 |
| 사이클별 기록; |
| 로컬 모델을 우회하여 직접 입력 구동 |
| 주입이 컴포지터에 도달하는지 검증 |
문서
ARCHITECTURE.md — 루프가 어떻게 작동하는지, 각 선택이 이루어진 이유, 시간이 어디에 소요되는지
PLAYBOOK.md — 작성 가이드
상태
디스크에 가중치 없이 가능한 한 멀리 빌드되고 검증되었습니다. 149개의 테스트가 버스트
DSL, 가드 샌드박스, 안전 거버너, Playbook 컴파일, GBNF 생성, uinput 와이어 인코딩,
그리고 실행 루프 자체를 다룹니다(스텁 모델로 구동 — 정적 화면에서 on_change 인식이
정말로 비전 모델을 건너뛰는지 확인하는 검사 포함).
MCP 서버는 실제 클라이언트에 의해 stdio를 통해 종단 간 구동되었습니다: 13개 도구,
올바른 스키마, execute_burst가 유효한 버스트를 수락하고 일치하는 두 규칙으로
sudo rm -rf /를 거부했습니다.
실행되지 않은 것은 라이브 모델입니다: llama.cpp 빌드와 가중치 다운로드가
필요하며, scripts/가 이를 설정합니다. 빌드 중 의도적으로 트리거하지 않은 두 가지도
있습니다 — 포털 권한 대화 상자와 실제 입력 주입 — 둘 다 데스크톱에 영향을 미치기
때문입니다.
여기서부터의 작업 순서:
./scripts/setup.sh # reports what needs sudo, doesn't run it
./scripts/build-llama.sh # ~15 min with CUDA
./scripts/fetch-models.sh lean
./scripts/serve.sh lean
.venv/bin/voltage doctor # should now say READY그런 다음 MCP 클라이언트에서: voltage_calibrate(커서가 실제로 움직이는지 확인),
voltage_observe(비전 모델이 라벨을 찾는지 확인), 그다음 dry_run Playbook을 실행하고
dry_run=false로 설정하기 전에 voltage_journal을 읽으세요.
저작자
Claude Opus 5(Anthropic)가 단일 세션에서 처음부터 끝까지 작성 — 아키텍처, 구현, 테스트, 문서화를 모두 포함합니다. 인간이 아이디어를 지정하고, 제약 조건을 설정했으며 (KDE Wayland, 6 GB VRAM, "컴퓨터 사용보다 빠르게"), 결과물을 검토했지만 코드를 작성하지는 않았습니다.
이 저장소에 반영된 플랫폼 조사 결과는 가정이 아니라 빌드 중 머신을 직접 탐색하여 얻은
것입니다 — KWin이 허용 목록에 없는 실행 파일에 ScreenShot2를 거부한다는 점, grim이
KWin에서 작동할 수 없다는 점, MCP 클라이언트가 세션 버스를 제거한다는 점 등이 그것입니다.
각 항목은 결정을 강제한 코드 지점에 문서화되어 있습니다.
LICENSE는 개인을 저작권 보유자로 명시하지 않으며, 그 이유가 해당 파일에 서술되어 있습니다.
라이선스
MIT. LICENSE를 참조하세요.
Available Tools
16 toolsvoltage_calibrateADestructive
Verify that input injection actually reaches the compositor.
Creates the virtual devices, moves the pointer to three known points, and captures after each to confirm the cursor moved. Reports whether absolute positioning works or whether the relative fallback is needed -- which cannot be known without trying, since it depends on how libinput classified the virtual device.
Run this once per machine before trusting a real (non-dry-run) Playbook.
| Name | Required | Description | Default |
|---|---|---|---|
| dry_run | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotations already declare destructiveHint=true and openWorldHint=true. The description adds valuable detail: it creates virtual devices, moves the pointer, and captures output—concrete side effects beyond the annotation. It also explains why these behaviors are unpredictable ('depends on how libinput classified the virtual device'), which aligns with openWorldHint. This goes beyond what annotations alone convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise and well-structured. It opens with the core purpose, immediately explains what the tool does, then provides the rationale and usage timing. Every sentence earns its place—no fluff or repetition. It is front-loaded with the most critical information and stays focused.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers the tool's purpose, mechanism, side effects, and when to run it. It also acknowledges an output summary (absolute vs relative fallback). However, it omits any explanation of the dry_run parameter, which is a key input that affects behavior. Given the presence of an output schema, return format doesn't need detailing, but the parameter gap leaves the description incomplete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The only parameter, dry_run, is entirely undocumented in the description. The schema gives its name, type, default, and requirement status but provides no semantic meaning. The description mentions 'non-dry-run' indirectly but never explains what dry_run does, when to set it to true, or what the difference is. With 0% schema description coverage, this is a significant gap—the description fails to compensate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's specific purpose: to verify that input injection reaches the compositor. It describes the concrete steps (creating virtual devices, moving pointer to three points, capturing) and the distinctive outcome (determining whether absolute positioning or relative fallback is needed). This distinguishes it from the many sibling tools, even without naming them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives a clear usage context: 'Run this once per machine before trusting a real (non-dry-run) Playbook.' It explains that this is a one-time calibration step and that the result cannot be known without trying, which implies this is the tool to use for that purpose. It doesn't explicitly mention alternatives or when not to use it, but the guidance is strong.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
voltage_captureARead-only
Take a screenshot and return it to you directly.
Use this to see the screen yourself -- before writing a Playbook, to pick coordinates for probes and click regions, or to work out why a run went wrong. This does not involve the local vision model.
region is [x, y, width, height] in desktop pixels; omit for the whole desktop.
| Name | Required | Description | Default |
|---|---|---|---|
| region | No | ||
| max_width | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=false, so the description's job is to add behavioral context. It does so by specifying 'return it to you directly' and explicitly noting it does not use the local vision model—information not in annotations. No contradictions exist.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is tightly written: a clear one-sentence purpose, a brief usage block, and a one-line parameter definition. No fluff or redundancy. The key purpose is front-loaded, and every sentence contributes value.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple screenshot tool, the description covers usage, region format, and direct return. However, the `max_width` parameter remains unexplained; an agent might not know how to set it or its effect on output. Given that the tool has no output schema, a bit more detail on max_width would make it fully complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description explains the `region` parameter fully: it is [x, y, width, height] in desktop pixels and can be omitted for the whole desktop. However, `max_width` is not described at all; the schema only shows it is an integer with default 1280. Since schema description coverage is 0%, the description should compensate for both parameters, but it only covers one.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening sentence 'Take a screenshot and return it to you directly' uses a specific verb and resource, and clearly states the result. It also distinguishes itself from the vision-model-based sibling by saying 'This does not involve the local vision model,' which makes its purpose unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It provides concrete use cases: 'before writing a Playbook, to pick coordinates for probes and click regions, or to work out why a run went wrong.' This tells the agent exactly when to invoke it. It does not explicitly mention alternatives or when not to use it, but the context is clear enough to guide selection.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
voltage_diagnoseARead-only
Explain why a run behaved as it did, and what to change.
Call this instead of reading the journal by hand. It computes what the journal
implies but does not state -- watch labels the vision model never once reported,
guards that never evaluated true, whether bursts actually moved the screen, whether
the actuator is chaining or emitting one action at a time -- and returns each with
the specific edit that fixes it, ordered blocker-first.
The distinction it exists for: a burst that never ran and a burst that ran and did nothing look identical in a summary and have unrelated causes. The first is policy or grammar; the second is window focus, pointer mode, or an application that ignores synthetic input.
Apply the highest-severity finding, re-run, diagnose again. Changing several things at once makes the next diagnosis uninterpretable.
| Name | Required | Description | Default |
|---|---|---|---|
| run_id | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=false, so the description doesn't need to restate safety. It adds valuable behavioral detail: it computes implicit journal information, returns specific edits ordered blocker-first, and distinguishes between a burst that never ran vs. ran but did nothing. This goes well beyond the annotation, providing non-obvious nuances.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with a crisp summary, then explains the key distinction and ends with an actionable workflow. Every sentence earns its place; there is no fluff or redundancy. Structure is clear and scannable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given that an output schema exists (the description doesn't need to detail return structure) and annotations cover safety, the description covers the essential context: the diagnostic purpose, the key distinction between two root causes, and the iterative workflow. The only minor gap is the run_id parameter semantics, which slightly detracts from completeness for an otherwise simple tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description carries the full burden for the single parameter run_id. It never mentions run_id, its format, how to obtain it, or whether it's required (though the schema marks it optional). The name 'run_id' is self-explanatory by convention, but the description provides no explicit guidance, and with only one parameter to cover, this is a noticeable gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Explain why a run behaved as it did, and what to change.' It then contrasts itself with reading the journal, making its purpose distinct from voltage_journal. No ambiguity about what the tool does.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly instructs to call this instead of reading the journal by hand, giving clear when-to-use context. It also provides a workflow (apply highest-severity finding, re-run, diagnose again). However, it doesn't name alternative siblings like voltage_doctor or voltage_observe, or describe conditions where those might be more appropriate, so it stops short of complete guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
voltage_doctorARead-only
Check that everything needed for a run is present and working.
Reports the session type, input-device permissions, which capture backends work, detected screen geometry, GPU memory versus the selected model profile, and whether both model backends respond. When something is missing it returns the exact command to fix it. Call this before the first run on a machine.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotations already declare readOnlyHint=true and openWorldHint=false, establishing a safe, closed-world read operation. The description adds valuable context beyond safety: it lists the specific components checked (session type, input-device permissions, capture backends, screen geometry, GPU memory vs model profile, both model backends) and states that it returns fix commands. This informs the agent about the scope of the check and the nature of the response, which is more than annotations provide. No contradiction found.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three sentences, tightly packed with information. The first sentence gives the core purpose, the second enumerates the checks and the fix-command behavior, and the third provides usage timing. Every sentence earns its place, and the most important information (purpose and when to use) is front-loaded. There is no fluff or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a no-argument diagnostic tool, the description is comprehensive. It clearly states what is checked, the response characteristic (returns fix commands), and when to call it. An output schema exists (though not shown), so the description doesn't need to detail return formats. Given the complexity of the checks and the existence of a schema, nothing essential is missing for the agent to use the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters and the schema shows none. The description implicitly confirms this by stating 'Call this before the first run on a machine' with no mention of inputs. Since there are no parameters to explain, the description effectively communicates that it requires no configuration. This is a perfect fit for the no-parameter case, and the baseline of 4 is exceeded because the description makes the absence of parameters obvious.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: 'Check that everything needed for a run is present and working.' It specifies a concrete action (check) and a distinct resource (run prerequisites). It differentiates from siblings like voltage_status and voltage_diagnose by enumerating the exact checklist items (session type, permissions, capture backends, geometry, GPU memory, model backends). This makes it unambiguous which tool to select for pre-flight validation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit timing guidance: 'Call this before the first run on a machine.' While it doesn't mention alternatives or when not to use it, the instruction is clear and actionable. It implies this is a single-use setup check, not a repeated monitoring tool. The guidance is sufficient for the agent to decide when to invoke it, though lacking explicit exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
voltage_execute_burstADestructive
Execute one input burst yourself, bypassing the local models entirely.
For moments that need your judgement rather than the actuator's: opening the right application, clicking a specific confirmed target, typing something exact. Also the fastest way to sanity-check that input injection works at all.
Syntax: m:640,360;c:l;w:120;t:"hello";k:enter. Call voltage_reference for the full
list. The safety policy still applies. Defaults to dry_run, so pass dry_run=false to
actually inject.
| Name | Required | Description | Default |
|---|---|---|---|
| burst | Yes | ||
| label | No | manual | |
| dry_run | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint: true, readOnlyHint: false, and openWorldHint: true. The description adds critical behavioral context: it defaults to dry_run ('pass dry_run=false to actually inject') and notes the safety policy. It also explains that this is a manual override path. These details go beyond the annotations and inform the agent about side effects and prerequisites.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-structured and efficient: it leads with the core action, then gives usage scenarios, then provides a syntax example and necessary caveats. Every sentence earns its place, and the dry_run warning is front-loaded within the critical context. No fluff or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (custom syntax, safety policy, dry_run default) and that an output schema exists, the description covers the essential aspects: purpose, when to use, how to construct the burst (via example and reference), and the dry_run behavior. The only gap is a full in-place explanation of the syntax and label, but the reference to voltage_reference and the presence of an output schema mitigate this. Overall, it is nearly complete for an agent to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It provides a concrete syntax example (`m:640,360;c:l;w:120;t:"hello";k:enter`) and explains the dry_run parameter clearly. However, burst syntax is not fully documented (only a pointer to voltage_reference) and the label parameter is not explained beyond its default. This is partial compensation—helpful but not exhaustive.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Execute one input burst yourself') and the resource (burst), and immediately differentiates from siblings by emphasizing 'bypassing the local models entirely' and 'moments that need your judgement rather than the actuator's'. It also names the exact use case (opening applications, clicking confirmed targets, typing exact text) and points to voltage_reference for full syntax, making the purpose unmistakable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly states when to use this tool ('For moments that need your judgement rather than the actuator's', 'the fastest way to sanity-check that input injection works at all'), implies alternatives by referencing voltage_reference for syntax, and reminds that 'the safety policy still applies'. This gives an agent clear decision-making guidance without ambiguity.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
voltage_journalARead-only
Read a run's cycle-by-cycle record: what was seen, decided, refused, executed.
only_refused=true filters to cycles the governor blocked, which is the fastest way
to see where a Playbook's policy and the actuator's intentions disagree.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| run_id | No | ||
| only_refused | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already mark readOnlyHint=true, and the description aligns with that by saying 'Read'. It adds value by explaining the behavioral semantics of the journal contents and the meaning of 'only_refused', which goes beyond the raw annotation. No contradiction exists.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two focused sentences with the core purpose front-loaded and the filter tip as a concise, well-formatted follow-up. No filler or repetition, and the code-styled parameter reference is efficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return format is covered. However, the description fails to explain the run_id parameter, which is central to selecting a run, and gives no mention of limit. The tool is simple with all optional params, but the missing parameter descriptions leave a gap in usability.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It explains only_refused in detail, but completely omits run_id and limit. run_id is critical for identifying which run to read, and limit is a common but still undocumented control. The description is inadequate for a zero-coverage schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Read') and resource ('a run's cycle-by-cycle record'), and lists the exact contents: what was seen, decided, refused, executed. This clearly distinguishes it from siblings like voltage_observe or voltage_diagnose by framing it as a chronological journal rather than a live observation or diagnostic.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives explicit context for using the 'only_refused' filter and explains the fastest way to see policy/actuator disagreement. While it doesn't mention sibling tools for comparison, the usage hint is concrete and actionable, and the description clearly implies this tool is for inspecting historical decisions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
voltage_learnADestructive
Record something worth carrying to the next run against this target.
Write these as concrete, reusable facts, not narration:
good "the health bar is at x=120..300, y=1010; region_mean on red channel works" good "vision reports 'hotbar' reliably but never 'crosshair' -- do not watch it" good "block placement needs w:100 after the right click or it does not register" bad "the run failed" bad "tried again and it worked better"
kind groups them: label (what the vision model does and does not recognise),
timing (waits that a specific application needs), policy (what the governor blocked
and whether that was right), burst (a sequence that works), observation (anything
else).
| Name | Required | Description | Default |
|---|---|---|---|
| kind | No | observation | |
| note | Yes | ||
| target | Yes | ||
| playbook | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already indicate a mutating, potentially destructive action (readOnlyHint=false, destructiveHint=true); the description does not contradict these and adds that notes are stored against a target. It does not describe side effects or permissions, but given annotation coverage it provides acceptable additional context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is efficiently structured: it opens with the core purpose, gives clear good/bad examples, and ends with a concise classification of kind values. Every sentence adds value, and the format is well-balanced for the tool's complexity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool that records notes, the description covers the purpose, content quality, and kind taxonomy, which is sufficient for basic use. Gaps remain around `playbook` and exact behavior (e.g., confirmation, persistence), but the presence of an output schema and annotations mitigates these. Overall it is fairly complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description compensates by explaining the meaning of `kind` (label, timing, policy, burst, observation) and prescribing the format for `note` via good/bad examples. It leaves `target` and `playbook` undefined, but `target` is self-evident and `playbook` remains ambiguous, so coverage is partial but effective.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool records reusable facts against a target, with concrete good/bad examples that make the purpose unmistakable. It does not explicitly differentiate from sibling tools like voltage_lessons, but the 'carrying to the next run' phrasing is specific enough to convey its unique role.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides strong guidance on what to record (concrete facts, not narration) and explains the kind grouping, but it never mentions alternative tools or conditions under which to avoid this tool. Usage context is implied rather than explicit, and no exclusions or comparisons are given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
voltage_lessonsARead-only
Recall what previous runs learned about driving something.
Call this before writing a Playbook for a target you have driven before. Lessons persist across sessions and are keyed by target ("minecraft", "roblox", "dolphin"), so a new Playbook can start from what the last one discovered -- which labels the vision model actually recognises, where the HUD probes are, what timing the game needs -- rather than rediscovering it.
Omit target to see everything recorded so far.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| target | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=true and the description aligns with that (no mutation implied). The description adds valuable behavioral context: lessons persist across sessions, are keyed by target, and include specific types of information (labels, HUD probes, timing). This goes beyond the annotation by describing persistence and content, which is useful for setting expectations about what the tool returns.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise and front-loaded with the purpose. It uses bold for emphasis ('before writing a Playbook') and keeps each sentence purposeful. There is no filler or redundant explanation. The structure guides the reader from what the tool does, to when to use it, to how to filter results.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has an output schema (as indicated by the context), so return values are documented elsewhere. The description provides sufficient context for an agent to decide when to call it: it explains the purpose, when it is appropriate (before writing a Playbook for a previously driven target), and how to control scope with the target parameter. No critical information is missing, given the read-only annotation and output schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 0% description coverage, so the description must compensate. It clearly explains the `target` parameter (keyed by target, omit to see everything) and gives examples of valid values. However, it does not mention the `limit` parameter at all, leaving its semantics to inference from the default value of 30. This is a partial compensation but not complete for both parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Recall what previous runs learned about driving something.' It then gives concrete examples of lesson content (labels, HUD probes, timing), which makes the tool's purpose unambiguous and distinct from any other sibling. The behavior is clearly scoped to recalling learned lessons, not a general-purpose query.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly instructs when to use the tool: 'Call this **before writing a Playbook** for a target you have driven before.' It also explains the benefit (start from previous discoveries rather than rediscovering) and provides parameter guidance: 'Omit `target` to see everything recorded so far.' This gives an agent clear, actionable context for choosing this tool over alternatives like voltage_learn.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
voltage_observeARead-only
Run one vision pass and return grounded elements in screen coordinates.
watch is the closed vocabulary the vision model may use -- it can only report
labels from this list, so name the things your Playbook's guards will test for.
Use this to check that the vision model can actually find what a state depends on
before committing to it in a Playbook. If an element does not come back here, a
sees(...) guard on it will never fire.
| Name | Required | Description | Default |
|---|---|---|---|
| watch | Yes | ||
| region | No | ||
| read_text | No | ||
| max_elements | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already convey read-only and closed-world hints. The description adds valuable behavioral context: it clarifies that 'watch' is a closed vocabulary, that the tool runs a single pass, and that missing elements imply guards never fire. This goes beyond the annotations and provides actionable insight into the tool's behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise, with two short paragraphs that are front-loaded with the core purpose. Every sentence adds distinct value—stating the action, vocabulary constraint, and practical implication. There is no fluff or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description captures the tool's primary purpose and a key behavioral consequence, and an output schema exists so return values are already documented. However, it does not explain non-required parameters (region, read_text, max_elements), which are likely needed for correct invocation. This gap reduces completeness, though the core use case is well covered.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It explains 'watch' as the closed vocabulary, which is essential, but it omits any explanation for 'region', 'read_text', and 'max_elements'. With only one parameter addressed, the description fails to adequately clarify the remaining parameters, leaving the agent with insufficient guidance.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states a specific verb and resource: 'Run one vision pass and return grounded elements in screen coordinates.' It also explains a distinct use case—checking if the vision model can find elements before committing to a Playbook. While it doesn't explicitly contrast with sibling tools, the purpose is specific and unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides explicit context for when to use the tool: 'Use this to check that the vision model can actually find what a state depends on before committing to it in a Playbook.' This is a clear directive without naming alternatives, but it effectively guides the agent on ideal usage scenarios.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
voltage_pauseBDestructive
Pause or resume a run. Held input is not released, so a paused run can continue.
| Name | Required | Description | Default |
|---|---|---|---|
| resume | No | ||
| run_id | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already flag destructiveHint=true, so the mutation nature is disclosed. The description adds the specific behavior that held input is retained, which goes beyond the annotations and gives the agent useful context about the pause/resume semantics. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with zero filler. The core action is front-loaded ('Pause or resume a run') and the clarifying detail about held input follows immediately. Every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the existence of an output schema and simple optional parameters, the description is far from complete. It lacks usage guidance, parameter semantics, and any mention of prerequisites or side effects beyond the held-input note. The agent would need to guess how to set 'resume' or when to pass 'run_id'.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% — neither 'resume' nor 'run_id' is explained in the schema. The description does not mention any parameters at all, so the agent has no idea that 'resume' likely indicates whether to resume or pause, or how 'run_id' selects the run. With two parameters and zero coverage, the description must compensate but fails completely.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a clear action (pause or resume), a specific resource (a run), and adds a key nuance (held input is not released). It distinguishes implicitly from voltage_stop but does not name sibling alternatives, so it falls short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No explicit guidance on when to use this tool versus alternatives like voltage_stop or voltage_run. The note about held input hints at a use case but does not state conditions or exclusions, leaving the agent to infer when pause is appropriate.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
voltage_referenceARead-only
Return everything needed to author and iterate on a run.
Call this before your first Playbook. Sections:
loop the learning loop -- how to go from a failed run to a working one, and what each failure mode actually means. Read this second. bursts the burst cookbook: how to chain inputs well, timing rules, ready-made patterns for desktop and for games, and the antipatterns that waste cycles. Read this if bursts are coming out one action at a time. burst the raw burst syntax playbook the state-machine JSON schema guards expression functions for transitions and reflexes example a complete working Playbook
| Name | Required | Description | Default |
|---|---|---|---|
| section | No | all |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, so the description need not restate safety. It adds value by explaining the content structure and the purpose of each section, which helps the agent understand what the tool actually returns. However, it does not disclose any potential caveats (e.g., response size, format specifics), though those may be covered by the output schema. The added context justifies a score slightly above baseline.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is efficiently organized: a one-line purpose, then a bulleted list of sections with clear labels and explanations. It front-loads the main instruction and uses formatting to allow fast scanning. No sentence is redundant; each adds useful detail about content or usage.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a reference tool, the description covers all essential information: what it returns, when to call it, what each section contains, and even contextual reading order. The read-only behavior is covered by annotations, and the output format is presumably defined by the output schema (present signal). Nothing necessary for an agent to select and invoke this tool is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must explain the 'section' parameter. It does so comprehensively by listing each enum value and its meaning, and even offers reading-order guidance (e.g., 'Read this second', 'Read this if...'). This fully compensates for the schema gap, making the parameter self-documenting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Return') and a resource ('everything needed to author and iterate on a run'), then enumerates the sections returned. It clearly distinguishes itself from sibling tools (e.g., voltage_execute_burst, voltage_validate_playbook) by being a reference/documentation tool, not an execution or validation tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly says 'Call this before your first Playbook,' giving a clear when-to-use directive. It also provides conditional reading order (e.g., 'Read this if bursts are coming out one action at a time') and labels like 'the learning loop,' which help an agent decide which section to request. This is strong, situation-specific guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
voltage_runADestructive
Start a Playbook. Returns immediately with a run_id; poll voltage_status.
dry_run overrides the Playbook's policy. Leave it unset for the Playbook's own
setting, which defaults to true. A dry run does everything except inject input, so
it is the correct way to check that your states, guards and transitions behave before
letting it touch the machine.
target_period_s is the loop period. 0.5 is a good default; lower it for games,
raise it for slow UI.
Stop a run with voltage_stop, adjust it live with voltage_steer. The run also stops on its own budget, on any physical keyboard or mouse input from the user, and on the panic file.
| Name | Required | Description | Default |
|---|---|---|---|
| dry_run | No | ||
| playbook | Yes | ||
| keep_frames | No | ||
| target_period_s | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already mark destructiveHint and openWorldHint, and the description complements these by explaining concrete behaviors: immediate return with run_id, polling requirement, dry_run overriding policy, and the specific conditions that terminate a run. It adds value beyond annotations without contradicting them.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is efficiently structured: the core action and return contract are front-loaded, followed by parameter guidance and termination behavior. Every sentence adds functional value, and no redundant or filler content is present.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers the essential lifecycle: starting, monitoring, adjusting, and stopping. It explains dry-run semantics and stopping triggers. However, it does not describe the structure of the `playbook` object or the meaning of `keep_frames`, which may be important for correct invocation. The presence of an output schema and related tools (voltage_validate_playbook) partially mitigates this.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must explain parameters. It does explain dry_run (including override semantics and default behavior) and target_period_s (with recommended values), but it does not explain `playbook` (the required parameter) or `keep_frames`. Since playbook is central and the schema offers no description, this leaves a gap for an agent constructing a valid call.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with 'Start a Playbook,' a specific verb-resource pairing that clearly states the tool's core function. It immediately distinguishes itself from siblings by mentioning polling with voltage_status, stopping with voltage_stop, and live adjustment with voltage_steer, so the agent can tell it apart without opening other schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides explicit context for when to use dry_run ('the correct way to check that your states, guards and transitions behave before letting it touch the machine'), recommends values for target_period_s, and explains how to stop or adjust a run using sibling tools. It also details automatic stopping conditions (budget, keyboard/mouse input, panic file), giving clear operational guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
voltage_statusARead-only
Poll a run: current state, variables, last burst, what the vision model sees.
Includes recent cycles, governor refusals, and per-stage timings so you can tell whether a slow loop is capture, vision, decision, or execution.
| Name | Required | Description | Default |
|---|---|---|---|
| run_id | No | ||
| journal_tail | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=false, so the read-only nature is covered. The description adds useful context beyond annotations: the specific data included (recent cycles, governor refusals, per-stage timings) and its diagnostic intent. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with no waste. The action is front-loaded ('Poll a run'), followed by a list of what it returns and the diagnostic purpose. Every phrase earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a monitoring tool with an output schema present, the description conveys enough about the returned data to be useful. However, the lack of parameter documentation is a notable gap that makes it slightly incomplete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description does not explain either parameter (run_id, journal_tail) at all. While run_id is somewhat inferable from its name, journal_tail is completely unexplained. The description fails to compensate for the schema's lack of documentation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb ('Poll') and resource ('a run'), then enumerates the returned data (state, variables, last burst, vision model view, cycles, refusals, timings). This clearly differentiates it from sibling tools like voltage_capture or voltage_execute_burst, which imply different actions.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It implies usage during a run to monitor state and diagnose slow loops ('so you can tell whether a slow loop is capture, vision, decision, or execution'). However, it doesn't explicitly state when not to use it or point to alternatives such as voltage_doctor or voltage_diagnose, leaving some ambiguity.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
voltage_steerADestructive
Correct a live run without restarting it.
hint is injected into the actuator's prompt as a supervisor note and persists until
changed -- use it when the actuator is doing something legal but wrong.
force_state jumps the machine on the next cycle. variables updates run variables.
dry_run can be flipped either way mid-run.
| Name | Required | Description | Default |
|---|---|---|---|
| hint | No | ||
| run_id | No | ||
| dry_run | No | ||
| variables | No | ||
| force_state | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already flag destructiveHint=true, so the description adds some context: hint persists, force_state jumps the machine, variables updates, dry_run can flip. However, it does not disclose potential side effects, irreversibility, or prerequisites despite the destructive nature. It does not contradict the annotations, but the coverage is not thorough.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise and well-structured: a one-sentence overview followed by per-parameter explanations. It is front-loaded with the main purpose, uses backticks for param names to aid scanning, and has no filler or redundant statements.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with 5 parameters, zero schema descriptions, a destructive annotation, and an output schema, the description covers the core actions but misses run_id semantics, any warning about destructive consequences, and what the output schema contains. It is usable but not fully complete for safe and correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description carries the full burden for parameter meaning. It explains hint, force_state, variables, and dry_run, but omits run_id entirely, leaving its role merely implied by the phrase 'a live run.' This is a partial but incomplete compensation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb-resource pair: 'Correct a live run without restarting it,' which clearly distinguishes this tool from siblings like voltage_stop, voltage_pause, or voltage_run. It also enumerates the effects of each parameter, leaving no ambiguity about what the tool does.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides a concrete usage scenario for `hint` ('when the actuator is doing something legal but wrong') and explains the function of each parameter (e.g., force_state jumps the machine, dry_run flips). It implies this tool is for mid-run corrections vs. restarting, but does not explicitly name alternatives or exclusion conditions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
voltage_stopADestructive
Stop a run and release every held key and button.
Safe to call at any time, including while a burst is mid-flight -- the burst is interrupted and anything held is released.
| Name | Required | Description | Default |
|---|---|---|---|
| reason | No | stopped by orchestrator | |
| run_id | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, and the description adds concrete behavior: releases every held key/button and interrupts bursts. This goes beyond the annotation's generic destroy flag without contradicting it, giving the agent a more precise model of consequences.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two concise sentences convey the purpose, safety, and edge-case behavior with zero filler. Information is front-loaded and every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple stop tool, the description covers the main behavior and safety profile. However, the lack of any parameter explanation means an agent might guess wrong about 'run_id' or 'reason' (e.g., whether run_id is required to target a specific run). Optional parameters with defaults mitigate, but the gap prevents full completeness.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the description does not explain either 'reason' or 'run_id.' The agent has no guidance on what these parameters control or when to provide them, though they are optional. With no parameter documentation anywhere, this is a notable gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Stop') and resource ('a run') while adding unique scope: 'release every held key and button.' This clearly distinguishes it from siblings like voltage_pause and voltage_run, and the mention of interrupting mid-flight bursts further clarifies its specific role.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context: 'Safe to call at any time' and explicitly covers the edge case of a mid-flight burst. However, it does not explicitly contrast with alternatives like voltage_pause or voltage_steer, leaving some ambiguity about when to choose this over a pause or a graceful stop.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
voltage_validate_playbookARead-only
Fully check a Playbook without running it.
Validates the schema, compiles every guard expression, parses every burst, checks that transition targets and probe references exist, and reports unreachable states and dead transitions. Errors come back as a complete list, not one at a time.
Always call this before voltage_run. Warnings are worth reading: "tests for X but X
is not in watch" means a transition that can never fire.
| Name | Required | Description | Default |
|---|---|---|---|
| playbook | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only mark readOnlyHint:true. The description adds substantial behavioral detail: it returns a complete list of errors rather than one at a time, reports unreachable states and dead transitions, and explains how to interpret warnings. This fully complements the annotation and does not contradict it.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is tightly written, starting with the primary purpose, then detailing checks, then error behavior, then usage guidance and a warning interpretation. Every sentence serves a purpose—no filler. It front-loads the action and clearly organizes information in short block format.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter validation tool with an output schema declared (though not shown explicitly), the description covers what it does, how it behaves, when to call it, and how to interpret results. With annotations covering read-only safety and the output schema expected to define return values, nothing essential is missing for an agent to decide and invoke correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema only defines a generic 'playbook' object with no description (0% coverage). The description compensates by making clear that the parameter is the Playbook being validated, and it describes what validation entails (schema, guards, bursts, references). This gives the agent enough context to pass the correct object, even without knowing its internal structure.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific, unambiguous statement: 'Fully check a Playbook without running it.' It enumerates the exact validations performed (schema, guards, bursts, transition targets, probe references) and reports unreachable states/dead transitions, distinguishing this validation tool from siblings like voltage_run and voltage_execute_burst. The verb 'validate' matches the tool name and clears its role.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly guides usage with 'Always call this before voltage_run,' which states when to use this tool relative to its primary sibling. It also adds a practical hint about interpreting warnings (e.g., 'tests for X but X is not in watch'). It does not list explicit exclusions, but the directive is clear and directly actionable.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
16 tool updates
v0.1.0- First observed
voltage_calibrate - First observed
voltage_capture - First observed
voltage_diagnose - First observed
voltage_doctor - First observed
voltage_execute_burst - First observed
voltage_journal - First observed
voltage_learn - First observed
voltage_lessons - First observed
voltage_observe - First observed
voltage_pause - First observed
voltage_reference - First observed
voltage_run - First observed
voltage_status - First observed
voltage_steer - First observed
voltage_stop - First observed
voltage_validate_playbook
TDQS
Scored across 16 tools
Each tool has a clearly distinct purpose: pre-flight checks, documentation, perception, input execution, validation, running, monitoring, control, and learning. Even similar tools like voltage_journal (raw data) and voltage_diagnose (analyzed explanation) are cleanly separated by their roles.
All tools follow a consistent voltage_ prefix with a verb or verb_noun pattern (capture, execute_burst, validate_playbook, etc.). No mixed conventions or ambiguous verbs; naming is predictable and intuitive.
16 tools is well-scoped for a comprehensive automation server covering setup, execution, monitoring, debugging, and learning. Each tool earns its place; the count supports the full workflow without bloat.
The tool surface covers the entire lifecycle: environment checks (doctor, calibrate), documentation (reference), perception (capture, observe), manual action (execute_burst), validation and execution (validate_playbook, run), live control (steer, stop, pause), monitoring (status, journal, diagnose), and cross-session learning (lessons, learn). No obvious gaps for the stated purpose.
Maintenance
Related MCP Connectors
Operate Linux, macOS and Windows from your LLM. Every action runs through an auditable allowlist.
Human-in-the-loop approval for agent actions, with verifiable action-bound receipts.
Control real Android and iOS devices with LLM agents — tap, swipe, type, automate flows.
Lets AI agents use a real human as a tool: visual checks, taste, phone calls, unblocking, approvals
Related MCP Servers
- AlicenseNot gradedqualityAmaintenanceA local autonomous AI agent that watches your screen, understands the visual layout, and executes native OS commands (clicking, typing) without cloud APIs.16MIT
- FlicenseNot gradedqualityBmaintenanceEnables AI agents to declaratively control web pages using real mouse and keyboard events via Chrome DevTools Protocol, without executing page JavaScript.10 npm1-
- AlicenseAqualityBmaintenanceEnables low-cost agent models to control Windows applications through a compact, state-safe proxy over Open Computer Use, reducing model-visible context by up to 99.8% with support for record/replay and reusable UI component memory.5MIT
- AlicenseNot gradedqualityAmaintenanceEnables AI agents to automate real desktop applications across Windows, Linux, and macOS using incremental screen perception, accessibility trees, OCR, and window management, dramatically reducing token usage compared to screenshot-per-step approaches.39 PyPI3MIT