hyprcu
# hyprcu
**Hyprland computer use.** Dialog-free desktop control for AI agents: an MCP
server and a CLI over the same primitives. A fork of
[hypruse](https://github.com/IlyasKhallouki/hypruse) (MIT, IlyasKhallouki) — see [Lineage](#lineage)
with the trust, journal, safety-beacon and CLI layers removed, keeping the
Wayland/Hyprland primitives intact. App knowledge and tooling live in `docs/` and `tools/`.
## Requires a TypeSafe API key (sends names to TypeSafe)
By default hyprcu resolves descriptions (`window="the file browser"`,
`click_ui "submit the form"`) with **TypeSafe's Jev** model, a cloud API.
Each such call **sends the candidate names to TypeSafe**: the titles and
classes of your open windows, or the names and values of the window's
accessible controls, plus your description. Addresses and exact names/
substrings resolve locally and send nothing.
- Get a key from [TypeSafe](https://typesafe.ai) and export `TYPESAFE_API_KEY`
in the environment hyprcu runs in. It is read from the environment only and
never written to disk.
- Without a key, descriptions fail with an explicit message; everything else
works.
- To keep everything on the machine, set `HYPRCU_CHOOSER=kev` and run the local
kev model instead (slower and less accurate; see below).
## Which model decides what
hyprcu separates two kinds of judgement, and uses different models for each.
**Text decisions** (which window is "the file browser", which control does
"submit the form"), made from window titles and the accessibility tree:
| model | where it runs | what leaves the machine | status |
|---|---|---|---|
| **Jev** (TypeSafe), default | cloud | window titles/classes or control names/values + your description | in hyprcu (`HYPRCU_CHOOSER=jev`) |
| **kev** (kev-4b) | local GPU (`kev.service`) | nothing | in hyprcu (`HYPRCU_CHOOSER=kev`) |
Jev is more accurate and faster (28/29 vs 20–25/29 on `tools/choice_bench.py`)
and can say "none of these"; kev keeps everything local.
**Visual decisions** (did it work, is a dialog open, what does this say,
which option is selected), made from a screenshot or a `--then changes` crop.
Jev cannot do these: it reads text only and never sees an image.
| model | where it runs | what leaves the machine | status |
|---|---|---|---|
| **kev-vision** (Qwen3-VL-4B-Instruct, 4-bit) | local GPU | nothing | benchmark only (`tools/vision_bench/`) |
| **Claude Haiku 4.5 / Sonnet 4.6** | cloud (via `pi -p`) | the image | benchmark only |
| Qwen describes the image, Jev decides from the text | local + cloud | the description | tested and dropped |
kev-vision answers without generating text: it reads the model's
probability for each lettered option in one pass (the readout of the
[Visual Jev](https://arxiv.org/abs/2609.25845) paper), and encodes each
screenshot once however many questions are asked about it.
Measured on 47 questions an agent asks after acting (32 about the nested
test session, in this repo; 15 about the owner's real desktop, kept
private), 2026-09-23:
| setup | correct | time per question |
|---|---|---|
| Sonnet 4.6 | 45/47 (96%) | ~1.4 s (incl. ~1 s pi startup) |
| Haiku 4.5 | 41/47 (87%) | ~1.3 s |
| kev-vision, local | 43/47 (91%) | 0.63 s (2 s to encode a full screenshot, then ~0.15 s per question) |
| Qwen describes, Jev decides | 36/47 (77%) | 5.9 s |
| **kev-vision, sending answers below 0.9 confidence to Haiku or Sonnet** | **45/47 (96%)** | 6% of questions leave the machine |
| kev-vision, below 0.995 to Sonnet | 46/47 (98%) | 15% leave the machine |
The local and cloud models fail in opposite ways (the cloud models guess
when the image cannot answer; kev-vision hedges on false claims), and
kev-vision's confidence drops on its mistakes, which is why escalating only
its uncertain answers works. The thresholds were chosen on these same 47
questions, so treat the combined rows as optimistic until the set grows.
kev-vision needs ~3.7 GB of GPU memory (4.5 GB peak) and does not fit
alongside kev on an 8 GB card.
Run the benchmark (kev-vision needs kev's venv, which has torch):
python3 tools/vision_bench/run.py pi:anthropic/claude-haiku-4-5 pi:anthropic/claude-sonnet-4-6
~/Work/kev/.venv/bin/python tools/vision_bench/run.py kev-vision
Next: a kev-vision server and a hyprcu tool that asks it about the latest
screenshot or change crop, escalating uncertain answers.
## What's kept from upstream
| module | what |
|---|---|
| `wire.py` | raw `zwlr_virtual_pointer_v1` client — motion+button+axis from one pointer, so drag works. hyprcu: the virtual keyboard puts every key on its real US keycode (wtype-style numbering made Chromium read '/' as Escape and '-' as Backspace) |
| `a11y.py` | AT-SPI tree. hyprcu: in-process D-Bus (jeepney) + Collection fast path instead of one `busctl` per call (Chromium `ui`: 5.2 s → 0.22 s); `busctl` remains the fallback |
| `events.py` | Hyprland socket2 listener → `wait_for`, event-driven `launch`/`close_window` |
| `hyprctl.py` | IPC layer; handles both hyprlang and Lua-dispatch Hyprland |
| `input.py`, `screenshot.py`, `clipboard.py`, `session.py` | as upstream, except typing/combos go through our virtual keyboard (wtype only as a fallback) |
| `server.py` | the MCP server; hyprcu adds window description resolution, `click_ui` by description, and per-toolkit accessibility coordinate mapping (Chromium UI vs web content) |
## What's replaced with no-op stubs
`trust.py`, `safety.py`, `skill.py` — same public names, every guard passes,
no log call does anything. `journal.py`
is NOT a stub: it is the training log (see below).
## What's removed
`waybar/`, `packaging/`, and upstream's `init` wizard, `journal`/`replay`
commands and skill installer.
## CLI (shell verbs) — kept
Every MCP tool is also a shell verb (the installed `hyprcu` binary starts
in ~20 ms cold; a few hundred ms under `uv run`) for agents that only
have bash (Codex, Claude Code in a terminal, scripts):
hyprcu # no args = MCP server on stdio
hyprcu desktop # one line per monitor/workspace/window
hyprcu screenshot --window 0x… # prints path + {"geometry","scale",…}
hyprcu hypr focus_window 0x…
hyprcu pointer click 800 60
hyprcu keyboard type "hello" --window 0x…
hyprcu sequence '[{"op":"keyboard","action":"type","text":"x"},…]'
hyprcu doctor # check binaries + session
hyprcu stop # kill any running server/verb
Exit codes: 0 delivered · 1 error · 2 usage · 3 refused (read-only) · 4 no result.
## Tests
`uv run pytest tests/ --ignore=tests/test_e2e.py` → 362 passed, 53 xfailed.
The xfails are enumerated in `tests/removed_guard_tests.txt`; each asserts
that a guard refuses something. They are `strict`, so a guard silently
coming back would fail the suite.
## Run
uv sync
uv run python -m hyprcu # MCP server on stdio
Tools: desktop, screenshot, zoom, ui, marks, binds, wait_for, pointer,
keyboard, click_ui, hypr, launch, use_bind, sequence.
## Measured on this machine (2026-09-20)
- `launch foot` → 167ms, returns address (event-driven, no sleep)
- `sequence` type+enter → 627ms, one round-trip
- `close_window` → 3ms, confirms destroy event
- `ui` on Strata: 1 element (sidebar toggle only — custom-drawn list)
- `ui` on Slack: 0 elements (Electron, needs `--force-renderer-accessibility`)
The a11y tools are only as good as the apps' trees. On this desktop, today,
they're mostly empty. `sequence`, `launch`, `wait_for` and the unified
pointer are the real gains.
## What hyprcu adds (the parts that are ours)
**`pick.py` — natural-language window targeting.** Every `window` argument
(hypr, pointer, keyboard, screenshot, ui, click_ui…) accepts an address, a
class/title substring, or a description. Resolution: exact address → unique
substring (free) → kev choice (~200ms, local). Ambiguous substrings are also
tie-broken by kev. Below `KEV_GATE` (0.5) it fails with "the app may not be
open — check desktop() or launch it", which has been right every time so far.
Results carry `[kev: 99% in 229ms]` so you can see when it was used.
hyprcu hypr focus_window "the file browser"
hyprcu keyboard type "hello" --window "the shell on workspace 2"
**Choosers: Jev (default) or kev.** By default these choices go to TypeSafe's
Jev (`TYPESAFE_API_KEY` from the environment) with an explicit "none of these"
option; this sends window titles and control names to TypeSafe (see
[Requires a TypeSafe API key](#requires-a-typesafe-api-key-sends-names-to-typesafe)).
`HYPRCU_CHOOSER=kev` uses the local model with a probability gate instead. `tools/choice_bench.py` (29 synthetic queries, nothing from the live
desktop): Jev 28/29 at ~0.6 s; kev-4b 20-25/29 at 1.7-2.8 s on a laptop GPU
(kev slows as the option list grows).
**`click_ui` by description.** When no control is *named* what you asked,
`click_ui` treats it as a description and the chooser picks among the
window's clickable controls (browser controls are labelled web page vs
toolbar), or refuses when none fits:
hyprcu click_ui "submit the form" --window "the web browser"
# clicked button 'Send message' ... [jev: 92% in 294ms]
**See only what changed: `--then changes`.** Every acting verb (pointer,
keyboard, click_ui, hypr, use_bind, sequence) can return its own effect.
hyprcu grabs a raw frame of the focused monitor before the action, waits
until the screen has been still for 0.3 s (up to 2 s), and compares:
nothing changed → one line of text and no image; a small change → one crop
per changed area; most of the screen → the whole screenshot. The mouse
pointer is masked out, so moving it is not a change. Reading an image is the
slow, costly part of computer use, and most effects are small:
hyprcu pointer click 638 696 --then changes
# changes: 4 area(s) changed on TEST (3.2% of it, settled after 1735 ms): ...
hyprcu pointer move 1400 950 --then changes
# changes: nothing changed on TEST (settled after 408 ms)
Measured on the nested test session (1600x1000): a raw frame is ~15 ms, a
typical call 0.5–0.9 s including the wait; typing a command in foot returned
two crops of 0.7% of the screen instead of a full screenshot.
**Window-relative clicks.** `pointer` takes `window` + `x_pct`/`y_pct`
(0.0–1.0); the window is focused first and the fraction is mapped to its
current geometry, so the click survives moves and resizes. CLI:
`hyprcu pointer click --in "Strata" --at 0.053 0.23`.
**`journal.py` — training log.** Acting tools append one JSONL row to
`~/.local/share/hyprcu/actions.jsonl` (`HYPRCU_LOG=0` disables): args,
the window list at the time, result, and — when kev chose — query and
probability. Rows are unlabelled; a `correct` field is meant to be added
later before anything is trained on them.
With `HYPRCU_CHOOSER=kev`, requires a Jev-compatible server at `KEV_URL` (default kev-4b on :8009).
The canonical launcher is the systemd user unit `kev.service`
(`systemctl --user start/stop/status kev`; runs on CUDA in NF4, offline-safe,
auto-starts at login) — a manual fallback is `tools/kev-serve.sh`. Without
one, substring and address targeting still work; descriptions fail with an
explicit message.
## No built-in judgement (2026-09-20)
hyprcu does what it's asked. Checks and controls belong in the caller, not
here. Removed from upstream's *acting* tools, beyond the trust layer:
| was | now |
|---|---|
| `sequence` aborts if the desktop changes between steps (default on) | runs every step; `stop_on_change=true` is opt-in |
| `sequence` capped at 20 steps / 30 s | effectively unbounded |
| `launch` wait clamped to 1–30 s | caller's value |
| `wait_for` timeout clamped to 1–60 s | caller's value |
The event stream still serves `wait_for` steps inside a sequence (so an
event between steps isn't missed) — that's plumbing, not a guard.
`tests/test_no_guards.py` pins this contract.
## Trust/approval language audit (2026-09-20)
Verified against the live MCP wire, not just the source:
- No code path can raise a refusal: `grep "raise TrustError"` → 0 hits outside
the stub's class definition.
- Server `instructions` (what the model reads at connect): 0 gating terms.
- Tool descriptions: the three `allow_auth=true overrides the refusal…`
sentences, "panic-kill guarantees", and "Refused while HYPRCU_CONFINE"
removed. Remaining "confirm" is "screenshot to confirm the click worked".
- `allow_auth` is still an accepted boolean on pointer/keyboard/click_ui —
it flows only into no-op stubs, FastMCP emits it with no description, and
8 upstream tests pass it. Left in to keep server.py logic identical to
upstream; harmless.
- `HYPRCU_READONLY=1` still works as an opt-in "observe only" mode (hides
acting tools). Nothing sets it by default.
## Lineage
hyprcu is a fork of [hypruse](https://github.com/IlyasKhallouki/hypruse) by
IlyasKhallouki (MIT). The Wayland/Hyprland primitives — `wire.py` (virtual
pointer), `a11y.py`, `events.py`, `hyprctl.py`, `input.py`, `screenshot.py`
and the MCP `server.py` — are his work, kept close to upstream so fixes can be
merged. hyprcu removes the trust/journal/safety layers and every built-in
refusal, and adds kev-based natural-language targeting, window-relative
coordinates, and a training log. Upstream reference checkout: `~/Work/hypruse`.
TDQS
Scored across 14 tools
Each tool targets a distinct capability: hypr for compositor window/workspace ops, desktop for state snapshots, screenshot/zoom for capture, ui/marks/click_ui for accessibility-driven interaction, pointer/keyboard for raw input, binds/use_bind for user-defined shortcuts, wait_for for events, launch for starting apps, and sequence for batching. Overlaps are minimal and clearly delineated (e.g., click_ui vs pointer).
Names are a mix of single-word nouns (desktop, ui, marks, binds, pointer, keyboard, sequence) and verbs (launch, zoom, screenshot), plus compound verb-noun forms (wait_for, click_ui, use_bind). The inconsistency is not chaotic but lacks a uniform pattern like verb_noun throughout.
14 tools is well-scoped for a desktop automation server covering input, capture, state, events, launching, and UI interaction. Each tool has a clear role with no redundancy, and the count is neither thin nor bloated.
The surface covers the full lifecycle of desktop automation: observing (desktop, screenshot, zoom), interacting (pointer, keyboard, click_ui, hypr), waiting for changes (wait_for), launching (launch), and executing user workflows (binds, use_bind). No obvious gaps for the stated purpose of controlling Hyprland desktops.