Skip to main content
Glama

jevdevice

License: MIT

jevdevice turns a plain-language goal into exactly one verified action on a real device.

Why

Scripted device automation uses a fixed map: exact coordinates, resource IDs, one decision tree per app. This map breaks the moment an app changes its layout, and it never covers an app nobody tested against. jevdevice reads the real device state before every action, so its choices adjust when the app changes. A typed judge model picks only from those real candidates, and every choice runs through a safety gate before execution.

Related MCP server: MCP Android Agent

Install

You need Python 3.12+ and uv.

uv sync

The local-shell device family needs nothing more. The Android phone family also needs adb, a phone reachable over USB or wireless debugging, and ANDROID_SERIAL set (from adb devices). The default hosted judge engine needs a TYPESAFE_AI_API key. Put both in a .env file one directory above the repo, or at the workspace root. jevdevice reads it automatically.

uv sync also fetches typesymbolic directly from GitHub, pinned to the revision in pyproject.toml's [tool.uv.sources]. See "The typesymbolic dependency" below for what it provides.

Usage

uv run python eval/phases/cli_family.py run   # dev goal set, local shell only, no phone or key
=== pc17: Which users are currently logged in?
picked pick: who (confidence 0.92, any_fit 0.91)
gate verdict: needs_approval (jev_uncertain, noul=0.65)
    -> escalated command=None satisfied=None
runs: 10, ok: 4, not-ok: 6

With a phone connected:

uv run python -m jevdevice.execution.dispatch "open the calculator"   # one goal
uv run python demo.py                                                 # guided interactive demo

Status

Active development, no fixed release. The phone and local-shell device families, the safety gate, and the decision journal work today with the hosted jev engine. The in-process laya engine swaps in without code changes but does not yet reach jev's accuracy. Run it only on CUDA, not on CPU. The calibration loop and recipe recall are new, and confirmed so far only on small live samples.

Tests

uv run pytest --deselect tests/test_mcp_approval_flow.py   # 257 offline tests, no device or network
uv run ruff check                                           # lint, must stay clean

test_mcp_approval_flow.py is the one live integration test. It needs a connected phone and a real API key.

Contributing

Issues and pull requests are welcome. For larger changes, open an issue first to discuss the approach.

License

MIT © 2026


Architecture

The core idea

Most automation tools work from a hand-written map: exact coordinates, resource IDs, a fixed decision tree per app. That map breaks when an app updates, and it never covers an app the author did not test. jevdevice hand-writes nothing device-specific. Every action follows the same pipeline.

  1. Enumerate the real, current device state.

  2. Narrow the real candidate list to one, with the judge engine.

  3. Gate the resulting command: a deny-list, then one direct safety judgment from the judge. Anything uncertain needs a human decision instead of running.

  4. Execute and verify against device truth (exit code, live screen state, real command output).

No step can select something that does not exist on the device at that moment, so the action set is correct by construction. The judge is only ever asked typed questions over closed option sets, so it can be calibrated, replayed, and swapped. That is how the engine below got exchanged for a second one without a single call site changing.

The typesymbolic dependency

jevdevice runs on typesymbolic, a separate repo that supplies the judge engine, question primitives, decision journal, gate types, and calibration math everything above rests on. jevdevice consumes it as a pinned git revision (pyproject.toml's [tool.uv.sources]) and never modifies it. A core change this repo needs becomes a change in typesymbolic itself, not an edit here.

What jevdevice takes from typesymbolic:

  • Question primitives: Noul, Choice, Score, Question, QuestionRef, Answer, re-exported unchanged from jev.py.

  • The engine and transport: JevEngine, ask_batch().

  • The journal: append-only decision and outcome rows, replay, and blob storage for large values.

  • Gate types: GateVerdict, GateResult, circuit.threshold_decision().

  • Calibration: current_threshold(), recalibrate() (tighten-only unless a human signs off).

  • Planner episodes: one Episode per resolve() call.

What stays domain-owned in jevdevice, with no typesymbolic counterpart: candidate narrowing (matching.py), per-engine budget and threshold profiles (budget.py), usage accounting (ledger.py), two-round narrowing and shadow-engine plumbing (judge/narrowing.py, judge/shadow.py), and the question_sets wording surface.

Judge engines

One judge contract, two interchangeable implementations, selected by JEV_ENGINE:

  • jev (default): TypeSafe's hosted System One model (docs.typesafe.ai). Needs TYPESAFE_AI_API. About 0.3 seconds per judgment.

  • laya: an in-process checkpoint, pinned to an exact revision, that answers the identical question contract locally (asyncio.to_thread predict). Needs no API key. JEV_DEVICE selects CPU or CUDA, and JEV_LAYA_REVISION pins the checkpoint (the default is the calibrated base revision).

Both engines see the same frozen question wordings (below), record usage into a shared ledger, and journal every decision. Runs from either engine are comparable, replayable, and safe to A/B.

Device families and the Device protocol

Device knowledge sits behind a five-member protocol (src/jevdevice/device/protocol.py): name (journal identity), dump_hierarchy() (the snapshot), run() (one shell command, exit code and stdout as device truth), run_binary() (raw stdout bytes), window_size() (geometry). Engine modules type against the protocol only.

  • AdbDevice: the Android phone, over a frozen adb/uiautomator2 transport. The snapshot is the live accessibility-tree XML. A warm read costs about 0.3 seconds because the uiautomator2 companion stays resident (a plain uiautomator dump costs about 2.2 seconds per call in process startup alone).

  • CliDevice: the local shell. The snapshot is plain text (pwd plus a listing), so the engine's XML parser rejects it and screen-grounded paths fail closed instead of assuming a screen exists. The CLI family's real path is the gated-command spine, with a frozen candidate command set per goal. It needs no phone at all, which makes it useful for exercising, calibrating, and training the judge on pure-text device data.

Adding a family means writing one adapter. The engine itself stays untouched. The second family was chosen on purpose to break the protocol's phone-shaped assumptions (screen geometry, "the snapshot is a UI tree") rather than confirm them.

What the judge is asked, and who writes it

Judge questions are frozen, versioned artifacts, never prose improvised at runtime: src/jevdevice/question_sets/v1.yaml (phone family) and v2.yaml (adds the local-shell family's wordings, and keeps every v1 entry byte-identical). JEV_QUESTION_SET selects the version. Call sites read their question text from the artifact. An unknown question id or a missing slot raises an error instead of falling back to improvised wording. Runtime-generated wording exists only as a journaled escape hatch, and can later be promoted into the next compiled version.

A compile validator (eval/phases/compile_questions.py) checks each frozen set against the journal: every real-path question instance must match a template exactly, and the validator stamps witness counts back into the artifact as provenance.

Gates and safety

Every command that changes device state passes the gate before it runs:

  • A deny-list rejects dangerous command shapes outright (rm, reboot, factory-wipe words, app uninstall or clear), and a read-only classification lets pure reads through without a judgment.

  • Then one safety judgment from the engine ("does this command do exactly what the chosen action says, and nothing more?") runs against a calibrated confidence threshold.

Below the threshold, the verdict is needs_approval and the command does not run. The MCP approval flow resolves it in a second tool call. The CLI prompts the terminal. Unattended runners (the planner, the eval harnesses) never auto-approve, and an approval-pending verdict counts as a failure. Gate thresholds are per-engine calibration knobs. Automatic recalibration may only ever tighten them. Loosening a threshold needs recorded human sign-off.

The decision journal

Everything is journaled, append-only, outside the repo (~/.jevdevice/tsjournal/ by default, JEV_JOURNAL_DIR overrides):

  • Decision rows record every judge ask: the full state, the questions, the answers, usage, latency, engine, and checkpoint revision.

  • Outcome rows record every execution: what ran, the verification result (verified, failed, unverified, or escalated), the device identity, and a before/after foreground edge for trajectory building.

  • Decision and outcome rows join by call_id, so an approval, its judgment, and its execution replay as one chain. Large values live in a content-addressed blob store.

The journal is the project's data flywheel. Calibration pulls distributions from it with zero model calls, the fine-tune exporter builds training sets from verified rows, the recipe builder aggregates verified chains, and the question compiler validates wordings against it. Every outcome row carries its device family's name, so data from different device families stays separate by construction.

Recipes and the tiered planner

Multi-step goals resolve through a tiered planner (src/jevdevice/execution/planner.py) that never generates a free-form plan.

  1. A stored recipe (a previously verified chain of actions for the same or a similar goal) runs first. The judge only fills variable slots.

  2. A stored recipe that mostly fits gets adapted: each step is checked against the live device, drop-only.

  3. Otherwise the planner selects steps itself, choosing among the engine's closed action vocabulary with a done-check after each step.

  4. A goal nothing covers goes back to the calling agent to break down. Each step still runs gated through the same engine.

  5. An unresolved goal escalates to a human.

Every tier transition is journaled. Verified runs feed back into the recipe store automatically.

Calibration

Each engine has its own budget and threshold profile (src/jevdevice/budget.py): option-count limits, state truncation, chunk sizes, abstain options, and the gate and confidence thresholds, all named knobs, never bare numbers. The calibration CLIs (src/jevdevice/calibrate/) refit profiles from journal data. A continuous rolling-window loop refits quality metrics (Brier score, expected calibration error, precision) from stored rows with zero model calls, and proposes threshold changes that run in shadow mode before any tighten-only promotion.

The MCP surface (three tools, kept minimal)

  • device_do(goal, verify, auto_approve): resolves one goal to one atomic action (kind selection, propose, gate, execute, verify). Sequencing multi-step tasks stays the calling agent's job.

  • device_approve(thread_id, decision): resolves a pending needs_approval action. A denied command never runs, in any mode.

  • device_screenshot(): returns the live screen.

Eval tooling and the data split

eval/phases/ holds the batch and analysis tooling (functional names, see its own README): A/B runners, shadow-agreement reports, threshold refits, gate audits, question compilation, the fine-tune export and Kaggle training script, a perturbation harness, recovery-pair mining, training-format exporters, recipe building, and METR Task Standard-shaped adapters over both device families. eval/goals.yaml holds the plain-language goal sets with a frozen dev/held-out split. The held-out half is never run, tuned on, or exported, and every batch tool checks this through one shared loader.

Why this design

  • Ground everything in device truth. Candidates, verification, and escalation all come from what the device actually reports.

  • Keep the judge's job small and typed. Closed option sets, frozen wordings, and calibrated thresholds make an engine swappable and its decisions auditable.

  • Fail closed. When confidence is short, the action does not run. A human decides, or the goal escalates.

  • Journal everything. The journal is the single source of labeled data, provenance, and replay for every later improvement: calibration, fine-tunes, recipes, question compilation.

Available Tools

3 tools
device_approveA

Resolve a pending action from any device_* tool. decision is "approve" or "deny". An optional command overrides the proposed command (e.g. a human-corrected variant), matching PLAN.md's device_approve(thread_id, decision, command?) signature. include_screenshot=True attaches a real screenshot, same off-by-default tradeoff as device_do.

ParametersJSON Schema
NameRequiredDescriptionDefault
commandNo
decisionYes
thread_idYes
include_screenshotNo

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does solid work: it explains that 'decision' is approve or deny, that command can override the proposed command as a human-corrected variant, and that include_screenshot=True attaches a real screenshot with an off-by-default tradeoff. It does not describe post-approval effects or authorization requirements, but it discloses the core behavioral semantics.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences deliver the essential behavioral content with no filler. The action is front-loaded, and each sentence earns its place: one for core resolution semantics, one for optional overrides and the screenshot tradeoff.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with no annotations and no output schema, the description covers the main invocation concerns: what the tool does, what decisions are valid, how to override the command, and what include_screenshot does. It is slightly incomplete on what happens after approval or denial and does not describe the return value, but nothing critical prevents an agent from calling the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, and it largely does: it defines decision values, explains command's override role, and clarifies include_screenshot's effect and default. thread_id is only implied through the PLAN.md signature rather than semantically described, but the context of resolving a pending action makes its role reasonably clear.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Resolve a pending action from any device_* tool.' It clearly differentiates this from its siblings, device_do and device_screenshot, by positioning it as the approve/deny gate for pending actions rather than an executor or capture tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description clearly implies use when a pending action exists from any device_* tool and states the decision values and optional command override. It references device_do's tradeoff, giving a useful contextual anchor, though it does not explicitly state when not to use this tool or name direct alternatives for exclusion.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

device_doB

One tool for any single real phone action. Jev first picks which ONE atomic kind this goal wants -- open an app, tap, long-press, type, swipe, scroll-to-find, press a key, toggle a service, set DND, take a screenshot, or read a system service (the same ACTION_KINDS the CLI toolkit uses) -- then runs exactly that one real action and returns its result. include_screenshot=True also attaches a real screenshot of the resulting screen -- off by default since an image costs real context; ask for one at a checkpoint, not after every single step in a sequence. Never plans or chains more than one action; sequencing multiple goals is still the calling agent's job.

ParametersJSON Schema
NameRequiredDescriptionDefault
goalYes
verifyNo
directionNodown
auto_approveNo
max_attemptsNo
include_screenshotNo

TDQS

B3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations present, the description carries the transparency burden. It honestly discloses that exactly one action is run, results are returned, screenshots are off by default due to context cost, and chaining is not performed by the tool. However, it does not disclose side-effect potential for mutating actions, what verification or auto_approval do, or what happens on failure or retry, leaving significant behavioral gaps.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is reasonably sized and front-loaded with the tool's core purpose, and the screenshot guidance is practical. But it contains redundancy ('one tool for any single... action' vs 'Never plans or chains more than one action') and a confusing 'Jev first picks' narrative that could be stated more directly and compactly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with no output schemahol, no annotations, and six parameters, the description is incomplete. It covers core single-action behavior and screenshot usage, but omits the return format, semantics of four parameters, failure/retry behavior, and any relation to sibling tools. An agent would likely need additional information or experimentation to invoke it correctly in all cases.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate for all six parameters. It adds meaning for goal (an atomic action from the enumerated list) and include_screenshot (attaches screenshot, off by default, expensive), but verify, direction, auto_approve, and max_attempts remain entirely unexplained in both schema and description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states that the tool executes one real phone action at a time and enumerates the supported atomic kinds (open app, tap, long-press, type, swipe, scroll-to-find, press key, toggle service, set DND, take screenshot, read system service). It is more specific than the generic name, but it does not directly distinguish itself from sibling tools device_screenshot and device_approve, especially since 'take a screenshot' is listed among its own action kinds.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives useful context: invoke one atomic action per call, never plan or chain, and request screenshots only at checkpoints because images cost context. However, it does not explicitly say when to prefer device_screenshot or device_approve, nor what conditions would make those siblings the better choice over this tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

device_screenshotA

Take a screenshot of the current screen and return the actual image -- never gated (read-only), no goal needed, no candidates to narrow.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It clearly discloses read-only nature, no gating, no prerequisites, and that the actual image is returned. It could mention edge cases like unavailable screens, but the core behavioral profile is well covered.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Single compact sentence front-loads the core action and return value, then uses short clauses to add gating and read-only context. Every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a parameterless screenshot tool with no output schema, the description covers what it does, its safety profile, and prerequisites. The only omission is output format or failure behavior, but 'actual image' sufficiently sets expectations.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters and schema coverage is 100%, so the description need not document args. It usefully reinforces that no goal or candidate narrowing is required, confirming the parameterless call is complete.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: takes a screenshot of the current screen and returns the actual image. It also distinguishes itself from sibling tools by noting it is never gated, read-only, and needs no goal or candidate narrowing.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly signals when it applies: whenever a current-screen image is needed, without a goal or candidate selection. The sibling contrast is implied but not named, and there is no explicit exclusion like 'use X instead'.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 3 tool updatesv0.1.0
    • First observeddevice_approve
    • First observeddevice_do
    • First observeddevice_screenshot

TDQS

A3.7/5.0

Scored across 3 tools

Disambiguation5/5

Each tool has a clearly distinct purpose: device_do executes atomic actions, device_screenshot captures the screen, and device_approve resolves pending approvals. No overlap or ambiguity between them, even though device_do covers many sub-actions internally.

Naming Consistency4/5

All tools share the 'device_' prefix, creating a clear namespace. However, the verb part is inconsistent: 'do' is vague, 'screenshot' is a noun-as-verb, and 'approve' is a verb. The pattern is recognizable but not perfectly uniform.

Tool Count4/5

Three tools is on the thin side for a device control server, but the design intentionally consolidates all atomic actions into device_do, making the count reasonable. Each tool has a distinct role and earns its place.

Completeness4/5

The surface covers core device interactions (actions, screenshots, approval resolution) without obvious dead ends. Minor gaps exist, such as no explicit listing of available action kinds or device status, but these are workable given the CLI-toolkit reference.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers