jevdevice
This server lets an agent safely perform single atomic actions on a real device, approve or deny pending actions, and capture screenshots.
device_do(goal, ...): resolve one plain-language goal to exactly one atomic phone action, such as opening an app, tapping, long-pressing, typing, swiping, scrolling to find, pressing a key, toggling a service, setting DND, taking a screenshot, or reading a system service.It gates the proposed action, executes only that one action, verifies it by default, and can attach a real screenshot of the resulting screen.
device_screenshot(): take a screenshot of the current screen and return the image; read-only and never gated.device_approve(thread_id, decision, command?, ...): resolve a pending action from any device_* tool by approving or denying it, optionally overriding the command and optionally attaching a screenshot.It does not plan or chain multiple actions; sequencing multi-step goals is the calling agent's responsibility.
Allows an AI agent to control a real Android phone, performing actions such as tapping, typing, swiping, long-pressing, scrolling, opening apps, pressing keys, toggling settings, reading system services, and taking screenshots.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@jevdevicetap the search button on my phone"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
jevdevice
jevdevice turns a plain-language goal into exactly one verified action on a real device.
Why
Scripted device automation uses a fixed map: exact coordinates, resource IDs, one decision tree per app. This map breaks the moment an app changes its layout, and it never covers an app nobody tested against. jevdevice reads the real device state before every action, so its choices adjust when the app changes. A typed judge model picks only from those real candidates, and every choice runs through a safety gate before execution.
Related MCP server: MCP Android Agent
Install
You need Python 3.12+ and uv.
uv syncThe local-shell device family needs nothing more. The Android phone family also needs adb, a
phone reachable over USB or wireless debugging, and ANDROID_SERIAL set (from adb devices).
The default hosted judge engine needs a TYPESAFE_AI_API key. Put both in a .env file one
directory above the repo, or at the workspace root. jevdevice reads it automatically.
uv sync also fetches typesymbolic directly from
GitHub, pinned to the revision in pyproject.toml's [tool.uv.sources]. See "The typesymbolic
dependency" below for what it provides.
Usage
uv run python eval/phases/cli_family.py run # dev goal set, local shell only, no phone or key=== pc17: Which users are currently logged in?
picked pick: who (confidence 0.92, any_fit 0.91)
gate verdict: needs_approval (jev_uncertain, noul=0.65)
-> escalated command=None satisfied=None
runs: 10, ok: 4, not-ok: 6With a phone connected:
uv run python -m jevdevice.execution.dispatch "open the calculator" # one goal
uv run python demo.py # guided interactive demoStatus
Active development, no fixed release. The phone and local-shell device families, the safety gate,
and the decision journal work today with the hosted jev engine. The in-process laya engine
swaps in without code changes but does not yet reach jev's accuracy. Run it only on CUDA, not
on CPU. The calibration loop and recipe recall are new, and confirmed so far only on small live
samples.
Tests
uv run pytest --deselect tests/test_mcp_approval_flow.py # 257 offline tests, no device or network
uv run ruff check # lint, must stay cleantest_mcp_approval_flow.py is the one live integration test. It needs a connected phone and a
real API key.
Contributing
Issues and pull requests are welcome. For larger changes, open an issue first to discuss the approach.
License
MIT © 2026
Architecture
The core idea
Most automation tools work from a hand-written map: exact coordinates, resource IDs, a fixed decision tree per app. That map breaks when an app updates, and it never covers an app the author did not test. jevdevice hand-writes nothing device-specific. Every action follows the same pipeline.
Enumerate the real, current device state.
Narrow the real candidate list to one, with the judge engine.
Gate the resulting command: a deny-list, then one direct safety judgment from the judge. Anything uncertain needs a human decision instead of running.
Execute and verify against device truth (exit code, live screen state, real command output).
No step can select something that does not exist on the device at that moment, so the action set is correct by construction. The judge is only ever asked typed questions over closed option sets, so it can be calibrated, replayed, and swapped. That is how the engine below got exchanged for a second one without a single call site changing.
The typesymbolic dependency
jevdevice runs on typesymbolic, a separate repo that
supplies the judge engine, question primitives, decision journal, gate types, and calibration
math everything above rests on. jevdevice consumes it as a pinned git revision
(pyproject.toml's [tool.uv.sources]) and never modifies it. A core change this repo needs
becomes a change in typesymbolic itself, not an edit here.
What jevdevice takes from typesymbolic:
Question primitives:
Noul,Choice,Score,Question,QuestionRef,Answer, re-exported unchanged fromjev.py.The engine and transport:
JevEngine,ask_batch().The journal: append-only decision and outcome rows, replay, and blob storage for large values.
Gate types:
GateVerdict,GateResult,circuit.threshold_decision().Calibration:
current_threshold(),recalibrate()(tighten-only unless a human signs off).Planner episodes: one
Episodeperresolve()call.
What stays domain-owned in jevdevice, with no typesymbolic counterpart: candidate narrowing
(matching.py), per-engine budget and threshold profiles (budget.py), usage accounting
(ledger.py), two-round narrowing and shadow-engine plumbing (judge/narrowing.py,
judge/shadow.py), and the question_sets wording surface.
Judge engines
One judge contract, two interchangeable implementations, selected by JEV_ENGINE:
jev(default): TypeSafe's hosted System One model (docs.typesafe.ai). NeedsTYPESAFE_AI_API. About 0.3 seconds per judgment.laya: an in-process checkpoint, pinned to an exact revision, that answers the identical question contract locally (asyncio.to_threadpredict). Needs no API key.JEV_DEVICEselects CPU or CUDA, andJEV_LAYA_REVISIONpins the checkpoint (the default is the calibrated base revision).
Both engines see the same frozen question wordings (below), record usage into a shared ledger, and journal every decision. Runs from either engine are comparable, replayable, and safe to A/B.
Device families and the Device protocol
Device knowledge sits behind a five-member protocol (src/jevdevice/device/protocol.py): name
(journal identity), dump_hierarchy() (the snapshot), run() (one shell command, exit code and
stdout as device truth), run_binary() (raw stdout bytes), window_size() (geometry). Engine
modules type against the protocol only.
AdbDevice: the Android phone, over a frozen adb/uiautomator2 transport. The snapshot is the live accessibility-tree XML. A warm read costs about 0.3 seconds because the uiautomator2 companion stays resident (a plainuiautomator dumpcosts about 2.2 seconds per call in process startup alone).CliDevice: the local shell. The snapshot is plain text (pwdplus a listing), so the engine's XML parser rejects it and screen-grounded paths fail closed instead of assuming a screen exists. The CLI family's real path is the gated-command spine, with a frozen candidate command set per goal. It needs no phone at all, which makes it useful for exercising, calibrating, and training the judge on pure-text device data.
Adding a family means writing one adapter. The engine itself stays untouched. The second family was chosen on purpose to break the protocol's phone-shaped assumptions (screen geometry, "the snapshot is a UI tree") rather than confirm them.
What the judge is asked, and who writes it
Judge questions are frozen, versioned artifacts, never prose improvised at runtime:
src/jevdevice/question_sets/v1.yaml (phone family) and v2.yaml (adds the local-shell family's
wordings, and keeps every v1 entry byte-identical). JEV_QUESTION_SET selects the version. Call sites
read their question text from the artifact. An unknown question id or a missing slot raises an
error instead of falling back to improvised wording. Runtime-generated wording exists only as a
journaled escape hatch, and can later be promoted into the next compiled version.
A compile validator (eval/phases/compile_questions.py) checks each frozen set against the
journal: every real-path question instance must match a template exactly, and the validator stamps
witness counts back into the artifact as provenance.
Gates and safety
Every command that changes device state passes the gate before it runs:
A deny-list rejects dangerous command shapes outright (
rm,reboot, factory-wipe words, app uninstall or clear), and a read-only classification lets pure reads through without a judgment.Then one safety judgment from the engine ("does this command do exactly what the chosen action says, and nothing more?") runs against a calibrated confidence threshold.
Below the threshold, the verdict is needs_approval and the command does not run. The MCP
approval flow resolves it in a second tool call. The CLI prompts the terminal. Unattended runners
(the planner, the eval harnesses) never auto-approve, and an approval-pending verdict counts as a
failure. Gate thresholds are per-engine calibration knobs. Automatic recalibration may only ever
tighten them. Loosening a threshold needs recorded human sign-off.
The decision journal
Everything is journaled, append-only, outside the repo (~/.jevdevice/tsjournal/ by default,
JEV_JOURNAL_DIR overrides):
Decision rows record every judge ask: the full state, the questions, the answers, usage, latency, engine, and checkpoint revision.
Outcome rows record every execution: what ran, the verification result (verified, failed, unverified, or escalated), the device identity, and a before/after foreground edge for trajectory building.
Decision and outcome rows join by
call_id, so an approval, its judgment, and its execution replay as one chain. Large values live in a content-addressed blob store.
The journal is the project's data flywheel. Calibration pulls distributions from it with zero model calls, the fine-tune exporter builds training sets from verified rows, the recipe builder aggregates verified chains, and the question compiler validates wordings against it. Every outcome row carries its device family's name, so data from different device families stays separate by construction.
Recipes and the tiered planner
Multi-step goals resolve through a tiered planner (src/jevdevice/execution/planner.py) that
never generates a free-form plan.
A stored recipe (a previously verified chain of actions for the same or a similar goal) runs first. The judge only fills variable slots.
A stored recipe that mostly fits gets adapted: each step is checked against the live device, drop-only.
Otherwise the planner selects steps itself, choosing among the engine's closed action vocabulary with a done-check after each step.
A goal nothing covers goes back to the calling agent to break down. Each step still runs gated through the same engine.
An unresolved goal escalates to a human.
Every tier transition is journaled. Verified runs feed back into the recipe store automatically.
Calibration
Each engine has its own budget and threshold profile (src/jevdevice/budget.py): option-count
limits, state truncation, chunk sizes, abstain options, and the gate and confidence thresholds,
all named knobs, never bare numbers. The calibration CLIs (src/jevdevice/calibrate/) refit
profiles from journal data. A continuous rolling-window loop refits quality metrics (Brier score,
expected calibration error, precision) from stored rows with zero model calls, and proposes
threshold changes that run in shadow mode before any tighten-only promotion.
The MCP surface (three tools, kept minimal)
device_do(goal, verify, auto_approve): resolves one goal to one atomic action (kind selection, propose, gate, execute, verify). Sequencing multi-step tasks stays the calling agent's job.device_approve(thread_id, decision): resolves a pendingneeds_approvalaction. A denied command never runs, in any mode.device_screenshot(): returns the live screen.
Eval tooling and the data split
eval/phases/ holds the batch and analysis tooling (functional names, see its own README): A/B
runners, shadow-agreement reports, threshold refits, gate audits, question compilation, the
fine-tune export and Kaggle training script, a perturbation harness, recovery-pair mining,
training-format exporters, recipe building, and METR Task Standard-shaped adapters over both
device families. eval/goals.yaml holds the plain-language goal sets with a frozen dev/held-out
split. The held-out half is never run, tuned on, or exported, and every batch tool checks this
through one shared loader.
Why this design
Ground everything in device truth. Candidates, verification, and escalation all come from what the device actually reports.
Keep the judge's job small and typed. Closed option sets, frozen wordings, and calibrated thresholds make an engine swappable and its decisions auditable.
Fail closed. When confidence is short, the action does not run. A human decides, or the goal escalates.
Journal everything. The journal is the single source of labeled data, provenance, and replay for every later improvement: calibration, fine-tunes, recipes, question compilation.
Available Tools
3 toolsdevice_approveA
Resolve a pending action from any device_* tool. decision is "approve"
or "deny". An optional command overrides the proposed command (e.g. a
human-corrected variant), matching PLAN.md's device_approve(thread_id,
decision, command?) signature. include_screenshot=True attaches a real screenshot,
same off-by-default tradeoff as device_do.
| Name | Required | Description | Default |
|---|---|---|---|
| command | No | ||
| decision | Yes | ||
| thread_id | Yes | ||
| include_screenshot | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does solid work: it explains that 'decision' is approve or deny, that command can override the proposed command as a human-corrected variant, and that include_screenshot=True attaches a real screenshot with an off-by-default tradeoff. It does not describe post-approval effects or authorization requirements, but it discloses the core behavioral semantics.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences deliver the essential behavioral content with no filler. The action is front-loaded, and each sentence earns its place: one for core resolution semantics, one for optional overrides and the screenshot tradeoff.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with no annotations and no output schema, the description covers the main invocation concerns: what the tool does, what decisions are valid, how to override the command, and what include_screenshot does. It is slightly incomplete on what happens after approval or denial and does not describe the return value, but nothing critical prevents an agent from calling the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, and it largely does: it defines decision values, explains command's override role, and clarifies include_screenshot's effect and default. thread_id is only implied through the PLAN.md signature rather than semantically described, but the context of resolving a pending action makes its role reasonably clear.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Resolve a pending action from any device_* tool.' It clearly differentiates this from its siblings, device_do and device_screenshot, by positioning it as the approve/deny gate for pending actions rather than an executor or capture tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly implies use when a pending action exists from any device_* tool and states the decision values and optional command override. It references device_do's tradeoff, giving a useful contextual anchor, though it does not explicitly state when not to use this tool or name direct alternatives for exclusion.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
device_doB
One tool for any single real phone action. Jev first picks which ONE atomic kind this goal wants -- open an app, tap, long-press, type, swipe, scroll-to-find, press a key, toggle a service, set DND, take a screenshot, or read a system service (the same ACTION_KINDS the CLI toolkit uses) -- then runs exactly that one real action and returns its result. include_screenshot=True also attaches a real screenshot of the resulting screen -- off by default since an image costs real context; ask for one at a checkpoint, not after every single step in a sequence. Never plans or chains more than one action; sequencing multiple goals is still the calling agent's job.
| Name | Required | Description | Default |
|---|---|---|---|
| goal | Yes | ||
| verify | No | ||
| direction | No | down | |
| auto_approve | No | ||
| max_attempts | No | ||
| include_screenshot | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations present, the description carries the transparency burden. It honestly discloses that exactly one action is run, results are returned, screenshots are off by default due to context cost, and chaining is not performed by the tool. However, it does not disclose side-effect potential for mutating actions, what verification or auto_approval do, or what happens on failure or retry, leaving significant behavioral gaps.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is reasonably sized and front-loaded with the tool's core purpose, and the screenshot guidance is practical. But it contains redundancy ('one tool for any single... action' vs 'Never plans or chains more than one action') and a confusing 'Jev first picks' narrative that could be stated more directly and compactly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with no output schemahol, no annotations, and six parameters, the description is incomplete. It covers core single-action behavior and screenshot usage, but omits the return format, semantics of four parameters, failure/retry behavior, and any relation to sibling tools. An agent would likely need additional information or experimentation to invoke it correctly in all cases.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for all six parameters. It adds meaning for goal (an atomic action from the enumerated list) and include_screenshot (attaches screenshot, off by default, expensive), but verify, direction, auto_approve, and max_attempts remain entirely unexplained in both schema and description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states that the tool executes one real phone action at a time and enumerates the supported atomic kinds (open app, tap, long-press, type, swipe, scroll-to-find, press key, toggle service, set DND, take screenshot, read system service). It is more specific than the generic name, but it does not directly distinguish itself from sibling tools device_screenshot and device_approve, especially since 'take a screenshot' is listed among its own action kinds.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives useful context: invoke one atomic action per call, never plan or chain, and request screenshots only at checkpoints because images cost context. However, it does not explicitly say when to prefer device_screenshot or device_approve, nor what conditions would make those siblings the better choice over this tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
device_screenshotA
Take a screenshot of the current screen and return the actual image -- never gated (read-only), no goal needed, no candidates to narrow.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It clearly discloses read-only nature, no gating, no prerequisites, and that the actual image is returned. It could mention edge cases like unavailable screens, but the core behavioral profile is well covered.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Single compact sentence front-loads the core action and return value, then uses short clauses to add gating and read-only context. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a parameterless screenshot tool with no output schema, the description covers what it does, its safety profile, and prerequisites. The only omission is output format or failure behavior, but 'actual image' sufficiently sets expectations.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters and schema coverage is 100%, so the description need not document args. It usefully reinforces that no goal or candidate narrowing is required, confirming the parameterless call is complete.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: takes a screenshot of the current screen and returns the actual image. It also distinguishes itself from sibling tools by noting it is never gated, read-only, and needs no goal or candidate narrowing.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly signals when it applies: whenever a current-screen image is needed, without a goal or candidate selection. The sibling contrast is implied but not named, and there is no explicit exclusion like 'use X instead'.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
3 tool updates
v0.1.0- First observed
device_approve - First observed
device_do - First observed
device_screenshot
TDQS
Scored across 3 tools
Each tool has a clearly distinct purpose: device_do executes atomic actions, device_screenshot captures the screen, and device_approve resolves pending approvals. No overlap or ambiguity between them, even though device_do covers many sub-actions internally.
All tools share the 'device_' prefix, creating a clear namespace. However, the verb part is inconsistent: 'do' is vague, 'screenshot' is a noun-as-verb, and 'approve' is a verb. The pattern is recognizable but not perfectly uniform.
Three tools is on the thin side for a device control server, but the design intentionally consolidates all atomic actions into device_do, making the count reasonable. Each tool has a distinct role and earns its place.
The surface covers core device interactions (actions, screenshots, approval resolution) without obvious dead ends. Minor gaps exist, such as no explicit listing of available action kinds or device status, but these are workable given the CLI-toolkit reference.
Maintenance
Related MCP Connectors
Control real Android and iOS devices with LLM agents — tap, swipe, type, automate flows.
Decision-assurance for AI agents: an auditable action boundary + receipt before it acts.
Human-in-the-loop approval for agent actions, with verifiable action-bound receipts.
Melaya is a remote MCP server. It gives an assistant hands on your own Android phone and browser: it reads the screen through the accessibility tree, then taps, types and navigates inside the apps and sites you allow-list, with no per-app API. It also builds, schedules and runs agent pipelines across 6k+ connected tools. OAuth 2.1, nothing to install.
Related MCP Servers
- AlicenseCqualityCmaintenanceA lightweight bridge enabling AI agents to perform real-world tasks on Android devices such as app navigation, UI interaction, and automated QA testing without requiring computer-vision pipelines or preprogrammed scripts.14872MIT
- AlicenseNot gradedqualityDmaintenanceEnables AI agents to control Android devices through natural language, supporting app management, UI interaction, gestures, and system operations via uiautomator2.MIT
- AlicenseNot gradedqualityBmaintenanceEnables AI agents to see and operate a real Android phone over adb by fusing live screenshots with UI-tree data, supporting look, tap, swipe, type, and screenshot actions through any MCP client.3MIT
- AlicenseNot gradedqualityBmaintenanceEnables LLM agents to drive real Android devices or emulators by reading the live accessibility tree, acting on UI nodes instead of coordinates, verifying every action, and typing Unicode text.34 npm1Apache 2.0