simfast
Provides tools for controlling the iOS Simulator, enabling AI agents to read the screen as compact text, tap elements by label or coordinates, type text, scroll, swipe, press hardware buttons, launch apps, and batch steps, with each action returning the updated screen state.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@simfastlaunch the Settings app, tap General, and tell me what's on screen"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
simfast
Fast, token-cheap iOS Simulator control for AI agents. Read the screen as ~300 tokens of text instead of a 1.5k-token screenshot, tap in ~100 ms through a persistent daemon, and get the new screen back from every action so an agent needs one tool call per step instead of two.
Works as a CLI (great with Claude Code and any agent that has a shell) and as an MCP server (Claude Desktop, Cursor, Codex, …).
$ simfast tap General
tapped "General" (Button) at (201,406)
[Settings] 12 elements · screen changed
0 Button Settings (38,84)
1 Heading General (201,84)
2 Heading General (72,254)
3 Heading Manage your overall setup and preferences for iPhone, such as softwar… (199,315)
4 Button About (201,434)
5 Button Screen Capture (201,521)
6 Button AutoFill & Passwords (201,608)
7 Button Dictionary (201,660)
8 Button Fonts (201,712)
9 Button Keyboard (201,764)
10 Button Language & Region (201,816)
11 Button Trackpad (201,868)
· 1.5s via simgadget (act 0.1 · settle 0.5 · read 0.9)
Click for the 47 s side-by-side video (real tool time; model latency simulated at 2 s per tool call). Reproduce with demo/record.mjs + demo/compose.py.
Why
Most agent-driven simulator loops look like this: tap → screenshot → model reads pixels → guess coordinates → tap. That is slow and expensive for three reasons:
Two tool calls per step (act, then look), and every call is a model round trip.
Images are token-heavy: a full-screen screenshot is ~1,500 tokens, and the model still has to guess coordinates from pixels.
Each tap spawns a process (~0.9 s with AXe) and the raw accessibility tree is huge (600 KB, ~175k tokens for one real app screen), so nobody feeds it to the model.
simfast fixes each of them:
Problem | What simfast does |
Reading the screen |
|
Slow taps | A per-simulator daemon keeps SimGadget's |
Two calls per step | Every action returns the new screen and whether it changed. One call = one step. |
Known paths |
|
Related MCP server: Shotter
Benchmark
Same 6-step Settings flow (General → About → back → Settings → Camera → back), iPhone 17 simulator, iOS 27.0, Xcode 27.1, Apple silicon. Median of 6 runs. bench/compare.mjs reproduces it.
Tool time | Tool calls (model round trips) | Observation tokens | |
Screenshot loop (tap, then screenshot) | 13.2 s | 12 | ~9,000 |
simfast, one call per step | 10.5 s | 6 | ~1,100 |
simfast, whole flow in one | 7.3 s | 1 | ~230 |
Reading it honestly:
Tool time only. LLM thinking time is not measured. It is dominated by the number of round trips and tokens, which is where simfast wins most: with a (made-up) 3 s per model turn the totals are ~49 s / ~29 s / ~13 s. Plug in your own with
--llm-seconds.The screenshot loop is given the benefit of the doubt: a reliable held touch for taps, and 0.4 s per step for the transition to finish. Image tokens are estimated as
w×h/750after resizing to a 1568 px long edge.Per-step simfast is only ~20 % faster in raw tool time: reading the accessibility tree (~0.5-1 s) costs about what a screenshot does. The savings are in tokens and round trips.
Variance is real: in one of six batched runs a slow SimGadget accessibility read pushed it to 27 s (see Known issues).
Install
Requirements: macOS on Apple silicon, Xcode with a simulator runtime, Node 18+, and AXe.
brew install cameroncooke/axe/axe
npm install -g github:iamxicor/simfast # once on npm: npm install -g simfast
simfast doctorsimfast doctor checks everything and tells you what to fix. SimGadget is an optional dependency: without it simfast still works, with taps through AXe (~0.9 s each) instead of ~0.1 s. On first use SimGadget downloads a pinned, SHA-256-verified idb_companion (~19 MB) from its GitHub releases into ~/Library/Caches/simgadget.
Use it from the shell (Claude Code, scripts, any agent with Bash)
Boot a simulator, then:
simfast launch com.apple.Preferences # launch an app, print its first screen
simfast see # what is on screen right now
simfast tap "General" # by label (case-insensitive, prefix/substring)
simfast tap 3 # by index from the latest listing
simfast tap 120,340 # by coordinates
simfast tap Search --type TextField # restrict to an element type
simfast type "hello" --submit # type, then press Return
simfast scroll down # reveal content below
simfast do "tap General" "tap About" # several steps, one screen read at the end
simfast shot /tmp/screen.png # screenshot: only for visual checksWith several simulators booted, pass --udid <UDID> or set SIMFAST_UDID.
Teach Claude Code to use it
Copy the bundled skill so Claude Code reaches for simfast instead of screenshots:
mkdir -p ~/.claude/skills && cp -r skills/simfast ~/.claude/skills/Step reference
Step | Meaning |
| Flags: |
| Type into the focused field; |
| Keyboard keys (or a HID keycode) |
| Content direction: |
| Raw finger direction |
| Hardware buttons |
| Launch an app / open a URL or deep link |
| Pause (useful inside |
In do, use labels, not indexes: the screen changes after the first step. A batch stops at the first failed step and prints the screen where it stopped, so the agent can recover. When the target is not on screen yet (a page still sliding in), tap waits up to ~1.2 s for it to appear.
Reading the output
[Settings] 12 elements · screen changed <- app, element count, did the last action change anything
0 Button Settings (38,84) <- index, type, label, centre point
...
off-screen: 8 down (swipe to reveal) <- how much is hidden and where
· 1.4s via simgadget (act 0.1 · settle 0.5 · read 0.8)screen UNCHANGED after an action is the most useful signal an agent gets: it means the tap did not do what was expected.
Use it as an MCP server
# Claude Code
claude mcp add simfast -- npx -y github:iamxicor/simfast mcpOther clients (Cursor, Claude Desktop, …):
{ "mcpServers": { "simfast": { "command": "npx", "args": ["-y", "github:iamxicor/simfast", "mcp"] } } }Tools: screen, tap, type_text, scroll, swipe, key, button, launch_app, open_url, batch, screenshot. The whole schema is ~5.5 KB (~1.6k tokens) of context. The server starts the daemon immediately so the warm-up overlaps with the model thinking.
How it works
agent ──(CLI or MCP)──> simfast client ──unix socket──> daemon (one per simulator)
│
read: axe describe-ui ──> flatten ──> compact text + changed flag
act : SimGadget gRPC ──> warm idb_companion ──> HID input (~0.1 s)
(falls back to AXe if SimGadget is not installed)src/snapshot.js: pure, unit-tested flattening of the AXe tree (npm test).src/engine.js: the step grammar, label resolution, settle and wait logic.src/daemon.js: idle for 30 min (SIMFAST_IDLE_MIN) then exits and releases the companion.
Label resolution is deliberately defensive. SimGadget's on-device "first match wins" lookup can pick the wrong element (an app named Settings also matches a Back button labelled Settings), so simfast resolves against the snapshot it last showed the agent, and uses the device lookup only for stale screens and only when the match is a real control. Switches go through SimGadget's accessibility activation, because a switch's frame spans its whole row.
Configuration
Variable | Default | |
| the only booted simulator | Target simulator |
|
| Wait after an action before reading (300 ms caught 1 of 24 screens mid-transition, 500 ms caught 0 of 24) |
|
|
|
|
| Daemon idle timeout |
| unset |
|
Daemon logs: /tmp/simfast-<uid>/<udid-prefix>.log. Stop one with simfast stop.
Known issues
These are real, observed while building this, not hypothetical.
The accessibility tree can be wrong. On iOS 27's Settings app, the Accessibility page reports the previous page's list behind it and an unlabelled back arrow. The text listing is a claim about the screen, not the screen. If it looks stale or does not match your expectation, take one screenshot (
simfast shot). The tool descriptions and bundled skill tell agents to do this.Plain
axe tapdid not register in my tests on iOS 27.0 / AXe 1.8.0: 0 of 26 taps landed (plainaxe tap0/16 with and without the companion running;--tap-style physical0/10), while a 0.1 s held touch landed 10 of 10. simfast always holds the touch ≥ 0.1 s (SimGadget's documented floor; the AXe fallback usesaxe touch --delay). Your results may differ with other versions.Cold starts. The first command after the daemon starts waits for the companion (~10 s). The very first run on a machine also downloads the companion and may take ~50 s while the accessibility bridge installs. The first label lookup in a freshly launched app takes ~3 s.
Intermittent slow reads. SimGadget's on-device accessibility reads occasionally take 3-4 s (I saw it once on a real app and once as a 27 s batched run). simfast reads the tree with AXe, so only label lookups on stale screens are exposed.
SIMFAST_NO_SIMGADGET=1avoids it.button homereturns a slow, near-empty listing: the SpringBoard tree is large and mostly unlabelled.Simulators only (no physical devices), macOS on Apple silicon only.
How it relates to other tools
AXe: the accessibility reader and fallback input. simfast is a layer on top.
SimGadget /
simgadget-mcp: the fast-tap engine, and a full MCP of its own. simfast adds the compact snapshot, the act-and-observe loop, batching and defensive label resolution. (In my tests SimGadget's ownui_describe_allreturned ~2k-token nested JSON and sometimes took 3 s+, which is why simfast reads with AXe.)idb, XcodeBuildMCP, Maestro: broader or flow-oriented tools; simfast is narrowly about the agent's inner loop.
Contributing
Issues and PRs welcome, especially: other iOS versions' tap/tree behaviour, more tree-flattening heuristics (with a fixture in test/), and an Android emulator backend. npm test runs the unit tests; npm run bench -- --udid <UDID> the benchmark.
License
MIT. Built on AXe (MIT, Cameron Cooke), SimGadget (MIT, zafnz) and idb_companion (MIT, Meta).
This server cannot be deployed
Maintenance
Related MCP Connectors
Control real Android and iOS devices with LLM agents — tap, swipe, type, automate flows.
Build, run, and inspect iOS apps in disposable hosted Simulators from cloud coding agents.
Drive real Android & iOS devices and web browsers from natural language for mobile + web QA. 290+ tools across device control, app management, automation sessions, browser automation, and flow recording / replay. Bearer-auth — get a token at robotactions.com → Profile → API Tokens.
Drive real devices from your AI Coding tool. Embed a client SDK (Unity, Godot, Flutter, iOS/macOS, Android, React Native, Web) in your app, then capture screenshots, traverse the UI tree, inject taps and key events, and run automated test tasks on the physical device over a secure relay.
Related MCP Servers
- AlicenseAqualityAmaintenanceEnables interaction with iOS simulators by providing tools to inspect UI elements, control UI interactions, and manage simulators through natural language commands.175,176 npm2,185MIT
- AlicenseNot gradedqualityDmaintenanceEnables AI assistants to automate iOS Simulator interactions including device management, UI element interaction (tap, swipe, type), screenshot capture, and execution of YAML-defined navigation workflows.5 npmMIT
- AlicenseNot gradedqualityFmaintenanceEnables AI agents to control real iPhones and simulators on macOS, allowing for UI interaction, testing, and automation.49 npm88MIT
- AlicenseBqualityDmaintenanceEnables AI assistants to control iOS simulators through WebDriverAgent, supporting taps, swipes, typing, screenshots, recording, and app actions with a real-time dashboard for visual feedback.36Apache 2.0