@doublespeed/mcp
Build, test, and run iOS apps on a remote Mac fleet without needing a Mac or Xcode locally — plus local project discovery and diagnostics.
Discover projects locally:
discover_projsscans for.xcworkspace,.xcodeproj, andPackage.swift;list_schemesreads shared schemes.Validate without a key:
check_projectlints schemes, Swift syntax, and plists on a real Mac (10/day keyless trial) — no compile.Build:
build_simcompiles for the iOS Simulator (file:line:colerrors, optional.appartifact).Test:
test_simruns XCTest / Swift Testing with per-test results and supports test plans,onlyTesting/skipTesting, and configuration.Run live:
build_run_simbuilds, installs, launches, and returns a browser URL to an interactive simulator a human can tap/type/swipe;stop_app_simends the session.Archive:
archive_appproduces an unsigned Release.xcarchive.Inspect simulators and jobs:
list_simsshows fleet devices/runtimes/Xcode versions;job_statusreturns status, diagnostics, tests, and log tail.Diagnose connectivity:
doctorverifies key, API reachability, fleet capabilities, and CLI version.Choose environment: specify scheme, Xcode version, configuration, simulator name, and
workspaceRoot(content-addressed uploads; only changed files move after first run).
Enables building, testing, archiving, and running Expo-based iOS apps on a remote Mac fleet, including simulator execution and live simulator links.
Provides full iOS app lifecycle management—building, testing, archiving, and running apps on iOS simulators—using a remote Mac fleet without needing a local Mac or Xcode.
Supports building, testing, and running React Native apps on iOS simulators and devices through the remote Mac fleet, with per-test results and live simulator sessions.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@@doublespeed/mcpBuild my iOS app and open the live simulator."
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
@doublespeed/mcp
An MCP server that gives a coding agent a Mac. Build, test, archive and run iOS apps — native, React Native, Expo — on the doublespeed Mac fleet from Linux, a container, CI, or a Mac without Xcode.
Tool names match XcodeBuildMCP (build_sim, test_sim, build_run_sim, list_sims, discover_projs, list_schemes, stop_app_sim, doctor) so prompts and skills written for a local Mac work unchanged when there is no Mac.
Install
Needs Node 20+. check_project works with no key at all; everything else needs a DS_API_KEY (free $100 of Mac time at https://mac.doublespeed.ai/dashboard → API keys).
Claude Code
claude mcp add doublespeed -e DS_API_KEY=dsx_... -- npx -y @doublespeed/mcpCursor / Windsurf / Codex / VS Code — add to the client's MCP config:
{ "mcpServers": { "doublespeed": { "command": "npx", "args": ["-y", "@doublespeed/mcp"], "env": { "DS_API_KEY": "dsx_..." } } } }Related MCP server: xcforge
Tools
Tool | What it does | Where it runs |
| find .xcworkspace / .xcodeproj / Package.swift | local |
| shared schemes of a project/workspace | local |
| simulators, runtimes and Xcode versions on the fleet | API |
| no key needed — schemes resolve, every Swift file parses, plists valid (lint, no compile) | Mac |
| compile for the iOS Simulator; | Mac |
| build + XCTest / Swift Testing; per-test results | Mac |
| build, install, launch; returns a browser URL to a live simulator | Mac |
| end a live session | API |
| unsigned | Mac |
| status, diagnostics, tests, log tail of any job | API |
| key, API reachability, fleet capabilities | local + API |
Only files under workspaceRoot (default: the project's directory) are uploaded, content-addressed — after the first run only changed files move. A small app's test run takes ~20–40 s of Mac time and costs a few cents; nothing is billed while queued.
Not yet
Screenshots / accessibility tree of the running simulator, code signing and TestFlight upload. The live build_run_sim link is meant for a human to look at.
Agents without a Mac build, test and run iOS apps through mac.doublespeed.ai (npx -y @doublespeed/ds). Docs: https://mac.doublespeed.ai/docs.md
Available Tools
11 toolsarchive_appArchive (.xcarchive)A
Produce an unsigned .xcarchive for the device SDK on the fleet (Release). Signing and TestFlight upload are not available yet.
| Name | Required | Description | Default |
|---|---|---|---|
| scheme | No | Scheme to build; inferred when the project has exactly one app scheme | |
| extraArgs | No | Extra `ds xcode` flags, e.g. ['--setting', 'CODE_SIGNING_ALLOWED=NO', '--timeout', '3600'] | |
| projectPath | No | Path to the .xcodeproj (omit for a workspace or a Swift package) | |
| xcodeVersion | No | Required Xcode major or exact version, e.g. '16' or '16.4' | |
| configuration | No | Debug or Release (default: what the scheme defines) | |
| simulatorName | No | Simulator device name, e.g. 'iPhone 16 Pro' (default iPhone 16) | |
| workspacePath | No | Path to the .xcworkspace (CocoaPods / React Native projects) | |
| workspaceRoot | No | Directory to upload as the project root (default: the directory containing the project/workspace, or the current directory). Only files under it are sent. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden; it does state the artifact is unsigned, targets the device SDK, uses Release, and that signing/TestFlight upload are unsupported. However, it does not disclose whether execution is asynchronous, how results or artifacts are returned, or what happens when signing is required.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two concise sentences, front-loaded with the core action and followed by the most important limitation. Every sentence adds information and there is no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given 8 parameters, no output schema, and no annotations, the description provides essential context (unsigned, device SDK, Release, fleet) but omits operational details like whether the call is async or returns a job id. The sibling job_status suggests asynchronous behavior, yet the description does not mention this workflow.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline of 3 applies; the description adds no per-parameter detail beyond the schema. It aligns with the configuration parameter by calling out Release, but it doesn't clarify potential ambiguities such as whether simulatorName is relevant for a device SDK archive.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a concrete deliverable (unsigned .xcarchive), target (device SDK), environment (fleet), and configuration (Release), so an agent can tell exactly what action this performs. It also distinguishes itself from simulator-oriented siblings like build_sim and test_sim by emphasizing device SDK.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies this tool is for producing device archives and explicitly rules out signing/TestFlight upload ('not available yet'), but it does not explain when to choose this over build_sim/test_sim or note prerequisites such as needing a project/workspace. Usage guidance is mostly implied rather than explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
build_run_simBuild and run in a live simulatorA
Build, install and launch the app in an iOS Simulator on the fleet and return a browser URL where a human can see and use it (tap, type, swipe). The session stays open until stop_app_sim or the idle timeout. Billed per second while open.
| Name | Required | Description | Default |
|---|---|---|---|
| scheme | No | Scheme to build; inferred when the project has exactly one app scheme | |
| extraArgs | No | Extra `ds xcode` flags, e.g. ['--setting', 'CODE_SIGNING_ALLOWED=NO', '--timeout', '3600'] | |
| projectPath | No | Path to the .xcodeproj (omit for a workspace or a Swift package) | |
| xcodeVersion | No | Required Xcode major or exact version, e.g. '16' or '16.4' | |
| configuration | No | Debug or Release (default: what the scheme defines) | |
| simulatorName | No | Simulator device name, e.g. 'iPhone 16 Pro' (default iPhone 16) | |
| workspacePath | No | Path to the .xcworkspace (CocoaPods / React Native projects) | |
| workspaceRoot | No | Directory to upload as the project root (default: the directory containing the project/workspace, or the current directory). Only files under it are sent. | |
| idleTimeoutSeconds | No | End the session after this long with no viewer activity (60–7200, default 900) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral burden and does it well: it discloses session lifetime, idle timeout behavior, and per-second billing. It also implies the app remains installable/launchable in a remote simulator. It stops short of describing failure modes or whether the call waits for the build, but the main behaviors are transparent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences deliver the core purpose, the interactive nature, the session lifecycle, and the cost implication. The most important information is front-loaded and there is no fluff or repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 9-parameter tool with no output schema and no annotations, the description covers the key non-schema facts: browser URL output, human interactivity, session lifecycle, and billing. It could add response details like session identification or build-status behavior, but the provided information is enough for an agent to select and invoke the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so every parameter is already documented in the schema. The description itself adds no parameter-level meaning beyond establishing the high-level workflow, which is appropriate at the baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb chain ('Build, install and launch') with a clear resource ('the app in an iOS Simulator on the fleet') and a concrete outcome ('return a browser URL where a human can see and use it'). This clearly distinguishes it from siblings like build_sim, test_sim, and stop_app_sim.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description conveys the intended use case: interactive human use in a live simulator, with a browser URL. It also points to stop_app_sim as the way to close the session. It does not explicitly state when to choose this over build_sim or test_sim, but the interactive-session context is clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
build_simBuild for iOS SimulatorA
Compile the project for the iOS Simulator on a Mac in the doublespeed fleet. Returns file:line:col errors. Use when there is no Xcode locally.
| Name | Required | Description | Default |
|---|---|---|---|
| scheme | No | Scheme to build; inferred when the project has exactly one app scheme | |
| extraArgs | No | Extra `ds xcode` flags, e.g. ['--setting', 'CODE_SIGNING_ALLOWED=NO', '--timeout', '3600'] | |
| collectApp | No | Also upload the built .app as an artifact | |
| projectPath | No | Path to the .xcodeproj (omit for a workspace or a Swift package) | |
| xcodeVersion | No | Required Xcode major or exact version, e.g. '16' or '16.4' | |
| configuration | No | Debug or Release (default: what the scheme defines) | |
| simulatorName | No | Simulator device name, e.g. 'iPhone 16 Pro' (default iPhone 16) | |
| workspacePath | No | Path to the .xcworkspace (CocoaPods / React Native projects) | |
| workspaceRoot | No | Directory to upload as the project root (default: the directory containing the project/workspace, or the current directory). Only files under it are sent. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral burden. It discloses remote execution ('on a Mac in the doublespeed fleet') and error format ('Returns file:line:col errors'), which is useful. However, it does not mention side effects like artifact creation, possible long build times, or what a successful result looks like.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three short sentences with no filler. The first sentence states the core action and environment, the second states the error output format, and the third gives a clear usage condition. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is complex (9 optional parameters) and has no output schema, so the description should cover success behavior and invocation flow more fully. It covers purpose, error format, and one usage condition, but it leaves gaps such as whether the build is asynchronous, whether a job ID is returned, and what happens on success.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the input schema already documents all nine parameters, their defaults, and examples. The description adds no additional parameter-level meaning, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Compile the project for the iOS Simulator.' This clearly distinguishes it from sibling tools like test_sim, build_run_sim, and archive_app by limiting scope to simulator compilation rather than testing, running, or archiving.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives an explicit triggering condition: 'Use when there is no Xcode locally.' However, it does not explicitly name sibling alternatives such as build_run_sim or test_sim, nor does it state when not to use this tool, so it stops short of full 5-level guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
check_projectCheck project on a Mac (no API key needed)A
Lint-only check on a real Mac: resolves the project/workspace/package and schemes with xcodebuild -list, parses every Swift file with swiftc (syntax, file:line:col), lints plists. Does NOT compile. Works without DS_API_KEY (keyless trial: 10 per day per address, 25 MB snapshot). Use this to validate structure before asking the user for a key; use build_sim/test_sim for a real build.
| Name | Required | Description | Default |
|---|---|---|---|
| projectPath | No | Path to the .xcodeproj (omit for a workspace or a Swift package) | |
| workspacePath | No | Path to the .xcworkspace (CocoaPods / React Native projects) | |
| workspaceRoot | No | Directory to upload as the project root (default: the directory containing the project/workspace, or the current directory). Only files under it are sent. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden, and it does so well: it discloses the lint-only nature, the absence of compilation, the lack of DS_API_KEY requirement, and the rate/size limits of the keyless trial. It also reveals internal actions like xcodebuild -list and swiftc parsing, which goes beyond a generic 'check' description.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded with the most important differentiator: 'Lint-only check on a real Mac.' Every sentence earns its place—scope, non-behavior, keyless constraints, and sibling routing—with no filler or repetition of schema content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no output schema and no annotations, the description covers purpose, limitations, rate limits, and alternative tools well. The only notable gap is that it does not describe the return format or whether this is a synchronous result versus a job to be polled via job_status, but it is otherwise complete enough for invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema already provides 100% description coverage for all three parameters, so the bar for added value is lower. The description reinforces how the parameters relate to resolving projects/workspaces/packages, but it does not add meaningful parameter-level detail beyond what the schema already states.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description is explicit about what the tool does: a lint-only check on a real Mac that resolves projects/workspaces/packages and schemes via xcodebuild -list, parses Swift files with swiftc reporting file:line:col, and lints plists. It also clearly states what it does NOT do ('Does NOT compile'), which distinguishes it from build_sim/test_sim and other build-related siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The tool gives direct usage guidance: 'Use this to validate structure before asking the user for a key; use build_sim/test_sim for a real build.' It also explains the keyless trial constraints (10 per day per address, 25 MB snapshot), so an agent knows when this alternative is viable.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
discover_projsDiscover Xcode projectsA
Scan a directory for .xcworkspace, .xcodeproj and Package.swift (skipping node_modules, Pods, .build, DerivedData). Local; no Mac needed.
| Name | Required | Description | Default |
|---|---|---|---|
| maxDepth | No | ||
| workspaceRoot | No | Directory to scan (default: cwd) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral disclosure burden. It usefully discloses that the scan is local, requires no Mac, and skips heavy directories. However, it does not state whether the operation is read-only, how deep it scans by default, or what the return shape will be, leaving some behavioral ambiguity.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single dense sentence that front-loads the core action, names the target artifacts, lists exclusions, and adds the key 'local' constraint. Every element earns its place with no redundancy or filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a low-complexity discovery tool, the description covers the essential scope and exclusions. However, with no output schema and no annotations, it leaves out what the tool returns and what maxDepth controls, so an agent may need to infer important call details.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 50%, and the description does not compensate. workspaceRoot is documented in the schema, but maxDepth is left completely unexplained in both the schema and the description. The description adds context about file types and exclusions but no meaningful parameter-level semantics.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Scan'), a clear resource type (directories containing .xcworkspace, .xcodeproj, Package.swift), and adds distinctive scope ('skipping node_modules, Pods, .build, DerivedData'). It clearly differentiates itself from siblings like list_schemes or check_project by focusing on project discovery.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly implies when to use the tool: when you need to find Xcode project markers in a directory. It does not explicitly compare to alternatives or state exclusions, but the context is unambiguous enough for an agent to select it appropriately.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
doctorDoctorA
Check that this machine can reach the doublespeed fleet: key present, API reachable, fleet capabilities, ds CLI version.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It discloses that the tool checks key presence, API reachability, fleet capabilities, and ds CLI version, which is transpararent about its input checks. However, it does not explicitly state that the operation is read-only, what happens on failure, or whether it requires specific credentials. Since the description does not contradict any annotations (none provided), a mid-range score is appropriate because it gives some behavioral detail but not complete disclosure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, well-structured sentence that front-loads the main purpose and enumerates key checks via a colon. There is zero redundancy; every word contributes meaning. It is appropriately concise for a tool with no inputs.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a diagnostic tool with no inputs and no output schema, the description lists what it checks but does not describe the expected output format or how results are conveyed. Since there are no annotations to fill this gap, an agent might not know whether the tool returns a status code, a report, or throws exceptions. The description is adequate for understanding intent but lacks enough detail to fully anticipate the tool's behavior, so a mid-range score is warranted.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has no parameters, so the description does not need to explain parameter semantics. The baseline for zero-parameter tools is 4, and the description does not introduce confusion. It correctly makes no mention of inputs since there are none.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb 'Check' and the resource 'this machine can reach the doublespeed fleet', then enumerates specific aspects (key, API, fleet capabilities, CLI version). This distinguishes it from sibling tools like discover_projs or list_schemes, which are about listing or managing projects/schemes. The purpose is unambiguous and specific.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage when connectivity to the fleet needs verification, and sibling tools do not overlap with diagnostic connectivity checks. However, there is no explicit guidance on when to use this tool versus alternatives, nor any stated prerequisites or exclusions. The context makes it obvious, but the description does not directly address usage guidance, so it merits an average score.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
job_statusJob status / logsA
Fetch a previous job (status, error, diagnostics, tests) and optionally its log tail.
| Name | Required | Description | Default |
|---|---|---|---|
| jobId | Yes | ||
| logLines | No | Return the last N log lines (default 0) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden of behavior. It discloses the payload categories (status, error, diagnostics, tests) and optional log tail, which suggests a read-only inspection tool. It does not mention retention, permission requirements, or behavior when the jobId is invalid.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, front-loaded sentence contains both the core purpose and the optional log-tail behavior, with no filler or repetition of the title.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has no output schema and no annotations, so the description should fully carry return-value expectations. It names the major result components and the optional log tail, which is useful, but it omits how to obtain jobId and when this endpoint is preferable to sibling tools like check_project or doctor.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 50%: logLines is documented in the schema, and the description ties it to 'log tail'. However, jobId is only typed as string in the schema and the description does not clarify its format or source, so the agent must infer that it comes from a previously run job.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Fetch') with a clear resource ('a previous job') and enumerates what is returned (status, error, diagnostics, tests, optional log tail). This distinguishes it from siblings like build_run_sim or test_sim, which start or test jobs rather than inspect prior ones.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is given about when to call this tool versus alternatives, how a jobId is obtained, or when logLines should be used. The description implies retrieval of a past job but leaves the selection context to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_schemesList schemesA
List shared schemes of a project or workspace by reading xcshareddata (local). Unshared schemes are not visible; if none are listed, ds infers the scheme from the app target on the Mac.
| Name | Required | Description | Default |
|---|---|---|---|
| projectPath | No | ||
| workspacePath | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the behavioral burden. It discloses that the tool reads local xcshareddata, only shows shared schemes, and has a fallback inference behavior when none are listed. These are meaningful behavioral traits beyond what the tool name alone conveys.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two concise sentences with no filler. It front-loads the primary action and then adds two relevant behavioral caveats, making every word earn its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has no annotations, no output schema, and zero parameter description coverage, so the description must compensate. It covers key behavioral caveats but leaves parameter semantics unresolved and does not describe what the returned list looks like. An agent still cannot fully determine how to construct the required invocation when parameters are optional.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the description does not explain projectPath or workspacePath beyond the phrase 'project or workspace'. It vaguely suggests the two parameters correspond to project vs workspace, but it does not clarify whether they are alternatives, whether one is required, or how both should be handled. Since the schema provides no descriptions, this is a significant gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action and resource: 'List shared schemes of a project or workspace' and adds the mechanism 'by reading xcshareddata (local)'. This clearly differentiates the tool from sibling tools like discover_projs, check_project, and build_sim, making its purpose unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage context: call this when you need shared schemes from a local project or workspace. It provides relevant caveats ('Unshared schemes are not visible; if none are listed, ds infers the scheme from the app target') but does not explicitly name alternatives or say when not to use this tool versus siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_simsList simulatorsA
List simulator devices, iOS runtimes and Xcode versions available on the doublespeed fleet.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the transparency burden. The verb 'List' clearly indicates a read-only inspection operation with no side effects, and the phrase 'available on the doublespeed fleet' adds scoping context. It does not mention output format or potential latency, but these are minor for a simple zero-parameter list tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single front-loaded sentence with no filler: it starts with the verb and lists exactly what is returned. Every word contributes to meaning.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple zero-parameter list tool, the description is almost complete: it states the object being listed, the categories included, and the deployment context. It does not describe the return shape in detail, but no output schema exists and an agent can reasonably infer a list of simulators/runtimes/Xcode versions.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has zero parameters and 100% schema description coverage, so there are no parameter semantics for the description to clarify. The baseline for zero-parameter tools is 4, and the description correctly does not invent or imply nonexistent arguments.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('List') and names the exact resource categories: simulator devices, iOS runtimes, and Xcode versions. It also scopes the operation to the doublespeed fleet, making it clearly distinct from sibling action tools like build_sim, test_sim, and stop_app_sim.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The usage context is implied: an agent would call this to discover available simulators/runtimes/Xcode versions before running build or test actions. However, the description does not explicitly state when to use it, when not to use it, or name any alternative tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
stop_app_simStop a simulator sessionA
End a live simulator session started by build_run_sim (frees the Mac; stops billing).
| Name | Required | Description | Default |
|---|---|---|---|
| jobId | Yes | Job id returned by build_run_sim |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the behavioral disclosure burden. It explicitly reveals important side effects: 'frees the Mac; stops billing'. It does not mention whether the action is reversible or what happens if the jobId is invalid, but the core destructive consequence is disclosed clearly.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, information-dense sentence. It front-loads the action, names the prerequisite session, and adds high-value consequences without wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple one-parameter tool with no output schema, the description is nearly complete. It identifies the target session, the originating operation, and the practical effects. It could optionally mention what happens after the call or error behavior, but nothing essential is missing for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema already fully documents jobId as 'Job id returned by build_run_sim', so baseline is 3. The description adds minimal parameter meaning beyond reinforcing that the session must be live and started by build_run_sim. It does not describe formatting or validation rules, but none are needed given the simple one-parameter schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb and resource: 'End a live simulator session'. It also names the originating tool, build_run_sim, which distinguishes it from sibling tools like test_sim or job_status. The parenthetical 'frees the Mac; stops billing' adds concrete purpose.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly implies when to use this tool: after a simulator session was started by build_run_sim. It gives context but does not explicitly state exclusions such as 'do not use for sessions not started by build_run_sim' or alternatives. This is a strong contextual hint, though not a full routing guide.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
test_simRun tests on iOS SimulatorA
Build and run XCTest / Swift Testing on an iOS Simulator on the doublespeed fleet. Returns per-test results and compile errors.
| Name | Required | Description | Default |
|---|---|---|---|
| scheme | No | Scheme to build; inferred when the project has exactly one app scheme | |
| testPlan | No | ||
| extraArgs | No | Extra `ds xcode` flags, e.g. ['--setting', 'CODE_SIGNING_ALLOWED=NO', '--timeout', '3600'] | |
| onlyTesting | No | XCTest identifiers to run, e.g. ['AppTests/CounterTests'] | |
| projectPath | No | Path to the .xcodeproj (omit for a workspace or a Swift package) | |
| skipTesting | No | ||
| xcodeVersion | No | Required Xcode major or exact version, e.g. '16' or '16.4' | |
| configuration | No | Debug or Release (default: what the scheme defines) | |
| simulatorName | No | Simulator device name, e.g. 'iPhone 16 Pro' (default iPhone 16) | |
| workspacePath | No | Path to the .xcworkspace (CocoaPods / React Native projects) | |
| workspaceRoot | No | Directory to upload as the project root (default: the directory containing the project/workspace, or the current directory). Only files under it are sent. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It discloses the main behavior and return payload ('Returns per-test results and compile errors'), which is useful. It does not mention side effects, prerequisites, upload behavior, or failure semantics, but the core operation is transparent enough.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two concise sentences with no filler. The action is front-loaded, and the second sentence adds a valuable outcome detail. Every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description is adequate for understanding what the tool does and what it returns, and the schema handles most parameter documentation. However, without annotations or an output schema, and with 11 parameters and several sibling tools, the description leaves gaps around when to choose this tool and what the full execution lifecycle looks like.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 82%, so most parameters already carry their own documentation. The tool description itself adds no parameter-level meaning beyond what the schema provides. With high schema coverage, the baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'Build and run XCTest / Swift Testing on an iOS Simulator'. It also names the output ('per-test results and compile errors') and the fleet context, making the tool's purpose clear. It does not explicitly differentiate from siblings like build_sim or build_run_sim, so it misses the top score.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by the phrase 'Build and run XCTest / Swift Testing', which signals this is for running tests. However, the description does not explicitly say when to use this tool versus build_sim or build_run_sim, and it provides no exclusions or alternative routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
11 tool updates
v0.1.1- First observed
archive_app - First observed
build_run_sim - First observed
build_sim - First observed
check_project - First observed
discover_projs - First observed
doctor - First observed
job_status - First observed
list_schemes - First observed
list_sims - First observed
stop_app_sim - First observed
test_sim
TDQS
Scored across 11 tools
Each tool targets a distinct action: listing, building, testing, running, archiving, checking, or diagnosing. Even build_sim, build_run_sim, and test_sim are clearly separated by purpose and described outcomes.
Most tools follow a verb_noun pattern (stop_app_sim, discover_projs, list_schemes, check_project, build_sim, test_sim, archive_app). Minor deviations exist with job_status (noun_phrase) and doctor (single-verb command), but the overall style remains predictable.
11 tools is well-scoped for a remote iOS simulator build/test/run service. Each tool addresses a specific step or capability with no obvious redundancy.
The core lifecycle is covered: discovery, scheme/simulator listing, structural checks, build, test, run, stop, archive, job status, and connectivity diagnosis. Minor gaps like job cancellation or signed export are mentioned as future work, but they don't break primary workflows.
Maintenance
Related MCP Connectors
Build, run, and inspect iOS apps in disposable hosted Simulators from cloud coding agents.
Cloud macOS with Xcode for AI agents: run commands, build, test and ship iOS apps.
- LimrunOAuthcom.limrun
Cloud iOS simulators and Android emulators your agent can create, drive, and throw away.
Secure access to a dedicated Otherlay Mac for files, terminals, Git, builds and UI inspection.
Related MCP Servers
AlicenseNot gradedqualityBmaintenanceHigh-performance MCP server for iOS development and test automation. Gives AI coding assistants direct access to iOS simulators with sub-20ms screenshots, UI interaction, building, testing, and an intelligent operator mode.8MIT- AlicenseNot gradedqualityDmaintenanceMCP server and CLI for iOS development — build, test, automate, and diagnose from any AI agent or terminal.1MIT
- FlicenseNot gradedqualityDmaintenanceEnables AI to control iOS simulators through the MCP protocol. Supports device management, UI automation, and network interception including screenshot capture, text input, and HTTP request mocking.-
- AlicenseBqualityBmaintenanceA comprehensive MCP server for iOS/macOS development that gives AI assistants full control over Xcode build, test, simulator management, and app deployment.657 npm2MIT