NoMac
Server Details
A real cloud Mac with Xcode for your AI agent. Start a macOS VM, run commands (git clone, xcodebuild, simulator tests), read the output and stop it; billed from prepaid credit at $0.80 an hour by the minute. Optional App Store Connect tools upload builds to TestFlight and manage the listing, with previews before any change. OAuth sign-in, nothing to paste.
- Status
- Healthy
- Last Tested
- Transport
- Streamable HTTP · MCP 2025-11-25
- URL
TDQS
Scored across 38 tools
Most tools have clearly distinct resource+action scopes, and the descriptions explicitly tell the agent when to use each (e.g. get_store_operation vs get_store_capabilities vs read_store). A few boundaries are softer: read_store/update_store/get_store_capabilities/get_store_operation all orbit Apple JSON:API, and upload_screenshots vs upload_store_asset overlap on asset uploads, but the docs disambiguate them well.
Strong, predictable snake_case with verb-led prefixes (get_*, set_*, list_*, start_*/stop_*, update_*, read_*, upload_*). A handful of bare names (build, status, publish, review_lint) deviate from the verb_noun pattern but remain readable and unambiguous in context.
38 tools is on the heavy side, though the server spans a legitimately broad surface (builds, TestFlight, App Store review, metadata, screenshots, Mac VMs, billing). Many tools earn their place, but the count is high enough that consolidation (e.g. the four Apple read/write/capability/store-operation tools) would improve scanability.
The surface covers a full lifecycle: push → check → build → status → failure → publish → review → TestFlight → feedback, plus billing, metadata, assets and Mac VM management, with recovery/status tools for lost responses. Minor gaps exist where Apple's public API can't act (reviewer message send/read are documented browser handoffs) and there's no explicit build cancel, but no significant dead ends.
Available Tools
38 toolsbuildStart a buildAInspect
Start a build for a project. workflow=release (signed → TestFlight; uses one build from the plan quota) or smoke (unsigned compile check). Poll with status.
| Name | Required | Description | Default |
|---|---|---|---|
| workflow | No | release (default): signed build uploaded to TestFlight, uses one build from the plan. smoke: unsigned compile check. | release |
| project_id | Yes | NoMac project ID (prj_…). push_project and connect_status list your projects. | |
| request_key | No | Reuse to reconcile a lost response. Use a new key only after confirmed no build started. |
Output Schema
| Name | Required | Description |
|---|---|---|
| error | No | Present only when the call failed |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare this is a non-read-only, non-idempotent, non-destructive write, and the description adds genuinely useful behavior beyond them: that a release build consumes one build from the plan quota and that release output is signed and lands in TestFlight. It omits failure/log-retrieval behavior, but the quota cost is meaningful context annotations cannot convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, front-loaded with the action and followed by the mode distinction and the polling hint. No filler and no redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need no explanation, and the description covers the two modes, their side effects, cost, and how to check progress. It leaves unanswered what happens on a failed build and which sibling exposes build failures, which would round it out.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the schema already documents project_id, workflow, and the request_key reconcile semantics in detail. The description largely restates the workflow definitions the schema already provides, adding little beyond the baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Start a build for a project') and immediately subdivides into two named modes, release vs smoke. It is clearly distinguishable from peers like push_project or publish, though it never explicitly contrasts itself with them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear selection context for the two workflows (signed TestFlight upload vs unsigned compile check) and routes the agent to the 'status' sibling for polling. It stops short of stating when not to use the tool or how it relates to publish/push_project in a full workflow.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
check_projectCheck project readinessAInspect
Free source-readiness check of the latest uploaded snapshot; no Apple calls or paid compute. Run push_project after source changes. Does not prove compilation, signing or TestFlight delivery.
| Name | Required | Description | Default |
|---|---|---|---|
| project_id | Yes | NoMac project ID (prj_…). push_project and connect_status list your projects. |
Output Schema
| Name | Required | Description |
|---|---|---|
| error | No | Present only when the call failed |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Adds real behavioral context beyond the annotations: it is free, makes no Apple calls or paid compute, and crucially discloses the boundary 'Does not prove compilation, signing or TestFlight delivery'. This negative scope is genuinely useful and not present in the structured fields. It does not address why readOnlyHint is false, leaving a minor gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences: the core purpose is front-loaded, then the follow-up action and the limitation. Every clause earns its place with no redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need no explanation. The description covers purpose, cost profile, boundary conditions, and the follow-up tool, which is sufficient for a one-parameter check. Minor omission is any note on whether it mutates state given readOnlyHint=false.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With a single parameter at 100% schema coverage, the schema already documents project_id fully (including where to find it). The description adds no syntax or format detail beyond that, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: a source-readiness check of the latest uploaded snapshot. The scope qualifiers ('free', 'no Apple calls or paid compute') further pin down what kind of check this is, distinguishing it from build/signing siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly routes the agent: 'Run push_project after source changes' names the sibling and the triggering condition. It stops short of stating when NOT to use check_project (e.g. versus build), so it lacks the full alternatives/exclusions treatment that would earn a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
connect_statusCheck Apple connectionARead-onlyInspect
Check the Apple App Store Connect connection: whether the API key works, when the signing certificate expires, and webhook health. Start here if any build, TestFlight or App Store tool fails; with no connection it explains the one-time setup for the human. Read-only.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| error | No | Present only when the call failed |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and destructiveHint=false, so the description's 'Read-only' adds no new safety signal. However, it does add useful context beyond the structured fields: what the check reports and, importantly, that with no connection it surfaces one-time human setup instructions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tightly written sentences, front-loaded with the core purpose before the usage trigger and the no-connection behavior. No filler or repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so the description is not obligated to describe return values, and it covers purpose, when to invoke, and the diagnostic-failure path. For a zero-parameter read-only check, nothing an agent needs to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so the baseline of 4 applies. There is no parameter syntax or format the description needs to compensate for, and it correctly focuses on behavior rather than inputs.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource (check the Apple App Store Connect connection) and enumerates exactly what is verified: API key validity, signing certificate expiry, and webhook health. This clearly distinguishes it from write-oriented siblings like set_project_connection.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit when-to-use guidance: 'Start here if any build, TestFlight or App Store tool fails.' It names concrete trigger conditions (build failure, missing connection) and describes the fallback path (explains one-time setup), leaving nothing to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
exec_macRun a command on a MacADestructiveInspect
Run arbitrary argv in your Mac workspace. Returns a durable job ID; poll get_mac_job and read_mac_output. Preserve request_key and exact arguments after lost responses. Use bash -lc explicitly for shell syntax. timeout_seconds defaults to 300 and caps at 600; raise it for slow work such as a first simulator boot, which needs about 90 seconds.
| Name | Required | Description | Default |
|---|---|---|---|
| argv | Yes | Command and arguments as an array, for example ["bash", "-lc", "xcodebuild -list"]. Runs as the agent user in the workspace directory. | |
| session_id | Yes | Mac session ID (ses_…) returned by start_mac or list_mac_sessions. | |
| request_key | Yes | Idempotency key you choose (letters, digits, dash, underscore). Retry a lost response with the same key and identical arguments; use a new key for new work. | |
| timeout_seconds | No | Stop the command after this many seconds (1-600, default 300). |
Output Schema
| Name | Required | Description |
|---|---|---|
| error | No | Present only when the call failed |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover destructive/openWorld, and the description adds the facts an agent actually needs: the call is asynchronous and returns a durable job ID rather than output, timeout defaults to 300 and caps at 600, ~90s is expected for a first simulator boot, and request_key + identical args is the recovery path after a lost response. This is genuinely additive behavioral context, not restatement.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four tightly packed sentences with the core action front-loaded, followed by the follow-up tools, correctness rule, and timeout behavior in priority order. No filler or repetition of the title.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The async job model, the full retrieval path (get_mac_job/read_mac_output), timeout bounds, retry semantics, and shell invocation guidance are all present, and an output schema exists so return-value detail is not needed here. Nothing required to call this tool correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3, but the description adds real meaning beyond it: the bash -lc guidance for shell syntax and the simulator-boot rationale for raising timeout_seconds go past what the schema fields say. The idempotency explanation for request_key largely duplicates the schema text, which keeps this from a 5.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a concrete verb+resource ("Run arbitrary argv in your Mac workspace") and immediately distinguishes itself from siblings by naming get_mac_job and read_mac_output as the polling/read path. An agent can tell this is the execution entry point in the Mac session family without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear operational routing (poll get_mac_job, read_mac_output) and a specific usage rule (use bash -lc explicitly for shell syntax) plus a sizing heuristic for timeout_seconds on slow work like a simulator boot. It does not state when NOT to use the tool or what alternative exists for non-argv operations, so it falls short of a full 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_billing_statusGet build plan statusARead-onlyInspect
Read confirmed payment, contracted price, quota/reset, renewal and guarantee. After backend payment confirmation, resume a blocked build once with a new request_key only if the earlier response said build_started:false. Reconcile unknown build outcomes first.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| error | No | Present only when the call failed |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, openWorldHint=false and destructiveHint=false, so the safety profile is covered structurally. The description adds which fields are read and a recovery precondition, but the 'resume a blocked build once with a new request_key' instruction is ambiguous on a read-only tool and is not clearly scoped to this tool rather than the build sibling.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The returned-data list is front-loaded and compact, which is good. However, the trailing workflow sentences (resume the build, reconcile outcomes) are dense and read as belonging to the build tool rather than a status read, which muddies the structure.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, return values need not be explained, and with zero parameters there is no input contract to document. The description supplies the field inventory and a workflow precondition, leaving it largely complete for a zero-param read tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There are zero input parameters, so the baseline is 4 and the schema cannot be under-served. The mention of a 'new request_key' is not a parameter here and could briefly confuse the agent, but no parameter semantics are actually missing.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening clause names a specific read verb and enumerates the resource fields it returns (payment, contracted price, quota/reset, renewal, guarantee), so the agent knows this is a billing/plan status read. It does not, however, distinguish itself from billing-adjacent siblings like get_checkout_link or get_refund_status, and the title ('Get build plan status') drifts from the name ('get_billing_status').
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The second sentence gives a conditional workflow ('resume a blocked build once ... only if the earlier response said build_started:false') and a precondition ('reconcile unknown build outcomes first'), which implies context. But it never tells the agent when to call this tool versus the checkout/refund/build siblings, and the guidance is tangled.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_checkout_linkGet build plan checkout linkAInspect
Return the org-bound checkout URL. The human chooses a plan and pays explicitly; this call cannot charge them. Reuses an existing checkout; existing subscribers receive account management.
| Name | Required | Description | Default |
|---|---|---|---|
| plan | No | Builder plan to buy: starter or pro. Omit to reuse an existing checkout. |
Output Schema
| Name | Required | Description |
|---|---|---|
| error | No | Present only when the call failed |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare non-read-only, non-idempotent, non-destructive, but the description adds genuinely non-obvious semantics: the call cannot charge the human, payment is explicit and human-driven, and an existing checkout is reused rather than recreated. That materially changes how an agent should treat the result, though it says nothing about expiration of the returned URL or auth requirements.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the return value, then behavior. Every sentence earns its place with no filler. Slightly clipped phrasing ('existing subscribers receive account management') costs it the top mark.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With only one optional enum parameter, a full schema, an output schema, and annotations present, the description carries what it needs to: what is returned and the non-charging, human-driven nature of the flow. The gap is routing among billing/checkout siblings, which an agent must guess.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the single enum parameter is fully documented, so the baseline is 3. The description's 'omit to reuse an existing checkout' merely restates what the schema already says about the optional plan parameter, adding no new meaning.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource — 'Return the org-bound checkout URL' — so the agent knows this yields a URL for a checkout flow. It is distinguishable from siblings like get_billing_status and get_mac_checkout by the 'org-bound' and plan-purchase framing, though it never names which sibling to prefer.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied rather than stated: the description explains that a human picks a plan and pays, and that existing subscribers get account management, which hints at the right scenario. But it never explicitly says when to call this instead of get_billing_status or get_mac_checkout, and gives no prerequisites.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_failureExplain a failed buildARead-onlyInspect
Structured failure for a failed build: {stage, code, kind, error_lines, hints}. kind=project → fix source and re-push; kind=account → relay the deep link to your human.
| Name | Required | Description | Default |
|---|---|---|---|
| build_id | Yes | NoMac build ID (bld_…) returned by build. |
Output Schema
| Name | Required | Description |
|---|---|---|
| error | No | Present only when the call failed |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, destructiveHint=false, openWorldHint=false, so the safety profile is covered. The description adds genuinely useful behavioral context the annotations lack: the exact response fields and an interpretation contract for kind values that tells the agent how to act on the output.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two dense sentences, front-loaded with the return shape followed by the action semantics. No filler, no restatement of the title.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present and one fully documented required parameter, the description is nearly complete for a read-only diagnostic call. Enumerating the return fields is mildly redundant with the output schema, but the kind→action mapping is the piece the structured data alone could not convey.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the single build_id parameter already documents the bld_… format from build. The description contributes no additional parameter syntax or format detail, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific resource and scope: a structured failure object {stage, code, kind, error_lines, hints} for a failed build. That is concrete and distinguishable from generic siblings like status. It does not explicitly name which sibling it replaces, keeping it just short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by 'for a failed build' and the description gives post-call routing for the two kind values (project → fix source, account → relay deep link). It never states when to call this versus siblings such as status or get_mac_job, so guidance is only partially covered.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_feedbackGet TestFlight feedback and crashesARead-onlyInspect
Read recent TestFlight crash reports and written tester feedback for the project's app. Use it after testers install a build, to find crashes and complaints from real devices before the next build. Read-only.
| Name | Required | Description | Default |
|---|---|---|---|
| project_id | Yes | NoMac project ID (prj_…). push_project and connect_status list your projects. |
Output Schema
| Name | Required | Description |
|---|---|---|
| error | No | Present only when the call failed |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, destructiveHint=false, and openWorldHint=false, so the safety profile is fully covered by structured data. The description's 'Read-only' merely repeats that, and the only extra behavioral signal is 'recent' (an implied time window) with no detail on volume, pagination, or limits. Modest added value over annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences are front-loaded with the resource and the usage trigger, with zero filler. The trailing 'Read-only.' is slightly redundant given readOnlyHint=true, a minor waste that keeps it just under a 5.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be documented, and the description covers what is read and when to call it. It is complete enough for an agent to invoke correctly, though it says nothing about the scope of 'recent' or result volume.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There is a single required parameter (project_id) and schema description coverage is 100%, with the schema itself pointing users to push_project/connect_status for discovery. Baseline 3 is appropriate since the description adds no parameter meaning beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a concrete verb ('Read') and two specific resources (TestFlight crash reports and written tester feedback) scoped to the project's app, which clearly separates it from write-oriented siblings like manage_testflight. It does not explicitly name the nearest alternative (get_testflight) to disambiguate, so it falls short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Use it after testers install a build, to find crashes and complaints from real devices before the next build' gives a clear trigger and intent for the tool. No explicit when-not conditions or named alternative tools are provided, so it is clear context without exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_mac_checkoutGet Mac credit payment linkAInspect
Get a payment link for Mac credit, for the human to open. Credit is $0.80 an hour, billed by the minute, sold in $10 packs (12.5 hours each), 1 to 10 packs per purchase, and never expires. Call this when get_mac_credits shows too little credit or start_mac returns insufficient_credit. Then give the human the returned url and ask them to open it and pay; if you can run a shell, also open it for them (macOS: open URL, Linux: xdg-open URL, Windows: start URL). Never enter payment details yourself. Poll get_mac_credits until the balance appears, then start the Mac. Reuse request_key after a lost response; use a new one for a new purchase.
| Name | Required | Description | Default |
|---|---|---|---|
| packs | No | Number of $10 packs, 1-10. Pick from the human's stated budget or expected hours; default 1. | |
| request_key | Yes | Idempotency key you choose (letters, digits, dash, underscore). Retry a lost response with the same key and identical arguments; use a new key for new work. |
Output Schema
| Name | Required | Description |
|---|---|---|
| error | No | Present only when the call failed |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Goes well beyond the annotations: states pricing and billing mechanics ($0.80/hr, per-minute, $10 packs, 1-10 packs, never expires), the idempotency contract ('reuse request_key after a lost response; use a new one for a new purchase'), platform-specific open commands, and an explicit safety rule ('never enter payment details yourself'). This is exactly the extra behavioral context annotations cannot carry.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the action and then packed with genuinely load-bearing facts (price, pack sizing, workflow, shell commands, retry rule); no filler sentences, though it is on the dense side with several clauses chained together.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists so return values need not be described, yet the description still tells the agent what to do with the returned url and how the operation fits between get_mac_credits, start_mac, and payment. Nothing needed to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3; the description adds economic meaning to the 'packs' parameter ('$10 packs (12.5 hours each), 1 to 10 packs per purchase, and never expires') that helps an agent size the request, and reinforces the request_key retry semantics beyond the schema wording.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Get a payment link for Mac credit') plus its intended recipient ('for the human to open'), which distinguishes it from the generic sibling get_checkout_link and from read-only credit tools like get_mac_credits.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit trigger conditions ('when get_mac_credits shows too little credit or start_mac returns insufficient_credit') and names the alternatives by name, then lays out the follow-up sequence (give url, poll get_mac_credits, then start the Mac). When-to-use and what-to-do-next are both fully specified.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_mac_creditsGet Mac credit balanceARead-onlyInspect
Read the prepaid Mac credit balance: spendable dollars, spendable_seconds (how long a Mac can run on it), amounts held by a running session, pricing and recent purchases. Credit never expires. If spendable_seconds is below what the task needs, call get_mac_checkout.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| error | No | Present only when the call failed |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover safety (readOnlyHint=true, destructiveHint=false, openWorldHint=false), so the description is not required to restate them. It still adds real context beyond structured data: credit never expires, and the balance includes amounts held by an in-flight session, which explains why spendable_seconds may differ from the raw balance.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the core purpose and the enumerated fields, then the business rule, then the escalation path. No restatement of the title and no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be fully specified, yet the description still names the key fields an agent should inspect. Combined with the never-expires rule and the get_mac_checkout fallback, an agent has everything needed to call and act on this tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so the baseline of 4 applies. The description correctly spends no space on argument semantics and instead documents the shape of the result.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a precise verb+resource ('Read the prepaid Mac credit balance') and enumerates the concrete payload fields: spendable dollars, spendable_seconds, amounts held by a running session, pricing, and recent purchases. It also names a sibling (get_mac_checkout), so an agent can distinguish it without opening another schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a clear conditional: if spendable_seconds is below what the task needs, call get_mac_checkout. That is genuine routing guidance, though it doesn't contrast against nearby read tools like get_billing_status or get_mac_session, so the 'when not to use this' dimension is only partially covered.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_mac_jobGet Mac command statusARead-onlyInspect
Check a command started by exec_mac: running, exited or failed, with exit code and output size. Poll it until the job finishes, then read the output with read_mac_output. Never reruns the command. Read-only.
| Name | Required | Description | Default |
|---|---|---|---|
| job_id | Yes | Command job ID (job_…) returned by exec_mac. | |
| session_id | Yes | Mac session ID (ses_…) returned by start_mac or list_mac_sessions. |
Output Schema
| Name | Required | Description |
|---|---|---|
| error | No | Present only when the call failed |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and destructiveHint=false; the description reinforces this and adds a genuinely useful behavioral fact: 'Never reruns the command.' It also implies safe repeated polling. idempotentHint=false is consistent with polling (results change from running to exited), so no contradiction. Lacks any note on pacing/backoff, which keeps it short of a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three compact sentences, front-loaded with what is checked, then the workflow, then the safety guarantee. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Output schema exists, so return-value details need not be explained, and the description still hints at exit code and output size. Combined with the hand-off to read_mac_output and the strict no-rerun guarantee, an agent has everything needed to call and sequence it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% with only two required params, so the schema already documents job_id and session_id fully, including their provenance formats. The description adds provenance context (job IDs come from exec_mac) but nothing beyond schema. Baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (check) plus the resource (a command started by exec_mac) and enumerates the observable states (running, exited, failed) plus returned fields (exit code, output size). This is clearly distinguishable from siblings like exec_mac and read_mac_output.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly instructs to poll until the job finishes and then hand off to read_mac_output, naming the alternative and the condition that selects it. No inference required.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_mac_sessionGet Mac sessionARead-onlyInspect
Read one Mac session: state, funded time, cost so far and whether cleanup is confirmed. Poll it after start_mac until state is ready (usually under a minute), and after stop_mac until cleanup_confirmed is true. queued or provisioning means wait; stopping is not yet confirmed cleanup. Read-only.
| Name | Required | Description | Default |
|---|---|---|---|
| session_id | Yes | Mac session ID (ses_…) returned by start_mac or list_mac_sessions. |
Output Schema
| Name | Required | Description |
|---|---|---|
| error | No | Present only when the call failed |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and destructiveHint=false, so "Read-only" is largely redundant. However, the description adds genuine behavioral context beyond the annotations: expected timing ("usually under a minute"), polling cadence, and the meaning of transitional states.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, front-loaded with what is returned, then the polling lifecycle, then state interpretation. No filler; every clause carries operational meaning.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return structure needn't be re-explained, yet the description still orients the agent on which fields matter (state, cleanup_confirmed). Combined with read-only annotations and clear polling guidance, nothing needed to invoke this correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the single session_id parameter is fully documented in the schema with pattern and origin (returned by start_mac or list_mac_sessions). The description adds no syntax or format detail beyond that, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ("Read one Mac session") and enumerates exactly what it returns: state, funded time, cost so far, cleanup confirmation. It is clearly distinguishable from list_mac_sessions and read_mac_output among the siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit when/when-not routing: poll after start_mac until state is ready, poll after stop_mac until cleanup_confirmed is true. It also decodes intermediate states (queued/provisioning = wait; stopping is not yet confirmed cleanup), so the agent knows how to interpret responses without guessing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_metadataGet App Store metadataARead-onlyInspect
Read the App Store listing for one iOS version: description, keywords, what's new, URLs, category and review contact. Defaults to the marketing version of the last pushed source. Use it before editing with set_metadata. Read-only.
| Name | Required | Description | Default |
|---|---|---|---|
| project_id | Yes | NoMac project ID (prj_…). push_project and connect_status list your projects. | |
| version_string | No | App Store version, for example 1.2.0. Defaults to the marketing version of the last pushed source. |
Output Schema
| Name | Required | Description |
|---|---|---|
| error | No | Present only when the call failed |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, destructiveHint=false, and openWorldHint=false, so the safety profile is covered; the description's 'Read-only' is largely redundant. The genuinely additive detail is the default-resolution behavior ('Defaults to the marketing version of the last pushed source'), which explains what version is actually read when the parameter is omitted.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short, front-loaded sentences with no filler: scope, default behavior, and the routing hint appear in that order. Every clause carries information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter read tool with a 100% covered schema and an output schema, the description supplies everything an agent needs. Listing the returned fields is mildly redundant given the output schema, but it helps the agent confirm the tool's scope.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and both parameters are documented in the schema, so the baseline is 3. The description's note about defaulting to the marketing version restates the version_string schema description rather than adding new semantics, and the field list refers to outputs rather than inputs.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Read) plus resource (the App Store listing for one iOS version) and even enumerates the fields returned (description, keywords, what's new, URLs, category, review contact). It is clearly separable from siblings set_metadata and get_metadata_schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly routes the agent: 'Use it before editing with set_metadata,' establishing the read-before-write pairing with a named sibling. It gives clear context but stops short of stating when this is not the right tool (e.g., vs. read_store or get_metadata_schema).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_metadata_schemaGet App Store metadata rulesARead-onlyInspect
Get the rules for App Store metadata: every editable field with its character limit, supported locales, and screenshot display types with their exact pixel sizes. Call it before set_metadata or upload_screenshots. Read-only.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| error | No | Present only when the call failed |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and destructiveHint=false, and the trailing 'Read-only' merely restates that, earning no credit. The description adds useful framing that this is a static rule set to consult prior to mutation, but says nothing about auth needs, caching/staleness, or rate limits beyond what annotations cover.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, zero waste, and the payload contents are front-loaded before the usage instruction. The closing 'Read-only' is the only mildly redundant clause.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, the description needn't enumerate return values in detail, and annotations already cover the safety profile for a zero-parameter call. Nothing an agent needs in order to invoke this correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so the schema carries no semantics to add; baseline 4 applies. The description correctly implies a parameterless read that needs no input.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb + resource ('Get the rules for App Store metadata') followed by an enumeration of exactly what is returned: editable fields, character limits, supported locales, and screenshot pixel sizes. This clearly distinguishes it from siblings like get_metadata, get_store_capabilities, and set_metadata.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Call it before set_metadata or upload_screenshots' gives an explicit precondition and names the two sibling tools it precedes, which is strong routing guidance. It stops short of stating when not to call it, so it isn't a full 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_refund_statusGet refund statusARead-onlyInspect
Read the progress of a refund the account owner requested at nomac.app/usage. Agents cannot request refunds; point the human to that page. Does not submit a refund or cancel anything.
| Name | Required | Description | Default |
|---|---|---|---|
| request_id | Yes | Refund request ID. Refunds are started by the account owner at nomac.app/usage. |
Output Schema
| Name | Required | Description |
|---|---|---|
| error | No | Present only when the call failed |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and destructiveHint=false, but the description reinforces this by explicitly disclaiming side effects ('Does not submit a refund or cancel anything') and stating the agent's capability boundary. It does not address the idempotentHint=false flag or any auth/rate-limit behavior, so it is strong but not exhaustive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the read action and scoping, followed by the negative capability and side-effect disclaimers. No filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need no explanation. For a single-parameter read tool, the description covers purpose, ownership, and the critical boundary that agents cannot initiate refunds, leaving nothing an agent needs to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The single request_id parameter is fully documented in the schema (100% coverage), including where the refund originates. The description repeats that origin context but adds no new format, validation, or lookup semantics beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Read the progress of a refund') plus precise scope ('the account owner requested at nomac.app/usage'). It is clearly distinguishable from billing siblings like get_billing_status or get_checkout_link without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly scopes who this is for and what it is not for: agents cannot request refunds and should redirect the human to nomac.app/usage. It gives clear context but never names a sibling alternative by name for related needs (e.g. where to check billing state instead).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_reviewGet App Review statusARead-onlyInspect
Read live App Review submissions and all item states, including submissions created outside nomac. Pass asc_submission_id for version/build/contact/attachment details. Apple does not expose reviewer correspondence through its public API; the result provides an explicit App Store Connect handoff.
| Name | Required | Description | Default |
|---|---|---|---|
| project_id | Yes | NoMac project ID (prj_…). push_project and connect_status list your projects. | |
| asc_submission_id | No | Apple App Review submission ID, from get_review. |
Output Schema
| Name | Required | Description |
|---|---|---|
| error | No | Present only when the call failed |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint/destructiveHint/openWorldHint, so the safety profile is covered. The description still adds real value beyond them: it discloses the scope ('submissions created outside nomac') and an important limitation ('Apple does not expose reviewer correspondence through its public API'), which the agent could not learn from annotations or schema alone.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, front-loaded with the core read action and scoping, with no filler. Slightly dense but every sentence carries distinct information (scope, parameter effect, API limitation).
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be described, and the description covers scope, the optional parameter's effect, and a known data limitation. The main missing piece is routing guidance against the review/testflight siblings, which keeps it short of a 5.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds meaning the schema lacks: it explains that asc_submission_id unlocks version/build/contact/attachment details, whereas the schema only labels it as an ID sourced 'from get_review'. This makes it clear why and when to pass the optional parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Read live App Review submissions and all item states') and even extends scope to submissions created outside nomac. It does not, however, name or distinguish itself from close siblings like get_testflight, update_review, or prepare_review_reply, so the agent must infer the boundary.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The line 'Pass asc_submission_id for version/build/contact/attachment details' implies when to supply the optional parameter, and 'read live' implies a status-check use case. But there is no explicit when-to-use/when-not guidance and no named alternative among the review-related siblings, so usage is only implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_screenshot_uploadGet screenshot upload statusARead-onlyInspect
Check screenshot replacement or restoration. Omit operation_id to list recent operations after a lost upload response. complete confirms Apple processed every image; rolled_back means the previous set was retained/restored. needs_attention requires support reconciliation.
| Name | Required | Description | Default |
|---|---|---|---|
| project_id | Yes | NoMac project ID (prj_…). push_project and connect_status list your projects. | |
| operation_id | No | Operation ID from upload_screenshots. Omit to list recent operations. |
Output Schema
| Name | Required | Description |
|---|---|---|
| error | No | Present only when the call failed |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and destructiveHint=false, so safety is covered; the description adds genuinely useful semantics for the outcome states (complete/rolled_back/needs_attention) and the escalation path for needs_attention. It does not address idempotentHint=false or listing limits, keeping it below 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences: purpose first, then the invocation-mode note, then the status glossary. Every clause carries information and nothing is padded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists so return values need not be described, and the description covers purpose, mode selection, and status meanings. Minor gap: no mention of pagination or how many 'recent operations' are returned, which matters for the listing mode.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both parameters are already documented, and the description's 'Omit operation_id to list recent operations' essentially restates the schema text. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (check) and resource (screenshot replacement/restoration), and the enumeration of status outcomes makes the operation-status nature unambiguous. It does not explicitly name or distinguish itself from upload_screenshots, which is the closest sibling, so it falls short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a concrete trigger ('after a lost upload response') and explains the omit-operation_id listing mode, which tells the agent when each invocation style applies. No explicit when-not-to-use clause or named alternative, so not a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_store_capabilitiesList App Store Connect operationsARead-onlyInspect
Discover supported App Store Connect operations for listing, paid pricing/availability, subscriptions, in-app purchases, public customer-review replies, phased releases, events, product pages, files, analytics and webhooks. Pass category/search to list; pass operation for its exact Apple JSON:API schema. Review messages/appeals remain an App Store Connect browser workflow.
| Name | Required | Description | Default |
|---|---|---|---|
| offset | No | Pagination offset for long lists. | |
| search | No | Free-text search over operation names and summaries. | |
| category | No | List operations in one area, for example pricing, subscriptions, testflight or analytics. | |
| operation | No | Exact operation name, for example apps_getInstance, to get its full request schema. |
Output Schema
| Name | Required | Description |
|---|---|---|
| error | No | Present only when the call failed |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, openWorldHint=false and destructiveHint=false, so the safety profile is covered. The description adds the useful scope exclusion about review messages/appeals, but says nothing about pagination limits or the closed nature of the capability list beyond what annotations imply.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, front-loaded with the purpose and then the key usage instruction. The long enumeration of supported areas is dense but informative; the only mild cost is the padded category list.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, return values need not be described, and the description covers purpose, usage modes and scope exclusions. What remains thin is pagination/rate behavior for the offset-based listing, but that is a minor gap for a read-only discovery tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3, but the description adds meaning beyond the schema by explaining the two modes of use — list mode (category/search) versus schema-retrieval mode (operation) — which clarifies the semantic role of each parameter rather than just restating its type.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource — discovering/listing supported App Store Connect operations — and enumerates the covered areas concretely. It is broadly distinguishable from siblings like get_store_operation, though the overlap around fetching an operation's schema is not explicitly disambiguated.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives explicit invocation guidance ('pass category/search to list; pass operation for its exact Apple JSON:API schema') and draws a clear out-of-scope boundary ('Review messages/appeals remain an App Store Connect browser workflow'). It does not name sibling alternatives, so it falls short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_store_operationGet saved App Store operationARead-onlyInspect
Read saved publish/metadata/review/TestFlight requests and recover a lost response or request_key. Resume pending work by repeating the original tool arguments with its returned request_key. This status call performs no Apple writes.
| Name | Required | Description | Default |
|---|---|---|---|
| project_id | Yes | NoMac project ID (prj_…). push_project and connect_status list your projects. | |
| request_key | No | Idempotency key you choose (letters, digits, dash, underscore). Retry a lost response with the same key and identical arguments; use a new key for new work. |
Output Schema
| Name | Required | Description |
|---|---|---|
| error | No | Present only when the call failed |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and destructiveHint=false, so the safety profile is covered; the description still adds value by framing it as a status call with no Apple writes and by explaining the lost-response recovery workflow. It doesn't describe the shape/pagination of returned operation records, but the output schema handles that.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences, front-loaded with the read scope, then the recovery recipe, then the reassurance about no writes. Each sentence carries distinct information with no redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present and full annotation coverage, the description supplies the one thing structured fields can't: the retry/recovery workflow that motivates the call. Only minor gaps remain, such as whether multiple stored operations can be listed versus a single lookup by key.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and both project_id and request_key are fully documented in the schema, so the description need not carry parameter detail. It reinforces request_key retry semantics but adds nothing the schema lacks, matching the baseline for complete schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (read) and resource (saved publish/metadata/review/TestFlight requests), enumerating the operation categories it retrieves. It is clear enough to differentiate this record-retrieval tool from artifact-readers like get_metadata, get_review, and get_testflight, though it never names those siblings to sharpen the distinction.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives concrete when-to-use conditions: recover a lost response or request_key, and resume pending work by re-invoking the original tool with the returned key. It stops short of explicit exclusions or naming the alternative tools to use for other cases.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_testflightGet TestFlight statusARead-onlyInspect
Read external/internal groups, test information and recent Apple builds. Pass an exact asc_build_id for beta-review state, group access and what to test. Internal readiness does not imply external approval.
| Name | Required | Description | Default |
|---|---|---|---|
| group_id | No | Read all testers in this app-owned group | |
| project_id | Yes | NoMac project ID (prj_…). push_project and connect_status list your projects. | |
| asc_build_id | No | Apple build ID, from get_testflight. |
Output Schema
| Name | Required | Description |
|---|---|---|
| error | No | Present only when the call failed |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and destructiveHint=false, so the safety profile is covered. The description adds a genuine domain caveat ('Internal readiness does not imply external approval') and enumerates the data scope returned, but says nothing about permissions, pagination, or rate limits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three compact sentences, front-loaded with the resource scope and followed by the parameter cue and the caveat. No filler, though the phrasing 'Read external/internal groups, test information and recent Apple builds' is slightly listy rather than crisp.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present and annotations covering safety, the description need only convey scope and the one non-obvious caveat, which it does. The main remaining gap is sibling routing (manage_testflight), which an agent must infer from names alone.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description earns an extra point by clarifying the semantics of asc_build_id — it must be exact and it unlocks beta-review state, group access, and test guidance — which is more than the schema's circular 'from get_testflight' note provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a clear verb (Read) and specific resources: external/internal groups, test information, and recent Apple builds. An agent can distinguish it from the mutating sibling manage_testflight by the read framing, though the description never names that sibling explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives one conditional usage cue — pass an exact asc_build_id to get beta-review state, group access and what-to-test — which tells the agent when a parameter matters. However there is no explicit routing guidance against alternatives like manage_testflight, so usage is only implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_mac_sessionsList Mac sessionsARead-onlyInspect
List this account's Mac sessions, newest first, with state, cost and cleanup status. Use it to find a session ID, to check nothing is left running, or to recover after a lost start_mac response. Read-only.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| error | No | Present only when the call failed |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, openWorldHint=false and destructiveHint=false, so 'Read-only' in the text is redundant. It does add the sort order and the fact that state/cost/cleanup status come back, which is modest behavioral value beyond the structured fields; no pagination or volume limits are disclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences, front-loading the resource and ordering before the usage scenarios. No filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be spelled out, and annotations carry the safety profile. The description covers purpose, ordering and recovery scenarios; only result-volume/pagination behavior is unaddressed for a listing tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Zero parameters, so there is nothing for the description to disambiguate; the baseline for an argument-less tool applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (List) and resource (this account's Mac sessions) plus scope, ordering ('newest first'), and returned fields ('state, cost and cleanup status'). An agent can separate it from the singular sibling get_mac_session purely on the plural list semantics.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives three concrete when-to-use cases: finding a session ID, verifying nothing is left running, and recovering from a lost start_mac response. No explicit when-not or named alternative (e.g. get_mac_session for a known ID), so it stops short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
manage_testflightManage TestFlight testingADestructiveInspect
Set test metadata/contact/demo credentials, create/update external groups and public links, invite/remove an explicit tester, submit/distribute a specific Apple build to selected external groups, notify after approval, expire testing, or relay a human-provided encryption answer. Preview without confirm; confirm:true applies. distribute may send to Beta App Review; auto_notify defaults false. Never guess encryption/demo-account attestations. On pending, repeat unchanged arguments with request_key; get_store_operation recovers it.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | Name of the TestFlight group to create or rename. | |
| No | Tester's email address, for invite_tester. | ||
| action | Yes | metadata, create_group, update_group, invite_tester, remove_tester, remove_build, distribute, notify, expire or encryption. | |
| locale | No | App Store locale code, for example en-US or de-DE. Defaults to en-US. | en-US |
| confirm | No | false (default) returns a preview and changes nothing. true performs the change; set it only after the user approves that exact preview. | |
| app_info | No | Beta app description, feedback email, marketing URL and privacy policy URL. | |
| group_id | No | TestFlight group ID from get_testflight, for update_group, invite_tester or remove_tester. | |
| group_ids | No | External TestFlight group IDs from get_testflight, for distribute. | |
| last_name | No | Tester's last name, for invite_tester. | |
| tester_id | No | Apple tester ID from get_testflight, for remove_tester. | |
| whats_new | No | What to Test notes shown to testers for this build. | |
| first_name | No | Tester's first name, for invite_tester. | |
| project_id | Yes | NoMac project ID (prj_…). push_project and connect_status list your projects. | |
| auto_notify | No | Notify testers automatically when the build becomes available. Defaults to false. | |
| request_key | No | Idempotency key you choose (letters, digits, dash, underscore). Retry a lost response with the same key and identical arguments; use a new key for new work. | |
| asc_build_id | No | Apple build ID, from get_testflight. | |
| review_contact | No | Contact for Apple's reviewers (name, phone, email) and an optional demo account. | |
| public_link_limit | No | Maximum number of testers who can join through the public link (1-10000). | |
| public_link_enabled | No | Turn the group's public TestFlight link on or off. | |
| uses_non_exempt_encryption | No | The human's answer to Apple's export-compliance question. Never guess. |
Output Schema
| Name | Required | Description |
|---|---|---|
| error | No | Present only when the call failed |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations cover destructive/openWorld/non-idempotent, but the description adds substantial context beyond them: a two-phase preview/confirm flow, that distribute may trigger Beta App Review, that auto_notify defaults false, and the request_key retry contract. This is genuinely useful behavioral disclosure for a destructive multi-action tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
It front-loads the action list, then layers the critical procedural rules (preview/confirm, review submission, no guessing, idempotency) in compact sentences. It is dense but appropriate for a 10-action tool; no sentence is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With annotations declaring the safety profile and an output schema covering returns, the description fills the remaining gaps an agent needs: the preview/confirm gate, the Beta App Review risk, the encryption-attestation rule, and the idempotency recovery path. Complete enough to call correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so every parameter is already documented in the schema (including auto_notify's default and confirm's meaning). The description adds action-level semantics but no parameter-level detail beyond what the schema provides, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description enumerates concrete operations (set metadata, create/update groups, invite/remove tester, distribute build, expire testing, relay encryption answer) that map cleanly to the action enum, so the agent understands exactly what the tool does. It doesn't explicitly distinguish itself from the read sibling get_testflight, which keeps it from a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives clear procedural guidance: preview without confirm, confirm:true applies after user approval, never guess encryption/demo attestations, and on a pending result retry with the same request_key while get_store_operation recovers it. It names a sibling for recovery but does not state when to prefer this over get_testflight or set_metadata.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
prepare_review_replyDraft a reply to App ReviewARead-onlyInspect
Prepare YOUR authored reply to Apple App Review with the correct app handoff. Returns sent:false because Apple's public API cannot send or read reviewer messages. The human must paste and send the reply in App Store Connect. This is not a public customer-review response or a change to reviewer notes.
| Name | Required | Description | Default |
|---|---|---|---|
| reply | Yes | Your reply to App Review, up to 4000 characters. The human sends it in App Store Connect. | |
| project_id | Yes | NoMac project ID (prj_…). push_project and connect_status list your projects. | |
| asc_submission_id | Yes | Apple App Review submission ID, from get_review. |
Output Schema
| Name | Required | Description |
|---|---|---|
| error | No | Present only when the call failed |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Goes well beyond the annotations: it discloses that the tool only drafts (returns sent:false), explains why (Apple's public API cannot send or read reviewer messages), and specifies the required human workflow. This is exactly the kind of behavioral context annotations cannot convey, and it is consistent with readOnlyHint=true and destructiveHint=false.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight, front-loaded sentences with no wasted words; the scope constraint and the not-this clauses are efficiently placed. Minor vagueness in 'the correct app handoff' keeps it from being flawless.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, return values need not be explained, yet the description still clarifies the sent:false behavior. Combined with annotations covering safety and a fully documented schema, nothing an agent needs in order to call this correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so each of the three parameters is already fully documented, and the description adds no syntax, format, or constraint detail beyond what the schema provides. The baseline of 3 applies when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Prepare YOUR authored reply to Apple App Review') and immediately names what it is NOT (not a public customer-review response, not a change to reviewer notes), which distinguishes it from siblings like update_review and get_review.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear negative guidance by excluding the public-review and reviewer-notes use cases, and states the human must paste and send in App Store Connect. It stops short of an explicit 'use X instead when Y' routing to a named sibling, so it falls just below the top band.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
publishSubmit for App Store reviewADestructiveInspect
Submit for App Store review. Requires confirm:true with authorization for this app/version; run without confirm first to see the staged result + Apple blockers. On pending, repeat unchanged arguments with the returned request_key; get_store_operation recovers a lost response.
| Name | Required | Description | Default |
|---|---|---|---|
| force | No | Proceed despite non-blocking warnings shown in the preview. Leave false unless the user accepts them. | |
| confirm | No | false (default) returns a preview and changes nothing. true performs the change; set it only after the user approves that exact preview. | |
| build_id | No | Exact nomac release build; use this when replacing a rejected binary | |
| project_id | Yes | NoMac project ID (prj_…). push_project and connect_status list your projects. | |
| request_key | No | Only supply to resume an interrupted invocation with unchanged arguments | |
| asc_submission_id | No | Exact Apple submission from get_review | |
| resolve_rejection | No | Explicitly assert the selected version's rejection was fixed and mark its item ready for review |
Output Schema
| Name | Required | Description |
|---|---|---|
| error | No | Present only when the call failed |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the annotations (destructive, non-idempotent, open-world), the description discloses the authorization prerequisite, the safe dry-run pattern, the pending state and request_key resumption, and a recovery tool for lost responses. These are exactly the traits an agent cannot get from annotations alone.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tightly packed sentences with the action front-loaded and the confirm workflow, pending state, and recovery pointer following in priority order. No filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values needn't be described. The description covers the destructive-confirm pattern, the pending/resume edge case, and failure recovery, which is complete for a mutation tool of this complexity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so confirm, request_key, and the other parameters are already documented in the schema descriptions. The description reinforces the confirm/request_key workflow but adds no syntax or format detail beyond what the schema already provides, making the 3 baseline correct.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Submit for App Store review') that is immediately distinguishable from read/update siblings like get_review, update_review, and review_lint. An agent knows exactly what action is being performed.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly prescribes the sequence: run without confirm first to see the staged result + Apple blockers, then repeat with confirm:true only with authorization. It also names the recovery path (get_store_operation) and the pending-resume condition, leaving no inference required.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
push_projectHow to push source codeBRead-onlyInspect
Hosted transport cannot read your working tree — push from where the code lives: run npx @nomac/cli login once, then npx @nomac/cli push in the project directory (or use the stdio server nomac mcp, whose push_project packs locally). Returns current projects instead.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| error | No | Present only when the call failed |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and destructiveHint=false, so the safety profile is covered. The description adds real behavioral context: the hosted transport cannot read the working tree, and the tool returns current projects instead of pushing. It never resolves the tension between a tool named 'push_project' and a read-only hint, but it does not contradict the annotations either.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
It is a single sentence, and the key constraint ('Hosted transport cannot read your working tree') is front-loaded. But it crams login steps, push commands, an alternative server mode, and the return value into one dense line without clean structure.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be spelled out, and the description does address the hosted-transport constraint and fallback. What remains incomplete is the core identity of the tool: an agent still cannot cleanly tell whether push_project moves code or just lists projects.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so the schema is trivially complete and there is no parameter semantics for the description to add. Baseline 4 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description ultimately states the tool's real behavior ('Returns current projects instead'), but its center of gravity is a CLI push walkthrough, so what push_project actually does is buried and ambiguous. It also never distinguishes itself from siblings like status, check_project, or build, which is confusing given the name implies a mutation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It routes the agent away from this tool for real pushes ('push from where the code lives') and names concrete alternatives (npx CLI, the stdio server nomac mcp), which is useful. However, it never says positively when to call push_project itself versus those alternatives, so the usage guidance is only implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
read_mac_outputRead Mac command outputARead-onlyInspect
Read stdout/stderr as base64 chunks; resume from next_cursor. If truncated, the final response includes tail_base64 and its absolute tail_offset. Retrieve promptly: only recent completed jobs retain output. output_expired preserves the job receipt and never means rerun. Retrieve before stopping the VM.
| Name | Required | Description | Default |
|---|---|---|---|
| cursor | No | Byte offset to read from. Start at 0, then pass next_cursor from the previous response. | |
| job_id | Yes | Command job ID (job_…) returned by exec_mac. | |
| stream | No | stdout (default) or stderr. | |
| session_id | Yes | Mac session ID (ses_…) returned by start_mac or list_mac_sessions. |
Output Schema
| Name | Required | Description |
|---|---|---|
| error | No | Present only when the call failed |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnly/destructive annotations, it discloses chunking with cursor-based resumption, the truncation contract (tail_base64 plus absolute tail_offset), a retention window, and the crucial non-error semantics of output_expired. This is exactly the kind of context annotations cannot convey and it prevents a needless re-execution.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four tight sentences ordered from core mechanic to truncation handling to retention risk to the VM-stopping warning. No filler and the most actionable warning (retrieve before stopping) is saved for emphasis at the end.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, return fields need not be explained, and the description still covers truncation, retries, and expiry semantics thoroughly. The only shortfall is that it never positions this tool against the sibling that returns job metadata, so an agent has no explicit cue for choosing between them.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents cursor, stream, job_id, and session_id, including the next_cursor resumption pattern and the stdout default. The description restates cursor resumption but adds no format or boundary details beyond the schema, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource (read stdout/stderr output as base64 chunks) with the resumption mechanism, and is clearly distinct from siblings like get_mac_job or exec_mac. An agent knows immediately this is the raw output reader rather than the job status tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives real timing guidance: retrieve promptly because only recent completed jobs retain output, and retrieve before stopping the VM. It also tells the agent how to interpret output_expired (do not rerun). However, it never names an alternative tool or a when-not condition, so the routing to siblings is implied rather than explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
read_storeRead App Store Connect dataARead-onlyInspect
Read a discovered Apple GET operation for the linked app. Begin with an apps_* relationship and scope:[]; copy each returned resource's scope into subsequent calls. Paginate with next_cursor and unchanged query. Raw IDs are accepted only for global lookups. This tool only reads. File reservations expose Apple uploadOperations; analytics segments expose report URLs.
| Name | Required | Description | Default |
|---|---|---|---|
| query | No | Query parameters for the Apple GET operation, such as filters, fields and limit. | |
| scope | No | Parent resource chain copied from earlier read_store results, as [{relationship, id}]. Start with [] for app-level reads. | |
| cursor | No | next_cursor from the previous page. Keep the other arguments unchanged. | |
| operation | Yes | App Store Connect operation name from get_store_capabilities, for example appInfos_getInstance. | |
| project_id | Yes | NoMac project ID (prj_…). push_project and connect_status list your projects. | |
| resource_id | No | Raw Apple resource ID, only for global lookups outside the linked app. |
Output Schema
| Name | Required | Description |
|---|---|---|
| error | No | Present only when the call failed |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With readOnlyHint already declared, the description goes beyond annotations by disclosing the scope-propagation workflow, the pagination contract ('next_cursor and unchanged query'), and what specific operations return (uploadOperations, report URLs). 'This tool only reads' merely restates the annotation, but the surrounding behavioral detail is genuinely additive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Five compact sentences, front-loaded with purpose and followed by actionable operational rules; nothing is wasted. It is slightly dense and terse, which costs a point on readability but not on economy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has an output schema and annotations covering safety, and the description still supplies the non-obvious pieces an agent needs: how to start, how to chain scope, how to paginate, and when raw IDs are allowed. Nothing essential is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3, but the description adds workflow meaning: scope must be seeded with [] and copied from prior results, cursor requires unchanged other arguments, and resource_id is valid only for global lookups. This clarifies how parameters interact rather than restating their types.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('read') and resource ('a discovered Apple GET operation') and scopes it to the linked app, so the agent knows exactly what the tool operates on. It does not explicitly differentiate itself from nearby siblings such as get_store_operation or get_store_capabilities, which would have earned a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives concrete entry conditions ('Begin with an apps_* relationship and scope:[]'), a chaining rule for follow-up calls, and a restriction ('Raw IDs are accepted only for global lookups'). It stops short of naming when to prefer a sibling tool, so it is strong context without explicit alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
report_issueReport an issue to NoMac supportAInspect
File an unknown/unfixable error with nomac support. For a Mac use session_id and optional job_id; for a build use ref_id. Attaches account-scoped context and deduplicates repeats. Optional email requests a direct resolution update; omit description to read replies.
| Name | Required | Description | Default |
|---|---|---|---|
| No | Optional address for a direct reply from NoMac support. Only when the user asks for one. | ||
| job_id | No | Command job ID (job_…) the issue is about. | |
| ref_id | No | Build ID or other NoMac reference the issue is about. | |
| session_id | No | Mac session ID (ses_…) the issue is about. | |
| description | No | What went wrong, including the tool, arguments and error you saw. Omit it to read replies to earlier reports. |
Output Schema
| Name | Required | Description |
|---|---|---|
| error | No | Present only when the call failed |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the annotations (which only declare readOnly=false, idempotent=false, destructive=false), the description discloses that account-scoped context is attached, that repeats are deduplicated, and that omitting description turns the call into a reply reader. That dual read/write behavior and the dedup guarantee are exactly the extra context annotations cannot convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four tight sentences, front-loaded with the core action, then routing, then side effects, then the reply mode. No filler and nothing repeated from the schema or annotations.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, return values need no explanation, and the description covers routing, side effects and the read mode for a zero-required-parameter tool. It is complete enough to invoke correctly, though it says nothing about auth requirements or submission limits, minor omissions for an openWorld=false support tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% so the baseline is 3, but the description adds meaning the schema lacks: it maps session_id (+ optional job_id) to the Mac case and ref_id to the build case, and clarifies that email is only for requesting a direct resolution update and that omitting description switches modes. This is useful selection logic beyond field definitions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (file) plus the resource (unknown/unfixable error) and the destination (NoMac support), and immediately gives the routing rule distinguishing a Mac session (session_id/job_id) from a build (ref_id). An agent can tell this apart from the many status/get_* siblings without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The scope qualifier 'unknown/unfixable error' tells the agent when this tool is warranted and implicitly when it is not (known/fixable failures). It also explains the two parameterization paths and the omit-description mode for reading replies. No sibling tool is named as an alternative, so it stops short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
review_lintRun App Review readiness checkAInspect
Review-readiness report (green/yellow/red). Red blocks publish. Findings carry evidence, fix hints, sometimes ready patches. 4.3-style findings are signals for YOU to judge.
| Name | Required | Description | Default |
|---|---|---|---|
| project_id | Yes | NoMac project ID (prj_…). push_project and connect_status list your projects. |
Output Schema
| Name | Required | Description |
|---|---|---|
| error | No | Present only when the call failed |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description adds real value beyond annotations: output is a tri-state report, red blocks publishing, findings include evidence and fix hints, and 4.3-style findings are advisory signals for the agent to judge. However, annotations declare readOnlyHint=false, and the description never explains what side effect makes this a non-read-only operation (the 'sometimes ready patches' phrase hints at it but is not clarified).
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four short, front-loaded sentences with no filler; the blocking rule and output semantics come early. The '4.3-style' shorthand is jargon that assumes shared context but costs no space.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, the description correctly focuses on decision semantics (severity levels, blocking behavior, advisory findings) rather than return fields. It is largely complete for a one-parameter check, though it omits any note on side effects or run frequency.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Only one parameter and schema coverage is 100%; the schema's own description already explains the prj_ format and points to push_project and connect_status for discovery. The description adds nothing about project_id, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The title and description establish a specific function: a review-readiness lint that returns a green/yellow/red report. It is clearly distinguishable from siblings like get_review, build, and check_project, though the verb 'lint' itself is never stated plainly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Red blocks publish' implies a pre-publish gate, which is useful routing context, but the description never states when to run this versus check_project or status, nor any prerequisite (e.g. a connected project). Usage is implied rather than explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
set_metadataUpdate App Store metadataADestructiveInspect
Write App Store metadata YOU authored (validated before any Apple call). fields: description/keywords/whats_new/support_url/…; plus primary_category, age_rating (human attestations), content_rights, copyright, review_contact, price:'FREE'. On pending, repeat unchanged arguments with the returned request_key; get_store_operation recovers a lost response.
| Name | Required | Description | Default |
|---|---|---|---|
| price | No | Only FREE is supported here. Use update_store for paid pricing. | |
| fields | No | Listing text by field name: description, keywords, whats_new, promotional_text, support_url, marketing_url and so on. get_metadata_schema lists fields and character limits. | |
| locale | No | App Store locale code, for example en-US or de-DE. Defaults to en-US. | en-US |
| copyright | No | Copyright line shown on the App Store, for example "2026 Example Inc." | |
| age_rating | No | Age-rating questionnaire answers. These are legal attestations: use only what the human told you. | |
| project_id | Yes | NoMac project ID (prj_…). push_project and connect_status list your projects. | |
| request_key | No | Reuse only to resume pending work with the original unchanged arguments | |
| content_rights | No | Whether the app uses third-party content, as the human answered. | |
| review_contact | No | Contact for Apple's reviewers (name, phone, email) and an optional demo account. | |
| version_string | No | Defaults to the pushed source marketing version; select a live version explicitly for promotional text | |
| primary_category | No | App Store primary category, for example UTILITIES or DEVELOPER_TOOLS. |
Output Schema
| Name | Required | Description |
|---|---|---|
| error | No | Present only when the call failed |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare destructive=true and idempotent=false, and the description usefully adds that inputs are 'validated before any Apple call' and how to safely resume an in-flight write via request_key. This materially explains the non-idempotent mutation behavior beyond the structured hints.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the core action, but the body is a dense, punctuated run-on using '…/' shorthand that is harder to parse than necessary. Information is useful but packed without clean structure.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need no explanation, and the description fills the key remaining gap (the pending/retry workflow). For an 11-param destructive mutation with human-attestation constraints, it is nearly complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents every parameter, including the FREE-only price enum, locale default, and the attestation caveat on age_rating. The description reiterates these (price:'FREE', human attestations) but adds no new syntax or format detail, so baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Write App Store metadata') and enumerates the field families it touches. It clearly differs from read-side siblings (get_metadata, get_metadata_schema) and the write-side update_store, though it doesn't name those alternatives directly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides concrete operational guidance for the pending/idempotency path: repeat unchanged arguments with the returned request_key, and use get_store_operation to recover a lost response. It also flags that age_rating/content_rights must come from a human. It stops short of explicitly saying when to prefer update_store (that lives in the schema).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
set_project_connectionChoose a project's Apple connectionAInspect
Link a project to an Apple connection, for example after the human rotates their App Store Connect key. Verifies that the connection can see the app's bundle ID. Release builds already running must finish first.
| Name | Required | Description | Default |
|---|---|---|---|
| project_id | Yes | NoMac project ID (prj_…). push_project and connect_status list your projects. | |
| connection_id | Yes | Apple connection ID, from connect_status. |
Output Schema
| Name | Required | Description |
|---|---|---|
| error | No | Present only when the call failed |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare a non-read-only, non-idempotent mutation, and the description adds genuine behavioral detail beyond that: it performs a bundle-ID visibility check and will block while release builds are in flight. This is real context an agent could not infer from the structured fields. No contradiction with the mutation-shaped annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight clauses with the action front-loaded, then the motivating example, then the precondition. Every sentence earns its place and nothing is redundant.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, return values need not be explained, and the description covers the mutation's side effects and precondition. It is essentially complete for a two-parameter link operation, with only the absence of an explicit alternative to connect_status as a minor gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both parameters (project_id, connection_id) are already fully documented, including where to obtain them. The description adds no syntax or format detail beyond the schema, which is the expected baseline when the schema carries the load.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: "Link a project to an Apple connection." The mutation is unambiguous and distinct from the read-only connect_status sibling, though the description never names an alternative explicitly. A clear, self-contained statement of purpose.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a concrete trigger ("after the human rotates their App Store Connect key") and a precondition ("Release builds already running must finish first"). It stops short of naming when NOT to use it or which sibling to prefer as an alternative, so it falls just below the top band.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
start_macStart a cloud MacAInspect
Allocate a full Mac VM, billed from prepaid credit at $0.80 an hour by the minute. If it returns insufficient_credit, call get_mac_checkout and have the human pay, then retry with the same request_key. public_key is optional for agents using HTTPS commands only. Reuse request_key and identical settings after a lost response. Poll get_mac_session until ready; inspect its offer, cost and limits. Stop deletes the VM and files. No predefined project pipeline.
| Name | Required | Description | Default |
|---|---|---|---|
| limits | No | Optional caps for this session: max_duration_seconds (60-86400), spend_cap_microunits (USD × 1,000,000) and idle_timeout_seconds (60-3600). Defaults: 2 hours, $5, 10 minutes idle. | |
| public_key | No | Optional SSH public key (ed25519, RSA 2048+ or P-256) for direct SSH access. Omit when you only use exec_mac. | |
| request_key | Yes | Idempotency key you choose (letters, digits, dash, underscore). Retry a lost response with the same key and identical arguments; use a new key for new work. |
Output Schema
| Name | Required | Description |
|---|---|---|
| error | No | Present only when the call failed |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Goes well beyond the annotations: prepaid-credit billing at a per-minute rate, the insufficient_credit error code, idempotent retry semantics via request_key, and critically that 'Stop deletes the VM and files' (destructive consequence not captured by destructiveHint=false). It also states there is no predefined project pipeline, preempting a wrong assumption about siblings like build/push_project.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the allocation action and cost, then sequences error handling, retry, and polling in roughly the order an agent would need them. It is dense but every clause carries operational information; the 'No predefined project pipeline' fragment is slightly terse and could be folded into a cleaner sentence.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a stateful, billable allocation tool with a nested limits object and an output schema, the description covers creation, cost, failure recovery, idempotent retry, readiness polling, and teardown consequences. Return-format details are legitimately delegated to the output schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds real meaning: public_key is optional specifically for agents using HTTPS/exec_mac only, and request_key must be reused only after a lost response with identical settings. It does not explain the limits object fields (idle_timeout, max_duration, spend cap), which the schema description covers but the prose does not reinforce.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Allocate a full Mac VM') and immediately bounds scope with cost ('billed from prepaid credit at $0.80 an hour by the minute'). This is clearly distinguishable from siblings like stop_mac, get_mac_session, and list_mac_sessions.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly routes the agent on the error path ('If it returns insufficient_credit, call get_mac_checkout...'), the lost-response path (reuse request_key with identical settings), and the readiness path ('Poll get_mac_session until ready'). It also states when public_key can be omitted ('agents using HTTPS commands only'), which is a genuine when/when-not condition.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
statusGet build statusARead-onlyInspect
Read a build's current state: from queued and building through uploading and processing to ready or failed. Poll it after build until ready or failed; on failed, call get_failure for the reason. Read-only.
| Name | Required | Description | Default |
|---|---|---|---|
| build_id | Yes | NoMac build ID (bld_…) returned by build. |
Output Schema
| Name | Required | Description |
|---|---|---|
| error | No | Present only when the call failed |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnlyHint, destructiveHint=false), so the bar is lower, and the description still adds real value by disclosing the state machine and the polling contract that an agent must follow. It stops short of stating rate limits or whether repeated polls are throttled, which would be the remaining useful detail.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, no filler; the lifecycle is front-loaded and the polling/alternative instruction follows. Every clause carries actionable information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return-value explanation is unnecessary; the description covers purpose, lifecycle, polling, failure routing, and read-only nature. For a single-param, read-only status tool this is fully sufficient.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the build_id schema text already explains it is the 'bld_…' ID returned by build, so the description adds nothing parameter-specific. Baseline 3 applies since the schema does all the work for this single required field.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Read') and resource (a build's current state) and enumerates the full lifecycle (queued, building, uploading, processing, ready/failed), so the agent immediately knows what is returned. It also implicitly separates itself from get_failure, the sibling that supplies the failure reason.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit when-to-use ('Poll it after build until ready or failed') and names the alternative and its triggering condition ('on failed, call get_failure for the reason'). This is a complete routing instruction, not an inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
stop_macStop and delete a MacADestructiveInspect
Stop a Mac session: shuts down the VM and permanently deletes it and all its files, which ends billing. Read any output you need first with read_mac_output. Call it whenever the task is done so credit is not spent on an idle Mac, then poll get_mac_session until cleanup_confirmed. Cannot be undone.
| Name | Required | Description | Default |
|---|---|---|---|
| session_id | Yes | Mac session ID (ses_…) returned by start_mac or list_mac_sessions. |
Output Schema
| Name | Required | Description |
|---|---|---|
| error | No | Present only when the call failed |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true and idempotentHint=false, so the safety profile is covered; the description goes beyond that by specifying exactly what is destroyed ('it and all its files'), that billing stops, that the action is irreversible, and that cleanup requires polling. That is meaningful added context, though it does not discuss failure modes or partial-state behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences, front-loaded with the destructive action and its billing consequence, followed by the required pre-step and post-step. No filler and no repetition of the title.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter destructive tool with an output schema available, the description covers what it does, what it destroys, the reading prerequisite, and the follow-up polling. Nothing needed to invoke it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Only one parameter, and the schema already documents session_id at 100% coverage including the ses_… pattern and its origin (start_mac/list_mac_sessions). The description adds nothing about the parameter, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Stop a Mac session: shuts down the VM and permanently deletes it') plus the consequence (ends billing). An agent can distinguish it from read_mac_output, get_mac_session, and list_mac_sessions without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly names the prerequisites and alternatives: read output first with read_mac_output, call whenever the task is done to avoid idle spend, then poll get_mac_session until cleanup_confirmed. When-to-use and sequencing are fully specified.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
update_reviewChange an App Review submissionADestructiveInspect
Resolve a fixed rejected item, remove an item, cancel review, resubmit all ready items, or release an approved app version. Uses exact Apple IDs from get_review. Preview without confirm; confirm:true performs the action. Removing items cannot be undone in that submission. Fix the actual issue before resolve_item. On pending, repeat unchanged arguments with request_key; get_store_operation recovers it.
| Name | Required | Description | Default |
|---|---|---|---|
| force | No | Proceed despite non-blocking warnings shown in the preview. Leave false unless the user accepts them. | |
| action | Yes | resolve_item (after fixing a rejection), remove_item, cancel, resubmit or release (an approved version). | |
| confirm | No | false (default) returns a preview and changes nothing. true performs the change; set it only after the user approves that exact preview. | |
| item_id | No | Apple review item ID from get_review. Needed for resolve_item and remove_item. | |
| project_id | Yes | NoMac project ID (prj_…). push_project and connect_status list your projects. | |
| request_key | No | Idempotency key you choose (letters, digits, dash, underscore). Retry a lost response with the same key and identical arguments; use a new key for new work. | |
| asc_submission_id | Yes | Apple App Review submission ID, from get_review. |
Output Schema
| Name | Required | Description |
|---|---|---|
| error | No | Present only when the call failed |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare destructive/non-idempotent/open-world, and the description adds material context beyond them: preview-by-default with confirm:true committing the change, irreversible item removal in that submission, and the request_key idempotency retry pattern. These are exactly the traits an agent needs to avoid an accidental destructive call.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The action list is front-loaded and the operational rules follow compactly with no filler. Sentences are dense and comma-heavy, so a small amount of parsing effort is required, but nothing is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With 7 parameters, an enum action, rich annotations and an output schema, the description still covers the whole workflow: preview-then-confirm, idempotent retry, irreversibility, and the prerequisite fix before resolve_item. An agent can act correctly without consulting anything else.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description goes further by anchoring item_id and asc_submission_id to get_review provenance and explaining the confirm/request_key interaction in prose. It adds real cross-tool meaning rather than restating the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The first sentence enumerates the five concrete operations (resolve a rejected item, remove an item, cancel review, resubmit ready items, release an approved version) performed against an App Review submission, so the verb+resource+scope are unambiguous. It is clearly separable from read-only siblings like get_review and prepare_review_reply.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It routes the agent to get_review for the exact Apple IDs and to get_store_operation for recovering a pending request, and it states the precondition 'Fix the actual issue before resolve_item.' It stops short of explicit when-not-to-use guidance (e.g. when to prefer update_store or prepare_review_reply), but the action-level conditions are clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
update_storeChange App Store Connect dataADestructiveInspect
Execute a discovered Apple write using its exact JSON:API body. Copy scope and relationship references from read_store. Defaults to a preview; confirm:true changes real Apple resources. Supports pricing/offers, product creation and submission, public customer replies, release settings and asset reservations/commit. Use owner-provided attestations and prices. On pending, retry identical arguments with the returned request_key; do not create a fresh key. Use publish/update_review/manage_testflight for their dedicated actions.
| Name | Required | Description | Default |
|---|---|---|---|
| body | No | Exact JSON:API request body, following the operation's schema from get_store_capabilities. | |
| scope | No | Parent resource chain copied from earlier read_store results, as [{relationship, id}]. Start with [] for app-level reads. | |
| confirm | No | false (default) returns a preview and changes nothing. true performs the change; set it only after the user approves that exact preview. | |
| operation | Yes | App Store Connect operation name from get_store_capabilities, for example appInfos_getInstance. | |
| project_id | Yes | NoMac project ID (prj_…). push_project and connect_status list your projects. | |
| references | No | Related resource references the body needs, copied from read_store results. | |
| request_key | No | Idempotency key you choose (letters, digits, dash, underscore). Retry a lost response with the same key and identical arguments; use a new key for new work. | |
| asset_sha256 | No | Set by upload_store_asset for exact file-reservation recovery. |
Output Schema
| Name | Required | Description |
|---|---|---|
| error | No | Present only when the call failed |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations flag destructiveHint=true and idempotentHint=false, but the description adds crucial behavior beyond them: the default is a non-mutating preview and only confirm:true touches real Apple resources, plus explicit idempotency/retry semantics via request_key. This is meaningful context an agent could not derive from the structured fields alone.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The core safety rule (preview default vs confirm:true) is front-loaded and every sentence carries operational weight, but the paragraph is dense with multiple distinct concerns (scope sourcing, supported domains, attestations, retry, routing) that could be more tightly grouped.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an open-world destructive write with an output schema present, the description covers everything an agent needs: preview/confirm gating, parameter sourcing, idempotency recovery, and sibling routing. Return values are correctly left to the output schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds sourcing guidance the schema lacks: scope/references come from read_store, operation/body follow get_store_capabilities, and request_key must be reused identically on retry rather than regenerated. That elevates it above a pure schema restatement.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Execute a discovered Apple write using its exact JSON:API body') and enumerates the covered domains (pricing/offers, product creation, replies, release settings, asset reservations). It explicitly distinguishes itself from publish/update_review/manage_testflight, so an agent can route without opening sibling schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit when-to-use ('execute a discovered Apple write'), when-to-defer ('Use publish/update_review/manage_testflight for their dedicated actions'), a dependency ('Copy scope and relationship references from read_store'), and a recovery rule ('On pending, retry identical arguments with the returned request_key'). Nothing is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
upload_screenshotsReplace App Store screenshotsADestructiveInspect
Start a screenshot replacement for one display type and iOS version. PNGs are validated and prior images are backed up. Reuse request_key and the original body when retrying. Poll get_screenshot_upload until complete, rolled_back or needs_attention; accepted means still in progress. images: [{filename, data: base64 PNG}].
| Name | Required | Description | Default |
|---|---|---|---|
| images | Yes | Screenshots in display order as [{filename, data}], where data is a base64 PNG of the exact size for display_type. | |
| locale | No | App Store locale code, for example en-US or de-DE. Defaults to en-US. | en-US |
| project_id | Yes | NoMac project ID (prj_…). push_project and connect_status list your projects. | |
| request_key | No | Unique per replacement; reuse this key and the same images on a retry | |
| display_type | Yes | Apple screenshot display type, for example APP_IPHONE_67. get_metadata_schema lists types and exact pixel sizes. | |
| version_string | No | Defaults to the pushed source marketing version |
Output Schema
| Name | Required | Description |
|---|---|---|
| error | No | Present only when the call failed |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Adds substantive behavior beyond the annotations: PNGs are validated, prior images are backed up (contextualizing the destructiveHint=true), and the async lifecycle states are enumerated with the meaning of "accepted." The retry guidance also meaningfully supplements idempotentHint=false.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the action, then the safety/retry facts, then the lifecycle states — good ordering with no filler sentences. Minor deduction because the trailing inline images spec duplicates the schema without adding new information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained; the description instead covers validation, backup, retry semantics, and terminal polling states. Nothing an agent needs to invoke and correctly follow up on this async tool is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all six parameters are already documented in the schema, and the description's "images: [{filename, data: base64 PNG}]" restates rather than extends that. Baseline 3 applies since the schema carries the parameter burden.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource — "Start a screenshot replacement" — and scopes it to "one display type and iOS version," which cleanly separates it from sibling upload_store_asset and from the read-side get_screenshot_upload. An agent can identify the operation without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear post-call workflow guidance: poll get_screenshot_upload until complete/rolled_back/needs_attention, and reuse request_key plus the original body on retry. It does not, however, say when to prefer this over the sibling upload_store_asset, so the alternative-selection guidance is incomplete.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
upload_store_assetUpload an App Store assetAInspect
Upload a new Apple screenshot, preview, purchase image, supplemental reviewer attachment or encryption document. Discover its createInstance operation/body with get_store_capabilities and copy the parent scope. Hosted calls accept data_base64 up to 2 MiB; use CLI file_path for larger files. Preview by default. confirm:true reserves, uploads and commits this file. Preserve request_key and identical file/body to resume; ready:true requires Apple processing COMPLETE. App Review message attachments remain a browser workflow.
| Name | Required | Description | Default |
|---|---|---|---|
| body | No | Exact JSON:API request body, following the operation's schema from get_store_capabilities. | |
| scope | No | Parent resource chain copied from earlier read_store results, as [{relationship, id}]. Start with [] for app-level reads. | |
| confirm | No | false (default) returns a preview and changes nothing. true performs the change; set it only after the user approves that exact preview. | |
| operation | Yes | App Store Connect operation name from get_store_capabilities, for example appInfos_getInstance. | |
| project_id | Yes | NoMac project ID (prj_…). push_project and connect_status list your projects. | |
| references | No | Related resource references the body needs, copied from read_store results. | |
| data_base64 | Yes | The file's contents, base64-encoded, up to 2 MiB. | |
| request_key | No | Idempotency key you choose (letters, digits, dash, underscore). Retry a lost response with the same key and identical arguments; use a new key for new work. | |
| asset_sha256 | No | Set by upload_store_asset for exact file-reservation recovery. |
Output Schema
| Name | Required | Description |
|---|---|---|
| error | No | Present only when the call failed |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare mutation (readOnlyHint false), open-world behavior, non-idempotency, and non-destructiveness. The description adds important context beyond that: preview/confirm staging, the 2 MiB hosted size limit, the request_key resume contract, and the ready/COMPLETE processing condition. It does not cover permissions or error behavior, so not a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with purpose and workflow, and every sentence adds distinct operational guidance. It is telegraphic in places but not wasteful.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given a complex 9-parameter mutation with an output schema and rich annotations, the description covers the staging workflow, size constraints, retry contract, and one explicit exclusion. Remaining gaps like permissions and failure handling are minor.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description reiterates and lightly supplements a few parameters (operation discovery, scope copy, data_base64 size, confirm, request_key) but adds little detail beyond the schema, and mentions 'ready:true' even though there is no top-level ready parameter in the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb 'upload' and resource 'App Store asset', enumerating asset types. It also distinguishes from the browser workflow for App Review attachments, but does not explicitly route screenshot uploads to or from the sibling upload_screenshots tool, so not a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
States prerequisite discovery via get_store_capabilities, scope copying, CLI file_path for files over 2 MiB, preview-by-default then confirm:true, and excludes App Review message attachments. Missing explicit guidance on when to prefer this over the sibling upload_screenshots.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
38 tool updates
- First observed
build - First observed
check_project - First observed
connect_status - First observed
exec_mac - First observed
get_billing_status - First observed
get_checkout_link - First observed
get_failure - First observed
get_feedback - First observed
get_mac_checkout - First observed
get_mac_credits - First observed
get_mac_job - First observed
get_mac_session - First observed
get_metadata - First observed
get_metadata_schema - First observed
get_refund_status - First observed
get_review - First observed
get_screenshot_upload - First observed
get_store_capabilities - First observed
get_store_operation - First observed
get_testflight - First observed
list_mac_sessions - First observed
manage_testflight - First observed
prepare_review_reply - First observed
publish - First observed
push_project - First observed
read_mac_output - First observed
read_store - First observed
report_issue - First observed
review_lint - First observed
set_metadata - First observed
set_project_connection - First observed
start_mac - First observed
status - First observed
stop_mac - First observed
update_review - First observed
update_store - First observed
upload_screenshots - First observed
upload_store_asset
Publisher details
- Operator
- NoMac · Publisher source
- Operator website
- https://nomac.app · Publisher source
- Vendor relationship
- First-party · Publisher source
- Documentation
- https://nomac.app/install · Publisher source
- Trust center
- Not available
- Restrictions
- Sign in with a NoMac account (OAuth). Starting a Mac needs prepaid credit, sold in $10 packs at $0.80 an hour. App Store and TestFlight tools also need the user's App Store Connect API key. · Publisher source
Related MCP Connectors
Build, run, and inspect iOS apps in disposable hosted Simulators from cloud coding agents.
- LimrunOAuthcom.limrun
Cloud iOS simulators and Android emulators your agent can create, drive, and throw away.
- app-managerOAuthapp.lance
App Store Connect operator for AI agents: icons, TestFlight builds, listings, IAP, rejection fixes.
Run App Store Connect from your IDE: pricing, listings, screenshots, releases, AI visibility.
Related MCP Servers
- AlicenseAqualityCmaintenanceEnables coding agents to build, test, archive, and run iOS apps on a remote Mac fleet from any environment, with XcodeBuildMCP-compatible tools and live simulator access.1139 npmMIT
- AlicenseAqualityCmaintenanceBuild, sign, and publish iOS and Android apps through AI agents. Integrates Codemagic CI/CD, App Store Connect, and Google Play in one server.6314 npm1MIT
- AlicenseNot gradedqualityBmaintenanceAutomate App Store Connect from your AI agent. Manage versions, metadata, builds, and submissions through natural language.16 npm8MIT
- AlicenseNot gradedqualityBmaintenanceConnects AI coding clients like OpenCode, Codex, and Claude Code to Xcode and Apple development tools, enabling builds, tests, simulator and device management, profiling, asset handling, and more through natural language.208 npm1MIT
Glama MCP Gateway
Add one secure layer between your agents and this server.