Skip to main content
Glama

Server Details

A real cloud Mac with Xcode for your AI agent. Start a macOS VM, run commands (git clone, xcodebuild, simulator tests), read the output and stop it; billed from prepaid credit at $0.80 an hour by the minute. Optional App Store Connect tools upload builds to TestFlight and manage the listing, with previews before any change. OAuth sign-in, nothing to paste.

Ownership verified
Status
Healthy
Last Tested
Transport
Streamable HTTP · MCP 2025-11-25
URL

TDQS

A3.8/5.0

Scored across 38 tools

Disambiguation4/5

Most tools have clearly distinct resource+action scopes, and the descriptions explicitly tell the agent when to use each (e.g. get_store_operation vs get_store_capabilities vs read_store). A few boundaries are softer: read_store/update_store/get_store_capabilities/get_store_operation all orbit Apple JSON:API, and upload_screenshots vs upload_store_asset overlap on asset uploads, but the docs disambiguate them well.

Naming Consistency4/5

Strong, predictable snake_case with verb-led prefixes (get_*, set_*, list_*, start_*/stop_*, update_*, read_*, upload_*). A handful of bare names (build, status, publish, review_lint) deviate from the verb_noun pattern but remain readable and unambiguous in context.

Tool Count3/5

38 tools is on the heavy side, though the server spans a legitimately broad surface (builds, TestFlight, App Store review, metadata, screenshots, Mac VMs, billing). Many tools earn their place, but the count is high enough that consolidation (e.g. the four Apple read/write/capability/store-operation tools) would improve scanability.

Completeness4/5

The surface covers a full lifecycle: push → check → build → status → failure → publish → review → TestFlight → feedback, plus billing, metadata, assets and Mac VM management, with recovery/status tools for lost responses. Minor gaps exist where Apple's public API can't act (reviewer message send/read are documented browser handoffs) and there's no explicit build cancel, but no significant dead ends.

Available Tools

38 tools
buildStart a buildAInspect

Start a build for a project. workflow=release (signed → TestFlight; uses one build from the plan quota) or smoke (unsigned compile check). Poll with status.

ParametersJSON Schema
NameRequiredDescriptionDefault
workflowNorelease (default): signed build uploaded to TestFlight, uses one build from the plan. smoke: unsigned compile check.release
project_idYesNoMac project ID (prj_…). push_project and connect_status list your projects.
request_keyNoReuse to reconcile a lost response. Use a new key only after confirmed no build started.

Output Schema

ParametersJSON Schema
NameRequiredDescription
errorNoPresent only when the call failed

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare this is a non-read-only, non-idempotent, non-destructive write, and the description adds genuinely useful behavior beyond them: that a release build consumes one build from the plan quota and that release output is signed and lands in TestFlight. It omits failure/log-retrieval behavior, but the quota cost is meaningful context annotations cannot convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, front-loaded with the action and followed by the mode distinction and the polling hint. No filler and no redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need no explanation, and the description covers the two modes, their side effects, cost, and how to check progress. It leaves unanswered what happens on a failed build and which sibling exposes build failures, which would round it out.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the schema already documents project_id, workflow, and the request_key reconcile semantics in detail. The description largely restates the workflow definitions the schema already provides, adding little beyond the baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Start a build for a project') and immediately subdivides into two named modes, release vs smoke. It is clearly distinguishable from peers like push_project or publish, though it never explicitly contrasts itself with them.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear selection context for the two workflows (signed TestFlight upload vs unsigned compile check) and routes the agent to the 'status' sibling for polling. It stops short of stating when not to use the tool or how it relates to publish/push_project in a full workflow.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

check_projectCheck project readinessAInspect

Free source-readiness check of the latest uploaded snapshot; no Apple calls or paid compute. Run push_project after source changes. Does not prove compilation, signing or TestFlight delivery.

ParametersJSON Schema
NameRequiredDescriptionDefault
project_idYesNoMac project ID (prj_…). push_project and connect_status list your projects.

Output Schema

ParametersJSON Schema
NameRequiredDescription
errorNoPresent only when the call failed

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Adds real behavioral context beyond the annotations: it is free, makes no Apple calls or paid compute, and crucially discloses the boundary 'Does not prove compilation, signing or TestFlight delivery'. This negative scope is genuinely useful and not present in the structured fields. It does not address why readOnlyHint is false, leaving a minor gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences: the core purpose is front-loaded, then the follow-up action and the limitation. Every clause earns its place with no redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need no explanation. The description covers purpose, cost profile, boundary conditions, and the follow-up tool, which is sufficient for a one-parameter check. Minor omission is any note on whether it mutates state given readOnlyHint=false.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With a single parameter at 100% schema coverage, the schema already documents project_id fully (including where to find it). The description adds no syntax or format detail beyond that, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: a source-readiness check of the latest uploaded snapshot. The scope qualifiers ('free', 'no Apple calls or paid compute') further pin down what kind of check this is, distinguishing it from build/signing siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly routes the agent: 'Run push_project after source changes' names the sibling and the triggering condition. It stops short of stating when NOT to use check_project (e.g. versus build), so it lacks the full alternatives/exclusions treatment that would earn a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

connect_statusCheck Apple connectionA
Read-only
Inspect

Check the Apple App Store Connect connection: whether the API key works, when the signing certificate expires, and webhook health. Start here if any build, TestFlight or App Store tool fails; with no connection it explains the one-time setup for the human. Read-only.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
errorNoPresent only when the call failed

TDQS

A4.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false, so the description's 'Read-only' adds no new safety signal. However, it does add useful context beyond the structured fields: what the check reports and, importantly, that with no connection it surfaces one-time human setup instructions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tightly written sentences, front-loaded with the core purpose before the usage trigger and the no-connection behavior. No filler or repetition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so the description is not obligated to describe return values, and it covers purpose, when to invoke, and the diagnostic-failure path. For a zero-parameter read-only check, nothing an agent needs to call it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so the baseline of 4 applies. There is no parameter syntax or format the description needs to compensate for, and it correctly focuses on behavior rather than inputs.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource (check the Apple App Store Connect connection) and enumerates exactly what is verified: API key validity, signing certificate expiry, and webhook health. This clearly distinguishes it from write-oriented siblings like set_project_connection.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit when-to-use guidance: 'Start here if any build, TestFlight or App Store tool fails.' It names concrete trigger conditions (build failure, missing connection) and describes the fallback path (explains one-time setup), leaving nothing to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

exec_macRun a command on a MacA
Destructive
Inspect

Run arbitrary argv in your Mac workspace. Returns a durable job ID; poll get_mac_job and read_mac_output. Preserve request_key and exact arguments after lost responses. Use bash -lc explicitly for shell syntax. timeout_seconds defaults to 300 and caps at 600; raise it for slow work such as a first simulator boot, which needs about 90 seconds.

ParametersJSON Schema
NameRequiredDescriptionDefault
argvYesCommand and arguments as an array, for example ["bash", "-lc", "xcodebuild -list"]. Runs as the agent user in the workspace directory.
session_idYesMac session ID (ses_…) returned by start_mac or list_mac_sessions.
request_keyYesIdempotency key you choose (letters, digits, dash, underscore). Retry a lost response with the same key and identical arguments; use a new key for new work.
timeout_secondsNoStop the command after this many seconds (1-600, default 300).

Output Schema

ParametersJSON Schema
NameRequiredDescription
errorNoPresent only when the call failed

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover destructive/openWorld, and the description adds the facts an agent actually needs: the call is asynchronous and returns a durable job ID rather than output, timeout defaults to 300 and caps at 600, ~90s is expected for a first simulator boot, and request_key + identical args is the recovery path after a lost response. This is genuinely additive behavioral context, not restatement.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four tightly packed sentences with the core action front-loaded, followed by the follow-up tools, correctness rule, and timeout behavior in priority order. No filler or repetition of the title.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The async job model, the full retrieval path (get_mac_job/read_mac_output), timeout bounds, retry semantics, and shell invocation guidance are all present, and an output schema exists so return-value detail is not needed here. Nothing required to call this tool correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3, but the description adds real meaning beyond it: the bash -lc guidance for shell syntax and the simulator-boot rationale for raising timeout_seconds go past what the schema fields say. The idempotency explanation for request_key largely duplicates the schema text, which keeps this from a 5.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a concrete verb+resource ("Run arbitrary argv in your Mac workspace") and immediately distinguishes itself from siblings by naming get_mac_job and read_mac_output as the polling/read path. An agent can tell this is the execution entry point in the Mac session family without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear operational routing (poll get_mac_job, read_mac_output) and a specific usage rule (use bash -lc explicitly for shell syntax) plus a sizing heuristic for timeout_seconds on slow work like a simulator boot. It does not state when NOT to use the tool or what alternative exists for non-argv operations, so it falls short of a full 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_billing_statusGet build plan statusA
Read-only
Inspect

Read confirmed payment, contracted price, quota/reset, renewal and guarantee. After backend payment confirmation, resume a blocked build once with a new request_key only if the earlier response said build_started:false. Reconcile unknown build outcomes first.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
errorNoPresent only when the call failed

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, openWorldHint=false and destructiveHint=false, so the safety profile is covered structurally. The description adds which fields are read and a recovery precondition, but the 'resume a blocked build once with a new request_key' instruction is ambiguous on a read-only tool and is not clearly scoped to this tool rather than the build sibling.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The returned-data list is front-loaded and compact, which is good. However, the trailing workflow sentences (resume the build, reconcile outcomes) are dense and read as belonging to the build tool rather than a status read, which muddies the structure.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, return values need not be explained, and with zero parameters there is no input contract to document. The description supplies the field inventory and a workflow precondition, leaving it largely complete for a zero-param read tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There are zero input parameters, so the baseline is 4 and the schema cannot be under-served. The mention of a 'new request_key' is not a parameter here and could briefly confuse the agent, but no parameter semantics are actually missing.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The opening clause names a specific read verb and enumerates the resource fields it returns (payment, contracted price, quota/reset, renewal, guarantee), so the agent knows this is a billing/plan status read. It does not, however, distinguish itself from billing-adjacent siblings like get_checkout_link or get_refund_status, and the title ('Get build plan status') drifts from the name ('get_billing_status').

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The second sentence gives a conditional workflow ('resume a blocked build once ... only if the earlier response said build_started:false') and a precondition ('reconcile unknown build outcomes first'), which implies context. But it never tells the agent when to call this tool versus the checkout/refund/build siblings, and the guidance is tangled.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_failureExplain a failed buildA
Read-only
Inspect

Structured failure for a failed build: {stage, code, kind, error_lines, hints}. kind=project → fix source and re-push; kind=account → relay the deep link to your human.

ParametersJSON Schema
NameRequiredDescriptionDefault
build_idYesNoMac build ID (bld_…) returned by build.

Output Schema

ParametersJSON Schema
NameRequiredDescription
errorNoPresent only when the call failed

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, destructiveHint=false, openWorldHint=false, so the safety profile is covered. The description adds genuinely useful behavioral context the annotations lack: the exact response fields and an interpretation contract for kind values that tells the agent how to act on the output.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two dense sentences, front-loaded with the return shape followed by the action semantics. No filler, no restatement of the title.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present and one fully documented required parameter, the description is nearly complete for a read-only diagnostic call. Enumerating the return fields is mildly redundant with the output schema, but the kind→action mapping is the piece the structured data alone could not convey.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the single build_id parameter already documents the bld_… format from build. The description contributes no additional parameter syntax or format detail, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific resource and scope: a structured failure object {stage, code, kind, error_lines, hints} for a failed build. That is concrete and distinguishable from generic siblings like status. It does not explicitly name which sibling it replaces, keeping it just short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied by 'for a failed build' and the description gives post-call routing for the two kind values (project → fix source, account → relay deep link). It never states when to call this versus siblings such as status or get_mac_job, so guidance is only partially covered.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_feedbackGet TestFlight feedback and crashesA
Read-only
Inspect

Read recent TestFlight crash reports and written tester feedback for the project's app. Use it after testers install a build, to find crashes and complaints from real devices before the next build. Read-only.

ParametersJSON Schema
NameRequiredDescriptionDefault
project_idYesNoMac project ID (prj_…). push_project and connect_status list your projects.

Output Schema

ParametersJSON Schema
NameRequiredDescription
errorNoPresent only when the call failed

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, destructiveHint=false, and openWorldHint=false, so the safety profile is fully covered by structured data. The description's 'Read-only' merely repeats that, and the only extra behavioral signal is 'recent' (an implied time window) with no detail on volume, pagination, or limits. Modest added value over annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences are front-loaded with the resource and the usage trigger, with zero filler. The trailing 'Read-only.' is slightly redundant given readOnlyHint=true, a minor waste that keeps it just under a 5.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be documented, and the description covers what is read and when to call it. It is complete enough for an agent to invoke correctly, though it says nothing about the scope of 'recent' or result volume.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There is a single required parameter (project_id) and schema description coverage is 100%, with the schema itself pointing users to push_project/connect_status for discovery. Baseline 3 is appropriate since the description adds no parameter meaning beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a concrete verb ('Read') and two specific resources (TestFlight crash reports and written tester feedback) scoped to the project's app, which clearly separates it from write-oriented siblings like manage_testflight. It does not explicitly name the nearest alternative (get_testflight) to disambiguate, so it falls short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Use it after testers install a build, to find crashes and complaints from real devices before the next build' gives a clear trigger and intent for the tool. No explicit when-not conditions or named alternative tools are provided, so it is clear context without exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_mac_checkoutGet Mac credit payment linkAInspect

Get a payment link for Mac credit, for the human to open. Credit is $0.80 an hour, billed by the minute, sold in $10 packs (12.5 hours each), 1 to 10 packs per purchase, and never expires. Call this when get_mac_credits shows too little credit or start_mac returns insufficient_credit. Then give the human the returned url and ask them to open it and pay; if you can run a shell, also open it for them (macOS: open URL, Linux: xdg-open URL, Windows: start URL). Never enter payment details yourself. Poll get_mac_credits until the balance appears, then start the Mac. Reuse request_key after a lost response; use a new one for a new purchase.

ParametersJSON Schema
NameRequiredDescriptionDefault
packsNoNumber of $10 packs, 1-10. Pick from the human's stated budget or expected hours; default 1.
request_keyYesIdempotency key you choose (letters, digits, dash, underscore). Retry a lost response with the same key and identical arguments; use a new key for new work.

Output Schema

ParametersJSON Schema
NameRequiredDescription
errorNoPresent only when the call failed

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well beyond the annotations: states pricing and billing mechanics ($0.80/hr, per-minute, $10 packs, 1-10 packs, never expires), the idempotency contract ('reuse request_key after a lost response; use a new one for a new purchase'), platform-specific open commands, and an explicit safety rule ('never enter payment details yourself'). This is exactly the extra behavioral context annotations cannot carry.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the action and then packed with genuinely load-bearing facts (price, pack sizing, workflow, shell commands, retry rule); no filler sentences, though it is on the dense side with several clauses chained together.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values need not be described, yet the description still tells the agent what to do with the returned url and how the operation fits between get_mac_credits, start_mac, and payment. Nothing needed to call it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3; the description adds economic meaning to the 'packs' parameter ('$10 packs (12.5 hours each), 1 to 10 packs per purchase, and never expires') that helps an agent size the request, and reinforces the request_key retry semantics beyond the schema wording.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Get a payment link for Mac credit') plus its intended recipient ('for the human to open'), which distinguishes it from the generic sibling get_checkout_link and from read-only credit tools like get_mac_credits.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit trigger conditions ('when get_mac_credits shows too little credit or start_mac returns insufficient_credit') and names the alternatives by name, then lays out the follow-up sequence (give url, poll get_mac_credits, then start the Mac). When-to-use and what-to-do-next are both fully specified.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_mac_creditsGet Mac credit balanceA
Read-only
Inspect

Read the prepaid Mac credit balance: spendable dollars, spendable_seconds (how long a Mac can run on it), amounts held by a running session, pricing and recent purchases. Credit never expires. If spendable_seconds is below what the task needs, call get_mac_checkout.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
errorNoPresent only when the call failed

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover safety (readOnlyHint=true, destructiveHint=false, openWorldHint=false), so the description is not required to restate them. It still adds real context beyond structured data: credit never expires, and the balance includes amounts held by an in-flight session, which explains why spendable_seconds may differ from the raw balance.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, front-loaded with the core purpose and the enumerated fields, then the business rule, then the escalation path. No restatement of the title and no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be fully specified, yet the description still names the key fields an agent should inspect. Combined with the never-expires rule and the get_mac_checkout fallback, an agent has everything needed to call and act on this tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so the baseline of 4 applies. The description correctly spends no space on argument semantics and instead documents the shape of the result.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a precise verb+resource ('Read the prepaid Mac credit balance') and enumerates the concrete payload fields: spendable dollars, spendable_seconds, amounts held by a running session, pricing, and recent purchases. It also names a sibling (get_mac_checkout), so an agent can distinguish it without opening another schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives a clear conditional: if spendable_seconds is below what the task needs, call get_mac_checkout. That is genuine routing guidance, though it doesn't contrast against nearby read tools like get_billing_status or get_mac_session, so the 'when not to use this' dimension is only partially covered.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_mac_jobGet Mac command statusA
Read-only
Inspect

Check a command started by exec_mac: running, exited or failed, with exit code and output size. Poll it until the job finishes, then read the output with read_mac_output. Never reruns the command. Read-only.

ParametersJSON Schema
NameRequiredDescriptionDefault
job_idYesCommand job ID (job_…) returned by exec_mac.
session_idYesMac session ID (ses_…) returned by start_mac or list_mac_sessions.

Output Schema

ParametersJSON Schema
NameRequiredDescription
errorNoPresent only when the call failed

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false; the description reinforces this and adds a genuinely useful behavioral fact: 'Never reruns the command.' It also implies safe repeated polling. idempotentHint=false is consistent with polling (results change from running to exited), so no contradiction. Lacks any note on pacing/backoff, which keeps it short of a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three compact sentences, front-loaded with what is checked, then the workflow, then the safety guarantee. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Output schema exists, so return-value details need not be explained, and the description still hints at exit code and output size. Combined with the hand-off to read_mac_output and the strict no-rerun guarantee, an agent has everything needed to call and sequence it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% with only two required params, so the schema already documents job_id and session_id fully, including their provenance formats. The description adds provenance context (job IDs come from exec_mac) but nothing beyond schema. Baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (check) plus the resource (a command started by exec_mac) and enumerates the observable states (running, exited, failed) plus returned fields (exit code, output size). This is clearly distinguishable from siblings like exec_mac and read_mac_output.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly instructs to poll until the job finishes and then hand off to read_mac_output, naming the alternative and the condition that selects it. No inference required.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_mac_sessionGet Mac sessionA
Read-only
Inspect

Read one Mac session: state, funded time, cost so far and whether cleanup is confirmed. Poll it after start_mac until state is ready (usually under a minute), and after stop_mac until cleanup_confirmed is true. queued or provisioning means wait; stopping is not yet confirmed cleanup. Read-only.

ParametersJSON Schema
NameRequiredDescriptionDefault
session_idYesMac session ID (ses_…) returned by start_mac or list_mac_sessions.

Output Schema

ParametersJSON Schema
NameRequiredDescription
errorNoPresent only when the call failed

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false, so "Read-only" is largely redundant. However, the description adds genuine behavioral context beyond the annotations: expected timing ("usually under a minute"), polling cadence, and the meaning of transitional states.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, front-loaded with what is returned, then the polling lifecycle, then state interpretation. No filler; every clause carries operational meaning.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return structure needn't be re-explained, yet the description still orients the agent on which fields matter (state, cleanup_confirmed). Combined with read-only annotations and clear polling guidance, nothing needed to invoke this correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the single session_id parameter is fully documented in the schema with pattern and origin (returned by start_mac or list_mac_sessions). The description adds no syntax or format detail beyond that, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ("Read one Mac session") and enumerates exactly what it returns: state, funded time, cost so far, cleanup confirmation. It is clearly distinguishable from list_mac_sessions and read_mac_output among the siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit when/when-not routing: poll after start_mac until state is ready, poll after stop_mac until cleanup_confirmed is true. It also decodes intermediate states (queued/provisioning = wait; stopping is not yet confirmed cleanup), so the agent knows how to interpret responses without guessing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_metadataGet App Store metadataA
Read-only
Inspect

Read the App Store listing for one iOS version: description, keywords, what's new, URLs, category and review contact. Defaults to the marketing version of the last pushed source. Use it before editing with set_metadata. Read-only.

ParametersJSON Schema
NameRequiredDescriptionDefault
project_idYesNoMac project ID (prj_…). push_project and connect_status list your projects.
version_stringNoApp Store version, for example 1.2.0. Defaults to the marketing version of the last pushed source.

Output Schema

ParametersJSON Schema
NameRequiredDescription
errorNoPresent only when the call failed

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, destructiveHint=false, and openWorldHint=false, so the safety profile is covered; the description's 'Read-only' is largely redundant. The genuinely additive detail is the default-resolution behavior ('Defaults to the marketing version of the last pushed source'), which explains what version is actually read when the parameter is omitted.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short, front-loaded sentences with no filler: scope, default behavior, and the routing hint appear in that order. Every clause carries information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a two-parameter read tool with a 100% covered schema and an output schema, the description supplies everything an agent needs. Listing the returned fields is mildly redundant given the output schema, but it helps the agent confirm the tool's scope.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and both parameters are documented in the schema, so the baseline is 3. The description's note about defaulting to the marketing version restates the version_string schema description rather than adding new semantics, and the field list refers to outputs rather than inputs.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Read) plus resource (the App Store listing for one iOS version) and even enumerates the fields returned (description, keywords, what's new, URLs, category, review contact). It is clearly separable from siblings set_metadata and get_metadata_schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly routes the agent: 'Use it before editing with set_metadata,' establishing the read-before-write pairing with a named sibling. It gives clear context but stops short of stating when this is not the right tool (e.g., vs. read_store or get_metadata_schema).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_metadata_schemaGet App Store metadata rulesA
Read-only
Inspect

Get the rules for App Store metadata: every editable field with its character limit, supported locales, and screenshot display types with their exact pixel sizes. Call it before set_metadata or upload_screenshots. Read-only.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
errorNoPresent only when the call failed

TDQS

A4.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false, and the trailing 'Read-only' merely restates that, earning no credit. The description adds useful framing that this is a static rule set to consult prior to mutation, but says nothing about auth needs, caching/staleness, or rate limits beyond what annotations cover.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, zero waste, and the payload contents are front-loaded before the usage instruction. The closing 'Read-only' is the only mildly redundant clause.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, the description needn't enumerate return values in detail, and annotations already cover the safety profile for a zero-parameter call. Nothing an agent needs in order to invoke this correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so the schema carries no semantics to add; baseline 4 applies. The description correctly implies a parameterless read that needs no input.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Specific verb + resource ('Get the rules for App Store metadata') followed by an enumeration of exactly what is returned: editable fields, character limits, supported locales, and screenshot pixel sizes. This clearly distinguishes it from siblings like get_metadata, get_store_capabilities, and set_metadata.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Call it before set_metadata or upload_screenshots' gives an explicit precondition and names the two sibling tools it precedes, which is strong routing guidance. It stops short of stating when not to call it, so it isn't a full 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_refund_statusGet refund statusA
Read-only
Inspect

Read the progress of a refund the account owner requested at nomac.app/usage. Agents cannot request refunds; point the human to that page. Does not submit a refund or cancel anything.

ParametersJSON Schema
NameRequiredDescriptionDefault
request_idYesRefund request ID. Refunds are started by the account owner at nomac.app/usage.

Output Schema

ParametersJSON Schema
NameRequiredDescription
errorNoPresent only when the call failed

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false, but the description reinforces this by explicitly disclaiming side effects ('Does not submit a refund or cancel anything') and stating the agent's capability boundary. It does not address the idempotentHint=false flag or any auth/rate-limit behavior, so it is strong but not exhaustive.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, front-loaded with the read action and scoping, followed by the negative capability and side-effect disclaimers. No filler or redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need no explanation. For a single-parameter read tool, the description covers purpose, ownership, and the critical boundary that agents cannot initiate refunds, leaving nothing an agent needs to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The single request_id parameter is fully documented in the schema (100% coverage), including where the refund originates. The description repeats that origin context but adds no new format, validation, or lookup semantics beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Read the progress of a refund') plus precise scope ('the account owner requested at nomac.app/usage'). It is clearly distinguishable from billing siblings like get_billing_status or get_checkout_link without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly scopes who this is for and what it is not for: agents cannot request refunds and should redirect the human to nomac.app/usage. It gives clear context but never names a sibling alternative by name for related needs (e.g. where to check billing state instead).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_reviewGet App Review statusA
Read-only
Inspect

Read live App Review submissions and all item states, including submissions created outside nomac. Pass asc_submission_id for version/build/contact/attachment details. Apple does not expose reviewer correspondence through its public API; the result provides an explicit App Store Connect handoff.

ParametersJSON Schema
NameRequiredDescriptionDefault
project_idYesNoMac project ID (prj_…). push_project and connect_status list your projects.
asc_submission_idNoApple App Review submission ID, from get_review.

Output Schema

ParametersJSON Schema
NameRequiredDescription
errorNoPresent only when the call failed

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint/destructiveHint/openWorldHint, so the safety profile is covered. The description still adds real value beyond them: it discloses the scope ('submissions created outside nomac') and an important limitation ('Apple does not expose reviewer correspondence through its public API'), which the agent could not learn from annotations or schema alone.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, front-loaded with the core read action and scoping, with no filler. Slightly dense but every sentence carries distinct information (scope, parameter effect, API limitation).

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be described, and the description covers scope, the optional parameter's effect, and a known data limitation. The main missing piece is routing guidance against the review/testflight siblings, which keeps it short of a 5.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds meaning the schema lacks: it explains that asc_submission_id unlocks version/build/contact/attachment details, whereas the schema only labels it as an ID sourced 'from get_review'. This makes it clear why and when to pass the optional parameter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Read live App Review submissions and all item states') and even extends scope to submissions created outside nomac. It does not, however, name or distinguish itself from close siblings like get_testflight, update_review, or prepare_review_reply, so the agent must infer the boundary.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The line 'Pass asc_submission_id for version/build/contact/attachment details' implies when to supply the optional parameter, and 'read live' implies a status-check use case. But there is no explicit when-to-use/when-not guidance and no named alternative among the review-related siblings, so usage is only implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_screenshot_uploadGet screenshot upload statusA
Read-only
Inspect

Check screenshot replacement or restoration. Omit operation_id to list recent operations after a lost upload response. complete confirms Apple processed every image; rolled_back means the previous set was retained/restored. needs_attention requires support reconciliation.

ParametersJSON Schema
NameRequiredDescriptionDefault
project_idYesNoMac project ID (prj_…). push_project and connect_status list your projects.
operation_idNoOperation ID from upload_screenshots. Omit to list recent operations.

Output Schema

ParametersJSON Schema
NameRequiredDescription
errorNoPresent only when the call failed

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false, so safety is covered; the description adds genuinely useful semantics for the outcome states (complete/rolled_back/needs_attention) and the escalation path for needs_attention. It does not address idempotentHint=false or listing limits, keeping it below 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences: purpose first, then the invocation-mode note, then the status glossary. Every clause carries information and nothing is padded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values need not be described, and the description covers purpose, mode selection, and status meanings. Minor gap: no mention of pagination or how many 'recent operations' are returned, which matters for the listing mode.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both parameters are already documented, and the description's 'Omit operation_id to list recent operations' essentially restates the schema text. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (check) and resource (screenshot replacement/restoration), and the enumeration of status outcomes makes the operation-status nature unambiguous. It does not explicitly name or distinguish itself from upload_screenshots, which is the closest sibling, so it falls short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives a concrete trigger ('after a lost upload response') and explains the omit-operation_id listing mode, which tells the agent when each invocation style applies. No explicit when-not-to-use clause or named alternative, so not a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_store_capabilitiesList App Store Connect operationsA
Read-only
Inspect

Discover supported App Store Connect operations for listing, paid pricing/availability, subscriptions, in-app purchases, public customer-review replies, phased releases, events, product pages, files, analytics and webhooks. Pass category/search to list; pass operation for its exact Apple JSON:API schema. Review messages/appeals remain an App Store Connect browser workflow.

ParametersJSON Schema
NameRequiredDescriptionDefault
offsetNoPagination offset for long lists.
searchNoFree-text search over operation names and summaries.
categoryNoList operations in one area, for example pricing, subscriptions, testflight or analytics.
operationNoExact operation name, for example apps_getInstance, to get its full request schema.

Output Schema

ParametersJSON Schema
NameRequiredDescription
errorNoPresent only when the call failed

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, openWorldHint=false and destructiveHint=false, so the safety profile is covered. The description adds the useful scope exclusion about review messages/appeals, but says nothing about pagination limits or the closed nature of the capability list beyond what annotations imply.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, front-loaded with the purpose and then the key usage instruction. The long enumeration of supported areas is dense but informative; the only mild cost is the padded category list.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, return values need not be described, and the description covers purpose, usage modes and scope exclusions. What remains thin is pagination/rate behavior for the offset-based listing, but that is a minor gap for a read-only discovery tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so baseline is 3, but the description adds meaning beyond the schema by explaining the two modes of use — list mode (category/search) versus schema-retrieval mode (operation) — which clarifies the semantic role of each parameter rather than just restating its type.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource — discovering/listing supported App Store Connect operations — and enumerates the covered areas concretely. It is broadly distinguishable from siblings like get_store_operation, though the overlap around fetching an operation's schema is not explicitly disambiguated.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives explicit invocation guidance ('pass category/search to list; pass operation for its exact Apple JSON:API schema') and draws a clear out-of-scope boundary ('Review messages/appeals remain an App Store Connect browser workflow'). It does not name sibling alternatives, so it falls short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_store_operationGet saved App Store operationA
Read-only
Inspect

Read saved publish/metadata/review/TestFlight requests and recover a lost response or request_key. Resume pending work by repeating the original tool arguments with its returned request_key. This status call performs no Apple writes.

ParametersJSON Schema
NameRequiredDescriptionDefault
project_idYesNoMac project ID (prj_…). push_project and connect_status list your projects.
request_keyNoIdempotency key you choose (letters, digits, dash, underscore). Retry a lost response with the same key and identical arguments; use a new key for new work.

Output Schema

ParametersJSON Schema
NameRequiredDescription
errorNoPresent only when the call failed

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false, so the safety profile is covered; the description still adds value by framing it as a status call with no Apple writes and by explaining the lost-response recovery workflow. It doesn't describe the shape/pagination of returned operation records, but the output schema handles that.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences, front-loaded with the read scope, then the recovery recipe, then the reassurance about no writes. Each sentence carries distinct information with no redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present and full annotation coverage, the description supplies the one thing structured fields can't: the retry/recovery workflow that motivates the call. Only minor gaps remain, such as whether multiple stored operations can be listed versus a single lookup by key.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and both project_id and request_key are fully documented in the schema, so the description need not carry parameter detail. It reinforces request_key retry semantics but adds nothing the schema lacks, matching the baseline for complete schema coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (read) and resource (saved publish/metadata/review/TestFlight requests), enumerating the operation categories it retrieves. It is clear enough to differentiate this record-retrieval tool from artifact-readers like get_metadata, get_review, and get_testflight, though it never names those siblings to sharpen the distinction.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives concrete when-to-use conditions: recover a lost response or request_key, and resume pending work by re-invoking the original tool with the returned key. It stops short of explicit exclusions or naming the alternative tools to use for other cases.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_testflightGet TestFlight statusA
Read-only
Inspect

Read external/internal groups, test information and recent Apple builds. Pass an exact asc_build_id for beta-review state, group access and what to test. Internal readiness does not imply external approval.

ParametersJSON Schema
NameRequiredDescriptionDefault
group_idNoRead all testers in this app-owned group
project_idYesNoMac project ID (prj_…). push_project and connect_status list your projects.
asc_build_idNoApple build ID, from get_testflight.

Output Schema

ParametersJSON Schema
NameRequiredDescription
errorNoPresent only when the call failed

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false, so the safety profile is covered. The description adds a genuine domain caveat ('Internal readiness does not imply external approval') and enumerates the data scope returned, but says nothing about permissions, pagination, or rate limits.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three compact sentences, front-loaded with the resource scope and followed by the parameter cue and the caveat. No filler, though the phrasing 'Read external/internal groups, test information and recent Apple builds' is slightly listy rather than crisp.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present and annotations covering safety, the description need only convey scope and the one non-obvious caveat, which it does. The main remaining gap is sibling routing (manage_testflight), which an agent must infer from names alone.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description earns an extra point by clarifying the semantics of asc_build_id — it must be exact and it unlocks beta-review state, group access, and test guidance — which is more than the schema's circular 'from get_testflight' note provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a clear verb (Read) and specific resources: external/internal groups, test information, and recent Apple builds. An agent can distinguish it from the mutating sibling manage_testflight by the read framing, though the description never names that sibling explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives one conditional usage cue — pass an exact asc_build_id to get beta-review state, group access and what-to-test — which tells the agent when a parameter matters. However there is no explicit routing guidance against alternatives like manage_testflight, so usage is only implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_mac_sessionsList Mac sessionsA
Read-only
Inspect

List this account's Mac sessions, newest first, with state, cost and cleanup status. Use it to find a session ID, to check nothing is left running, or to recover after a lost start_mac response. Read-only.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
errorNoPresent only when the call failed

TDQS

A4.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, openWorldHint=false and destructiveHint=false, so 'Read-only' in the text is redundant. It does add the sort order and the fact that state/cost/cleanup status come back, which is modest behavioral value beyond the structured fields; no pagination or volume limits are disclosed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences, front-loading the resource and ordering before the usage scenarios. No filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be spelled out, and annotations carry the safety profile. The description covers purpose, ordering and recovery scenarios; only result-volume/pagination behavior is unaddressed for a listing tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Zero parameters, so there is nothing for the description to disambiguate; the baseline for an argument-less tool applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (List) and resource (this account's Mac sessions) plus scope, ordering ('newest first'), and returned fields ('state, cost and cleanup status'). An agent can separate it from the singular sibling get_mac_session purely on the plural list semantics.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives three concrete when-to-use cases: finding a session ID, verifying nothing is left running, and recovering from a lost start_mac response. No explicit when-not or named alternative (e.g. get_mac_session for a known ID), so it stops short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

manage_testflightManage TestFlight testingA
Destructive
Inspect

Set test metadata/contact/demo credentials, create/update external groups and public links, invite/remove an explicit tester, submit/distribute a specific Apple build to selected external groups, notify after approval, expire testing, or relay a human-provided encryption answer. Preview without confirm; confirm:true applies. distribute may send to Beta App Review; auto_notify defaults false. Never guess encryption/demo-account attestations. On pending, repeat unchanged arguments with request_key; get_store_operation recovers it.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameNoName of the TestFlight group to create or rename.
emailNoTester's email address, for invite_tester.
actionYesmetadata, create_group, update_group, invite_tester, remove_tester, remove_build, distribute, notify, expire or encryption.
localeNoApp Store locale code, for example en-US or de-DE. Defaults to en-US.en-US
confirmNofalse (default) returns a preview and changes nothing. true performs the change; set it only after the user approves that exact preview.
app_infoNoBeta app description, feedback email, marketing URL and privacy policy URL.
group_idNoTestFlight group ID from get_testflight, for update_group, invite_tester or remove_tester.
group_idsNoExternal TestFlight group IDs from get_testflight, for distribute.
last_nameNoTester's last name, for invite_tester.
tester_idNoApple tester ID from get_testflight, for remove_tester.
whats_newNoWhat to Test notes shown to testers for this build.
first_nameNoTester's first name, for invite_tester.
project_idYesNoMac project ID (prj_…). push_project and connect_status list your projects.
auto_notifyNoNotify testers automatically when the build becomes available. Defaults to false.
request_keyNoIdempotency key you choose (letters, digits, dash, underscore). Retry a lost response with the same key and identical arguments; use a new key for new work.
asc_build_idNoApple build ID, from get_testflight.
review_contactNoContact for Apple's reviewers (name, phone, email) and an optional demo account.
public_link_limitNoMaximum number of testers who can join through the public link (1-10000).
public_link_enabledNoTurn the group's public TestFlight link on or off.
uses_non_exempt_encryptionNoThe human's answer to Apple's export-compliance question. Never guess.

Output Schema

ParametersJSON Schema
NameRequiredDescription
errorNoPresent only when the call failed

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations cover destructive/openWorld/non-idempotent, but the description adds substantial context beyond them: a two-phase preview/confirm flow, that distribute may trigger Beta App Review, that auto_notify defaults false, and the request_key retry contract. This is genuinely useful behavioral disclosure for a destructive multi-action tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It front-loads the action list, then layers the critical procedural rules (preview/confirm, review submission, no guessing, idempotency) in compact sentences. It is dense but appropriate for a 10-action tool; no sentence is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With annotations declaring the safety profile and an output schema covering returns, the description fills the remaining gaps an agent needs: the preview/confirm gate, the Beta App Review risk, the encryption-attestation rule, and the idempotency recovery path. Complete enough to call correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so every parameter is already documented in the schema (including auto_notify's default and confirm's meaning). The description adds action-level semantics but no parameter-level detail beyond what the schema provides, so baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description enumerates concrete operations (set metadata, create/update groups, invite/remove tester, distribute build, expire testing, relay encryption answer) that map cleanly to the action enum, so the agent understands exactly what the tool does. It doesn't explicitly distinguish itself from the read sibling get_testflight, which keeps it from a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives clear procedural guidance: preview without confirm, confirm:true applies after user approval, never guess encryption/demo attestations, and on a pending result retry with the same request_key while get_store_operation recovers it. It names a sibling for recovery but does not state when to prefer this over get_testflight or set_metadata.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

prepare_review_replyDraft a reply to App ReviewA
Read-only
Inspect

Prepare YOUR authored reply to Apple App Review with the correct app handoff. Returns sent:false because Apple's public API cannot send or read reviewer messages. The human must paste and send the reply in App Store Connect. This is not a public customer-review response or a change to reviewer notes.

ParametersJSON Schema
NameRequiredDescriptionDefault
replyYesYour reply to App Review, up to 4000 characters. The human sends it in App Store Connect.
project_idYesNoMac project ID (prj_…). push_project and connect_status list your projects.
asc_submission_idYesApple App Review submission ID, from get_review.

Output Schema

ParametersJSON Schema
NameRequiredDescription
errorNoPresent only when the call failed

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well beyond the annotations: it discloses that the tool only drafts (returns sent:false), explains why (Apple's public API cannot send or read reviewer messages), and specifies the required human workflow. This is exactly the kind of behavioral context annotations cannot convey, and it is consistent with readOnlyHint=true and destructiveHint=false.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight, front-loaded sentences with no wasted words; the scope constraint and the not-this clauses are efficiently placed. Minor vagueness in 'the correct app handoff' keeps it from being flawless.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, return values need not be explained, yet the description still clarifies the sent:false behavior. Combined with annotations covering safety and a fully documented schema, nothing an agent needs in order to call this correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so each of the three parameters is already fully documented, and the description adds no syntax, format, or constraint detail beyond what the schema provides. The baseline of 3 applies when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Prepare YOUR authored reply to Apple App Review') and immediately names what it is NOT (not a public customer-review response, not a change to reviewer notes), which distinguishes it from siblings like update_review and get_review.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear negative guidance by excluding the public-review and reviewer-notes use cases, and states the human must paste and send in App Store Connect. It stops short of an explicit 'use X instead when Y' routing to a named sibling, so it falls just below the top band.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

publishSubmit for App Store reviewA
Destructive
Inspect

Submit for App Store review. Requires confirm:true with authorization for this app/version; run without confirm first to see the staged result + Apple blockers. On pending, repeat unchanged arguments with the returned request_key; get_store_operation recovers a lost response.

ParametersJSON Schema
NameRequiredDescriptionDefault
forceNoProceed despite non-blocking warnings shown in the preview. Leave false unless the user accepts them.
confirmNofalse (default) returns a preview and changes nothing. true performs the change; set it only after the user approves that exact preview.
build_idNoExact nomac release build; use this when replacing a rejected binary
project_idYesNoMac project ID (prj_…). push_project and connect_status list your projects.
request_keyNoOnly supply to resume an interrupted invocation with unchanged arguments
asc_submission_idNoExact Apple submission from get_review
resolve_rejectionNoExplicitly assert the selected version's rejection was fixed and mark its item ready for review

Output Schema

ParametersJSON Schema
NameRequiredDescription
errorNoPresent only when the call failed

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the annotations (destructive, non-idempotent, open-world), the description discloses the authorization prerequisite, the safe dry-run pattern, the pending state and request_key resumption, and a recovery tool for lost responses. These are exactly the traits an agent cannot get from annotations alone.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tightly packed sentences with the action front-loaded and the confirm workflow, pending state, and recovery pointer following in priority order. No filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values needn't be described. The description covers the destructive-confirm pattern, the pending/resume edge case, and failure recovery, which is complete for a mutation tool of this complexity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so confirm, request_key, and the other parameters are already documented in the schema descriptions. The description reinforces the confirm/request_key workflow but adds no syntax or format detail beyond what the schema already provides, making the 3 baseline correct.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Submit for App Store review') that is immediately distinguishable from read/update siblings like get_review, update_review, and review_lint. An agent knows exactly what action is being performed.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly prescribes the sequence: run without confirm first to see the staged result + Apple blockers, then repeat with confirm:true only with authorization. It also names the recovery path (get_store_operation) and the pending-resume condition, leaving no inference required.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

push_projectHow to push source codeB
Read-only
Inspect

Hosted transport cannot read your working tree — push from where the code lives: run npx @nomac/cli login once, then npx @nomac/cli push in the project directory (or use the stdio server nomac mcp, whose push_project packs locally). Returns current projects instead.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
errorNoPresent only when the call failed

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false, so the safety profile is covered. The description adds real behavioral context: the hosted transport cannot read the working tree, and the tool returns current projects instead of pushing. It never resolves the tension between a tool named 'push_project' and a read-only hint, but it does not contradict the annotations either.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is a single sentence, and the key constraint ('Hosted transport cannot read your working tree') is front-loaded. But it crams login steps, push commands, an alternative server mode, and the return value into one dense line without clean structure.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be spelled out, and the description does address the hosted-transport constraint and fallback. What remains incomplete is the core identity of the tool: an agent still cannot cleanly tell whether push_project moves code or just lists projects.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so the schema is trivially complete and there is no parameter semantics for the description to add. Baseline 4 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description ultimately states the tool's real behavior ('Returns current projects instead'), but its center of gravity is a CLI push walkthrough, so what push_project actually does is buried and ambiguous. It also never distinguishes itself from siblings like status, check_project, or build, which is confusing given the name implies a mutation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It routes the agent away from this tool for real pushes ('push from where the code lives') and names concrete alternatives (npx CLI, the stdio server nomac mcp), which is useful. However, it never says positively when to call push_project itself versus those alternatives, so the usage guidance is only implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

read_mac_outputRead Mac command outputA
Read-only
Inspect

Read stdout/stderr as base64 chunks; resume from next_cursor. If truncated, the final response includes tail_base64 and its absolute tail_offset. Retrieve promptly: only recent completed jobs retain output. output_expired preserves the job receipt and never means rerun. Retrieve before stopping the VM.

ParametersJSON Schema
NameRequiredDescriptionDefault
cursorNoByte offset to read from. Start at 0, then pass next_cursor from the previous response.
job_idYesCommand job ID (job_…) returned by exec_mac.
streamNostdout (default) or stderr.
session_idYesMac session ID (ses_…) returned by start_mac or list_mac_sessions.

Output Schema

ParametersJSON Schema
NameRequiredDescription
errorNoPresent only when the call failed

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the readOnly/destructive annotations, it discloses chunking with cursor-based resumption, the truncation contract (tail_base64 plus absolute tail_offset), a retention window, and the crucial non-error semantics of output_expired. This is exactly the kind of context annotations cannot convey and it prevents a needless re-execution.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four tight sentences ordered from core mechanic to truncation handling to retention risk to the VM-stopping warning. No filler and the most actionable warning (retrieve before stopping) is saved for emphasis at the end.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, return fields need not be explained, and the description still covers truncation, retries, and expiry semantics thoroughly. The only shortfall is that it never positions this tool against the sibling that returns job metadata, so an agent has no explicit cue for choosing between them.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents cursor, stream, job_id, and session_id, including the next_cursor resumption pattern and the stdout default. The description restates cursor resumption but adds no format or boundary details beyond the schema, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource (read stdout/stderr output as base64 chunks) with the resumption mechanism, and is clearly distinct from siblings like get_mac_job or exec_mac. An agent knows immediately this is the raw output reader rather than the job status tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives real timing guidance: retrieve promptly because only recent completed jobs retain output, and retrieve before stopping the VM. It also tells the agent how to interpret output_expired (do not rerun). However, it never names an alternative tool or a when-not condition, so the routing to siblings is implied rather than explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

read_storeRead App Store Connect dataA
Read-only
Inspect

Read a discovered Apple GET operation for the linked app. Begin with an apps_* relationship and scope:[]; copy each returned resource's scope into subsequent calls. Paginate with next_cursor and unchanged query. Raw IDs are accepted only for global lookups. This tool only reads. File reservations expose Apple uploadOperations; analytics segments expose report URLs.

ParametersJSON Schema
NameRequiredDescriptionDefault
queryNoQuery parameters for the Apple GET operation, such as filters, fields and limit.
scopeNoParent resource chain copied from earlier read_store results, as [{relationship, id}]. Start with [] for app-level reads.
cursorNonext_cursor from the previous page. Keep the other arguments unchanged.
operationYesApp Store Connect operation name from get_store_capabilities, for example appInfos_getInstance.
project_idYesNoMac project ID (prj_…). push_project and connect_status list your projects.
resource_idNoRaw Apple resource ID, only for global lookups outside the linked app.

Output Schema

ParametersJSON Schema
NameRequiredDescription
errorNoPresent only when the call failed

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With readOnlyHint already declared, the description goes beyond annotations by disclosing the scope-propagation workflow, the pagination contract ('next_cursor and unchanged query'), and what specific operations return (uploadOperations, report URLs). 'This tool only reads' merely restates the annotation, but the surrounding behavioral detail is genuinely additive.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Five compact sentences, front-loaded with purpose and followed by actionable operational rules; nothing is wasted. It is slightly dense and terse, which costs a point on readability but not on economy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has an output schema and annotations covering safety, and the description still supplies the non-obvious pieces an agent needs: how to start, how to chain scope, how to paginate, and when raw IDs are allowed. Nothing essential is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3, but the description adds workflow meaning: scope must be seeded with [] and copied from prior results, cursor requires unchanged other arguments, and resource_id is valid only for global lookups. This clarifies how parameters interact rather than restating their types.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('read') and resource ('a discovered Apple GET operation') and scopes it to the linked app, so the agent knows exactly what the tool operates on. It does not explicitly differentiate itself from nearby siblings such as get_store_operation or get_store_capabilities, which would have earned a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives concrete entry conditions ('Begin with an apps_* relationship and scope:[]'), a chaining rule for follow-up calls, and a restriction ('Raw IDs are accepted only for global lookups'). It stops short of naming when to prefer a sibling tool, so it is strong context without explicit alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

report_issueReport an issue to NoMac supportAInspect

File an unknown/unfixable error with nomac support. For a Mac use session_id and optional job_id; for a build use ref_id. Attaches account-scoped context and deduplicates repeats. Optional email requests a direct resolution update; omit description to read replies.

ParametersJSON Schema
NameRequiredDescriptionDefault
emailNoOptional address for a direct reply from NoMac support. Only when the user asks for one.
job_idNoCommand job ID (job_…) the issue is about.
ref_idNoBuild ID or other NoMac reference the issue is about.
session_idNoMac session ID (ses_…) the issue is about.
descriptionNoWhat went wrong, including the tool, arguments and error you saw. Omit it to read replies to earlier reports.

Output Schema

ParametersJSON Schema
NameRequiredDescription
errorNoPresent only when the call failed

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the annotations (which only declare readOnly=false, idempotent=false, destructive=false), the description discloses that account-scoped context is attached, that repeats are deduplicated, and that omitting description turns the call into a reply reader. That dual read/write behavior and the dedup guarantee are exactly the extra context annotations cannot convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four tight sentences, front-loaded with the core action, then routing, then side effects, then the reply mode. No filler and nothing repeated from the schema or annotations.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, return values need no explanation, and the description covers routing, side effects and the read mode for a zero-required-parameter tool. It is complete enough to invoke correctly, though it says nothing about auth requirements or submission limits, minor omissions for an openWorld=false support tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% so the baseline is 3, but the description adds meaning the schema lacks: it maps session_id (+ optional job_id) to the Mac case and ref_id to the build case, and clarifies that email is only for requesting a direct resolution update and that omitting description switches modes. This is useful selection logic beyond field definitions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (file) plus the resource (unknown/unfixable error) and the destination (NoMac support), and immediately gives the routing rule distinguishing a Mac session (session_id/job_id) from a build (ref_id). An agent can tell this apart from the many status/get_* siblings without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The scope qualifier 'unknown/unfixable error' tells the agent when this tool is warranted and implicitly when it is not (known/fixable failures). It also explains the two parameterization paths and the omit-description mode for reading replies. No sibling tool is named as an alternative, so it stops short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

review_lintRun App Review readiness checkAInspect

Review-readiness report (green/yellow/red). Red blocks publish. Findings carry evidence, fix hints, sometimes ready patches. 4.3-style findings are signals for YOU to judge.

ParametersJSON Schema
NameRequiredDescriptionDefault
project_idYesNoMac project ID (prj_…). push_project and connect_status list your projects.

Output Schema

ParametersJSON Schema
NameRequiredDescription
errorNoPresent only when the call failed

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description adds real value beyond annotations: output is a tri-state report, red blocks publishing, findings include evidence and fix hints, and 4.3-style findings are advisory signals for the agent to judge. However, annotations declare readOnlyHint=false, and the description never explains what side effect makes this a non-read-only operation (the 'sometimes ready patches' phrase hints at it but is not clarified).

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four short, front-loaded sentences with no filler; the blocking rule and output semantics come early. The '4.3-style' shorthand is jargon that assumes shared context but costs no space.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, the description correctly focuses on decision semantics (severity levels, blocking behavior, advisory findings) rather than return fields. It is largely complete for a one-parameter check, though it omits any note on side effects or run frequency.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Only one parameter and schema coverage is 100%; the schema's own description already explains the prj_ format and points to push_project and connect_status for discovery. The description adds nothing about project_id, so baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The title and description establish a specific function: a review-readiness lint that returns a green/yellow/red report. It is clearly distinguishable from siblings like get_review, build, and check_project, though the verb 'lint' itself is never stated plainly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Red blocks publish' implies a pre-publish gate, which is useful routing context, but the description never states when to run this versus check_project or status, nor any prerequisite (e.g. a connected project). Usage is implied rather than explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

set_metadataUpdate App Store metadataA
Destructive
Inspect

Write App Store metadata YOU authored (validated before any Apple call). fields: description/keywords/whats_new/support_url/…; plus primary_category, age_rating (human attestations), content_rights, copyright, review_contact, price:'FREE'. On pending, repeat unchanged arguments with the returned request_key; get_store_operation recovers a lost response.

ParametersJSON Schema
NameRequiredDescriptionDefault
priceNoOnly FREE is supported here. Use update_store for paid pricing.
fieldsNoListing text by field name: description, keywords, whats_new, promotional_text, support_url, marketing_url and so on. get_metadata_schema lists fields and character limits.
localeNoApp Store locale code, for example en-US or de-DE. Defaults to en-US.en-US
copyrightNoCopyright line shown on the App Store, for example "2026 Example Inc."
age_ratingNoAge-rating questionnaire answers. These are legal attestations: use only what the human told you.
project_idYesNoMac project ID (prj_…). push_project and connect_status list your projects.
request_keyNoReuse only to resume pending work with the original unchanged arguments
content_rightsNoWhether the app uses third-party content, as the human answered.
review_contactNoContact for Apple's reviewers (name, phone, email) and an optional demo account.
version_stringNoDefaults to the pushed source marketing version; select a live version explicitly for promotional text
primary_categoryNoApp Store primary category, for example UTILITIES or DEVELOPER_TOOLS.

Output Schema

ParametersJSON Schema
NameRequiredDescription
errorNoPresent only when the call failed

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare destructive=true and idempotent=false, and the description usefully adds that inputs are 'validated before any Apple call' and how to safely resume an in-flight write via request_key. This materially explains the non-idempotent mutation behavior beyond the structured hints.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the core action, but the body is a dense, punctuated run-on using '…/' shorthand that is harder to parse than necessary. Information is useful but packed without clean structure.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need no explanation, and the description fills the key remaining gap (the pending/retry workflow). For an 11-param destructive mutation with human-attestation constraints, it is nearly complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents every parameter, including the FREE-only price enum, locale default, and the attestation caveat on age_rating. The description reiterates these (price:'FREE', human attestations) but adds no new syntax or format detail, so baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Write App Store metadata') and enumerates the field families it touches. It clearly differs from read-side siblings (get_metadata, get_metadata_schema) and the write-side update_store, though it doesn't name those alternatives directly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides concrete operational guidance for the pending/idempotency path: repeat unchanged arguments with the returned request_key, and use get_store_operation to recover a lost response. It also flags that age_rating/content_rights must come from a human. It stops short of explicitly saying when to prefer update_store (that lives in the schema).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

set_project_connectionChoose a project's Apple connectionAInspect

Link a project to an Apple connection, for example after the human rotates their App Store Connect key. Verifies that the connection can see the app's bundle ID. Release builds already running must finish first.

ParametersJSON Schema
NameRequiredDescriptionDefault
project_idYesNoMac project ID (prj_…). push_project and connect_status list your projects.
connection_idYesApple connection ID, from connect_status.

Output Schema

ParametersJSON Schema
NameRequiredDescription
errorNoPresent only when the call failed

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare a non-read-only, non-idempotent mutation, and the description adds genuine behavioral detail beyond that: it performs a bundle-ID visibility check and will block while release builds are in flight. This is real context an agent could not infer from the structured fields. No contradiction with the mutation-shaped annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight clauses with the action front-loaded, then the motivating example, then the precondition. Every sentence earns its place and nothing is redundant.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, return values need not be explained, and the description covers the mutation's side effects and precondition. It is essentially complete for a two-parameter link operation, with only the absence of an explicit alternative to connect_status as a minor gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both parameters (project_id, connection_id) are already fully documented, including where to obtain them. The description adds no syntax or format detail beyond the schema, which is the expected baseline when the schema carries the load.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: "Link a project to an Apple connection." The mutation is unambiguous and distinct from the read-only connect_status sibling, though the description never names an alternative explicitly. A clear, self-contained statement of purpose.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives a concrete trigger ("after the human rotates their App Store Connect key") and a precondition ("Release builds already running must finish first"). It stops short of naming when NOT to use it or which sibling to prefer as an alternative, so it falls just below the top band.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

start_macStart a cloud MacAInspect

Allocate a full Mac VM, billed from prepaid credit at $0.80 an hour by the minute. If it returns insufficient_credit, call get_mac_checkout and have the human pay, then retry with the same request_key. public_key is optional for agents using HTTPS commands only. Reuse request_key and identical settings after a lost response. Poll get_mac_session until ready; inspect its offer, cost and limits. Stop deletes the VM and files. No predefined project pipeline.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitsNoOptional caps for this session: max_duration_seconds (60-86400), spend_cap_microunits (USD × 1,000,000) and idle_timeout_seconds (60-3600). Defaults: 2 hours, $5, 10 minutes idle.
public_keyNoOptional SSH public key (ed25519, RSA 2048+ or P-256) for direct SSH access. Omit when you only use exec_mac.
request_keyYesIdempotency key you choose (letters, digits, dash, underscore). Retry a lost response with the same key and identical arguments; use a new key for new work.

Output Schema

ParametersJSON Schema
NameRequiredDescription
errorNoPresent only when the call failed

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well beyond the annotations: prepaid-credit billing at a per-minute rate, the insufficient_credit error code, idempotent retry semantics via request_key, and critically that 'Stop deletes the VM and files' (destructive consequence not captured by destructiveHint=false). It also states there is no predefined project pipeline, preempting a wrong assumption about siblings like build/push_project.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the allocation action and cost, then sequences error handling, retry, and polling in roughly the order an agent would need them. It is dense but every clause carries operational information; the 'No predefined project pipeline' fragment is slightly terse and could be folded into a cleaner sentence.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a stateful, billable allocation tool with a nested limits object and an output schema, the description covers creation, cost, failure recovery, idempotent retry, readiness polling, and teardown consequences. Return-format details are legitimately delegated to the output schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds real meaning: public_key is optional specifically for agents using HTTPS/exec_mac only, and request_key must be reused only after a lost response with identical settings. It does not explain the limits object fields (idle_timeout, max_duration, spend cap), which the schema description covers but the prose does not reinforce.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Allocate a full Mac VM') and immediately bounds scope with cost ('billed from prepaid credit at $0.80 an hour by the minute'). This is clearly distinguishable from siblings like stop_mac, get_mac_session, and list_mac_sessions.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly routes the agent on the error path ('If it returns insufficient_credit, call get_mac_checkout...'), the lost-response path (reuse request_key with identical settings), and the readiness path ('Poll get_mac_session until ready'). It also states when public_key can be omitted ('agents using HTTPS commands only'), which is a genuine when/when-not condition.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

statusGet build statusA
Read-only
Inspect

Read a build's current state: from queued and building through uploading and processing to ready or failed. Poll it after build until ready or failed; on failed, call get_failure for the reason. Read-only.

ParametersJSON Schema
NameRequiredDescriptionDefault
build_idYesNoMac build ID (bld_…) returned by build.

Output Schema

ParametersJSON Schema
NameRequiredDescription
errorNoPresent only when the call failed

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover the safety profile (readOnlyHint, destructiveHint=false), so the bar is lower, and the description still adds real value by disclosing the state machine and the polling contract that an agent must follow. It stops short of stating rate limits or whether repeated polls are throttled, which would be the remaining useful detail.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, no filler; the lifecycle is front-loaded and the polling/alternative instruction follows. Every clause carries actionable information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return-value explanation is unnecessary; the description covers purpose, lifecycle, polling, failure routing, and read-only nature. For a single-param, read-only status tool this is fully sufficient.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the build_id schema text already explains it is the 'bld_…' ID returned by build, so the description adds nothing parameter-specific. Baseline 3 applies since the schema does all the work for this single required field.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Read') and resource (a build's current state) and enumerates the full lifecycle (queued, building, uploading, processing, ready/failed), so the agent immediately knows what is returned. It also implicitly separates itself from get_failure, the sibling that supplies the failure reason.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit when-to-use ('Poll it after build until ready or failed') and names the alternative and its triggering condition ('on failed, call get_failure for the reason'). This is a complete routing instruction, not an inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

stop_macStop and delete a MacA
Destructive
Inspect

Stop a Mac session: shuts down the VM and permanently deletes it and all its files, which ends billing. Read any output you need first with read_mac_output. Call it whenever the task is done so credit is not spent on an idle Mac, then poll get_mac_session until cleanup_confirmed. Cannot be undone.

ParametersJSON Schema
NameRequiredDescriptionDefault
session_idYesMac session ID (ses_…) returned by start_mac or list_mac_sessions.

Output Schema

ParametersJSON Schema
NameRequiredDescription
errorNoPresent only when the call failed

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true and idempotentHint=false, so the safety profile is covered; the description goes beyond that by specifying exactly what is destroyed ('it and all its files'), that billing stops, that the action is irreversible, and that cleanup requires polling. That is meaningful added context, though it does not discuss failure modes or partial-state behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences, front-loaded with the destructive action and its billing consequence, followed by the required pre-step and post-step. No filler and no repetition of the title.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter destructive tool with an output schema available, the description covers what it does, what it destroys, the reading prerequisite, and the follow-up polling. Nothing needed to invoke it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Only one parameter, and the schema already documents session_id at 100% coverage including the ses_… pattern and its origin (start_mac/list_mac_sessions). The description adds nothing about the parameter, so baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Stop a Mac session: shuts down the VM and permanently deletes it') plus the consequence (ends billing). An agent can distinguish it from read_mac_output, get_mac_session, and list_mac_sessions without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly names the prerequisites and alternatives: read output first with read_mac_output, call whenever the task is done to avoid idle spend, then poll get_mac_session until cleanup_confirmed. When-to-use and sequencing are fully specified.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update_reviewChange an App Review submissionA
Destructive
Inspect

Resolve a fixed rejected item, remove an item, cancel review, resubmit all ready items, or release an approved app version. Uses exact Apple IDs from get_review. Preview without confirm; confirm:true performs the action. Removing items cannot be undone in that submission. Fix the actual issue before resolve_item. On pending, repeat unchanged arguments with request_key; get_store_operation recovers it.

ParametersJSON Schema
NameRequiredDescriptionDefault
forceNoProceed despite non-blocking warnings shown in the preview. Leave false unless the user accepts them.
actionYesresolve_item (after fixing a rejection), remove_item, cancel, resubmit or release (an approved version).
confirmNofalse (default) returns a preview and changes nothing. true performs the change; set it only after the user approves that exact preview.
item_idNoApple review item ID from get_review. Needed for resolve_item and remove_item.
project_idYesNoMac project ID (prj_…). push_project and connect_status list your projects.
request_keyNoIdempotency key you choose (letters, digits, dash, underscore). Retry a lost response with the same key and identical arguments; use a new key for new work.
asc_submission_idYesApple App Review submission ID, from get_review.

Output Schema

ParametersJSON Schema
NameRequiredDescription
errorNoPresent only when the call failed

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare destructive/non-idempotent/open-world, and the description adds material context beyond them: preview-by-default with confirm:true committing the change, irreversible item removal in that submission, and the request_key idempotency retry pattern. These are exactly the traits an agent needs to avoid an accidental destructive call.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The action list is front-loaded and the operational rules follow compactly with no filler. Sentences are dense and comma-heavy, so a small amount of parsing effort is required, but nothing is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With 7 parameters, an enum action, rich annotations and an output schema, the description still covers the whole workflow: preview-then-confirm, idempotent retry, irreversibility, and the prerequisite fix before resolve_item. An agent can act correctly without consulting anything else.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description goes further by anchoring item_id and asc_submission_id to get_review provenance and explaining the confirm/request_key interaction in prose. It adds real cross-tool meaning rather than restating the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The first sentence enumerates the five concrete operations (resolve a rejected item, remove an item, cancel review, resubmit ready items, release an approved version) performed against an App Review submission, so the verb+resource+scope are unambiguous. It is clearly separable from read-only siblings like get_review and prepare_review_reply.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It routes the agent to get_review for the exact Apple IDs and to get_store_operation for recovering a pending request, and it states the precondition 'Fix the actual issue before resolve_item.' It stops short of explicit when-not-to-use guidance (e.g. when to prefer update_store or prepare_review_reply), but the action-level conditions are clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update_storeChange App Store Connect dataA
Destructive
Inspect

Execute a discovered Apple write using its exact JSON:API body. Copy scope and relationship references from read_store. Defaults to a preview; confirm:true changes real Apple resources. Supports pricing/offers, product creation and submission, public customer replies, release settings and asset reservations/commit. Use owner-provided attestations and prices. On pending, retry identical arguments with the returned request_key; do not create a fresh key. Use publish/update_review/manage_testflight for their dedicated actions.

ParametersJSON Schema
NameRequiredDescriptionDefault
bodyNoExact JSON:API request body, following the operation's schema from get_store_capabilities.
scopeNoParent resource chain copied from earlier read_store results, as [{relationship, id}]. Start with [] for app-level reads.
confirmNofalse (default) returns a preview and changes nothing. true performs the change; set it only after the user approves that exact preview.
operationYesApp Store Connect operation name from get_store_capabilities, for example appInfos_getInstance.
project_idYesNoMac project ID (prj_…). push_project and connect_status list your projects.
referencesNoRelated resource references the body needs, copied from read_store results.
request_keyNoIdempotency key you choose (letters, digits, dash, underscore). Retry a lost response with the same key and identical arguments; use a new key for new work.
asset_sha256NoSet by upload_store_asset for exact file-reservation recovery.

Output Schema

ParametersJSON Schema
NameRequiredDescription
errorNoPresent only when the call failed

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations flag destructiveHint=true and idempotentHint=false, but the description adds crucial behavior beyond them: the default is a non-mutating preview and only confirm:true touches real Apple resources, plus explicit idempotency/retry semantics via request_key. This is meaningful context an agent could not derive from the structured fields alone.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The core safety rule (preview default vs confirm:true) is front-loaded and every sentence carries operational weight, but the paragraph is dense with multiple distinct concerns (scope sourcing, supported domains, attestations, retry, routing) that could be more tightly grouped.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For an open-world destructive write with an output schema present, the description covers everything an agent needs: preview/confirm gating, parameter sourcing, idempotency recovery, and sibling routing. Return values are correctly left to the output schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds sourcing guidance the schema lacks: scope/references come from read_store, operation/body follow get_store_capabilities, and request_key must be reused identically on retry rather than regenerated. That elevates it above a pure schema restatement.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Execute a discovered Apple write using its exact JSON:API body') and enumerates the covered domains (pricing/offers, product creation, replies, release settings, asset reservations). It explicitly distinguishes itself from publish/update_review/manage_testflight, so an agent can route without opening sibling schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit when-to-use ('execute a discovered Apple write'), when-to-defer ('Use publish/update_review/manage_testflight for their dedicated actions'), a dependency ('Copy scope and relationship references from read_store'), and a recovery rule ('On pending, retry identical arguments with the returned request_key'). Nothing is left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

upload_screenshotsReplace App Store screenshotsA
Destructive
Inspect

Start a screenshot replacement for one display type and iOS version. PNGs are validated and prior images are backed up. Reuse request_key and the original body when retrying. Poll get_screenshot_upload until complete, rolled_back or needs_attention; accepted means still in progress. images: [{filename, data: base64 PNG}].

ParametersJSON Schema
NameRequiredDescriptionDefault
imagesYesScreenshots in display order as [{filename, data}], where data is a base64 PNG of the exact size for display_type.
localeNoApp Store locale code, for example en-US or de-DE. Defaults to en-US.en-US
project_idYesNoMac project ID (prj_…). push_project and connect_status list your projects.
request_keyNoUnique per replacement; reuse this key and the same images on a retry
display_typeYesApple screenshot display type, for example APP_IPHONE_67. get_metadata_schema lists types and exact pixel sizes.
version_stringNoDefaults to the pushed source marketing version

Output Schema

ParametersJSON Schema
NameRequiredDescription
errorNoPresent only when the call failed

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Adds substantive behavior beyond the annotations: PNGs are validated, prior images are backed up (contextualizing the destructiveHint=true), and the async lifecycle states are enumerated with the meaning of "accepted." The retry guidance also meaningfully supplements idempotentHint=false.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the action, then the safety/retry facts, then the lifecycle states — good ordering with no filler sentences. Minor deduction because the trailing inline images spec duplicates the schema without adding new information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained; the description instead covers validation, backup, retry semantics, and terminal polling states. Nothing an agent needs to invoke and correctly follow up on this async tool is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all six parameters are already documented in the schema, and the description's "images: [{filename, data: base64 PNG}]" restates rather than extends that. Baseline 3 applies since the schema carries the parameter burden.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource — "Start a screenshot replacement" — and scopes it to "one display type and iOS version," which cleanly separates it from sibling upload_store_asset and from the read-side get_screenshot_upload. An agent can identify the operation without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear post-call workflow guidance: poll get_screenshot_upload until complete/rolled_back/needs_attention, and reuse request_key plus the original body on retry. It does not, however, say when to prefer this over the sibling upload_store_asset, so the alternative-selection guidance is incomplete.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

upload_store_assetUpload an App Store assetAInspect

Upload a new Apple screenshot, preview, purchase image, supplemental reviewer attachment or encryption document. Discover its createInstance operation/body with get_store_capabilities and copy the parent scope. Hosted calls accept data_base64 up to 2 MiB; use CLI file_path for larger files. Preview by default. confirm:true reserves, uploads and commits this file. Preserve request_key and identical file/body to resume; ready:true requires Apple processing COMPLETE. App Review message attachments remain a browser workflow.

ParametersJSON Schema
NameRequiredDescriptionDefault
bodyNoExact JSON:API request body, following the operation's schema from get_store_capabilities.
scopeNoParent resource chain copied from earlier read_store results, as [{relationship, id}]. Start with [] for app-level reads.
confirmNofalse (default) returns a preview and changes nothing. true performs the change; set it only after the user approves that exact preview.
operationYesApp Store Connect operation name from get_store_capabilities, for example appInfos_getInstance.
project_idYesNoMac project ID (prj_…). push_project and connect_status list your projects.
referencesNoRelated resource references the body needs, copied from read_store results.
data_base64YesThe file's contents, base64-encoded, up to 2 MiB.
request_keyNoIdempotency key you choose (letters, digits, dash, underscore). Retry a lost response with the same key and identical arguments; use a new key for new work.
asset_sha256NoSet by upload_store_asset for exact file-reservation recovery.

Output Schema

ParametersJSON Schema
NameRequiredDescription
errorNoPresent only when the call failed

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare mutation (readOnlyHint false), open-world behavior, non-idempotency, and non-destructiveness. The description adds important context beyond that: preview/confirm staging, the 2 MiB hosted size limit, the request_key resume contract, and the ready/COMPLETE processing condition. It does not cover permissions or error behavior, so not a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with purpose and workflow, and every sentence adds distinct operational guidance. It is telegraphic in places but not wasteful.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given a complex 9-parameter mutation with an output schema and rich annotations, the description covers the staging workflow, size constraints, retry contract, and one explicit exclusion. Remaining gaps like permissions and failure handling are minor.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description reiterates and lightly supplements a few parameters (operation discovery, scope copy, data_base64 size, confirm, request_key) but adds little detail beyond the schema, and mentions 'ready:true' even though there is no top-level ready parameter in the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Specific verb 'upload' and resource 'App Store asset', enumerating asset types. It also distinguishes from the browser workflow for App Review attachments, but does not explicitly route screenshot uploads to or from the sibling upload_screenshots tool, so not a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

States prerequisite discovery via get_store_capabilities, scope copying, CLI file_path for files over 2 MiB, preview-by-default then confirm:true, and excludes App Review message attachments. Missing explicit guidance on when to prefer this over the sibling upload_screenshots.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 38 tool updates
    • First observedbuild
    • First observedcheck_project
    • First observedconnect_status
    • First observedexec_mac
    • First observedget_billing_status
    • First observedget_checkout_link
    • First observedget_failure
    • First observedget_feedback
    • First observedget_mac_checkout
    • First observedget_mac_credits
    • First observedget_mac_job
    • First observedget_mac_session
    • First observedget_metadata
    • First observedget_metadata_schema
    • First observedget_refund_status
    • First observedget_review
    • First observedget_screenshot_upload
    • First observedget_store_capabilities
    • First observedget_store_operation
    • First observedget_testflight
    • First observedlist_mac_sessions
    • First observedmanage_testflight
    • First observedprepare_review_reply
    • First observedpublish
    • First observedpush_project
    • First observedread_mac_output
    • First observedread_store
    • First observedreport_issue
    • First observedreview_lint
    • First observedset_metadata
    • First observedset_project_connection
    • First observedstart_mac
    • First observedstatus
    • First observedstop_mac
    • First observedupdate_review
    • First observedupdate_store
    • First observedupload_screenshots
    • First observedupload_store_asset

Publisher details

Operator
NoMac · Publisher source
Vendor relationship
First-party · Publisher source
Trust center
Not available
Restrictions
Sign in with a NoMac account (OAuth). Starting a Mac needs prepaid credit, sold in $10 packs at $0.80 an hour. App Store and TestFlight tools also need the user's App Store Connect API key. · Publisher source

Related MCP Connectors

Related MCP Servers

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.

Resources