App Store Connect MCP
Connects to Apple's App Store Connect API (api.appstoreconnect.apple.com) using an App Store Connect API key (team or individual) to read and modify app data, with read-only access by default and writes gated behind ASC_WRITE=1.
Manages apps in App Store Connect: uploads and distributes TestFlight builds, manages beta groups and testers, edits store listings (What's New, age rating), runs the full screenshot lifecycle (upload, replace, reorder, delete), manages subscriptions and intro offers, replies to customer reviews, downloads sales/financial reports, and prepares versions for App Review including a submission pre-flight and dry-run review submission.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@App Store Connect MCPWhat's the status of my app? Is the latest build ready for testers?"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
App Store Connect MCP
An MCP server that lets coding agents (Claude Code, Cursor, Codex and others) manage your apps in App Store Connect: TestFlight builds and testers, the store listing, screenshots, subscriptions, App Review submission, customer reviews and sales reports.
Apple doesn't ship an MCP server for App Store Connect. Xcode's xcrun mcpbridge covers building, testing and previews, not App Store Connect. Most community servers expose one tool per REST endpoint and leave the agent to chain them correctly. This one is built differently:
Workflows as tools. "Ship this build to these groups with these notes" is one call, and so is "replace screenshot 3". The server handles the ordering, the waiting and the edge cases that agents get wrong.
Safe by default. The server is read-only until you set
ASC_WRITE=1. Anything destructive is a dry run until the agent confirms it, and the docs recommend a least-privilege API key.Reliable. Every step checks the current state first, so you can re-run a tool after a timeout or a dropped connection. Tools poll Apple's asynchronous processing, back off on rate limits, and return errors that say what to do next.
Status: 0.1. Every tool is tested against an in-memory App Store Connect that checks each request against Apple's OpenAPI spec (v4.5). The workflows have also been run against a real App Store Connect account on a sandbox app: build upload through the API, distribution, groups and testers, listing edits, the whole screenshot lifecycle (including a full set of 10 and an image Apple rejects), version prep and the review pre-flight. Submitting to App Review and Beta App Review was only dry-run, because those send the app to Apple. Please report anything that behaves differently for you.
Quick start
1. Create an API key with the App Manager role. In App Store Connect, go to Users and Access > Integrations > App Store Connect API > Team Keys and generate a key with App Manager access. Never use Admin. Download the .p8 file (Apple only lets you download it once), and note the Key ID and the Issuer ID. The full walkthrough, including where to keep the key, is in docs/api-key.md.
2. Download and build it. You need Node.js 20 or later and git.
git clone https://github.com/swlittles/app-store-connect-mcp.git ~/app-store-connect-mcp
cd ~/app-store-connect-mcp
npm ci # installs the exact locked dependencies and builds dist/index.jsYou can clone it anywhere. The examples below assume ~/app-store-connect-mcp.
3. Check the key:
ASC_KEY_ID=YOUR_KEY_ID ASC_ISSUER_ID=YOUR_ISSUER_ID ASC_KEY_PATH=~/.appstoreconnect/AuthKey_YOUR_KEY_ID.p8 \
node ~/app-store-connect-mcp/dist/index.js --checkIt lists the apps the key can see.
4. Add it to your agent. Point the agent at dist/index.js, using absolute paths.
Claude Code:
claude mcp add asc --scope user \
-e ASC_KEY_ID=YOUR_KEY_ID \
-e ASC_ISSUER_ID=YOUR_ISSUER_ID \
-e ASC_KEY_PATH=$HOME/.appstoreconnect/AuthKey_YOUR_KEY_ID.p8 \
-- node $HOME/app-store-connect-mcp/dist/index.jsAdd -e ASC_WRITE=1 when you want the agent to make changes, and -e ASC_APP_ID=com.example.app to set a default app.
Cursor (~/.cursor/mcp.json) and other clients that use the mcpServers JSON format:
{
"mcpServers": {
"asc": {
"command": "node",
"args": ["/Users/you/app-store-connect-mcp/dist/index.js"],
"env": {
"ASC_KEY_ID": "YOUR_KEY_ID",
"ASC_ISSUER_ID": "YOUR_ISSUER_ID",
"ASC_KEY_PATH": "/Users/you/.appstoreconnect/AuthKey_YOUR_KEY_ID.p8"
}
}
}
}Codex (~/.codex/config.toml):
[mcp_servers.asc]
command = "node"
args = ["/Users/you/app-store-connect-mcp/dist/index.js"]
env = { ASC_KEY_ID = "YOUR_KEY_ID", ASC_ISSUER_ID = "YOUR_ISSUER_ID", ASC_KEY_PATH = "/Users/you/.appstoreconnect/AuthKey_YOUR_KEY_ID.p8" }Updates are automatic. Each time your agent starts the server, the server checks GitHub for a newer release (at most every 30 minutes) and installs it in the background. The update is used from the next time the agent starts the server; the running server isn't interrupted. See Updates.
Trying it without cloning: npx -y github:swlittles/app-store-connect-mcp --check (with the same environment variables) downloads and builds it in npm's cache. That's fine for a quick try. For everyday use, clone it: npx rebuilds on first launch, which can be slower than an agent waits for a server to start.
This project is distributed only through GitHub. It isn't published to npm, so a package with this name on npm isn't this project.
Then ask for things in plain language:
"What's the status of my app? Is the latest build ready for testers?"
"Ship build 202610041200 to the Friends and Public Beta groups with the notes 'New daily puzzles, please try the hint button'."
"Replace the third iPhone screenshot with ~/Desktop/shot-3.png."
"Upload everything in ./fastlane/screenshots/en-US/iphone as the 6.9-inch iPhone screenshots."
"Remove the free trial from the yearly subscription in every country."
"Reply to this week's 1-star reviews."
"Create version 1.2, attach the latest build, update What's New, and show me what's missing before we submit."
Related MCP server: App Store Connect MCP Server
Configuration
Variable | |
| API key ID. Required. |
| Issuer ID for team keys. Leave it unset for an individual key, which signs with |
| Path to the |
| The PEM contents instead of a path. Literal |
|
|
| Default app, as an app ID, bundle ID or exact name. If it's unset and the key can see only one app, that app is used. |
| Vendor number for |
| Only offer these tools or groups. Default: all. See Turning tools off. |
| Never offer these tools or groups. |
|
|
|
|
The key is only used to sign tokens. The server never logs it or puts it in tool output or error messages, and it only sends tokens to api.appstoreconnect.apple.com. That host is fixed and can't be overridden.
If the configuration is incomplete, the server still starts, and every tool returns an error explaining what's missing. Your client won't just show "failed to connect".
Updates
Clones of this repository update themselves:
When: on each server start, the server checks GitHub at most every 30 minutes. It does this in the background after starting, so the agent never waits on it.
What: with the default
releasechannel, the newest version tag (v0.3.0,v0.4.0, …). WithASC_UPDATE_CHANNEL=main, the newest commit onmain.How:
git fetch, then switch to the new version. Then it rebuilds, which takes seconds and works offline.npm ciruns instead when dependencies changed. The new version is used from the next start.Safe to leave on:
It only moves forward to the newer version, and never touches a copy with local edits or its own commits, such as a development clone.
If installing or building fails, it switches back to the previous version and rebuilds it. It won't retry the failed version until a newer one is published.
Two servers starting at once can't both update; one waits for the next check.
Check status:
node ~/app-store-connect-mcp/dist/index.js --checkprints the version and the last update result.Update now:
node ~/app-store-connect-mcp/dist/index.js --update.Turn it off or pin a version: set
ASC_AUTO_UPDATE=0, then pick a version yourself, for examplegit checkout v0.3.0 && npm ci.
Copies installed before v0.3.0 don't have the updater yet. Update those once by hand: cd ~/app-store-connect-mcp && git checkout -- package-lock.json && git fetch --tags && git checkout v0.3.0 && npm ci.
Turning tools off
There's no settings screen: the server runs in the background of your agent app. Choose its tools with two environment variables, set in the same place as the key:
ASC_TOOLSis an allowlist. Only the tools listed are offered. Without it, every tool is offered.ASC_DISABLED_TOOLSis a denylist. It's applied afterASC_TOOLS.
Both take tool names and group names, separated by commas or spaces. Turned-off tools aren't shown to the agent at all, so it can't call them.
Group | Tools |
| Every read-only tool: |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| Everything that deletes or can't be taken back: |
| Everything |
Examples:
-e ASC_TOOLS=read # look, never touch
-e ASC_TOOLS=read,testflight # TestFlight only, plus reading
-e ASC_DISABLED_TOOLS=destructive # everything except deleting and submitting
-e ASC_DISABLED_TOOLS=submit_for_review,raw # no App Review submissions, no raw API callsIf an entry isn't a known tool or group, every tool refuses to run until it's fixed. That way a typo in ASC_DISABLED_TOOLS can't leave a tool switched on. --check prints which tools are on.
ASC_WRITE still applies on top: without it, even enabled write tools only return dry-run plans. Your agent app may also let you block tools; for example, Claude Code's permissions.deny takes names like mcp__asc__submit_for_review.
Tools
Tools take an app argument, which can be an app ID, a bundle ID or a name. Results are short, readable text that includes the IDs a follow-up call needs.
Read (always available):
Tool | What it does |
| Apps the key can see, with IDs, bundle IDs and SKUs. |
| One-call overview: versions and their states, the attached build, the latest builds and their TestFlight states, uploads still processing, and review submissions. |
| Recent builds with processing and TestFlight states. |
| One build: compliance, beta review, groups, What to Test, App Store version. Use |
| TestFlight groups: internal or external, public link, tester count. |
| Testers of an app or a group, with status and groups. |
| Name, subtitle, categories, age rating, localized description, keywords, promotional text, What's New and URLs (with character counts), and App Review details. |
| Screenshot sets in display order, with positions, sizes, states and IDs. |
| Subscription groups, subscriptions, price in a territory, introductory offers summarized across territories, and in-app purchases. |
| Customer reviews with your replies. |
| Sales or finance report: unzipped, totaled, optionally saved as TSV. |
Workflows (these change things, so they need ASC_WRITE=1; any of them can be called with dry_run: true to see the plan):
Tool | Writes | What it does |
| yes | Uploads an |
| yes | Waits for processing, answers export compliance (if you say how), sets What to Test, adds the build to groups, and submits for beta review when an external group needs it. |
| yes | Creates a group (or returns the existing one with that name), optionally with a public link. |
| yes | Invites new testers and adds existing ones to a group; people who were already invited elsewhere aren't an error. |
| destructive | Removes testers from a group or from the app. |
| yes | Edits localized metadata, categories and game subcategories, showing old → new for each field and changing only what differs. |
| yes | App Store "What's New" or TestFlight "What to Test". |
| yes | Changes answers in the age rating questionnaire. |
| destructive | Makes a screenshot set match a folder or list of files ( |
| destructive | Replaces one screenshot (by position or ID), without a gap in the listing. |
| yes | Reorders a set. |
| destructive | Deletes screenshots by position or ID. |
| yes | Creates or renames the version being prepared, attaches a build, sets the release type. |
| yes | App Review contact, demo account and notes. The password is never echoed back. |
| destructive | Pre-flight checks, then submits through |
| destructive | Withdraws an active submission. |
| destructive | Bulk-deletes introductory offers (e.g. a free trial) across territories. |
| yes | Bulk-adds a free trial in every territory that lacks an introductory offer. |
| destructive | Publishes a public reply to a customer review; |
Escape hatch: asc_request(method, path, query, body) calls any App Store Connect endpoint. GET works in read-only mode. POST, PATCH and DELETE need ASC_WRITE=1, and DELETE also needs confirm: true.
How it stays safe
Read-only by default. Without
ASC_WRITE=1, write tools refuse to change anything, but they still return a plan if called withdry_run: true.Destructive means dry run first. Tools that delete things (screenshots, offers, testers) or can't be taken back (submitting for review) default to
dry_run: true. The agent gets the plan, shows it to you, and calls again withdry_run: false. The tools carry MCPdestructiveHintandreadOnlyHintannotations, so clients can ask before running them.Order that never leaves a gap. When
replace_screenshotreplaces an image, it uploads the new one, waits until Apple has processed it, puts it in place, and only then deletes the old one. If Apple rejects the image, nothing in the live listing changes. The one exception is a full set of 10. Apple won't allow an 11th screenshot even briefly, so the old one has to be deleted first, and the dry run says so.Least privilege. See below.
Choosing a key role
Use App Manager unless you have a reason not to. Step-by-step instructions are in docs/api-key.md.
Role | Enough for |
App Manager | Everything here except sales and finance reports. Recommended. |
Marketing | Listing metadata and screenshots, if that's all the agent should touch. |
Customer Support | Reading and replying to customer reviews. |
Finance or Sales |
|
Admin | Never needed. Don't use it: an Admin key can manage users and other keys. |
If a tool needs more access than the key has, the error says so. Keep keys out of repositories: store the .p8 somewhere like ~/.appstoreconnect/ with chmod 600, and revoke any key you think has leaked under Users and Access.
Long-running jobs
Apple processes builds (usually 5 to 30 minutes) and images (usually seconds) asynchronously. Tools that wait take wait_minutes, and they send MCP progress notifications while they wait. When time runs out, they don't fail. They return a status such as "still processing; run distribute_build again with the same arguments". Re-running is always safe:
distribute_buildskips groups the build is already in, notes that already match, and beta reviews already submitted. It treats Apple's "already submitted" response as success.Screenshot tools match files to existing screenshots by MD5, so an image uploaded by an interrupted run isn't uploaded again. Each delete, reorder and commit is its own request. If a step fails partway, the error lists the steps that finished and gives the exact call that finishes the job. For
replace_screenshot, that call includes thescreenshot_id, because positions can shift between runs. Leftovers from broken uploads never block a set; replacing or uploading to it cleans them up.Bulk jobs (
remove_intro_offers,add_free_trial,invite_testers) keep going past individual failures and report what succeeded, what failed and what wasn't tried. They stop early when the hourly rate limit (about 3,600 requests per key) runs low. Running them again finishes the job.
HTTP retries: the server retries GETs, PATCHes and DELETEs on 5xx errors and connection resets, with jittered exponential backoff. It retries POSTs only when Apple says the request didn't happen (429 or 503), because a reset POST may have gone through. It honors Retry-After and refreshes the token once on a 401.
Uploading builds
upload_build uploads an exported .ipa. Make one with:
xcodebuild archive -scheme MyApp -archivePath build/MyApp.xcarchive -destination 'generic/platform=iOS'
xcodebuild -exportArchive -archivePath build/MyApp.xcarchive -exportPath build -exportOptionsPlist ExportOptions.plistThe ExportOptions.plist needs method set to app-store-connect and destination set to export. (With destination set to upload, Xcode uploads the build itself, and you can skip upload_build.)
method: "api"(the default) uses App Store Connect's build upload API (buildUploadsandbuildUploadFiles, added in API 4.1). It needs no Xcode, so it works on Linux CI too.method: "altool"runsxcrun altool --upload-appwith the same API key, on a Mac.
xcodebuild -exportArchive registers a placeholder upload with App Store Connect even when it only exports. upload_build recognizes these placeholders and ignores them.
Every upload needs a new, higher CFBundleVersion. A timestamp such as YYYYMMDDHHMM works well. Add ITSAppUsesNonExemptEncryption = NO to Info.plist if it applies to your app, and TestFlight will never ask about export compliance.
Apple's rules worth knowing
These came up when testing against the real API. The tools handle each one, but they explain why a change is refused:
What's New can't be set on an app's first version.
The age rating questionnaire starts with every answer empty, and Apple won't accept any change until all 21 questions are answered.
update_age_ratingwithfill_unanswered: trueanswers the rest NONE/false. Check those answers with whoever owns the app.Once App Review details exist, every edit needs the full contact: first and last name, email, and phone with
+and the country code.Fields are cleared with
null, not an empty string. Pass""to the tools, and they sendnull.Apple reports a wrong-size screenshot only as
ASSET_FAILED, so the tools check image sizes before uploading. They also delete the rejected upload, so it doesn't take up one of the set's 10 slots.External testers aren't emailed until a build passes Beta App Review.
What the API can't do
These still need the App Store Connect website. This list was checked against API spec 4.5:
The App Privacy questionnaire ("nutrition labels"). The public API has no endpoints for it.
Creating a new app record.
POST /v1/appsdoesn't exist. Create the app on the website, then manage it from here.Agreements, tax and banking, including accepting updated Paid Apps agreements. When an agreement is missing, API calls fail with an error that this server explains.
Export compliance documentation for apps that use non-exempt encryption.
Paid introductory offers, which need a price point per territory. This server only covers free trials, so use
asc_requestfor paid offers. App pricing, Game Center and in-app purchase editing are also in the API but not wrapped yet;asc_requestreaches them.
Development
npm install
npm test # unit and workflow tests against a spec-checked fake App Store Connect
npm run typecheck
npm run build # dist/index.jsAPI types are generated from Apple's OpenAPI spec, which is pinned in spec/openapi.oas.json.gz. Its version and SHA-256 are in spec/VERSION.json.
npm run spec:updatedownloads the latest spec from Apple and regenerates the types insrc/generated/asc-api.ts.npm run spec:generateregenerates the types from the pinned spec.
Only the schemas listed in scripts/spec-schemas.json (and what they reference) are generated.
Tests run against test/helpers/fake-asc.ts, an in-memory App Store Connect. It supports JSON:API includes, filters, sorting, pagination, relationship endpoints and fault injection (connection resets, 409s, 429s). It also checks every request's path, method and query parameters against the pinned spec, so a typo in an endpoint fails a test. All fixture data is synthetic.
scripts/call-tool.mjs calls one tool on the built server over stdio, which is handy against a real account:
npm run build && node scripts/call-tool.mjs get_app_status '{"app": "com.example.app"}'Live tests are opt-in. They run against a real account, ideally a sandbox app:
ASC_LIVE_TEST=1 ASC_KEY_ID=… ASC_ISSUER_ID=… ASC_KEY_PATH=… ASC_LIVE_APP_ID=… npm run test:liveThey're read-only and dry runs unless you also set ASC_WRITE=1 and ASC_LIVE_WRITE=1. Then they also set and restore the promotional text of the live version.
License
MIT. See LICENSE. Not affiliated with Apple. App Store Connect and TestFlight are trademarks of Apple Inc.
Available Tools
31 toolsadd_free_trialAdd a free trialAIdempotent
Adds a free-trial introductory offer to a subscription in every territory where it's sold (or the ones you list), skipping territories that already have an introductory offer. Runs as a resumable bulk job with progress. Paid introductory offers need a price per territory and aren't covered; use asc_request for those.
Changes App Store Connect (needs ASC_WRITE=1). Safe to re-run: steps already done are skipped.
| Name | Required | Description | Default |
|---|---|---|---|
| app | No | App ID, bundle ID or exact app name. Defaults to ASC_APP_ID, or to the only app the key can see. | |
| dry_run | No | If true, return the plan without changing anything. | |
| duration | Yes | ||
| end_date | No | ||
| start_date | No | ||
| territories | No | Only these territories. Default: every territory the subscription is available in. | |
| subscription | Yes | Subscription ID or product ID (e.g. com.example.app.yearly). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare idempotentHint=true, destructiveHint=false, openWorldHint=true and readOnlyHint=false; the description adds value beyond them by naming the required credential (ASC_WRITE=1), disclosing that it runs as a resumable bulk job with progress, and explaining the mechanism behind idempotency ('steps already done are skipped'). It stops short of noting rate limits, partial-failure reporting, or what a resumable job returns, so it is strong but not exhaustive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, front-loaded with the core action and scope, then the paid-offer exclusion, then the environment/permission caveat. Dense and largely waste-free, though the parenthetical '(or the ones you list)' is mildly redundant with the territories parameter.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 7-parameter mutation with no output schema, the description covers the essential agent-facing facts: write permission needed, idempotent/resumable behavior, territory defaults, and the sibling to use instead. The remaining gap is that start_date/end_date semantics and their interaction with the offer window are left entirely to an undocumented schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 57%: app, territories, dry_run and subscription carry their own descriptions, and the description's territory scoping largely restates the schema's 'Default: every territory the subscription is available in.' However, start_date and end_date have no description in either place, and the duration enum values are unexplained, so the description does not compensate for the remaining coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Adds a free-trial introductory offer to a subscription') plus the scope (every territory where it's sold, or listed ones). It is clearly distinguishable from remove_intro_offers and from the generic asc_request sibling it names.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly routes the agent: free trials go here, paid introductory offers 'aren't covered; use asc_request for those'. It also states the skip condition (territories that already have an introductory offer), so both when-to-use and when-not-to-use are covered.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
asc_requestRaw App Store Connect requestADestructive
Escape hatch for anything the workflow tools don't cover: calls the App Store Connect API directly and returns the JSON. path is relative to https://api.appstoreconnect.apple.com, e.g. /v1/apps/123/appInfos. GET works in read-only mode; POST, PATCH and DELETE need ASC_WRITE=1, and DELETE also needs confirm: true. Prefer the workflow tools: they check state, retry safely and wait for Apple's processing.
| Name | Required | Description | Default |
|---|---|---|---|
| body | No | JSON:API request body for POST and PATCH. | |
| path | Yes | API path, e.g. /v1/apps or /v1/builds/{id}. Must start with /v1/, /v2/ or /v3/. | |
| query | No | Query parameters, e.g. {"filter[app]": "123", "limit": 50}. | |
| method | No | GET | |
| confirm | No | Must be true for DELETE. | |
| all_pages | No | For GET collections: follow links.next and return every page (at most 2,000 items). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare destructive/openWorld/non-idempotent, and the description adds substantial context beyond them: the base URL the path resolves against, the ASC_WRITE=1 environment requirement for POST/PATCH/DELETE, and the extra confirm: true gate on DELETE. That is exactly the auth/prerequisite detail annotations cannot express.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three dense sentences, each carrying distinct information (purpose, endpoint/auth mechanics, sibling preference), with the escape-hatch framing front-loaded and no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an open-world raw-API tool with no output schema, the description covers the base URL, method/auth gating, delete confirmation, and return type (raw JSON). Pagination (all_pages) is only in the schema, and rate-limit behavior on Apple's API is unmentioned, leaving minor gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is high (83%), so the baseline is 3, but the description adds real meaning: path is relative to https://api.appstoreconnect.apple.com with a concrete example, and it explains that confirm: true is required for DELETE (linking a parameter to a method).
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (calls the App Store Connect API directly) and resource (raw API path), explicitly positioning itself as an 'escape hatch' for gaps the workflow tools don't cover. An agent can immediately distinguish it from the ~30 workflow siblings like list_apps or submit_for_review.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says when not to use it ('Prefer the workflow tools') and why ('they check state, retry safely and wait for Apple's processing'), and names the alternative. It also gates usage by method: GET works read-only, writes need ASC_WRITE=1.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
cancel_review_submissionCancel App Review submissionADestructiveIdempotent
Withdraws the app's active review submission (waiting for or in review), so the version can be edited again.
Destructive: dry_run defaults to true and returns the plan. Show it to the user, then call again with dry_run: false. Needs ASC_WRITE=1.
| Name | Required | Description | Default |
|---|---|---|---|
| app | No | App ID, bundle ID or exact app name. Defaults to ASC_APP_ID, or to the only app the key can see. | |
| dry_run | No | Defaults to true: return the plan without changing anything. Call again with dry_run: false to apply it. | |
| platform | No | Platform. Defaults to IOS. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructive=true, but the description adds genuinely new behavioral context: the dry_run default-true confirmation pattern and the required ASC_WRITE=1 environment flag. These are the exact details an agent needs to avoid an unintended destructive call, and they go beyond what the annotations expose.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tightly written sentences with zero filler; the action and scope come first, followed by the destructive/dry-run warning and auth requirement. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive mutation with no output schema, the description covers the essential concerns: what it changes, that it is destructive, the dry-run confirmation workflow, and the auth prerequisite. There is no output schema to explain, and nothing critical is missing for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so app, dry_run, and platform are already fully documented in the schema (including defaults and the platform enum). The description only reinforces the dry_run default and does not add syntax or format detail beyond the schema, so the baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb (withdraws/cancels) and resource (the app's active review submission), and narrows scope to submissions 'waiting for or in review'. The stated effect ('so the version can be edited again') tells the agent exactly what state change to expect, distinguishing it cleanly from siblings like submit_for_review or prepare_version.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives a clear precondition (an active submission in waiting-for or in-review state) and an explicit two-step invocation flow: call with the default dry_run, show the plan, then re-call with dry_run: false. It also states the ASC_WRITE=1 prerequisite. It does not name alternative tools, so it falls just short of the 5 bar.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_beta_groupCreate a beta groupAIdempotent
Creates a TestFlight group, or returns the existing one with the same name. External groups can have a public link anyone can join. Internal groups can only contain people who are already App Store Connect users on the team.
Changes App Store Connect (needs ASC_WRITE=1). Safe to re-run: steps already done are skipped.
| Name | Required | Description | Default |
|---|---|---|---|
| app | No | App ID, bundle ID or exact app name. Defaults to ASC_APP_ID, or to the only app the key can see. | |
| name | Yes | ||
| dry_run | No | If true, return the plan without changing anything. | |
| feedback | No | Let testers send feedback and screenshots. | |
| internal | No | ||
| all_builds | No | Internal groups only: automatically get every new build. | |
| public_link | No | External groups only: turn on a public invite link. | |
| public_link_limit | No | Maximum testers who can join through the public link. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, idempotentHint=true and destructiveHint=false; the description adds genuinely new context on top of that: it requires ASC_WRITE=1, mutates App Store Connect, and concretely explains idempotency ('steps already done are skipped'). Return format is unaddressed, but no output schema exists to cover it.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose is front-loaded in the first sentence and the supporting detail is broken into short, scannable blocks with no filler. The trailing paragraph is slightly fragmentary ('Changes App Store Connect (needs ASC_WRITE=1). Safe to re-run...') but each clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an eight-parameter mutation tool with no output schema, the description supplies the key operational facts: auth requirement, idempotency, and the semantic difference between the two group types. Only per-parameter behavior for the optional flags is left entirely to the schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 75% and eight parameters exist, so the schema carries most of the load. The description adds conceptual meaning for the external/internal split and the public-link concept, but says nothing about feedback, all_builds, or public_link_limit, so it does not compensate for the remaining gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Creates a TestFlight group') and immediately discloses upsert semantics ('or returns the existing one with the same name'), which cleanly separates it from list_beta_groups and invite_testers in the sibling set.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explains the internal vs external group distinction and what each can contain, giving the agent a clear basis for choosing a mode. It does not name an alternative tool or state when not to use this one, so it stops short of full routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
delete_screenshotsDelete screenshotsADestructiveIdempotent
Deletes screenshots from a set by position (1-based) or ID. The rest keep their order.
Destructive: dry_run defaults to true and returns the plan. Show it to the user, then call again with dry_run: false. Needs ASC_WRITE=1.
| Name | Required | Description | Default |
|---|---|---|---|
| app | No | App ID, bundle ID or exact app name. Defaults to ASC_APP_ID, or to the only app the key can see. | |
| locale | No | Localization, e.g. en-US. Defaults to the app's primary locale. | |
| dry_run | No | Defaults to true: return the plan without changing anything. Call again with dry_run: false to apply it. | |
| version | No | App Store version string (e.g. "1.2"), "editable" (the version being prepared; the default) or "live". | |
| platform | No | Platform. Defaults to IOS. | |
| screenshots | Yes | Positions or IDs to delete. | |
| display_type | Yes | Screenshot display type, e.g. APP_IPHONE_67 (6.9" iPhone, 1320×2868), APP_IPHONE_65, APP_IPAD_PRO_3GEN_129 (13" iPad, 2064×2752), APP_DESKTOP, APP_APPLE_TV, APP_APPLE_VISION_PRO, APP_WATCH_ULTRA. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, idempotentHint=true and openWorldHint=true, yet the description still adds material context beyond them: the two-step dry-run confirmation workflow, the ASC_WRITE=1 authorization requirement, and the fact that surviving screenshots retain their relative order. That is exactly the extra behavioral payload the rubric credits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short paragraphs with zero filler: the operation and its addressing modes come first, then the destructive/confirmation protocol. Every sentence carries information an agent needs.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 7-parameter destructive mutation with no output schema, the definition covers what an agent needs: what is deleted, how targets are addressed, that it is destructive, the required confirmation loop, and the auth gate. Nothing essential is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents app, locale, version, platform, dry_run, screenshots and display_type. The description adds only the '1-based' indexing detail for positions, which is genuinely absent from the schema, but otherwise restates what is already structured (e.g. the dry_run default). Baseline 3 is the right level.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Deletes screenshots') plus the two addressing modes (1-based position or ID), and the ordering side effect on remaining items. An agent can distinguish this immediately from replace_screenshot, reorder_screenshots, upload_screenshots and list_screenshots without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a concrete operational protocol: dry_run defaults to true, show the plan to the user, then re-call with dry_run: false, and notes the ASC_WRITE=1 prerequisite. It does not explicitly route between this tool and its nearest siblings (replace_screenshot / reorder_screenshots), so it stops short of full alternative selection guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
distribute_buildDistribute a TestFlight buildAIdempotent
Ships a build to TestFlight testers in one call: waits for processing, answers export compliance if you say how, sets the What to Test notes, adds the build to beta groups, and submits it for beta app review when an external group needs it. Every step checks the current state first, so re-running after a timeout or error picks up where it stopped.
Changes App Store Connect (needs ASC_WRITE=1). Safe to re-run: steps already done are skipped.
| Name | Required | Description | Default |
|---|---|---|---|
| app | No | App ID, bundle ID or exact app name. Defaults to ASC_APP_ID, or to the only app the key can see. | |
| build | No | Build number (CFBundleVersion, e.g. "202610010021"), build resource ID, or "latest" (the default). | |
| notes | No | "What to Test" text shown to testers. | |
| groups | No | Beta group names or IDs. Internal groups with access to all builds get it automatically. | |
| dry_run | No | If true, return the plan without changing anything. | |
| version | No | Marketing version (CFBundleShortVersionString, e.g. "1.2"). Needed only if build numbers repeat across versions. | |
| platform | No | Platform. Defaults to IOS. | |
| notes_locale | No | en-US | |
| wait_minutes | No | Minutes to wait for Apple's processing before returning (0-30, default 5). Processing usually takes 5-30 minutes; if it isn't done in time, the result says how to resume. | |
| submit_for_beta_review | No | Submit for beta app review. Defaults to true when any target group is external, since external testers need an approved build. | |
| uses_non_exempt_encryption | No | Export compliance answer, used only if the build is waiting on it. false means the app uses no encryption or only exempt encryption such as HTTPS. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the annotations (idempotentHint/destructiveHint/readOnlyHint), it discloses the required environment flag (ASC_WRITE=1), that it mutates App Store Connect, and the mechanism behind idempotency: every step checks current state first, so already-done steps are skipped on re-run. It also explains the resume behavior when Apple processing exceeds wait_minutes. This is context an agent cannot get from the annotations alone.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action, then the operational caveats (auth flag, safe re-run) in a short second paragraph. The first sentence is a long comma-chained list of steps, which is dense but still scannable; nothing is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an 11-parameter mutation with no output schema, the description covers the workflow, auth prerequisite, idempotency, and partial-failure resume path, which is most of what an agent needs. It stops short of describing the shape of the returned plan/result (relevant for dry_run), leaving that to inference.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 91%, so the schema already documents nearly every parameter; per the rubric the baseline is 3. The description adds only marginal framing (that the export-compliance answer is used "only if the build is waiting on it", and the conditional default for beta review), and does not address the undocumented notes_locale parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb+resource ("Ships a build to TestFlight testers") and then enumerates the exact sub-steps performed in the call: wait for processing, export compliance, What to Test notes, beta group assignment, beta app review submission. This clearly distinguishes it from siblings like upload_build (which only uploads) and submit_for_review (which only submits).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It states the conditions under which optional steps fire ("answers export compliance if you say how", "submits it for beta app review when an external group needs it") and explicitly endorses re-running after a timeout or error. It never names a sibling alternative or an explicit when-not-to-use case, so it falls short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
download_reportDownload a sales or finance reportARead-onlyIdempotent
Downloads a sales or finance report (Apple returns gzipped TSV), optionally saves it, and returns the row count, columns, simple totals and the first rows. Needs the vendor number (App Store Connect > Payments and Financial Reports) as vendor_number or ASC_VENDOR_NUMBER, and a key with the Finance or Sales role.
| Name | Required | Description | Default |
|---|---|---|---|
| kind | No | sales | |
| rows | No | How many rows to show. | |
| save_to | No | Absolute path to save the uncompressed TSV. | |
| version | No | Report format version, e.g. 1_0 or 1_4. Apple's error says which version it wants if this is wrong. | |
| frequency | No | DAILY | |
| region_code | No | Finance only: region, e.g. US, EU, or ZZ for all (Z1 for FINANCE_DETAIL). | ZZ |
| report_date | Yes | Sales: YYYY-MM-DD (daily/weekly), YYYY-MM (monthly) or YYYY (yearly). Finance: fiscal month YYYY-MM. | |
| report_type | No | Default SALES for sales, FINANCIAL for finance. | |
| vendor_number | No | ||
| report_sub_type | No | SUMMARY |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint/openWorld/idempotent/non-destructive, so the safety profile is covered. The description adds real value beyond that: the payload is gzipped TSV, saving is optional, and the response includes row count, columns, totals and leading rows, plus the API key role requirement.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with purpose and return behavior, followed by prerequisites. Dense but every clause carries information; only slight cost is the run-on packing of output fields.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 10-parameter tool with no output schema, the description covers purpose, return shape, auth and the vendor-number sourcing fallback. Remaining gaps (report_type/frequency interplay, save path semantics) are largely handled by the schema's enums and field descriptions.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 60%, and the description compensates by disclosing that vendor_number can alternatively come from the ASC_VENDOR_NUMBER environment variable – a fact absent from the schema – and by naming the required key role. It does not explain kind/frequency/report_sub_type semantics, but those carry enums and defaults.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (downloads) and resource (sales or finance report), plus the return format (Apple gzipped TSV) and what the tool yields. No sibling tool overlaps with report retrieval, so the agent can route to it unambiguously.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit prerequisites: the vendor number (with where to find it in App Store Connect) and a key carrying the Finance or Sales role. It does not spell out when-not-to-use or name alternatives, but no sibling tool competes for this task, so the guidance is effectively complete for context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_app_statusGet app statusARead-onlyIdempotent
One-call overview of an app: App Store versions and their states, the build attached to each, the latest builds and their TestFlight states, uploads still processing, and open review submissions.
| Name | Required | Description | Default |
|---|---|---|---|
| app | No | App ID, bundle ID or exact app name. Defaults to ASC_APP_ID, or to the only app the key can see. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, non-destructive, and open-world behavior. The description adds valuable context by enumerating the categories of data returned, which matters especially because there is no output schema. It still does not describe return format, pagination, or authentication expectations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single front-loaded sentence that begins with the purpose and then efficiently enumerates the data categories. Every listed item earns its place by clarifying the scope of the overview.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only aggregate tool with full annotation coverage and a fully documented single parameter, the description gives enough context about what will be returned. It could be slightly more complete by hinting at output structure or how to interpret the states, but no critical information is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There is a single optional 'app' parameter with 100% schema description coverage, including the default fallback behavior. The description adds no additional parameter semantics, so the baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific purpose: a one-call overview of an app covering versions, builds, TestFlight states, processing uploads, and open submissions. It clearly distinguishes itself as an aggregate status tool, though it does not explicitly name granular siblings like list_apps or list_builds.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'One-call overview' implies the tool should be used when an agent wants a combined snapshot rather than multiple list/get calls. However, there is no explicit when-to-use guidance, no exclusions, and no named alternative to route to for narrower queries.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_buildGet buildARead-onlyIdempotent
Shows one build in detail: processing and TestFlight states, export compliance, beta review, groups, What to Test notes and the App Store version it's attached to. Set wait_minutes to wait for Apple to finish processing a fresh upload.
| Name | Required | Description | Default |
|---|---|---|---|
| app | No | App ID, bundle ID or exact app name. Defaults to ASC_APP_ID, or to the only app the key can see. | |
| build | No | Build number (CFBundleVersion, e.g. "202610010021"), build resource ID, or "latest" (the default). | |
| version | No | Marketing version (CFBundleShortVersionString, e.g. "1.2"). Needed only if build numbers repeat across versions. | |
| platform | No | Platform. Defaults to IOS. | |
| wait_minutes | No | Minutes to wait for Apple's processing before returning (0-30, default 0). Processing usually takes 5-30 minutes; if it isn't done in time, the result says how to resume. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint true and destructiveHint false, so safety is covered by structured data. The description adds non-obvious behavior: the tool can block for a bounded period (wait_minutes), processing typically takes 5-30 minutes, and a timeout result explains how to resume. It doesn't mention auth/permission scope or result truncation, keeping it short of a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, zero filler, with the core purpose and return contents front-loaded before the operational note about wait_minutes. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, and the description usefully compensates by listing what the response contains, which is the key thing an agent needs to decide whether to call it. With 100% schema coverage and read-only annotations, the remaining gap is minor: no guidance on the singular-vs-list choice against list_builds.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% for all five parameters, including defaults (latest build, ASC_APP_ID, IOS), so the schema already carries the semantics. The description's only parameter-related addition is restating the purpose of wait_minutes, which the schema already explains; baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb and resource ('Shows one build in detail') and enumerates the returned facets: processing/TestFlight states, export compliance, beta review, groups, What to Test notes and the attached App Store version. The word 'one' implicitly separates it from list_builds, so an agent can route between them without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The second sentence gives a real usage cue for the wait_minutes parameter ('wait for Apple to finish processing a fresh upload'), but there is no explicit when-to-use/when-not framing and no alternative named (e.g. list_builds for enumeration, get_app_status for readiness checks). Usage is implied rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_listingGet store listingARead-onlyIdempotent
Shows the App Store listing: app name, subtitle, privacy URL and categories, the version's localized description, keywords, promotional text, what's new and URLs (with character counts against Apple's limits), the age rating, and the App Review contact details.
| Name | Required | Description | Default |
|---|---|---|---|
| app | No | App ID, bundle ID or exact app name. Defaults to ASC_APP_ID, or to the only app the key can see. | |
| full | No | Show full descriptions instead of the first 300 characters. | |
| locale | No | Localization, or "all". Defaults to the primary locale. | |
| version | No | App Store version string (e.g. "1.2"), "editable" (the version being prepared; the default) or "live". | |
| platform | No | Platform. Defaults to IOS. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly/idempotent/non-destructive/openWorld, so the safety profile is covered. The description adds genuine behavioral value beyond that: it discloses derived data ('character counts against Apple's limits') that the agent cannot infer from the schema, plus the breadth of what is retrieved.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence that leads with the verb and resource before listing contents. It is dense but every element earns its place; slightly long, but no padding or repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With five optional parameters and no output schema, the description usefully compensates by enumerating the return contents. The only minor gap is that it doesn't restate the default version selection (editable/live), though that is fully handled in the schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all five parameters (app, full, locale, version, platform) with their defaults and the enum are fully documented structurally. The description adds no parameter-level detail beyond the schema, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Shows') and resource ('the App Store listing'), then enumerates the returned fields (name, subtitle, privacy URL, categories, description, keywords, age rating, review contacts). This clearly differentiates it from read siblings like list_apps/get_app_status and from the write-side siblings like update_listing and set_whats_new.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit when-to-use or when-not-to-use guidance, and no alternative is named. However, the exhaustive content enumeration implicitly tells an agent this is the tool to fetch the current listing, so usage is inferable rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_reviewsGet customer reviewsARead-onlyIdempotent
Lists App Store customer reviews, newest first by default, with your published replies and review IDs (for reply_to_review).
| Name | Required | Description | Default |
|---|---|---|---|
| app | No | App ID, bundle ID or exact app name. Defaults to ASC_APP_ID, or to the only app the key can see. | |
| sort | No | -createdDate | |
| limit | No | ||
| rating | No | ||
| territory | No | ISO alpha-3, e.g. USA. | |
| unanswered | No | Only reviews without a published reply. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, openWorldHint and non-destructive, so the safety profile is covered. The description adds that replies and review IDs are returned and that ordering defaults to newest-first, but says nothing about pagination (limit max 200), rate limits, or auth requirements. Useful context, not rich behavioral disclosure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One sentence, front-loaded with the verb and resource, with the default ordering and the sibling pointer packed in without waste.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description usefully sketches the return payload (reviews, published replies, review IDs), which is the key thing an agent needs. It omits pagination behavior and how the filters (rating, territory, unanswered) compose, leaving a small gap for a 6-parameter listing tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 50% and the description only touches the sort default ('newest first'), which merely restates the schema's default value. It adds no syntax or meaning for app, limit, rating, territory, or unanswered beyond what the schema already documents, so it neither compensates for the gap nor adds value.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb ('Lists') plus resource ('App Store customer reviews') and scope ('newest first by default'). It also names what comes back (published replies, review IDs) and cross-references the sibling reply_to_review, so an agent can place it precisely in the review workflow.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description establishes the default retrieval context ('newest first by default') and routes the agent to reply_to_review for the downstream reply action. There is no alternative review-listing sibling to contrast against and no explicit when-not condition, so it stops short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
invite_testersInvite testersAIdempotent
Adds testers to a beta group, inviting them if they're new. People already invited elsewhere are added to the group instead of failing. External groups accept any email; internal groups only accept App Store Connect users on the team. Safe to re-run with the same list.
Changes App Store Connect (needs ASC_WRITE=1). Safe to re-run: steps already done are skipped.
| Name | Required | Description | Default |
|---|---|---|---|
| app | No | App ID, bundle ID or exact app name. Defaults to ASC_APP_ID, or to the only app the key can see. | |
| group | Yes | Group name or ID. | |
| dry_run | No | If true, return the plan without changing anything. | |
| testers | Yes | Emails, or objects like {"email": "ada@example.com", "first_name": "Ada"}. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, idempotentHint=true, destructiveHint=false and openWorldHint=true, so the safety profile is largely covered. The description adds the concrete auth prerequisite ASC_WRITE=1 and the external-vs-internal group constraint, which is useful, but the rerun-safety claim mostly restates idempotentHint and no return/pagination behavior is described.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the core behavior well, but the rerun-safety idea is stated twice ('Safe to re-run with the same list' and 'Safe to re-run: steps already done are skipped'), which is redundant rather than additive. Structure is otherwise fine but the duplication costs it.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and full parameter coverage, the description carries the mutation warning (ASC_WRITE=1), idempotency, and group-type constraints, which is close to complete for a write tool. It could mention return/permission-error behavior but nothing critical for correct invocation is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3, and the description adds meaning beyond it by explaining how the 'group' parameter behaves differently for external vs internal groups (email acceptance rules). The 'testers' dedupe/merge behavior ('added to the group instead of failing') also clarifies semantics not stated in the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Adds testers to a beta group') plus the scope of the operation, which cleanly distinguishes it from siblings like create_beta_group, remove_testers and list_testers. An agent can identify the operation without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear operating context: external groups accept any email while internal groups only accept App Store Connect team users, and reruns are safe. It does not name an explicit alternative tool or an exclusion condition, but the group-type rule plus idempotency note gives strong selection context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_appsList appsARead-onlyIdempotent
Lists the apps this API key can see, with their IDs, bundle IDs and SKUs. Start here to find an app ID.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | Only apps whose name contains this text. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, non-destructive and open-world, so safety is covered; the description adds the meaningful extra context that results are scoped to the API key's permissions and enumerates the returned fields. It omits pagination/result-count behavior, which is the only notable gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences with no filler, front-loading what is listed and following with the practical next-step hint. Every clause carries information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description usefully enumerates the return fields and positions the tool as the entry point. For a simple, no-required-param list call this is nearly complete, lacking only pagination or ordering details.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the single 'name' filter is fully documented there, so the baseline of 3 applies. The description never mentions that results can be filtered by name, adding no meaning beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Lists the apps') plus the visibility scope ('this API key can see') and the returned identifiers (IDs, bundle IDs, SKUs), which separates it from the app-specific siblings. It never names a sibling directly, but the scope and payload make the distinction inferable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Start here to find an app ID' gives clear entry-point guidance and implies that downstream tools (get_app_status, list_builds) require an app ID discovered here. No explicit when-not or alternatives are named, so it stops short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_beta_groupsList beta groupsBRead-onlyIdempotent
Lists an app's TestFlight groups: internal or external, access to all builds, public link, feedback, tester count and ID.
| Name | Required | Description | Default |
|---|---|---|---|
| app | No | App ID, bundle ID or exact app name. Defaults to ASC_APP_ID, or to the only app the key can see. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false and openWorldHint, so the safety profile is covered. The description adds real value by enumerating what each returned group carries (internal/external, build access, public link, feedback, tester count, ID), but says nothing about pagination, ordering, or rate limits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One dense sentence, front-loaded with the operation and resource, then listing the returned fields. Every clause earns its place with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description usefully compensates by telling the agent what fields the listing exposes. Combined with full parameter coverage in the schema, an agent has enough to call it correctly; only pagination/edge behavior is unstated.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With a single optional parameter at 100% schema description coverage, the schema already documents the app identifier and its ASC_APP_ID default. The description adds no semantic detail about the app argument, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Lists an app's TestFlight groups') and scopes it to TestFlight groups, which cleanly separates it from list_testers or list_apps. It does not, however, explicitly name or contrast with the sibling create_beta_group, so it stops short of full sibling differentiation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description is purely declarative about what is returned; it gives no when-to-use guidance, no prerequisites, and no pointer to alternatives such as create_beta_group for mutation or list_testers for the tester roster. Usage is only inferable from the tool name.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_buildsList buildsBRead-onlyIdempotent
Lists an app's most recent builds with processing state, TestFlight states and IDs.
| Name | Required | Description | Default |
|---|---|---|---|
| app | No | App ID, bundle ID or exact app name. Defaults to ASC_APP_ID, or to the only app the key can see. | |
| limit | No | ||
| version | No | Only builds of this marketing version, e.g. "1.2". | |
| platform | No | Platform. Defaults to IOS. | |
| include_expired | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, openWorldHint and destructiveHint=false, so the safety profile is covered. The description adds the meaningful behavioral detail that results are the app's most recent builds and include processing/TestFlight state, but it omits the default limit (10), what include_expired does, and whether results are paginated. With annotations carrying safety, this is adequate but not rich.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single compact sentence with the resource stated first and the payload contents second. Nothing is padded or redundant.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description partially compensates by listing the returned fields (processing state, TestFlight states, IDs). However, for a five-parameter listing tool it says nothing about default limit, expired-build inclusion, or result ordering, leaving real gaps for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 60%: app, version and platform are documented in the schema, while limit (default 10, max 50) and include_expired (default false) have no description anywhere. The description adds no parameter-level meaning at all, so the two undocumented parameters remain opaque to the agent.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Lists an app's most recent builds') and names what the result contains (processing state, TestFlight states, IDs), so the agent knows this is a read-only build listing. It never names a sibling such as get_build for a single build or upload_build for creation, so the differentiation from neighbors is left to inference.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use/when-not guidance and no mention of alternatives like get_build for a single build. The phrase 'most recent builds' weakly implies a default recency-ordered listing, but nothing tells the agent when this is the right call versus another listing tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_screenshotsList screenshotsARead-onlyIdempotent
Lists App Store screenshots for a version, per localization and display type, in display order with positions (1-based), file names, sizes, states and IDs.
| Name | Required | Description | Default |
|---|---|---|---|
| app | No | App ID, bundle ID or exact app name. Defaults to ASC_APP_ID, or to the only app the key can see. | |
| locale | No | Localization, e.g. en-US, or "all". Defaults to the primary locale. | |
| version | No | App Store version string (e.g. "1.2"), "editable" (the version being prepared; the default) or "live". | |
| platform | No | Platform. Defaults to IOS. | |
| display_type | No | Screenshot display type, e.g. APP_IPHONE_67 (6.9" iPhone, 1320×2868), APP_IPHONE_65, APP_IPAD_PRO_3GEN_129 (13" iPad, 2064×2752), APP_DESKTOP, APP_APPLE_TV, APP_APPLE_VISION_PRO, APP_WATCH_ULTRA. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false and openWorldHint=true, so safety is covered. The description adds genuinely useful behavioral context not in the annotations: results come back in display order with 1-based positions, plus file names, sizes, states and IDs. This is valuable since there is no output schema, though it says nothing about pagination or result-size limits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single well-formed sentence that front-loads the verb and resource, then appends the returned field list. Every clause earns its place; the trailing enumeration is dense but directly useful given the absence of an output schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only list tool whose annotations carry the full safety profile and whose schema is fully documented with defaults, the description is nearly complete: it conveys the scope, grouping, ordering and returned fields. The only gap is the lack of any usage routing or result-volume context, which is minor here.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all five parameters (app, locale, version, platform, display_type) are already documented with defaults and enum values. The description only echoes that results are grouped per localization and display type, adding no syntax or format detail beyond the schema. Baseline 3 applies when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb (Lists) and resource (App Store screenshots) with scope (for a version, per localization and display type), and enumerates the returned fields. It is clearly a read operation, distinguishable from siblings like replace_screenshot, reorder_screenshots and delete_screenshots, though it never names them explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance, no prerequisites, and no mention of alternatives. The only implied usage is that this is the read step before mutating screenshot tools such as reorder_screenshots, which the agent must infer on its own.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_subscriptionsList subscriptions and in-app purchasesARead-onlyIdempotent
Lists subscription groups and their subscriptions (product ID, period, state, price in one territory, introductory offers summarized across territories), plus one-time in-app purchases.
| Name | Required | Description | Default |
|---|---|---|---|
| app | No | App ID, bundle ID or exact app name. Defaults to ASC_APP_ID, or to the only app the key can see. | |
| territory | No | Territory for the price shown (ISO 3166-1 alpha-3). | USA |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, non-destructive and openWorld, so the safety profile is covered structurally. The description goes beyond that by disclosing the shape and scope of returned data, notably that price is shown in one territory while introductory offers are summarized across territories, which is genuinely useful behavioral context not present in the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with the verb and resource first, followed by a parenthetical enumeration of returned fields. No filler, no restatement of the title, every clause carries information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description must convey return contents, and it does so reasonably well by listing the fields returned. It is slightly incomplete on pagination/result-size behavior and on whether results span a single app or all visible apps, but nothing critical to invoking the tool is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both parameters ('app' default resolution and 'territory' ISO code/default) are fully documented in the schema. The description's phrase 'price in one territory' loosely echoes the territory parameter but adds no format or default semantics beyond what the schema already provides, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (lists) and resource (subscription groups, their subscriptions, plus one-time in-app purchases), and enumerates the returned fields, so the agent knows exactly what domain this covers. It does not name or differentiate from any sibling (e.g. remove_intro_offers or add_free_trial, which touch the same subscription objects), so it falls short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance, no prerequisites, and no mention of alternatives or when not to use it. Usage is only inferable from the verb 'Lists'. This matches the MID calibration where a read tool with no routing guidance scored 2.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_testersList testersARead-onlyIdempotent
Lists TestFlight testers for an app, or for one group, with email, name, status, invite type and groups.
| Name | Required | Description | Default |
|---|---|---|---|
| app | No | App ID, bundle ID or exact app name. Defaults to ASC_APP_ID, or to the only app the key can see. | |
| No | Only this tester. | ||
| group | No | Group name or ID. Omit for every tester of the app. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false, so the safety profile is covered. The description adds that results carry email, name, status, invite type and groups, but says nothing about pagination, result limits, or auth requirements.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler. It is efficient, though it packs the return-field list into the same sentence rather than treating it as a distinct concern.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read-only list tool with fully documented parameters and annotations covering the safety profile, the description is nearly sufficient. Because there is no output schema, the returned-field enumeration adds real value; only pagination/limit behavior is left unaddressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so app, email and group are all documented in the schema itself, including the ASC_APP_ID default. The description's mention of app/group scope adds no syntax or format detail beyond what the schema provides, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Lists') and resource ('TestFlight testers') with clear scope: for an app or for one group. The verb alone separates it from sibling mutations like invite_testers and remove_testers, though no sibling is named explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'for an app, or for one group' implies the two usage modes, and the schema notes 'Omit for every tester of the app.' However, there is no explicit when-to-use guidance or named alternative for the group-scoped vs app-scoped case.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
prepare_versionPrepare an App Store versionAIdempotent
Makes sure an App Store version with this version string is being prepared: reuses it if it exists, renames the version currently being prepared, or creates a new one (Apple copies the metadata from the previous version). Optionally attaches a build and sets the release type.
Changes App Store Connect (needs ASC_WRITE=1). Safe to re-run: steps already done are skipped.
| Name | Required | Description | Default |
|---|---|---|---|
| app | No | App ID, bundle ID or exact app name. Defaults to ASC_APP_ID, or to the only app the key can see. | |
| build | No | Build to attach: number, ID or "latest". Must be processed (VALID). | |
| dry_run | No | If true, return the plan without changing anything. | |
| platform | No | Platform. Defaults to IOS. | |
| release_type | No | ||
| version_string | Yes | ||
| earliest_release_date | No | For SCHEDULED: ISO 8601 date-time. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, idempotentHint=true and destructiveHint=false, but the description adds real behavioral detail beyond them: metadata is copied from the previous version, an in-progress version may be renamed, and the write flag is required. The one gap is that it doesn't say what the call returns or how the rename branch affects an already-submitted version.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core behavior, then a separate sentence for auth and idempotency. The parenthetical about Apple copying metadata is high-value and every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 7-parameter mutation tool with no output schema, the description covers prerequisites, side effects and re-run safety well. It stops short of describing the result of a successful preparation (e.g. the created/renamed version identifier), which an agent may need to chain into submit_for_review.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 71%, so most parameters are already documented. The description adds semantics for the build attachment and release type, but says nothing about earliest_release_date, platform defaults, or dry_run beyond what the schema states, so it only marginally exceeds the schema baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb (prepare) and resource (App Store version) and enumerates the three concrete branches the tool takes: reuse an existing version, rename the currently prepared one, or create a new one. No sibling tool (submit_for_review, get_listing, etc.) overlaps, so the agent can route to it without ambiguity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It states the operating prerequisite (changes App Store Connect, needs ASC_WRITE=1) and that re-running is safe, which tells the agent when invocation is allowed. It doesn't explicitly name a sibling alternative or a when-not-to-use case, but no sibling competes for this job, so the guidance is adequate.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
remove_intro_offersRemove introductory offersADestructiveIdempotent
Deletes a subscription's introductory offers (for example a free trial) in every territory or only some. Apple stores one offer per territory, so this is a bulk job: it reports progress, keeps going past individual failures, stops early if the hourly rate limit runs low, and can be re-run to finish.
Destructive: dry_run defaults to true and returns the plan. Show it to the user, then call again with dry_run: false. Needs ASC_WRITE=1.
| Name | Required | Description | Default |
|---|---|---|---|
| app | No | App ID, bundle ID or exact app name. Defaults to ASC_APP_ID, or to the only app the key can see. | |
| dry_run | No | Defaults to true: return the plan without changing anything. Call again with dry_run: false to apply it. | |
| offer_mode | No | Only offers of this kind. | |
| territories | No | Only these territories (ISO alpha-3, e.g. USA, GBR). Default: all. | |
| subscription | Yes | Subscription ID or product ID (e.g. com.example.app.yearly). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Goes well beyond what annotations declare (destructive/idempotent/openWorld) by disclosing the bulk-job semantics: reports progress, continues past individual failures, stops early when the hourly rate limit runs low, and is safely re-runnable. The dry_run-defaults-to-true workflow and the ASC_WRITE=1 auth requirement are exactly the operational context an agent otherwise couldn't infer.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the verb and scope, then a tight second paragraph on the destructive/dry_run contract. Every clause carries useful weight; the bulk-behavior sentence is dense but justified for a rate-limited batch operation.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive bulk operation with no output schema, the description supplies the missing pieces: progress/failure/rate-limit behavior, re-runnability, auth prerequisite, and the mandatory dry_run preview flow. Nothing essential to calling it correctly is left unexplained.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so every parameter is already documented, including dry_run's default and the territory/offer_mode enums. The description reinforces the dry_run workflow but adds no syntax or format detail beyond the schema, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (deletes) and resource (a subscription's introductory offers), with the scope explicitly qualified (every territory or only some). An agent can distinguish this from the inverse sibling add_free_trial without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is strongly implied by the verb and by the dry_run workflow ('Show it to the user, then call again with dry_run: false'), which is clear procedural guidance. However, no explicit when-not condition or named alternative (e.g. add_free_trial as the inverse) is given, so routing is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
remove_testersRemove testersADestructiveIdempotent
Removes testers from one beta group, or from the app entirely (every group, and their access to builds) when group is omitted. Removing someone from their last group leaves them listed as a tester of the app with no groups; omit group to remove them completely.
Destructive: dry_run defaults to true and returns the plan. Show it to the user, then call again with dry_run: false. Needs ASC_WRITE=1.
| Name | Required | Description | Default |
|---|---|---|---|
| app | No | App ID, bundle ID or exact app name. Defaults to ASC_APP_ID, or to the only app the key can see. | |
| group | No | Group name or ID. Omit to remove the testers from the app entirely. | |
| emails | Yes | ||
| dry_run | No | Defaults to true: return the plan without changing anything. Call again with dry_run: false to apply it. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already flag destructiveHint=true and idempotentHint=true, and the description goes well beyond them: it discloses the dry_run default of true, the mandatory show-the-plan-then-reapply sequence, the ASC_WRITE=1 environment prerequisite, and the non-obvious side effect that removing someone from their last group leaves them listed as a groupless tester.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight paragraphs, front-loaded with the scope distinction and followed by the operational workflow. Every sentence carries information the agent needs; there is no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive mutation with no output schema, the description covers the safety workflow, the auth requirement, the scope semantics, and the surprising residual state. Nothing an agent needs to call this correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 75% and the schema already documents group and dry_run. The description still adds meaning the schema does not: that omitting 'group' removes the tester from every group and revokes build access, and that the app-level removal is distinct from merely leaving them group-less.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (removes) and resource (testers) and precisely distinguishes its two operating modes: scoped to one beta group versus app-wide when 'group' is omitted. An agent can tell immediately what this tool does without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit selection logic for the group parameter (omit it to remove completely) and a clear two-step dry_run workflow with a user-confirmation gate. It does not name a sibling alternative such as invite_testers or list_testers, but the conditions for using this tool are unambiguous.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
reorder_screenshotsReorder screenshotsAIdempotent
Changes the order of a screenshot set. order lists current positions (1-based) or screenshot IDs in the new order; any not listed keep their relative order after them. Example: [4] moves screenshot 4 to the front.
Changes App Store Connect (needs ASC_WRITE=1). Safe to re-run: steps already done are skipped.
| Name | Required | Description | Default |
|---|---|---|---|
| app | No | App ID, bundle ID or exact app name. Defaults to ASC_APP_ID, or to the only app the key can see. | |
| order | Yes | ||
| locale | No | Localization, e.g. en-US. Defaults to the app's primary locale. | |
| dry_run | No | If true, return the plan without changing anything. | |
| version | No | App Store version string (e.g. "1.2"), "editable" (the version being prepared; the default) or "live". | |
| platform | No | Platform. Defaults to IOS. | |
| display_type | Yes | Screenshot display type, e.g. APP_IPHONE_67 (6.9" iPhone, 1320×2868), APP_IPHONE_65, APP_IPAD_PRO_3GEN_129 (13" iPad, 2064×2752), APP_DESKTOP, APP_APPLE_TV, APP_APPLE_VISION_PRO, APP_WATCH_ULTRA. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, idempotentHint=true, and destructiveHint=false. The description adds real context beyond them: the ASC_WRITE=1 auth requirement and the fact that already-applied steps are skipped on re-run. The 'safe to re-run' line partially restates idempotentHint, which keeps it from a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action, then the parameter semantics, then the safety/auth note. Every sentence carries weight and the example is compact. Slightly compressed wording ('after them') costs a little clarity, but nothing is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 7-parameter mutation tool with no output schema, the description covers the non-obvious parameter, the required write scope, and re-run safety. It does not describe the return payload of dry_run or a normal run, which is a minor gap given no output schema exists to carry that information.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 86%, so the baseline is 3, but the description meaningfully explains the one parameter the schema leaves undocumented: 'order' accepts 1-based positions or screenshot IDs, unlisted items retain relative order, and '[4] moves screenshot 4 to the front.' That is genuine semantic value beyond the schema, including the otherwise-unexplained mixed integer/string union.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource ('Changes the order of a screenshot set'), which clearly separates it from sibling mutators like upload_screenshots, delete_screenshots, and replace_screenshot. It is clear but never explicitly names or contrasts with those siblings, so it stops short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by the ordering semantics and the worked example, so an agent can infer how to call it. However, there is no explicit when-to-use guidance, no prerequisites beyond the write scope, and no routing to alternatives such as replace_screenshot when the goal is substitution rather than reordering.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
replace_screenshotReplace one screenshotADestructiveIdempotent
Replaces one screenshot with a new image, keeping its place. Identify it by position (1-based, from list_screenshots) or by screenshot_id. The new image is uploaded and processed before the old one is removed, so the listing never has a gap (except in a full set of 10, where the old one has to go first). If a run stops partway, it says how to continue: call again with screenshot_id so the right screenshot is replaced even if positions shifted.
Destructive: dry_run defaults to true and returns the plan. Show it to the user, then call again with dry_run: false. Needs ASC_WRITE=1.
| Name | Required | Description | Default |
|---|---|---|---|
| app | No | App ID, bundle ID or exact app name. Defaults to ASC_APP_ID, or to the only app the key can see. | |
| file | Yes | Absolute path of the new image. | |
| locale | No | Localization, e.g. en-US. Defaults to the app's primary locale. | |
| dry_run | No | Defaults to true: return the plan without changing anything. Call again with dry_run: false to apply it. | |
| version | No | App Store version string (e.g. "1.2"), "editable" (the version being prepared; the default) or "live". | |
| platform | No | Platform. Defaults to IOS. | |
| position | No | Position to replace, 1-based, as shown by list_screenshots. | |
| display_type | Yes | Screenshot display type, e.g. APP_IPHONE_67 (6.9" iPhone, 1320×2868), APP_IPHONE_65, APP_IPAD_PRO_3GEN_129 (13" iPad, 2064×2752), APP_DESKTOP, APP_APPLE_TV, APP_APPLE_VISION_PRO, APP_WATCH_ULTRA. | |
| wait_minutes | No | Minutes to wait for Apple to process uploaded images (usually under a minute). If time runs out, re-run the same call to continue. | |
| screenshot_id | No | ID of the screenshot to replace. Safer than position when continuing an interrupted run. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Goes well beyond the annotations (destructiveHint, idempotentHint, openWorldHint): it explains the upload-before-delete ordering, the exception when a full set of 10 forces the old image out first, the dry_run default flow, resume semantics, and the ASC_WRITE=1 permission requirement.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action, then layered with the ordering nuance, resume behavior, and a clearly separated destructive/dry-run paragraph. The 10-item-set edge case is a bit granular but still earns its place as a real exception.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 10-parameter destructive tool with no output schema, it covers permissions, dry-run flow, ordering side effects, and how to recover a partial run. It could say more about what the dry-run plan or final result contains, but nothing an agent needs to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds real meaning the schema does not: the position-vs-screenshot_id tradeoff ('safer than position when continuing an interrupted run'), the intent of the dry_run default, and where positions come from.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Replaces one screenshot with a new image, keeping its place') and pins the scope to a single item, which distinguishes it from the bulk siblings upload_screenshots, reorder_screenshots and delete_screenshots without needing to open a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear operational guidance: identify by 1-based position from list_screenshots or by screenshot_id, run dry_run first and confirm, and re-call after an interrupted run. It does not explicitly name an alternative tool ('to add instead of replace, use upload_screenshots'), so it stops short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
reply_to_reviewReply to a customer reviewADestructiveIdempotent
Publishes a developer reply to a customer review. Replies are public, so this is a dry run until confirmed. A review has at most one reply; to change an existing one pass replace: true, which deletes the old reply and posts the new one.
Destructive: dry_run defaults to true and returns the plan. Show it to the user, then call again with dry_run: false. Needs ASC_WRITE=1.
| Name | Required | Description | Default |
|---|---|---|---|
| text | Yes | ||
| dry_run | No | Defaults to true: return the plan without changing anything. Call again with dry_run: false to apply it. | |
| replace | No | ||
| review_id | Yes | Review ID from get_reviews. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint and idempotentHint, but the description goes well beyond them: dry_run defaults to true and returns a plan, a review holds at most one reply, replace deletes the old reply, replies are public, and ASC_WRITE=1 is required. This is exactly the mutation-specific context an agent needs before acting.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four tight sentences, front-loaded with purpose, then the safety workflow. There is mild redundancy: the dry_run default and the 'call again with false' instruction are repeated almost verbatim in the schema's dry_run description.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive, 4-parameter mutation tool with no output schema, the description covers the safety protocol, the permission prerequisite, and the one-reply-per-review edge case. Nothing an agent needs in order to invoke it safely is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 50% (text and replace have no schema descriptions), and the description compensates well for replace ('deletes the old reply and posts the new one') and the dry_run workflow. It says nothing extra about the text body, though maxLength/minLength are already encoded in the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource ('Publishes a developer reply to a customer review'), which is unmistakably distinct from the read-side sibling get_reviews. The added semantics (single reply per review, replace behavior) make the operation's scope unambiguous without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives explicit when-to-use guidance: use replace: true to change an existing reply, and the dry_run two-step workflow ('Show it to the user, then call again with dry_run: false'). The prerequisite ASC_WRITE=1 is also stated, leaving little to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
set_review_detailsSet App Review detailsAIdempotent
Sets the App Review information for the version being prepared: contact name, phone and email, demo account, and notes for the reviewer. Only fields you pass change. The password is never echoed back.
Changes App Store Connect (needs ASC_WRITE=1). Safe to re-run: steps already done are skipped.
| Name | Required | Description | Default |
|---|---|---|---|
| app | No | App ID, bundle ID or exact app name. Defaults to ASC_APP_ID, or to the only app the key can see. | |
| notes | No | ||
| dry_run | No | If true, return the plan without changing anything. | |
| version | No | App Store version string (e.g. "1.2"), "editable" (the version being prepared; the default) or "live". | |
| platform | No | Platform. Defaults to IOS. | |
| contact_email | No | ||
| contact_phone | No | Include the country code, e.g. +1 555 010 0000. | |
| contact_last_name | No | ||
| demo_account_name | No | ||
| contact_first_name | No | ||
| demo_account_password | No | ||
| demo_account_required | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Adds real information beyond the annotations: it names the auth prerequisite (ASC_WRITE=1), clarifies partial-update semantics ('Only fields you pass change'), notes the password is never returned, and explains the idempotent behavior ('steps already done are skipped'). This goes past the idempotentHint/destructiveHint hints, though failure behavior is unstated.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded: the first sentence gives purpose, fields, partial-update behavior, and the password note; the second carries the write requirement and idempotency. Tight and well ordered, with only minor redundancy around re-run safety.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 12-parameter write tool with no output schema, the description covers the key decision-relevant facts: auth requirement, partial updates, and idempotency. It omits any description of the response shape and error handling, which keeps it short of fully complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Description names the principal settable fields (contact name, phone, email, demo account, notes), which maps to roughly half the 12 parameters. With only 42% schema coverage, the app/version/platform selection and dry_run semantics are left largely to the schema, so it does not fully compensate for the gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Sets the App Review information for the version being prepared') and enumerates the fields it controls (contact name/phone/email, demo account, notes). This clearly separates it from other version-config siblings. It stops short of explicitly naming an alternative, so it does not fully meet the 5 bar.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied (App Review prep for the version being prepared) and the re-run safety is noted, but there is no explicit when-to-use vs when-not guidance and no alternatives named among siblings like prepare_version or submit_for_review. Requires inference to place in a workflow.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
set_whats_newSet what's newAIdempotent
Sets release notes. target "app_store" sets the version's "What's New in This Version"; target "testflight" sets a build's "What to Test".
Changes App Store Connect (needs ASC_WRITE=1). Safe to re-run: steps already done are skipped.
| Name | Required | Description | Default |
|---|---|---|---|
| app | No | App ID, bundle ID or exact app name. Defaults to ASC_APP_ID, or to the only app the key can see. | |
| text | Yes | ||
| build | No | Build number (CFBundleVersion, e.g. "202610010021"), build resource ID, or "latest" (the default). | |
| locale | No | Defaults to the primary locale (app_store) or en-US (testflight). | |
| target | Yes | ||
| dry_run | No | If true, return the plan without changing anything. | |
| version | No | App Store version string (e.g. "1.2"), "editable" (the version being prepared; the default) or "live". | |
| platform | No | Platform. Defaults to IOS. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare idempotentHint=true and destructiveHint=false, and the description adds real context beyond them: it names the external system mutated (App Store Connect), the required write credential (ASC_WRITE=1), and clarifies that re-runs skip completed steps. Return/result behavior is not described, keeping it out of 5 territory.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the purpose and target semantics, followed by the mutation target and safety note. Nothing redundant.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an 8-parameter write tool with no output schema and 75% schema coverage, the description covers purpose, both target modes, auth requirements, and idempotency. It omits only what a successful result contains and how the text length limit interacts with target selection.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 75%, and the description supplies the one genuinely undocumented enum meaning: what 'app_store' vs 'testflight' do to text. That adds value beyond the schema, though the interplay of locale/platform/version with each target is left implicit.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Sets release notes') and immediately disambiguates the two modes: app_store writes the version's What's New, testflight writes a build's What to Test. No sibling tool covers release notes, so the agent can identify the capability without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The target semantics effectively tell the agent which mode to pick for a version vs a build, and it flags the ASC_WRITE=1 prerequisite. It stops short of naming alternatives or stating when not to use it (e.g. versus update_listing), but the context is clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
submit_for_reviewSubmit for App ReviewADestructiveIdempotent
Submits the version being prepared to App Review. First checks what the API can see (build attached and processed, descriptions, keywords, support URL, screenshots, privacy policy URL, review contact) and stops if something is missing. Then creates or reuses a review submission, adds the version and submits it. Things the API can't check: the App Privacy questionnaire, pricing and availability, agreements and tax forms.
Destructive: dry_run defaults to true and returns the plan. Show it to the user, then call again with dry_run: false. Needs ASC_WRITE=1.
| Name | Required | Description | Default |
|---|---|---|---|
| app | No | App ID, bundle ID or exact app name. Defaults to ASC_APP_ID, or to the only app the key can see. | |
| dry_run | No | Defaults to true: return the plan without changing anything. Call again with dry_run: false to apply it. | |
| version | No | App Store version string (e.g. "1.2"), "editable" (the version being prepared; the default) or "live". | |
| platform | No | Platform. Defaults to IOS. | |
| skip_checks | No | Submit even if the pre-flight checks find problems; Apple will still validate. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the destructiveHint/openWorldHint/idempotentHint annotations, it discloses the safe default (dry_run=true returns a plan), the required human confirmation loop, the ASC_WRITE=1 authorization requirement, and the boundaries of what the API can verify (privacy questionnaire, pricing, agreements). This is unusually rich behavioral context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the core action, then the check sequence, then the caveats and dry_run workflow. Each sentence carries actionable information, though the 'things the API can't check' list is dense enough to be slightly long.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutating, no-output-schema tool, the description covers the mutation's effect, the returned plan artifact, the safety workflow, and auth prerequisites. Nothing essential for a correct invocation is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3; the description nonetheless adds meaning by explaining the dry_run contract, the pre-flight check behavior that skip_checks bypasses, and the default 'editable' version semantics. It slightly enriches rather than merely restating the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (submits) and resource (the version being prepared, to App Review), and scopes it to the editable version by default. It is clearly distinguishable from siblings like prepare_version (which readies a version) and cancel_review_submission (which reverses this).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explains the workflow context well: it runs pre-flight checks first and stops on missing items, and lays out the two-step dry_run confirmation. It does not explicitly name alternative tools, but the sequencing guidance is clear enough for an agent to know when to invoke it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
update_age_ratingUpdate age ratingAIdempotent
Answers the age rating questionnaire of the app info being prepared. Pass the answers to change, using Apple's attribute names, e.g. {"violenceCartoonOrFantasy": "INFREQUENT_OR_MILD", "gambling": false}. Content questions take NONE, INFREQUENT_OR_MILD or FREQUENT_OR_INTENSE: alcoholTobaccoOrDrugUseOrReferences, contests, gamblingSimulated, gunsOrOtherWeapons, horrorOrFearThemes, matureOrSuggestiveThemes, medicalOrTreatmentInformation, profanityOrCrudeHumor, sexualContentGraphicAndNudity, sexualContentOrNudity, violenceCartoonOrFantasy, violenceRealistic, violenceRealisticProlongedGraphicOrSadistic. Yes/no questions take true/false: advertising, ageAssurance, gambling, healthOrWellnessTopics, lootBox, parentalControls, unrestrictedWebAccess, userGeneratedContent, plus messagingAndChat, socialMedia. Apple needs every question answered before it accepts any change; fill_unanswered: true answers the rest NONE/false. Confirm those answers with the user.
Changes App Store Connect (needs ASC_WRITE=1). Safe to re-run: steps already done are skipped.
| Name | Required | Description | Default |
|---|---|---|---|
| app | No | App ID, bundle ID or exact app name. Defaults to ASC_APP_ID, or to the only app the key can see. | |
| answers | No | ||
| dry_run | No | If true, return the plan without changing anything. | |
| fill_unanswered | No | Answer every question that has no answer yet with NONE or false. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Adds value beyond annotations by disclosing the ASC_WRITE=1 permission requirement and explaining the safety model ('steps already done are skipped'), which substantiates the idempotentHint. The prerequisite that every question must be answered before any change is accepted is useful behavioral context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads purpose, then the input format, then the attribute taxonomy, then the completion constraint. It is dense and the long inline enumeration is hard to scan, but nearly every sentence carries necessary information an agent needs to construct valid input.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with no output schema and a nested free-form parameter, the description covers the operation, auth requirement, idempotency, and a prerequisite well. It is complete enough to call correctly, with only return-value behavior left implicit.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The `answers` parameter is an undocumented free-form object in the schema, and the description compensates fully by enumerating Apple's attribute names and their value domains (the three-level content scale and the true/false yes-no set). This is exactly the compensation needed for a nested, enum-less object.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: it answers/updates the age rating questionnaire for the app info being prepared. This is clearly distinguishable from siblings like update_listing or set_review_details, which handle different metadata.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear operative context: it must be completed before Apple accepts any change, and instructs the agent to use fill_unanswered and confirm answers with the user. No sibling alternatives are named or excluded, but none are obviously needed for this tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
update_listingUpdate store listingAIdempotent
Updates App Store metadata for one locale. Only fields you pass change, and only if they differ; the result shows each change as old → new. Name, subtitle and privacy URLs live on the app info; description, keywords, promotional text, what's new and URLs on the version. Adds the localization if it doesn't exist yet. Only versions being prepared can change, except promotional text, which can change on the live version (pass version: "live").
Changes App Store Connect (needs ASC_WRITE=1). Safe to re-run: steps already done are skipped.
| Name | Required | Description | Default |
|---|---|---|---|
| app | No | App ID, bundle ID or exact app name. Defaults to ASC_APP_ID, or to the only app the key can see. | |
| name | No | ||
| locale | No | Defaults to the primary locale. | |
| dry_run | No | If true, return the plan without changing anything. | |
| version | No | App Store version string (e.g. "1.2"), "editable" (the version being prepared; the default) or "live". | |
| keywords | No | Comma-separated, 100 characters in total. | |
| platform | No | Platform. Defaults to IOS. | |
| subtitle | No | ||
| copyright | No | Version copyright, e.g. "2026 Example Inc." | |
| whats_new | No | ||
| description | No | ||
| support_url | No | ||
| marketing_url | No | ||
| primary_category | No | Category ID such as GAMES, UTILITIES, PRODUCTIVITY, HEALTH_AND_FITNESS. | |
| promotional_text | No | ||
| privacy_policy_url | No | ||
| secondary_category | No | ||
| privacy_choices_url | No | ||
| primary_subcategories | No | Games only: up to two subcategories of the primary category, e.g. ["GAMES_PUZZLE", "GAMES_BOARD"]. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the annotations (readOnlyHint=false, idempotentHint=true), it discloses that only fields passed change and only if they differ, that results are shown as old → new, that localization is auto-created, that edits are gated on editable versions, and that re-runs are safe. This is rich context an agent needs before mutating metadata.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Dense but front-loaded: mutation scope and diff behavior lead, followed by field grouping and constraints. Every sentence carries information, though the field-placement sentence is a little list-heavy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 19-param mutation tool with no output schema, it covers auth, idempotency, editability gating, and the old → new return shape. Remaining gaps (exact response structure, per-field format) are minor given the annotations and schema already present.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With schema coverage at 47% across 19 params, the description compensates meaningfully by grouping fields by location (app info vs version) and explaining the version:"live" and default-locale semantics. It doesn't clarify the remaining undocumented params' formats, but most are self-evident fields.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb+resource: 'Updates App Store metadata for one locale.' It scopes the operation (one locale) and effectively distinguishes itself from read-oriented siblings like get_listing and the narrower set_whats_new by covering the full metadata surface.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives clear conditions for use: only versions being prepared can change, except promotional text on the live version via version:"live". It also notes the write requirement (ASC_WRITE=1). It stops short of naming explicit alternatives, so it falls just below the top tier.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
upload_buildUpload a buildAIdempotent
Uploads an exported .ipa (or macOS .pkg) to App Store Connect. method "api" uses Apple's build upload API (no Xcode needed); "altool" shells out to xcrun altool on a Mac. Make the .ipa first with xcodebuild archive + xcodebuild -exportArchive (method app-store-connect, destination export). Every upload needs a new, higher build number (CFBundleVersion). After it's processed, use distribute_build to send it to testers.
Changes App Store Connect (needs ASC_WRITE=1). Safe to re-run: steps already done are skipped.
| Name | Required | Description | Default |
|---|---|---|---|
| app | No | App ID, bundle ID or exact app name. Defaults to ASC_APP_ID, or to the only app the key can see. | |
| file | Yes | Absolute path to the .ipa or .pkg. | |
| method | No | api | |
| dry_run | No | If true, return the plan without changing anything. | |
| version | No | CFBundleShortVersionString, e.g. 1.2. Read from the .ipa if omitted (macOS only). | |
| platform | No | Platform. Defaults to IOS. | |
| build_number | No | CFBundleVersion. Read from the .ipa if omitted (macOS only). | |
| wait_minutes | No | Minutes to wait for Apple's processing before returning (0-30, default 0). Processing usually takes 5-30 minutes; if it isn't done in time, the result says how to resume. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare idempotentHint=true and destructiveHint=false, but the description adds value beyond them: the ASC_WRITE=1 auth requirement and the concrete idempotency mechanism ('steps already done are skipped'). It also notes Apple processing typically takes 5-30 minutes and that the result explains how to resume, which the annotations do not convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action and each sentence carries distinct information (method choice, build prerequisite, versioning rule, downstream step, auth/re-run note). It is dense and slightly long, but no sentence is filler, so it stays efficient rather than bloated.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with 8 params, no output schema and open-world behavior, the description covers prerequisites, auth, idempotency, the 5-30 minute processing wait, and the resume path. Nothing an agent needs to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is high (88%), so the baseline is 3, but the description adds real meaning: it clarifies the method enum tradeoff, the .ipa/.pkg file expectation, and the critical constraint that build_number (CFBundleVersion) must be new and higher on every upload. That last point is not derivable from the schema alone.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Uploads an exported .ipa (or macOS .pkg) to App Store Connect') and distinguishes itself from the sibling distribute_build by naming it as the downstream step. An agent can identify exactly what this tool does and where it sits in the pipeline without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly explains when to use each method ('api' uses Apple's build upload API, no Xcode; 'altool' shells out to xcrun altool on a Mac) and gives prerequisites: build the .ipa first with xcodebuild archive/-exportArchive, and use a new higher build number. It also names the follow-up action (distribute_build), leaving no inference required.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
upload_screenshotsUpload screenshotsADestructiveIdempotent
Uploads images to one screenshot set of the version being prepared. mode "replace" makes the set exactly these files in this order (the old ones are deleted only after the new ones are processed); mode "append" adds them at the end. Files already in the set (same MD5) aren't uploaded again, so re-running is safe.
Destructive: dry_run defaults to true and returns the plan. Show it to the user, then call again with dry_run: false. Needs ASC_WRITE=1.
| Name | Required | Description | Default |
|---|---|---|---|
| app | No | App ID, bundle ID or exact app name. Defaults to ASC_APP_ID, or to the only app the key can see. | |
| mode | No | replace | |
| files | No | Absolute image paths, in display order. | |
| folder | No | Folder of .png/.jpg files, used in natural filename order (1, 2, …, 10). Use instead of files. | |
| locale | No | Localization, e.g. en-US. Defaults to the app's primary locale. | |
| dry_run | No | Defaults to true: return the plan without changing anything. Call again with dry_run: false to apply it. | |
| version | No | App Store version string (e.g. "1.2"), "editable" (the version being prepared; the default) or "live". | |
| platform | No | Platform. Defaults to IOS. | |
| display_type | Yes | Screenshot display type, e.g. APP_IPHONE_67 (6.9" iPhone, 1320×2868), APP_IPHONE_65, APP_IPAD_PRO_3GEN_129 (13" iPad, 2064×2752), APP_DESKTOP, APP_APPLE_TV, APP_APPLE_VISION_PRO, APP_WATCH_ULTRA. | |
| wait_minutes | No | Minutes to wait for Apple to process uploaded images (usually under a minute). If time runs out, re-run the same call to continue. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, idempotentHint=true, and readOnlyHint=false, but the description adds real behavioral context beyond them: old images are deleted only after the new ones are processed (atomicity), MD5-based dedup makes re-runs safe (substantiates idempotency), and it requires ASC_WRITE=1. It also flags the destructive nature and the two-step dry_run confirmation pattern. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action, then behavior, then the safety workflow. Two tight paragraphs with little waste. Minor redundancy: dry_run's default and behavior are restated from the schema, and the destructive sentence partially overlaps the annotation.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 10-parameter mutation tool with no output schema, the description covers the essentials: what it does, ordering/atomicity, dedup, auth requirement, and the required dry_run confirmation loop. It leaves some parameters (display_type, wait_minutes, version/platform values) to the well-covered schema, which is acceptable given 90% coverage, but does not describe the plan object returned by dry_run.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 90%, so the baseline is 3, but the description adds genuine meaning: it explains mode's replace-vs-append semantics (replace makes the set exactly these files in order; append adds at the end) and why re-running is safe (MD5 dedup on the files parameter). This meaningfully complements the enum without restating it.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource+scope: "Uploads images to one screenshot set of the version being prepared." An agent can tell it operates on a whole set rather than a single slot, which helps versus replace_screenshot. However, it never names or contrasts with the closest siblings (replace_screenshot, reorder_screenshots, list_screenshots), so full sibling differentiation is left to inference.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives concrete workflow guidance: dry_run defaults to true, show the plan to the user, then re-call with dry_run: false, and explains that re-running is safe. It also documents when to use replace vs append. It lacks an explicit "use this instead of replace_screenshot when..." routing, so it stays at clear-context rather than explicit-alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
31 tool updates
v0.3.0- First observed
add_free_trial - First observed
asc_request - First observed
cancel_review_submission - First observed
create_beta_group - First observed
delete_screenshots - First observed
distribute_build - First observed
download_report - First observed
get_app_status - First observed
get_build - First observed
get_listing - First observed
get_reviews - First observed
invite_testers - First observed
list_apps - First observed
list_beta_groups - First observed
list_builds - First observed
list_screenshots - First observed
list_subscriptions - First observed
list_testers - First observed
prepare_version - First observed
remove_intro_offers - First observed
remove_testers - First observed
reorder_screenshots - First observed
replace_screenshot - First observed
reply_to_review - First observed
set_review_details - First observed
set_whats_new - First observed
submit_for_review - First observed
update_age_rating - First observed
update_listing - First observed
upload_build - First observed
upload_screenshots
TDQS
Scored across 31 tools
Most tools target a clearly distinct resource+action (builds vs screenshots vs testers vs reviews vs subscriptions), and the destructive/write tools are well-differentiated. Minor overlap exists between set_whats_new and update_listing (which also sets what's new), and asc_request is a deliberate catch-all that could tempt use where a workflow tool fits better.
Every tool is snake_case with a verb-first pattern (list_*, get_*, create_*, update_*, remove_*, upload_*, etc.), matching resource nouns consistently. asc_request is the only mild deviation but is still readable and intentional.
31 tools is on the heavy side and exceeds the comfortable 3-15 range, though the App Store Connect domain is genuinely broad (builds, testers, screenshots, listings, subscriptions, reviews, reports, review submission). The surface is large but mostly earns its place rather than being redundant.
Strong lifecycle coverage: app discovery, build upload/distribution, TestFlight groups and testers, listing/metadata editing, screenshots, age rating, subscriptions/offers, reviews, and review submission. Known gaps (pricing/availability, App Privacy questionnaire, paid intro offers) are explicitly documented and bridged by the asc_request escape hatch.
Maintenance
Related MCP Connectors
- app-managerOAuthapp.lance
App Store Connect operator for AI agents: icons, TestFlight builds, listings, IAP, rejection fixes.
- metarunOAuthdev.metarun
Run App Store Connect from your IDE: pricing, listings, screenshots, releases, AI visibility.
Create App Store screenshots, app icons and store copy, then push them to App Store Connect.
Read products, sales, subscribers and offer codes; verify, enable and disable product licenses.
Related MCP Servers
- AlicenseNot gradedqualityCmaintenanceEnables AI agents to manage Apple App Store Connect through the official API, including apps, metadata, reviews, TestFlight, provisioning, users, and reports.MIT
- AlicenseNot gradedqualityCmaintenanceEnables analysis and management of iOS/macOS apps via the App Store Connect API, including app management, reviews, sales reports, analytics, performance metrics, and TestFlight.16 npm2MIT
- AlicenseNot gradedqualityFmaintenanceEnables AI agents to manage App Store Connect apps, including registering bundle IDs, uploading metadata and screenshots, setting age ratings, managing TestFlight groups and testers, and submitting apps for review.MIT
- AlicenseCqualityCmaintenanceEnables management of App Store Connect resources including app reviews, TestFlight crashes, analytics reports, and Xcode Cloud workflows through natural language or AI agents.401MIT