AppFactory
Integrates with App Store Connect to manage apps, in-app purchases, metadata, and TestFlight distributions via the ASC API.
Allows setting up Firebase Analytics for the app.
Creates and manages a GitHub repository and issues for each app.
Provides backend integration with Supabase for database, edge functions, and credits management.
Enables local iOS builds, simulator runs, and archives via Xcode.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@AppFactoryfind 3 underserved iOS app niches in Health & Fitness and validate the best one"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
AppFactory
An MCP server that lets your coding agent take a SwiftUI iOS subscription app from idea to the App Store.
What it does: idea harvesting and validation, design, SwiftUI code from a template, Supabase backend, App Store Connect + RevenueCat setup, localization, analytics, ASO, screenshots and TestFlight. Every stage has a gate the agent must pass. You do the final App Store submission.
Quickstart
Requires a Mac, Python 3.12+ and uv.
1. Add the MCP to your agent
One line, no installer. The server command is uvx appfactory.
Agent | How |
Claude Code |
|
Codex CLI |
|
Gemini CLI |
|
Cursor |
|
2. Ask your agent: "Set up AppFactory"
The agent calls setup_status, asks which services you want, and turns them on. Keys never go through the chat:
for secrets it opens a local page in your browser (127.0.0.1 only, one-time token) where you type them; they are
stored in ~/.appfactory/config.toml (mode 0600). appfactory doctor shows what is configured.
Then try this first prompt. It needs no accounts:
Find 3 underserved iOS app niches in Health & Fitness and validate the best one
Related MCP server: agentloop
Works with
Claude Code, Codex CLI, Gemini CLI and Cursor. The run playbook is served by the MCP server itself (tool
playbook(), prompt run, resource appfactory://playbook), so any MCP-capable agent can follow it.
Choose what you use
Every service is optional. Tools for a disabled service do nothing and the pipeline skips those stages
(appfactory services enable|disable NAME).
Service | What it enables | Needs |
research | Idea harvesting, ASO research (always on) | nothing |
apple | App Store Connect: apps, in-app purchases, metadata, TestFlight | ASC API key, |
xcode | Local builds, simulator, archives | Xcode, |
supabase | Backend: database, edge functions, credits | access token, |
revenuecat | Subscriptions and entitlements | RevenueCat secret key |
ai | AI features through a server-side proxy | optional provider keys |
firebase | Firebase Analytics setup |
|
github | A repo and issues per app |
|
design | Screen design through the claude-design MCP | claude-design MCP |
maestro | End-to-end UI tests on the simulator | Maestro, Java 17+ |
lottie | Lottie animations |
|
Details and credential steps: docs/SETUP.md.
Safety
Approvals. Irreversible or outward actions (submission, store/backend/RevenueCat writes, GitHub pushes, uploads) do not run when an agent calls them. They wait for
appfactory approve <id>in your terminal.Dry run by default for deploys, repo creation, uploads and issue writes.
Secrets stay local and are scrubbed from every tool result.
Untrusted content (App Store listings, reviews) is labelled as data, not instructions.
Full threat model and limits: SECURITY.md.
How a run works
Interview. Before any work the agent calls
run_options(), asks you about every optional part, and saves your answers. Anything you turn off is skipped.Pipeline with gates.
scaffold, design, features, backend, store setup, localize, analytics, ASO, metadata, screenshots, TestFlight. A stage is done only whenpipeline_markpasses its gate; blockers stop the run with aNEEDS_HUMAN.md.You submit. App Store submission is always human, in the App Store Connect web UI.
Docs
docs/SETUP.md: credentials and per-app choices
docs/PIPELINE_PLAYBOOK.md: the why, order and gotchas of every stage
docs/TOOLS.md: every tool (generated)
License
Available Tools
125 toolsai_configureC
AI provider/model selection (for now fal.ai/Flux schnell — the cheapest).
| Name | Required | Description | Default |
|---|---|---|---|
| providers | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden, and it discloses almost nothing: not whether this persists a setting, applies immediately, is global or per-project, or what happens when the value is null (the schema default). The parenthetical cost remark is the only extra context, and it is incomplete (unclosed).
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One short sentence that is front-loaded and easy to scan, but it is under-specified rather than efficient, and the trailing parenthetical is malformed (missing close parenthesis and a fragmentary dash clause).
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained, but for a configuration-mutation tool with a 0%-documented nullable parameter and no annotations, the description is far too thin to let an agent invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% and the single parameter is named 'providers' (an array, nullable, default null), yet the description talks about 'provider/model selection' in the singular. The relationship between the array parameter and the described model choice is never clarified, so the description fails to compensate for the undocumented schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The phrase 'AI provider/model selection' names the resource but supplies no real verb describing what the tool does to it (set? read? apply?). It does not distinguish itself from siblings like config_set, ai_deploy_proxy, or setup_services, leaving the agent to guess the actual operation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use, when-not-to-use, or alternative-tool guidance at all. With config_set and config_doctor among siblings, the agent has no signal about which configuration entry point to pick.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ai_deploy_proxyC
Deploy the AI backend. With app.spec.json in subscription mode this is backend_deploy (the full function set + migrations + secrets + auth); credits mode keeps the legacy ai-proxy deploy. AI keys go only to server-side secrets. LIVE side effects: human approval.
| Name | Required | Description | Default |
|---|---|---|---|
| app_dir | Yes | ||
| approval_id | No | ||
| daily_limit | No | ||
| project_ref | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description must carry the full behavioral burden. It does disclose useful traits: 'AI keys go only to server-side secrets' and 'LIVE side effects: human approval.' However, it omits critical operational detail: what the human-approval workflow looks like, how approval_id is obtained, what gets mutated (functions, migrations, secrets, auth), and whether deployment is reversible. For a live-effect deploy tool with zero annotations, this is partial transparency.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three compact sentences with no filler. The key live-effect warning is placed last, which is a reasonable emphasis, though the conditional 'this is backend_deploy' framing makes the first sentence harder to parse than necessary.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Although an output schema exists, the tool has a live-effect deploy with human approval, four undocumented parameters, and no annotations. The description leaves the approval flow, parameter meanings, and override behavior under-specified. Given the complexity and risk, this is insufficient.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so all four parameters (app_dir, project_ref, approval_id, daily_limit) are undocumented in both schema and description. The description hints at approvals ('human approval') but does not explain approval_id, does not define project_ref or app_dir, and does not mention daily_limit at all, which is a meaningful limitation knob. The description does not compensate for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific action ('Deploy the AI backend') but then immediately complicates it by saying 'this is backend_deploy' in one mode and 'legacy ai-proxy deploy' in another. It is not clear whether ai_deploy_proxy is its own tool, an alias for backend_deploy, or a mode switch. The sibling list contains both backend_deploy and ai_configure, making the distinction ambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives a mode-dependent condition ('with app.spec.json in subscription mode this is backend_deploy... credits mode keeps the legacy ai-proxy deploy'), which implies when each behavior applies. However, it does not tell the agent when to call ai_deploy_proxy versus calling backend_deploy directly, nor what prerequisites besides app.spec.json are needed. Usage is implied rather than explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
animation_fetch_recolorA
Fetch a CUTE app-specific Lottie animation + recolor it to the palette → Resources/Animations/.json.
Source: the free LottieFiles library (Simple License = commercial OK, no attribution). The most-viewed
match is picked and its colors are mapped to the app palette (dominant→primary, others→accent, black/white kept).
Build-time recolor — NO runtime color code in Swift; lottie-spm renders it with .named(slot).
slot: loading | success | empty | onboarding_hero. keyword: app-themed override (e.g. 'happy dog celebration').
primary_hex/accent_hex: DS.Palette light hex ('#' optional, e.g. '006A63'). Called for every slot of every app.
| Name | Required | Description | Default |
|---|---|---|---|
| slot | Yes | ||
| app_dir | Yes | ||
| keyword | No | ||
| accent_hex | Yes | ||
| primary_hex | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does so well: it discloses the asset source and license, the selection rule (most-viewed match), the deterministic color-mapping policy (dominant→primary, others→accent, black/white kept), and critically that recolor is build-time with no runtime Swift code. It omits failure modes (no match found), network/rate-limit behavior, and auth needs.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The first sentence front-loads the action and destination, followed by compact source, recolor, and parameter notes. Usage of arrows and line-break fragments is dense but each sentence adds real information; only minor trimming is possible.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists so return values need not be explained, and the description compensates for 0% schema coverage across the non-obvious parameters. What remains missing is error/failure handling and any dependency on network access to LottieFiles.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate and largely does: slot is enumerated (loading | success | empty | onboarding_hero), keyword gets an example, and primary_hex/accent_hex get format rules ('#' optional, format hint '006A63'). app_dir is left undocumented but is self-evident from the output path.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific two-part action (fetch + recolor a Lottie animation) and names the exact output artifact (Resources/Animations/<slot>.json), which no sibling tool produces. It is immediately distinguishable from nearby tools like mascot_assets, icon_generate, or design_generate.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear context: the tool is called for every slot of every app, and the keyword parameter is for app-themed overrides. It does not explicitly name competing tools (e.g. mascot_assets, icon_generate) or state when NOT to use it, but the usage context is unambiguous.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
app_inject_configB
Fill in the scaffolded app's AppConfig + StoreKit tokens.
paywall_strategy: "hard_only" (hard paywall only) | "hard_and_offer" (default, hard paywall + a discounted offer paywall on dismiss). Credit fields fall back to defaults if omitted (15/10/10/30/60) — credit pack amounts must be kept in sync with the server-side ai-proxy PACK_MAP.
| Name | Required | Description | Default |
|---|---|---|---|
| posthog_key | No | ||
| privacy_url | No | ||
| project_dir | Yes | ||
| ai_proxy_url | No | ||
| supabase_url | No | ||
| support_email | No | ||
| product_weekly | No | ||
| product_yearly | No | ||
| revenuecat_key | No | ||
| weekly_credits | No | ||
| yearly_credits | No | ||
| paywall_strategy | No | ||
| credit_pack_large | No | ||
| credit_pack_small | No | ||
| supabase_anon_key | No | ||
| credit_pack_medium | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It usefully discloses paywall defaults, credit fallback values, and the need to keep pack amounts in sync with the server-side PACK_MAP, but it omits mutation side effects, overwrite behavior, permissions, and auth requirements expected of a config-injection tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loads the core purpose, then provides the most important parameter-specific details. The paywall strategy and credit default notes are reasonably concise, though the enum formatting could be slightly clearer.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 16-parameter configuration mutation with no annotations and no schema property descriptions, the description is materially incomplete. It partially covers paywall and credit settings but does not explain what most fields represent or how the injection affects the scaffolded app; the output schema only covers return values.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across 16 parameters, so the description must compensate and largely does not. It explains paywall_strategy values and some credit defaults, but leaves most parameters such as posthog_key, privacy_url, supabase_url, product identifiers, and revenuecat_key undocumented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: filling in a scaffolded app's AppConfig and StoreKit tokens. It is clear what the tool does, though it does not explicitly distinguish itself from siblings like config_set, setup_services, or app_scaffold beyond the 'scaffolded app' context.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'scaffolded app's' implies the prerequisite that scaffolding must already exist, and the paywall_strategy explanation gives some inline guidance. However, it does not say when to use this tool versus alternatives such as config_set or setup_services, nor does it state when not to use it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
app_scaffoldA
Scaffold a new SwiftUI app from app.spec.json (+xcodegen) and init the pipeline manifest.
spec: optional full spec dict, a dict of overrides on the defaults, or a path to an app.spec.json; otherwise the defaults for name/bundle_id. The spec is written to the app dir and drives products, StoreKit file, placements, locales, onboarding length, monetization mode (subscription strips the credit economy) and the Supabase backend mode.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | ||
| spec | No | ||
| dest_dir | No | ||
| bundle_id | No | ||
| display_name | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does disclose meaningful side effects: it writes the spec into the app dir, runs xcodegen, initializes the pipeline manifest, and controls monetization mode where 'subscription strips the credit economy'. It omits whether existing directories are overwritten, auth requirements, or failure modes.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action, and the spec explanation is dense but earns its place given the polymorphic parameter. No filler sentences, though the side-effect list is a long run-on.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists so return values need no explanation, and the main 'spec' parameter is covered. However, for a 5-parameter mutation with no annotations, the silent dest_dir and display_name parameters and the absence of overwrite/prerequisite behavior leave gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It explains the 'spec' parameter well (dict of overrides, full dict, or path to app.spec.json) and mentions name/bundle_id defaults, but dest_dir and display_name are never addressed, leaving the agent guessing where output lands.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Scaffold a new SwiftUI app') plus a second effect ('init the pipeline manifest'), which clearly separates it from siblings like app_sync_spec or design_screens_skeleton. An agent can identify the operation without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied (start a new app before building), and the description explains the spec fallback behavior ('otherwise the defaults for name/bundle_id'), but it never says when to prefer this over app_sync_spec or design_screens_skeleton, nor any prerequisites (e.g., must dest_dir be empty).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
app_sync_specB
Re-generate AppSpec.swift, PaywallSource.swift and Configuration.storekit (and prune the String Catalogs to locales.app) after editing app.spec.json.
| Name | Required | Description | Default |
|---|---|---|---|
| app_dir | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It usefully discloses that the operation overwrites generated files and prunes String Catalogs to locales.app, hinting at a destructive/reconciling mutation, but it says nothing about required permissions, reversibility, or what happens to manual edits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One dense sentence, front-loaded with the primary action and parenthetical detail for the secondary effect. Every clause earns its place with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need no explanation, and the description covers the generated artifacts well. It is still incomplete for a mutating tool with zero annotations and an undocumented required parameter.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The single parameter app_dir has 0% schema description coverage and is never mentioned in the description; the agent can only infer from 'app.spec.json' that a project directory is involved. The description does not compensate for the coverage gap, so it falls below the 0-param baseline of 4.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific action ('Re-generate') and enumerates the exact artifacts produced: AppSpec.swift, PaywallSource.swift and Configuration.storekit, plus a catalog-pruning side effect. This makes the tool's job unambiguous, though it stops short of explicitly distinguishing itself from close siblings like storekit_generate.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'after editing app.spec.json' supplies a clear triggering condition, which is genuine usage context. However, it names no alternative tool and gives no when-not guidance, so an agent must infer that this is the post-edit sync step.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
asc_add_subscription_group_localizationD
Subscription group display name (required for MISSING_METADATA).
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | ||
| locale | No | en-US | |
| group_id | Yes | ||
| approval_id | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It does disclose one useful fact — that this localization is needed to resolve a MISSING_METADATA condition — but says nothing about permissions, idempotency, what happens to existing localizations, or approval_id behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
It is short, but brevity here comes from under-specification rather than efficiency — a single parenthetical fragment that is not even a complete sentence. There is no front-loaded statement of the tool's action or scope.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 4-parameter mutation tool with no annotations and 0% schema description coverage, the description is far too thin. An output schema exists so return values need not be explained, but the inputs, side effects, and usage context are all absent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and all four parameters (group_id, locale, approval_id, name) are undocumented in the schema. The description adds only a partial hint about 'name' (its role in MISSING_METADATA) and leaves group_id, locale, and approval_id entirely unexplained, failing to compensate for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description reads as a parameter label ('Subscription group display name') rather than a statement of what the tool does. It never uses a verb like 'adds' or 'creates', so the agent must infer the action from the tool name alone. It only incidentally distinguishes it from siblings like asc_localize_group.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The parenthetical 'required for MISSING_METADATA' hints at one triggering condition, but there is no when-to-use guidance, no mention of alternatives (e.g., asc_localize_group vs. this tool), and no prerequisites. The agent gets a fragment of context, not routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
asc_append_subscription_disclosureA
APPEND subscription disclosure + Terms/EULA + Privacy link to the description (per language) (Apple 3.1.2). Idempotent (skips if the marker is present). If terms_url is omitted, Apple's standard EULA is used. Truncated to ≤4000.
| Name | Required | Description | Default |
|---|---|---|---|
| bundle_id | Yes | ||
| terms_url | No | ||
| approval_id | No | ||
| privacy_url | Yes | ||
| disclosure_by_locale | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does reasonably well: it discloses idempotency (skips if marker present), the default fallback (Apple's standard EULA when terms_url omitted), and a truncation limit (≤4000). It omits auth/permission requirements and exactly which description field is mutated, so it is not exhaustive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight clauses, front-loaded with the verb and resource, with the two behavioral caveats (idempotency, truncation) placed after. Nothing is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be described. The description covers the key behavioral facts (append target, idempotency, default EULA, truncation), but leaves the targeted description field and any authorization prerequisite unstated for a mutation tool with no annotations.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It usefully explains terms_url (defaults to Apple's EULA) and implies per-locale behavior via disclosure_by_locale, but bundle_id, privacy_url, and approval_id carry no added meaning beyond their names.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb (APPEND) and resource (subscription disclosure + Terms/EULA + Privacy link to the description), plus the compliance context (Apple 3.1.2). This distinguishes it clearly from siblings like asc_localize_subscription or asc_finalize_subscription.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It implies the usage context (App Store subscription disclosure compliance for Apple 3.1.2) and notes idempotent re-runs, but never states when to use this versus alternatives or what prerequisites apply. Usage is inferable but not spelled out.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
asc_create_appA
Create an app SHELL in App Store Connect (fastlane produce — requires Apple ID). AUTOMATIC.
The one ASC write that does not go through the asc CLI: Apple's public API cannot create apps and
asc web apps create needs the same interactive 2FA web session, so produce stays.
App names are globally unique. Pass candidate_names → the MCP AUTOMATICALLY picks
the first available one (on a "name taken" error it moves to the next) and returns app_id.
app_name (single name) is for backward compatibility; it's converted to a one-element candidate list.
| Name | Required | Description | Default |
|---|---|---|---|
| sku | No | ||
| app_name | No | ||
| bundle_id | Yes | ||
| approval_id | No | ||
| candidate_names | No | ||
| primary_language | No | en-US |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden well: it discloses the Apple ID requirement, the interactive 2FA dependency, that this is a write that creates an app shell, and the automatic name-selection fallback on 'name taken' errors. It stops short of stating permissions, rate limits, or what the partial shell contains.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short paragraphs, front-loaded with the purpose and verb. The middle paragraph justifying why produce stays is slightly tangential but provides useful context; overall efficient with no egregious padding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 6-param write tool with no annotations, it covers the critical creation flow, name-uniqueness behavior, and returns app_id (output schema also exists). The main gap is the undocumented sku/approval_id/primary_language parameters, but the core invocation path is complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate for all 6 params. It thoroughly explains candidate_names and app_name (backward-compat, converted to a one-element list) and the returned app_id, but sku, approval_id, primary_language, and bundle_id go entirely undocumented in both schema and description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Create an app SHELL in App Store Connect') with the underlying mechanism (fastlane produce). It distinguishes itself from siblings by noting it is 'the one ASC write that does not go through the asc CLI' and is separate from asc_create_bundle_id or subscription creation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explains why this tool exists versus the asc CLI and `asc web apps create` (the latter needs an interactive 2FA web session), which effectively routes the agent here. It lacks an explicit when-not statement or a named sibling to prefer for related tasks, but the context for selecting it is clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
asc_create_bundle_idA
Register a bundle id in App Store Connect (LIVE write), then its capabilities.
The App ID capabilities (IN_APP_PURCHASE, APPLE_ID_AUTH/PRIMARY_APP_CONSENT, HEALTHKIT if spec.health.enabled) must exist BEFORE any key, profile or signing step. With app_dir the spec's capabilities are checked right away; apply_capabilities=true also applies them. Human approval.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | ||
| app_dir | No | ||
| identifier | Yes | ||
| approval_id | No | ||
| apply_capabilities | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations present, the description carries the burden well: it flags 'LIVE write', states 'Human approval' (explaining the approval gate), and explains the capability preconditions and the effect of apply_capabilities/app_dir. It still omits failure/reversibility behavior and the relationship of approval_id to the approval flow.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The core action is front-loaded in the first clause, and the dense parenthetical enumerates capabilities efficiently. The second paragraph is information-rich but slightly crammed, preventing a 5.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained. Given no annotations and a mutation tool with a human-approval gate, the description supplies the essential safety and ordering context, though it does not cover error or reversal behavior.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate. It explains app_dir (checks capabilities right away), apply_capabilities (also applies them), and the capability enum values, but identifier, name, and approval_id receive no explicit mapping in the text. Partial compensation only.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Register a bundle id in App Store Connect') and marks it as a LIVE write, which distinguishes it from read-oriented siblings like asc_list_apps and asc_get_app. An agent can identify the operation without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear sequencing context: the App ID capabilities must exist BEFORE any key, profile or signing step, and explains the app_dir vs apply_capabilities behavior. It does not name a competing sibling tool or state explicit when-not-to-use conditions, keeping it below a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
asc_create_subscriptionC
Create a subscription product. period: ONE_WEEK/ONE_MONTH/.../ONE_YEAR (LIVE write).
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | ||
| period | No | ONE_YEAR | |
| group_id | Yes | ||
| product_id | Yes | ||
| approval_id | No | ||
| family_shareable | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. The '(LIVE write)' tag is a genuinely useful signal that this mutates production state, but the description says nothing about required permissions, reversibility, side effects on existing products, or the approval_id workflow implied by the schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two compact clauses with the core purpose front-loaded and zero filler. It is appropriately sized, though the ellipsis in the period list is a minor loss of precision rather than padding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be described. However, for a six-parameter mutation tool with no annotations, this is too thin: the relationship between group_id and product_id, the role of the optional approval_id, and the effect of family_shareable are all undisclosed, leaving an agent under-equipped to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for 6 parameters, so the description must compensate. It only documents one parameter (period), and only partially, by listing sample values (ONE_WEEK/ONE_MONTH/.../ONE_YEAR) that are not encoded as a schema enum. The meaning of group_id, product_id, approval_id, and family_shareable is left entirely unexplained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a clear verb+resource ('Create a subscription product'), which is distinguishable from the sibling asc_create_subscription_group that operates on a group rather than a product. It does not name that sibling or otherwise route the agent, but the resource noun is specific enough to identify the tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this tool versus alternatives such as asc_create_subscription_group, asc_ensure_subscription_prices, or asc_finalize_subscription, nor any stated prerequisites. The only context is the inline '(LIVE write)' flag, which is behavioral rather than usage-routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
asc_create_subscription_groupC
Create a subscription group for the app (LIVE write).
| Name | Required | Description | Default |
|---|---|---|---|
| app_id | Yes | ||
| approval_id | No | ||
| reference_name | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. '(LIVE write)' usefully warns that this affects live data, but nothing is said about required permissions, idempotency, whether the change is reversible, or what happens if the group already exists — meaningful gaps for a live mutation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with zero waste and the mutation warning in parentheses. It is efficient, though it is terse to the point of under-specification rather than maximally informative.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be described. However, for a live write with no annotations and three completely undocumented parameters, the description is far too thin — no prerequisites, no side-effect detail, no disambiguation from the many subscription-related siblings.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the description adds no parameter meaning at all. 'For the app' only weakly implies app_id; reference_name and the optional approval_id (which likely relates to an approval-gated live write) are entirely unexplained in both the schema and the description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource ('Create a subscription group') with a scope ('for the app'), and 'group' distinguishes it from the sibling asc_create_subscription and from the localization tools. It stops short of explicitly contrasting itself with asc_add_subscription_group_localization or asc_localize_group, so it is clear but not sibling-differentiating.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The '(LIVE write)' tag signals this is a real mutation rather than a dry run, but there is no guidance on when to call it, what prerequisites (e.g. an existing app record) are needed, or which sibling to use instead. An agent gets no routing information beyond the resource name.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
asc_ensure_subscription_pricesA
Walk ALL of the app's subscriptions; set a price on those that have NONE (clears MISSING_METADATA). Idempotent (those with a price are skipped). Price comes from usd_price_by_product[productId] or default_usd.
| Name | Required | Description | Default |
|---|---|---|---|
| bundle_id | Yes | ||
| approval_id | No | ||
| default_usd | No | ||
| usd_price_by_product | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does reasonably well: it discloses scope (all subscriptions), mutation semantics (sets a price only where none exists), idempotency, the metadata defect it resolves (MISSING_METADATA), and the price-source precedence. It omits permission/approval requirements and any rate or failure behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences, front-loaded with the operation and scope, followed by idempotency and price-source resolution. No redundant or filler content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with no annotations and an output schema (so return values need not be described), the description covers scope, idempotency, and price resolution well but leaves approval_id unexplained and omits any authorization or failure-mode context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate. It meaningfully explains usd_price_by_product (keyed by productId) and default_usd as the per-product vs. fallback price source, but says nothing about approval_id, leaving one of four parameters undocumented across both schema and description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb and resource ('set a price on those that have NONE') and constrains the scope ('ALL of the app's subscriptions'). It is clearly distinguishable in effect, though it does not explicitly name or differentiate from any sibling tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied: the reference to clearing MISSING_METADATA and the idempotency note suggest it is a fix-up/backfill tool safe to re-run. There is no explicit statement of when to prefer this over other setup tools, nor any when-not guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
asc_finalize_submission_requirementsA
Fill submit-blocking app-level fields in one call: Content Rights + Copyright + Age Rating (4+) + Free Price + App Review Contact + App Review Notes (7-item test flow). App Privacy EXCLUDED (not in Apple's API → set via ASC web UI; appDataUsages endpoints all 404). Called before submit, per app. RULE: if copyright is omitted, it's read from config ("copyright" key — set your own via config_set); if contact_email is omitted, support_email is used (NOT apple_id, which is only the developer-account login). If review_notes is omitted, a generic 7-item template is filled with app_name/ai_services (Apple 2.1).
| Name | Required | Description | Default |
|---|---|---|---|
| free | No | ||
| app_name | No | ||
| bundle_id | Yes | ||
| copyright | No | ||
| ai_services | No | ||
| approval_id | No | ||
| contact_last | No | ||
| review_notes | No | ||
| contact_email | No | ||
| contact_first | No | ||
| contact_phone | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It discloses important fallback behavior: copyright read from config, contact_email falling back to support_email, and a generic 7-item review_notes template. However, it does not state auth requirements, whether existing values are overwritten, idempotency, or other side effects of this mutation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with purpose, then exclusions, then omitted-parameter rules. Dense but efficient with no filler, though a few parentheticals are terse. Appropriate for the tool's complexity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained. The description covers purpose, timing, exclusions, and key defaults, but with 11 parameters at 0% schema coverage and no annotations, missing parameter and mutation-safety details leave it short of fully self-contained.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across 11 parameters, so the description must compensate. It groups and names most fields and gives specific fallback rules for copyright, contact_email, and review_notes, but approval_id and the required bundle_id are not explained. Compensation is partial.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Fill') and enumerates the exact submit-blocking fields (Content Rights, Copyright, Age Rating, Free Price, App Review Contact, App Review Notes). Distinguishes itself from generic submit tools by scope ('before submit, per app') and by explicitly excluding App Privacy.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
States when to call it (before submit, per app), what is excluded (App Privacy, must be set via ASC web UI), and dependent fallback configuration (config_set for the copyright key). It names both the alternative path for an excluded field and the source of defaults for omitted parameters.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
asc_finalize_subscriptionA
FINALIZE one subscription (ORDER: localization → availability (all but CHN) → price → intro offer).
The intro offer follows the SPEC: pass the spec product's intro ({"type":"free","duration":"P3D"});
it is created per territory. Offer products: pass nothing (no intro offer). period is ignored (the old
embedded offers were wrong). Prefer store_setup, which does every product idempotently.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | ||
| intro | No | ||
| period | No | ||
| sub_id | Yes | ||
| usd_price | Yes | ||
| approval_id | No | ||
| description | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full load and does disclose non-obvious behavior: the exact step ordering, that intro offers are created per territory, that `period` is silently ignored (legacy bug), and the implication from 'store_setup does every product idempotently' that this tool is not fully idempotent. It still omits permission/approval requirements and re-run safety.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the action and the key ordering, and the sentences earn their place, but the arrow notation, nested parentheses, and inline JSON make it dense and harder to parse quickly than it needs to be.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists so return values need not be explained, and the description covers ordering, the alternative tool, and the trickiest parameters. However, for a 7-parameter mutation with no annotations it leaves gaps around approval_id, auth requirements, and failure/re-run behavior.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It richly documents the difficult parameters (intro format with a concrete example, period being ignored, offer-product null behavior) but leaves sub_id, name, description, usd_price, and especially approval_id entirely unexplained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (FINALIZE) and resource (one subscription), plus the internal step order (localization → availability → price → intro offer). It partially distinguishes itself from the crowded sibling set by naming store_setup as the preferred alternative and calling out how offer products differ.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly routes the agent: 'Prefer store_setup, which does every product idempotently,' and gives the branching rule for spec products vs offer products. Clear when-to-use guidance, though it lacks exclusions for other finalize-adjacent siblings like asc_submit_for_review.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
asc_get_appC
Find the App Store Connect app by bundle id (LIVE).
| Name | Required | Description | Default |
|---|---|---|---|
| bundle_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. 'Find' implies read-only access and '(LIVE)' hints at live data, but the description does not explicitly state read-only safety, authentication requirements, error behavior, or rate limits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no wasted words. The parenthetical '(LIVE)' is compact and relevant.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained. However, with no annotations and no usage guidance, the description is only minimally complete for an agent choosing between this and sibling app tools.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It names the bundle id lookup key but adds no format, example, or constraint beyond the schema property name, so it does not meaningfully enrich parameter semantics.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb 'Find', resource 'App Store Connect app', and lookup key 'bundle id'. It clearly implies a single-app lookup, though it does not explicitly distinguish itself from sibling tools like asc_list_apps.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use, prerequisites, or alternative guidance. The phrase 'by bundle id' implies the lookup context, but the description does not tell the agent when to choose this over asc_list_apps or asc_create_app.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
asc_list_appsA
List the apps in App Store Connect (LIVE; requires issuer_id).
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full disclosure burden. It usefully signals that this hits the live App Store Connect API (not a mock/dry run) and that an issuer_id credential must be configured first, but it says nothing about pagination, rate limits, or what the LIVE designation implies for side effects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single 12-word sentence with the resource named first and the LIVE/auth caveat second; every clause carries information and there is no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be described, and a zero-parameter read tool has little else to specify. The remaining gap is the relationship to asc_get_app and how large result sets are returned.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so the baseline is 4. The mention of issuer_id is an auth prerequisite rather than a call argument, so it does not confuse the empty schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('List the apps in App Store Connect'), so the agent knows exactly what it returns. It does not differentiate itself from the sibling asc_get_app (singular fetch) or explain what 'apps' scopes to, so it stops short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The parenthetical '(LIVE; requires issuer_id)' implies a precondition for calling it, which is more than nothing. However, there is no guidance on when to pick this over asc_get_app or asc_token_check, and no statement of exclusions, so usage is only implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
asc_localize_groupC
Localize the subscription group name in multiple languages.
| Name | Required | Description | Default |
|---|---|---|---|
| group_id | Yes | ||
| approval_id | No | ||
| name_by_locale | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It does not say whether this overwrites or merges existing localizations, what permissions or approval flow is involved (the presence of approval_id strongly implies one), or what happens to locales not supplied. For a mutation tool with zero annotation coverage this is a substantial gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
It is a single, front-loaded sentence with no filler or restatement of the title. The brevity is appropriate structurally, though it borders on under-specification rather than true economy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be described, but the description still omits the approval workflow implied by approval_id, the merge/overwrite semantics needed for a mutation, and any differentiation from sibling localization tools. Incomplete for a 3-parameter mutation with no annotations.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across three parameters, including a nested locale-to-string map. The phrase 'in multiple languages' loosely implies name_by_locale, but group_id is unexplained and approval_id — clearly a meaningful workflow parameter — is not mentioned at all, nor is the expected locale key format.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a clear verb+resource ('Localize the subscription group name') plus scope ('in multiple languages'), so the action is unambiguous. However, it does not distinguish itself from the near-identical sibling asc_add_subscription_group_localization or asc_localize_subscription, which an agent would need to route between.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance, no prerequisites, and no mention of alternatives despite at least two strongly overlapping siblings (asc_add_subscription_group_localization, asc_localize_subscription). The agent is left to guess which localization tool applies.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
asc_localize_subscriptionB
Localize the subscription in MULTIPLE LANGUAGES. items={locale:{name,description}}. IAP-unsupported languages are skipped.
| Name | Required | Description | Default |
|---|---|---|---|
| items | Yes | ||
| sub_id | Yes | ||
| approval_id | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations the description carries the full burden. It usefully discloses that IAP-unsupported languages are skipped during localization, a genuine behavior not derivable from the schema, but says nothing about the approval_id-driven approval flow, permissions, idempotency, or whether existing localizations are overwritten.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short, front-loaded sentences with no filler. The items notation is terse but functional given the nested schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The output schema covers return values, but for a mutation tool with no annotations and an unexplained approval_id parameter the description is thin. It covers the items payload and skip behavior but omits the approval/auth context an agent would need to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It does clarify the nested items shape (locale:{name,description}), which is valuable, but leaves sub_id and the optional approval_id undocumented in both schema and description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb (localize) and resource (subscription) with a scope qualifier (MULTIPLE LANGUAGES). It is clearly distinguishable from the sibling asc_localize_group, which targets groups rather than subscriptions, though it does not name that sibling explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this tool versus the other localization or subscription-sibling tools, and no prerequisites are stated. The only contextual note, that unsupported languages are skipped, is a behavior, not a usage condition.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
asc_sbp_checkB
Small Business Program proof: latest SUBSCRIPTION/SUMMARY sales report (gzip TSV) → US proceeds/price ratio (≈0.85 SBP, ≈0.70 standard). Needs a Finance-role key: asc_finance_key_id + asc_finance_key_filepath (+ asc_vendor_number); 403 → clear error.
| Name | Required | Description | Default |
|---|---|---|---|
| days_back | No | ||
| report_date | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full disclosure burden and does well: it states the data source (latest SUBSCRIPTION/SUMMARY sales report, gzip TSV), the auth requirement (Finance-role key and the specific config parameters), and error behavior (403 → clear error). It does not state rate limits or caching, but auth and failure handling are the traits an agent most needs to invoke this correctly, and both are covered.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The definition is front-loaded with the purpose, then packs auth and error details into compact parentheticals without filler. It is dense and telegraphic but every clause contributes actionable information. A slightly more readable structure would be the only improvement.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values needn't be described, and the description covers purpose, auth, and error handling. The material gap is the two input parameters, which are neither documented in the schema nor the description. For a tool with a real auth dependency and a computed-ratio output, this is adequate but leaves invocation details under-specified.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for both parameters, and it does not. Although 'latest ... sales report' hints that days_back controls the lookback window and report_date selects a specific report, neither parameter is named or explained, leaving the agent to guess how they interact or which takes precedence.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific operation: fetching the latest SUBSCRIPTION/SUMMARY sales report and computing a US proceeds/price ratio to prove Small Business Program status. The verb+resource is concrete, and the ratio thresholds (≈0.85 SBP, ≈0.70 standard) make the intent unambiguous. It does not distinguish itself from siblings, but no sibling shares this function, so the gap is minor.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied ('Small Business Program proof') rather than stated explicitly as a when-to-use, and no alternatives or when-not conditions are named. It does supply the prerequisite (Finance-role key via asc_finance_key_id + asc_finance_key_filepath + asc_vendor_number), which is real guidance. This lands at implied usage with useful prerequisite context, not a full routing statement.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
asc_submit_for_reviewA
SUBMIT the app to App Store review. Always needs out-of-band human approval (appfactory approve <id>), even with approvals off. Without a valid approval_id nothing is submitted.
| Name | Required | Description | Default |
|---|---|---|---|
| app_id | Yes | ||
| approval_id | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden, and it discloses a critical non-obvious trait: human approval is always required even when approvals are disabled. That is genuinely valuable and not derivable from the schema. It stops short of describing what happens after a successful submission (irreversibility, follow-up state), leaving a minor gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences with the action front-loaded, followed by the gating rule and its consequence. Nothing is padded and every sentence adds a distinct, actionable fact.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with no annotations, the description adequately covers the preconditions an agent needs to avoid a failed call, and an output schema exists so return values need not be explained. The only real shortfall is the undocumented app_id parameter.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It partially does — it conveys that 'approval_id' must be valid and comes from `appfactory approve <id>`, giving that parameter real meaning. But 'app_id' remains entirely undocumented in both schema and description, so compensation is incomplete.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific, unambiguous verb+resource: 'SUBMIT the app to App Store review.' An agent can distinguish this from adjacent siblings like asc_finalize_submission_requirements or deliver_metadata without further reading, since the scope (App Store review submission) is stated explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives a strong precondition — a valid approval_id obtained out-of-band via `appfactory approve <id>` — and even states a when-not condition ('Without a valid approval_id nothing is submitted'). It does not, however, compare itself to any alternative tool or explain sequencing relative to finalize/preflight siblings, so it falls short of explicit routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
asc_token_checkA
Check App Store Connect auth through the asc CLI: binary present, which credentials it uses (config.toml asc_* keys as env, else the asc keychain profile) and one live read-only call.
A successful asc apps list --limit 1 proves key id, issuer and .p8 together.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does well: it discloses the exact credential resolution order (config.toml asc_* keys as env, else the asc keychain profile) and explicitly states the live call is read-only, plus how success is proven. It omits failure modes and what a failed check reports, keeping it below 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences with the core action front-loaded and the credential-resolution caveat and success proof following. The parenthetical is dense but earns its place by describing real resolution behavior; nothing is wasted, though the second sentence is somewhat compressed.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be described, and there are no parameters to document. The description covers what is checked and how credentials resolve, which is adequate for a no-arg diagnostic tool; only failure/edge-case behavior is unaddressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, and the schema confirms an empty properties object, so there is nothing for the description to clarify. Baseline 4 applies; no parameter meaning is missing.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a precise verb ('Check') and a specific resource (App Store Connect auth via the asc CLI), then enumerates exactly what is verified: binary presence, credential source resolution, and one live read-only call. An agent can distinguish it from asc_list_apps or the various doctor tools without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The purpose implies when to use it ('check ASC auth'), and the success criterion (`asc apps list --limit 1` proving key id, issuer and .p8) gives a concrete validation context. However, it names no alternatives despite many overlapping siblings (asc_sbp_check, config_doctor, env_doctor, setup_credentials), leaving the routing decision to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
aso_check_nameA
Is the App Store name (globally unique) available — iTunes Search exact-name collision. Call BEFORE asc_create_app.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | ||
| country | No | us |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden, and it does disclose useful behavior: the name is globally unique and the check is an exact-name collision against iTunes Search. It does not cover read-only vs mutating nature, network/auth requirements, or rate-limit behavior, which is a real gap for an unannotated tool, though the output schema covers return values.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two compact sentences, front-loaded with the core question and followed by the mechanism and the prerequisite call. Every clause earns its place with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple two-parameter read check with an output schema, the description is nearly sufficient, and it correctly omits return-value details. It is still thin on parameter meaning and on how the two name-checking siblings relate, which an agent needs to invoke this correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the description never mentions either parameter. The 'country' parameter in particular is left unexplained, and there is a subtle tension between 'globally unique' and a country-scoped iTunes lookup that the description could have resolved.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb ('check available') and resource ('App Store name'), and even discloses the mechanism ('iTunes Search exact-name collision'), so the operation is unambiguous. It does not differentiate itself from the sibling aso_find_available_name, which appears to cover similar ground, so it falls short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives an explicit sequencing instruction ('Call BEFORE asc_create_app'), which tells the agent exactly where this fits in a workflow. However, it offers no when-not guidance and never mentions aso_find_available_name, leaving the choice between the two name-checking siblings to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
aso_competitor_iapA
Competitor's subscription/IAP price ladder (from the App Store product page). Input to the pricing decision.
iTunes lookup does not expose IAPs; the product page embeds them. app_id = iTunes trackId (returned by aso_fetch_competitors). FRAGILE (undocumented): on failure ok:False + prices:None — ABSENCE means 'unknown', NOT 'free'.
| Name | Required | Description | Default |
|---|---|---|---|
| app_id | Yes | ||
| country | No | us |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does so well: it discloses the failure contract (ok:False + prices:None), warns that absence means 'unknown' NOT 'free', and explains why iTunes lookup can't substitute. It stops short of documenting auth or rate-limit behavior, but the fragility disclosure is the critical trait for correct interpretation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the core purpose, then the sourcing rationale, then the critical fragility warning. Tight and purposeful, though the fragmented parenthetical asides make it slightly choppy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values needn't be explained, and the description adds the one thing an agent most needs — the failure/absence semantics — plus the app_id source. Only the country parameter and any auth expectations are left uncovered.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate. It explains app_id precisely ('iTunes trackId returned by aso_fetch_competitors'), but the country parameter (default 'us') is left entirely undocumented in both schema and description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource: retrieves a competitor's subscription/IAP price ladder from the App Store product page. This is clearly distinguishable from siblings like aso_fetch_competitors (which supplies the app_id) and aso_unit_economics (which consumes pricing data).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives implied usage ('Input to the pricing decision') and a useful chaining hint that app_id comes from aso_fetch_competitors. However, it never states explicit when-to-use/when-not conditions or names an alternative tool for the same task.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
aso_completeA
Validate the ASO output + mark the 'aso' stage done (opens the metadata/deliver gate).
| Name | Required | Description | Default |
|---|---|---|---|
| app_dir | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, and it does disclose the key state mutation: marking the 'aso' stage done and unblocking the metadata/deliver gate. It omits whether validation failure is fatal, whether the call is idempotent, and any permission requirements, leaving meaningful gaps for a state-changing tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single tight sentence that front-loads the two actions and appends the downstream effect in parentheses. No filler, every clause carries information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be described, but the near-total absence of annotations and the completely undocumented parameter leave the tool only minimally specified. For a tool that mutates pipeline state, more on failure handling and sequencing relative to aso_run/pipeline_mark would be expected.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There is one parameter, app_dir, with 0% schema description coverage, and the description says nothing about it. The name is largely self-explanatory (the app's root directory), but no format or scope detail is added, so it neither compensates for nor meaningfully worsens the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific compound verb+resource: 'Validate the ASO output + mark the 'aso' stage done'. It also names the downstream consequence ('opens the metadata/deliver gate'), which helps separate it from aso_run and aso_validate_metadata. It stops short of explicitly contrasting those siblings, so not a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The 'opens the metadata/deliver gate' clause implies this is the step that follows ASO generation and precedes metadata/deliver work, giving implied sequencing. However, it never states when to prefer this over aso_validate_metadata or pipeline_mark, nor any preconditions (e.g. ASO output must already exist).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
aso_fetch_competitorsC
Fetch competitor apps via iTunes Search (LIVE, read-only).
| Name | Required | Description | Default |
|---|---|---|---|
| term | Yes | ||
| limit | No | ||
| country | No | us |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It does disclose two useful traits: the call is live (network-dependent, potentially slow or rate-limited) and read-only (safe, no mutation). It omits auth requirements, rate limits, error behavior, and result ordering.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler; the parenthetical efficiently conveys the live/read-only traits. It is terse to the point of under-specification, which caps it below a 5.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained, but for a 3-parameter tool with zero schema coverage and no annotations the description is far too thin — it provides no parameter meaning, no usage routing, and no operational caveats.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the description mentions no parameters at all. Three parameters (term, limit, country) are left entirely undocumented, so an agent gets no guidance on what 'term' should contain or what 'country' accepts.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb ('Fetch'), a specific resource ('competitor apps'), and the data source ('iTunes Search'). It does not differentiate itself from sibling tools like aso_competitor_iap or aso_search_hints, which an agent would have to disambiguate on its own.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The '(LIVE, read-only)' tag hints that this hits live external data rather than a cache, but there is no explicit statement of when to use this tool versus the other ASO siblings, no prerequisites, and no exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
aso_find_available_nameC
Pick the first AVAILABLE App Store name from the candidates (a taken name gets the submit rejected).
| Name | Required | Description | Default |
|---|---|---|---|
| country | No | us | |
| candidates | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It discloses the selection rule (first available candidate) and the consequence of a taken name, but does not say what happens if every candidate is taken, whether it checks the live App Store, what credentials or permissions are needed, or whether the check is cached or rate-limited.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single front-loaded sentence with no wasted words. The availability constraint and submission consequence are stated efficiently, though the extreme brevity leaves gaps that a longer description could close.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with two parameters, no annotations, and a live availability-checking purpose, the description is under-specified. It omits the country parameter, does not explain the return semantics beyond what the output schema may show, and says nothing about the no-available-name case.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for both parameters. It gives meaning to candidates as App Store name options, but completely omits the country parameter, which defaults to 'us' and could materially affect availability results.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: picking the first available App Store name from candidates. It is clear what the tool does, but it does not explicitly contrast itself with the sibling tool aso_check_name, which likely has an overlapping name-checking role.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The parenthetical implies usage by explaining why availability matters: a taken name gets submission rejected. However, there is no explicit when-to-use or when-not-to-use guidance, and no alternative tool is named for cases where none of the candidates are available.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
aso_niche_scoreA
Idea go/no-go score: demand + competitor weakness + saturation penalty + monetization (0-100).
HARD GATE: 0 autocomplete suggestions = REJECT. Competitor metrics from iTunes Search top 50 (country PINNED — userRatingCount is per storefront). Saturation penalty: the score drops if the median competitor is strong. genre: any App Store category (aso.GENRES) — the category's AI density is INFORMATIONAL only (AI is an optional edge, never required nor penalized; not part of the score). verdict: GO(≥60) | MAYBE(≥40) | WEAK | REJECT(median>50k or no demand).
| Name | Required | Description | Default |
|---|---|---|---|
| term | Yes | ||
| genre | No | ||
| country | No | us |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and mostly succeeds: it discloses the hard-gate rule, that competitor metrics come from the iTunes Search top 50, that country is PINNED and userRatingCount is per-storefront, that the saturation penalty grows when the median competitor is strong, and that genre AI density is informational and never scored. Gaps remain around read-only/network/cost behavior and error modes.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The scoring formula is front-loaded in the first sentence, followed by the hard gate and verdict bands. It is dense but every clause conveys a rule; a few telegraphic fragments ('HARD GATE:', 'country PINNED') cost a little readability but waste no space.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained, and the description supplies the scoring model, gating rule, and verdict cutoffs an agent needs to interpret results. For a 3-param tool it is nearly complete, missing only explicit term semantics and any note on failure conditions.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It adds real meaning for genre (any App Store category from aso.GENRES, AI density informational) and country (PINNED, per-storefront rating counts), but the required 'term' parameter is only implied and its format/constraints are never stated.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb+resource: it computes a 0-100 go/no-go score for a keyword/idea niche from demand, competitor weakness, saturation, and monetization. That is far more concrete than the bare name aso_niche_score. It does not, however, name the nearest siblings (idea_evaluate, aso_run) to disambiguate, so an agent must infer the boundary from the formula alone.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied through the HARD GATE ('0 autocomplete suggestions = REJECT') and the verdict thresholds (GO>=60, MAYBE>=40, WEAK, REJECT), which tell the agent what output to expect but not when to call this instead of idea_evaluate or aso_run. No explicit when-not or prerequisite guidance is given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
aso_runC
Start the ASO stage (mandatory step): outputs skeleton + skill instructions.
| Name | Required | Description | Default |
|---|---|---|---|
| app_dir | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. 'Start' implies a state-changing operation that likely writes skeleton files, but the description never states side effects, whether it requires prior setup, or whether it is idempotent. The only behavioral hint is that it is mandatory and emits instructions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single tight sentence with no waste, and the mandatory status plus outputs are front-loaded. It is appropriately sized, though extremely terse for a stage-entry tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The output schema covers return values, so those need not be described. But with no annotations, an undocumented required parameter, and no explanation of side effects or ordering against sibling ASO tools, the definition leaves real gaps for an agent invoking a stage-start operation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There is one required parameter, app_dir, and schema description coverage is 0%, so neither the schema nor the description explains it. The description adds no meaning about what app_dir should point to (repo root, app folder) or its expected format.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a verb ('Start') and a resource ('the ASO stage') and mentions the outputs (skeleton + skill instructions). However, 'start the ASO stage' is vague and does not distinguish this from sibling ASO tools such as aso_scaffold_outputs, aso_validate_metadata, or aso_complete, so an agent cannot easily tell which ASO entry point to pick.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It labels itself a 'mandatory step', which implies it should be run at the beginning of the ASO flow. It names no alternatives and gives no when-not-to-use conditions or prerequisites relative to the many other aso_* siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
aso_scaffold_outputsC
Create the outputs// skeleton (including apple-metadata.md, 32 languages).
| Name | Required | Description | Default |
|---|---|---|---|
| locales | No | ||
| app_name | Yes | ||
| base_dir | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It discloses what gets created but not whether existing files are overwritten, whether the operation is idempotent, whether it requires credentials/permissions, or what happens if the directory already exists — important for a file-creating scaffold.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler. Efficient, though the parenthetical about '32 languages' is terse to the point of ambiguity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists so return values need not be described, but for a file-writing scaffold with zero annotation coverage and 0% parameter documentation, the description leaves the agent without enough to call it safely or confidently.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and none of app_name, base_dir, or locales is explained in prose. The phrase '32 languages' faintly implies the locale default, but the required base_dir and app_name are left entirely to the schema's bare types.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Create) and resource (the outputs/<App>/ skeleton) and names concrete artifacts it produces (apple-metadata.md, 32 languages). Clear enough for an agent to know what it does, though it never distinguishes itself from plausible siblings like app_scaffold or design_screens_skeleton.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance, no prerequisites, and no mention of alternatives among the many scaffolding/metadata siblings. Usage is only weakly implied by the word 'skeleton'.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
aso_search_hintsA
App Store autocomplete = REAL search demand (free, no auth). IDEA-STAGE HARD GATE.
Returns suggestions in Apple's own popularity order. 0 suggestions for a real term = no demand → REJECT the idea. expand=True appends a–z to the seed and collects the keyword universe (ASO keyword research). An endpoint error returns ok:False (NOT an empty list) — a broken endpoint must not be read as 'no competitors'.
| Name | Required | Description | Default |
|---|---|---|---|
| term | Yes | ||
| expand | No | ||
| country | No | us |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does well: it discloses no-auth/free access, that results come in Apple's own popularity order, that expand=True appends a–z to build a keyword universe, and critically that an endpoint error returns ok:False rather than an empty list. Rate limits and the effect of the country parameter are not covered, keeping it below 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Dense and front-loaded: the core identity ('App Store autocomplete = REAL search demand') leads, followed by role, then behavior and error semantics. Every sentence adds information with no padding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists so return values need not be explained, and the description covers the gate logic and error-handling caveat well. The remaining gap is the country parameter's impact on results, which an agent would need to know for non-US stores.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It explains expand=True thoroughly (a–z expansion, keyword universe collection) and term is self-evident, but the country parameter and its effect on locale-specific results are entirely undocumented in both schema and description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('App Store autocomplete') and its semantic meaning ('REAL search demand'), plus its role as an 'IDEA-STAGE HARD GATE'. An agent can distinguish this from sibling ASO tools like aso_find_available_name or aso_niche_score without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a clear decision rule for when to use it (idea-stage validation) and how to interpret results (0 suggestions → REJECT the idea). It does not explicitly name competing sibling tools or exclusions, so it stops short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
aso_top_grossingA
Top-grossing apps (revenue proxy) — proven-idea hunting. genre: any App Store category name from aso.GENRES (books, business, developer_tools, education, entertainment, finance, food_drink, games, graphics_design, health, lifestyle, medical, music, navigation, news, photo_video, productivity, reference, shopping, social, sports, travel, utilities, weather) or its genre id; None = overall chart. An ignored genre filter returns ok:False, never the overall chart.
| Name | Required | Description | Default |
|---|---|---|---|
| genre | No | ||
| limit | No | ||
| country | No | us |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It does disclose a valuable failure mode (invalid genre returns ok:False, never the overall chart), which goes beyond the schema. But it says nothing about auth requirements, rate limits, pagination, or scoping, leaving meaningful behavioral gaps.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core purpose before the parameter detail, and the long genre enumeration is justified because the schema has no enum. Slightly verbose but every segment earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained here. Combined with the genre documentation and the explicit filter-failure rule, the agent has enough to call it correctly; only the undocumented limit/country params keep it from a 5.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate for all three params. It documents genre thoroughly (category name from aso.GENRES or its id, None = overall chart), but limit and country are left entirely unexplained. It covers one of three parameters well.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb-and-resource plus a use framing: 'Top-grossing apps (revenue proxy) — proven-idea hunting.' An agent immediately knows this returns a revenue-ranked chart. It does not name or contrast a sibling (e.g., aso_fetch_competitors, aso_niche_score), so it stops short of 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'proven-idea hunting' implies a use context, and the description warns that an ignored genre filter fails rather than silently returning the overall chart. However, there is no explicit when-to-use / when-not / alternative-tool routing against the many ASO siblings, so guidance remains implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
aso_unit_economicsB
Unit economics: credit count × AI cost vs subscription revenue → margin, break-even, warnings.
Grounds the price + credit decision in data at the idea/scaffold stage (the yearly_credits/weekly_credits passed to app_scaffold are validated here). apple_cut: Small Business 15% (default), otherwise 0.30.
| Name | Required | Description | Default |
|---|---|---|---|
| apple_cut | No | ||
| weekly_price | No | ||
| yearly_price | No | ||
| weekly_credits | No | ||
| ai_cost_per_credit | No | ||
| yearly_monthly_credits | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It does disclose what the tool produces (margin, break-even, warnings), the apple_cut default rule (0.15 Small Business, else 0.30), and its advisory role in validating scaffold credits. It does not state whether it is purely read-only, whether it persists anything, or whether it depends on project state, which is a gap given zero annotation coverage.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The computation is front-loaded in a compact formula line, followed by one sentence of stage/relationship context and a default-value note. Little waste; the parenthetical credit note is slightly dense but earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be described, which the description correctly omits. But for a 6-parameter, zero-schema-coverage, annotation-free tool the definition is only partially complete: parameter meanings and side-effect/safety behavior are thin, and sibling differentiation from pricing_unit_economics is absent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across 6 parameters, so the description must compensate. It explains apple_cut's default semantics and references the weekly/yearly credit inputs, but leaves weekly_price, yearly_price, ai_cost_per_credit, and yearly_monthly_credits entirely unexplained — no units, granularity, or accepted ranges.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a concrete computation: 'credit count × AI cost vs subscription revenue → margin, break-even, warnings.' An agent can tell this is a unit-economics calculator rather than a mutating setup tool. However, it never distinguishes itself from the near-identical sibling 'pricing_unit_economics', which is the most likely source of misselection.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It anchors usage to a phase: 'Grounds the price + credit decision in data at the idea/scaffold stage,' and ties it to app_scaffold via the credits it validates. That is implied when-to-use context, but there is no explicit when-not guidance and no routing away from the sibling pricing_unit_economics.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
aso_validate_metadataC
Validate outputs//02-metadata/apple-metadata.md.
| Name | Required | Description | Default |
|---|---|---|---|
| app_name | Yes | ||
| base_dir | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the entire behavioral burden. It does not say what is validated, whether a failure is fatal, whether it mutates anything, or what errors look like. An output schema exists, which covers return structure, but the validation behavior itself is undisclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no padding, which is good. But it is terse to the point of under-specification given everything it leaves unaddressed.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema the return values need no prose, but two undocumented required parameters, zero annotations, and a crowded sibling set of validators make this far too thin. An agent lacks enough to select or invoke it confidently.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and neither of the two required parameters (app_name, base_dir) is documented anywhere. The path fragment "outputs/<App>" loosely hints that app_name slots into a directory, but base_dir's role and expected format are entirely unexplained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ("Validate") and a concrete resource (the apple-metadata.md file), which is more than a tautology. However, it does nothing to distinguish itself from sibling validators like metadata_check, metadata_listing_check, or pipeline_validate, so an agent cannot tell which validator applies.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance, no exclusions, and no mention of the many sibling validation tools. The agent must guess at the workflow position and prerequisites on its own.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
backend_deployA
Deploy the spec's Supabase backend (LIVE unless dry_run): migrations (tracked), secrets (AI_MODEL, AI_FALLBACK_MODEL, caps, REQUIRE_CONSENT, RC_PROJECT_ID from app outputs, RC_SECRET_KEY
FAL_KEY from config — reported by name only), auth (anonymous + Sign in with Apple + manual linking), legal build, and the function set (analyze, usage-status, delete-account, legal, rc-webhook). dry_run=true (default) returns the plan without any call; a live deploy needs human approval.
| Name | Required | Description | Default |
|---|---|---|---|
| app_dir | Yes | ||
| dry_run | No | ||
| approval_id | No | ||
| project_ref | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden and does well: it discloses that dry_run makes no calls, that live deployment requires human approval, that migrations are tracked, and that secret values are reported by name only rather than echoed. It stops short of covering idempotency, rollback, or permission requirements for live runs.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the action and resource, and the parenthetical lists are dense but each item earns its place by telling the agent what will actually be provisioned. The single run-on sentence is information-heavy but slightly unwieldy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a high-complexity deployment tool with 4 params, an output schema present, and no annotations, the description covers scope, safety gating, and dry-run behavior well. The unexplained app_dir/project_ref semantics are the main residual gap, but return values are correctly left to the output schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for 4 parameters, so the description must compensate. It explains dry_run's semantics and implies the approval/approval_id mechanism, but app_dir and project_ref receive no explanation in either place, leaving half the surface documented only by name.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb (Deploy) plus a precise resource (the spec's Supabase backend), with an explicit enumeration of what is deployed: migrations, secrets, auth, legal build, and a named function set. An agent can distinguish this from siblings like supabase_create_project or supabase_run_sql without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Clear context: dry_run=true is the default and returns only a plan, while a live deploy requires human approval via the approval path. It does not explicitly name alternative tools or state when-not to use this versus supabase_* setup tools, so it falls short of the explicit-alternatives bar.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
backend_renderA
Make the scaffolded backend follow app.spec.json (offline, idempotent): prune supabase/ to the monetization mode, render functions/_shared/app.gen.ts + config.toml (Apple client id) + the CONTRACT.md quota block, resolve/fill the legal sources, and sync listing.json locales/legal URLs.
| Name | Required | Description | Default |
|---|---|---|---|
| app_dir | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden and does well: 'offline' tells the agent no network or credentials are needed, and 'idempotent' signals it is safe to re-run. 'Prune supabase/' also discloses that files may be removed, which is important destructive context. It stops short of saying what happens on failure or whether existing files are overwritten.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose is front-loaded, then a compact colon-delimited list of side effects. It is one dense sentence, but every clause names a real artifact or trait, so there is little waste.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be described, and the description covers offline/idempotent behavior plus the concrete mutation set. The main gap is placement in the broader setup pipeline relative to siblings like app_scaffold and backend_deploy.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% and the single app_dir parameter is never named or described. The file paths and the app.spec.json reference imply app_dir is the project root containing those artifacts, which gives some implicit semantic value, but the description never compensates fully for the undocumented parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific action (rendering the backend to match app.spec.json) and enumerates the concrete artifacts produced: pruned supabase/, functions/_shared/app.gen.ts, config.toml, the CONTRACT.md quota block, legal sources, and listing.json URLs. This is far more specific than a sibling like backend_deploy, though it never explicitly names the alternative it contrasts with.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Calling it a step that makes 'the scaffolded backend' follow the spec implies it runs after app_scaffold, so prerequisite order is inferable. However, there is no explicit when/when-not guidance and no mention of how it relates to backend_deploy, app_inject_config, or app_sync_spec, so the agent must infer sequencing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
build_archiveC
xcodebuild archive (generic iOS) → .xcarchive.
| Name | Required | Description | Default |
|---|---|---|---|
| scheme | Yes | ||
| project | Yes | ||
| archive_path | Yes | ||
| configuration | No | Release |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It does add the useful detail that it emits a .xcarchive artifact, but gives no information about required signing/credentials, execution time, where the archive is written, or whether the path directory is created. For a build mutation with zero annotation coverage this is a large gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
It is short but this is under-specification rather than true conciseness. The single fragment is front-loaded, yet it omits information that would earn its place (prerequisites, artifact location). A one-line arrow notation is too sparse for a four-parameter build tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained. But given four undocumented parameters, no annotations, and a complex build operation, the description leaves out signing requirements, project/scheme expectations, and how the result feeds downstream shipping steps. Incomplete for the tool's complexity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and there are four parameters, so the description must compensate but does not. 'archive_path' is only obliquely implied by '→ .xcarchive'; project, scheme, and the Release-defaulted configuration are entirely unaddressed. No format hints (workspace vs project) or path conventions are given.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a concrete verb+resource (xcodebuild archive) and identifies the output artifact (.xcarchive), so the core action is inferable. However, it does not distinguish this from siblings like build_export_ipa or build_for_sim, and the '(generic iOS)' fragment is the only scoping hint. It states what happens but not how it differs from adjacent build tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance whatsoever on when to use this tool versus the many other build/export siblings (build_for_sim, build_export_ipa, build_test). No prerequisites, no mention of needing a configured scheme or signing setup. The agent must guess where archive fits in the pipeline.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
build_boot_simC
Boot the simulator and open Simulator.app.
| Name | Required | Description | Default |
|---|---|---|---|
| udid | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It discloses the UI side effect of opening Simulator.app, which is useful, but says nothing about state mutation (booting changes device state), whether an already-booted device errors, or any auth/permission needs.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with zero filler. It is efficient, though its brevity contributes to the guidance gaps rather than being purely economical.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained. However, for a state-changing tool with no annotations and a fully undocumented parameter, the description is thinner than ideal, leaving the udid semantics and prerequisites unaddressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% for the single required parameter, so the description must compensate and it does not. It never mentions 'udid', its format, or where to obtain it; the only hint is the parameter name itself.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Boot the simulator') and adds the side effect of opening Simulator.app. This distinguishes it from build_for_sim and build_list_simulators, though it never names a sibling explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance, no alternatives, and no prerequisites. It does not mention that the udid must come from build_list_simulators or how this relates to build_for_sim, which are the natural adjacent tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
build_export_ipaC
xcodebuild -exportArchive → .ipa.
| Name | Required | Description | Default |
|---|---|---|---|
| export_dir | Yes | ||
| archive_path | Yes | ||
| export_options_plist | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden, and it discloses nothing: not that it shells out to xcodebuild, not the host/toolchain requirements, not that it writes files into export_dir, not what happens on failure or whether output is overwritten. A single command line is far short of adequate for a 3-required-parameter mutating build step.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
It is undeniably front-loaded with zero filler, but the extreme brevity here reflects under-specification rather than disciplined conciseness. There is no wasted sentence simply because there are almost no sentences.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be described, but for a tool with three required, undocumented parameters and no annotations, the description should at minimum explain the inputs and preconditions. It leaves an agent unable to call the tool correctly without outside knowledge.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for three required parameters. The '-exportArchive' flag faintly implies archive_path and export_options_plist, and '.ipa' faintly implies export_dir, but no parameter is named or explained, and key facts (plist is a file path, expected plist keys, directory must exist) are absent from both schema and description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description maps the tool to a concrete command ('xcodebuild -exportArchive') and its output ('.ipa'), so an informed agent can infer it exports an existing Xcode archive into a signed .ipa. However, it never states this in plain terms and offers no differentiation from the sibling 'build_archive', which an agent could easily confuse with it.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this versus 'build_archive' or any other build tool, no stated prerequisites (an archive must already exist, xcodebuild/macOS required), and no mention of ordering in a pipeline. The agent must guess entirely.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
build_for_simC
Build the project for the simulator (verification). project = .xcodeproj/.xcworkspace path.
| Name | Required | Description | Default |
|---|---|---|---|
| scheme | Yes | ||
| project | Yes | ||
| device_name | No | iPhone 16 |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It identifies the operation as a simulator build but does not disclose side effects, failure behavior, permission needs, or whether it installs/launches anything. Only the project path format is clarified.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Very concise and front-loaded: the core action appears first, followed by a brief parameter hint. The second sentence is slightly fragmentary but still efficient and useful.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no annotations, 0% schema description coverage, and two required parameters, the definition is missing key context about scheme semantics, device selection, and the build workflow. An output schema exists, so return values need not be explained, but the invocation context remains thin.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for all parameters. It explains the 'project' parameter as a .xcodeproj/.xcworkspace path, but says nothing about the required 'scheme' parameter or the optional 'device_name' default.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific action (build), target (project), and platform context (simulator verification). Clear enough for an agent to understand what the tool does, but it does not explicitly differentiate itself from sibling build tools such as build_test or build_archive.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The parenthetical '(verification)' hints that this is for simulator build verification, but there is no explicit when-to-use guidance, no prerequisites, and no mention of alternatives like build_test or build_archive. Usage must be inferred.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
build_list_simulatorsA
List the available iOS simulators (iPhones are surfaced first).
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, and it only discloses one behavioral trait: iPhone simulators are ordered first. For a read-only enumeration that is mild but real added context; permissions, side effects, and whether only booted/installed devices appear are unaddressed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One sentence, front-loaded with the core action, and the parenthetical ordering note is the only supplementary content. Nothing extraneous.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained, and a zero-parameter read-only tool requires little else. The only remaining gap is the unstated relationship to build_boot_sim/build_for_sim, which would help the agent sequence calls.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so there is nothing for the description to disambiguate. Baseline 4 applies; no schema semantics are missing.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: list iOS simulators, with the extra detail that iPhones are surfaced first. It is clearly distinguishable from siblings like build_boot_sim or build_for_sim by name, but the description never explicitly contrasts itself with those tools, so it stops short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied rather than stated: an agent can infer this is a discovery step before build_boot_sim, but the description never says when to call it or name the follow-up tool. No exclusions or alternatives are offered.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
build_screenshotC
Capture a screenshot from the booted simulator → out_path (PNG).
| Name | Required | Description | Default |
|---|---|---|---|
| udid | Yes | ||
| out_path | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It does not disclose whether an existing file at out_path is overwritten, what permissions are needed, or what happens if no simulator is booted. For a file-writing tool with zero annotation coverage this is a real gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, front-loaded sentence with zero wasted words and the output target placed last for emphasis. It is arguably too terse, but nothing is padded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be described. However, with no annotations and 0% schema coverage, the definition leaves overwrite behavior and UDID semantics unexplained for a two-required-parameter tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for two undocumented parameters. It clarifies only that out_path is a PNG destination and implies udid identifies the simulator; it never explains the UDID format or confirms which simulator to target.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Capture) and resource (screenshot), plus the source (booted simulator) and output (out_path, PNG). It is clearer than most siblings, though it never distinguishes itself from the nearby screenshot_capture sibling.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'booted simulator' implies a prerequisite (a simulator must already be running, likely via build_boot_sim), but this is left implicit. There is no explicit when-to-use guidance and no mention of the screenshot_capture alternative.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
build_testC
xcodebuild test (on the simulator).
| Name | Required | Description | Default |
|---|---|---|---|
| scheme | Yes | ||
| project | Yes | ||
| device_name | No | iPhone 16 |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden, and it delivers almost nothing: no indication that test execution is long-running, whether a simulator must already be booted (a sibling `build_boot_sim` exists), how failures are reported, or whether artifacts/logs are produced. The only disclosed trait is the simulator environment.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
It is a single front-loaded phrase with zero filler, so there is no verbosity problem. But it is a fragment rather than a structured sentence, and its brevity comes at the cost of under-specification rather than economy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained, but for a tool that spawns a test run the description omits everything else an agent needs: simulator prerequisites, expected runtime, and the meaning of the required project/scheme arguments. Given no annotations and 0% param coverage, the description is well below sufficient.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for all three parameters with no annotations to compensate. The description never clarifies whether `project` is an .xcodeproj or .xcworkspace path or what form `scheme` takes; the phrase 'on the simulator' only loosely gestures at `device_name`. This is far short of compensating for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific command and environment: it runs `xcodebuild test` against a simulator. That implicitly separates it from `build_for_sim` (build only), `maestro_test` (different runner), and `build_archive`, even though no sibling is named explicitly. It is a clear verb+resource, just extremely terse.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance at all: nothing says when to pick this over `maestro_test`, `build_for_sim`, or `cpp_build_all`, and no prerequisites (e.g. simulator boot, scheme/project format) are given. The single parenthetical is a runtime detail, not usage guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
build_xcode_versionA
Return the Xcode version (the first validation tool of the build_* group).
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. 'Return the Xcode version' implies a harmless read with no side effects, but it does not disclose failure behaviour (e.g. what happens when Xcode is not installed) or whether it is read-only. The 'validation' framing adds some context about intent, keeping this above a 2.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single short sentence with no waste, front-loading the action and resource. Nothing could be removed without losing information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so the return value need not be explained, and there are no parameters to document. For a trivial version query the description is nearly sufficient, with only error/failure semantics left unaddressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so there is no parameter syntax for the description to explain. The baseline of 4 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb+resource: 'Return the Xcode version.' The parenthetical 'the first validation tool of the build_* group' positions it within a family of tools, which aids recognition, though it does not explicitly name a sibling it should be chosen over.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Calling it 'the first validation tool of the build_* group' implies it is a preliminary preflight check to run before other build_* tools, but there is no explicit when-to-use, when-not-to-use, or named alternative. Usage is inferable rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
config_doctorA
Report config status: which keys are present/missing (secrets are masked).
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It usefully discloses that secrets are masked (a real safety-relevant trait) and that output is a presence/absence report, but says nothing about permissions, whether it reads repo or runtime config, or whether it mutates anything.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single sentence, front-loaded with the action and resource, with the masking caveat appended where it matters. Nothing is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so the description needn't enumerate return values, and it still sketches the report's shape (present/missing keys, masked secrets). For a zero-parameter diagnostic this is close to complete, though a hint about which config source it inspects would help.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so per the rubric the baseline is 4; there is nothing for the description to disambiguate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb ('Report') plus resource ('config status') and concrete output content (present/missing keys). It is clear on its own, but it never distinguishes itself from related siblings like config_set, env_doctor, or setup_status.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The diagnostic framing implies you run this to check config health, but there is no explicit when-to-use or when-not-to-use guidance, and no routing to alternatives such as env_doctor for environment issues or config_set for fixing the gaps it finds.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
config_setA
Write NON-secret settings to ~/.appfactory/config.toml (0600). Fields left empty are unchanged.
Secrets (passwords, sessions, tokens, API keys, phone) are refused here so they never pass through the model: the human enters them on the page opened by setup_credentials().
| Name | Required | Description | Default |
|---|---|---|---|
| team_id | No | ||
| apple_id | No | ||
| copyright | No | ||
| asc_key_id | No | ||
| asc_issuer_id | No | ||
| support_email | No | ||
| asc_key_filepath | No | ||
| legal_controller | No | ||
| asc_vendor_number | No | ||
| asc_finance_key_id | No | ||
| asc_finance_key_filepath | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does well: it discloses patch semantics ('Fields left empty are unchanged'), the file written and its 0600 mode, and a security constraint (secrets never pass through the model). It does not mention validation, error behavior, or that changes are immediate/persistent beyond the file write.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, front-loaded with the action and target file, with the security caveat following. No filler; every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained, and the behavioral essentials are present. The gap is parameter meaning for an 11-field tool with zero schema descriptions, which keeps it from being fully complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
11 parameters with 0% schema description coverage, and the description documents none of them individually (team_id, apple_id, asc_key_id, etc. are never explained). The only parameter-level meaning added is the general 'empty = unchanged' rule, which is insufficient to compensate for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource: 'Write NON-secret settings to ~/.appfactory/config.toml (0600)'. It explicitly scopes itself against the secret-handling sibling by naming setup_credentials(), so an agent can distinguish it without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a clear exclusion ('Secrets ... are refused here') and names the correct alternative (setup_credentials()), which is genuine when-not/when-to guidance. It does not, however, position itself against the other setup_* siblings, so routing among non-secret config tools is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
cpp_build_allB
Create one CPP per Search Ads theme (create→version→localization) + return the deep-link URLs. themes=[{name, locale?}] from the theme/keyword clusters in the ASO output. The URLs feed Ad Ops. LIVE writes: human approval.
| Name | Required | Description | Default |
|---|---|---|---|
| themes | Yes | ||
| bundle_id | Yes | ||
| approval_id | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the safety burden, and 'LIVE writes: human approval' is genuinely important disclosure about reversibility and gating. However, it does not explain what approval_id must contain, what happens if it is omitted, whether reruns duplicate CPPs, or failure behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Very compact and front-loaded: the core action and the returned URLs come first, context and the approval constraint follow. The telegraphic fragment style is efficient though slightly clipped.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return value detail is not required, and the key constraint (live write + approval) is present. Still missing the approval_id contract and any note on idempotency or partial failure, which matters for a multi-step live write tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate, and it partially does by documenting the themes element shape as {name, locale?} where the schema only says 'array of objects'. bundle_id and especially approval_id (the human-approval token) remain unexplained, leaving a real gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Create one CPP per Search Ads theme') and even sketches the internal sequence (create→version→localization) plus the side output (deep-link URLs). It does not name a sibling for contrast, so it falls short of a 5, but the operation is unmistakable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives the input-context hint that themes come from 'the theme/keyword clusters in the ASO output', which implies when to reach for this tool. It stops short of any when-not guidance or explicit alternatives (e.g. aso_scaffold_outputs), so usage is implied rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
deliver_metadataC
Upload metadata to the ASC draft — DIRECT API (bypasses the fastlane 'No data' bug). ASO-gated.
| Name | Required | Description | Default |
|---|---|---|---|
| app_dir | Yes | ||
| bundle_id | Yes | ||
| approval_id | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It does disclose two useful traits — that it uses a direct API path and bypasses a known fastlane 'No data' bug — and hints at gating ('ASO-gated'), but it says nothing about permissions, what gets overwritten in the draft, failure modes, or reversibility.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Extremely terse and front-loaded — the core action leads and the parenthetical rationale follows. It earns most of its words, though 'ASO-gated' is a compressed term that adds ambiguity rather than clarity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be documented. But for a write/gated operation with three undocumented parameters and no annotations, the description omits prerequisite conditions, parameter meaning, and error/side-effect behavior — substantial gaps remain.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the description explains none of the three parameters. 'approval_id' in particular is an opaque, non-obvious identifier whose role (which approval, issued where) is never clarified, and app_dir vs bundle_id precedence is left ambiguous. The description does not compensate for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Upload metadata to the ASC draft'), which is clear enough to distinguish from read-oriented siblings like metadata_check or metadata_export. However, it does not name or contrast with the many nearby metadata tools (metadata_render_listing, metadata_listing_check, aso_validate_metadata), and the trailing 'ASO-gated' is a cryptic qualifier rather than a clarifying one.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No explicit when-to-use or when-not-to-use guidance. 'ASO-gated' hints at a prerequisite stage/approval but never states the condition or names an alternative tool. The agent must infer from the presence of an 'approval_id' parameter that some approval workflow precedes invocation.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
deliver_screenshotsA
Upload screenshots to the ASC draft — checksum-based INCREMENTAL sync (NOT fastlane). Skips when local MD5 == ASC sourceFileChecksum; uploads only missing/changed ones → fast + reliable (fastlane was randomly dropping images). Gated on screenshots: run screenshot_build_all first.
| Name | Required | Description | Default |
|---|---|---|---|
| app_dir | Yes | ||
| bundle_id | Yes | ||
| approval_id | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden and does well: it discloses the checksum-based incremental strategy, the skip condition (local MD5 == ASC sourceFileChecksum), and what gets uploaded (missing/changed only). It lacks details on permissions, rate limits, or failure handling, but the core sync behavior is well explained.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The core behavior and prerequisite are front-loaded, but the parenthetical about fastlane randomly dropping images is marketing-style justification that does not help invocation. Trimming that clause would make the description tighter without losing actionable information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be described, and the behavior and prerequisite are covered. However, with zero annotation coverage and zero schema description coverage, the omission of any parameter explanation and lack of auth/permission context leave gaps for a mutation tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across three parameters, so the description must compensate for parameter meaning but does not. It never explains what app_dir, bundle_id, or approval_id expect, leaving the agent to infer their formats and roles.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Upload screenshots to the ASC draft') and specifies the sync mechanism (checksum-based incremental). It clearly distinguishes its behavior from fastlane, but does not differentiate itself from close siblings like screenshot_sync or deliver_screenshots_audit, leaving some ambiguity about which upload path to choose.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly names a prerequisite ('run screenshot_build_all first') and implies when not to use fastlane. It gives clear context for invocation but stops short of naming alternative tools or exclusions among siblings such as screenshot_sync.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
deliver_screenshots_auditA
Read-only screenshot readiness audit: per locale and display type (iPhone 6.9" = APP_IPHONE_67, Watch = APP_WATCH_ULTRA) the ASC set holds exactly the local fastlane/screenshots files, in order, all COMPLETE, each with the local file's MD5; lists the preview sets. Run after deliver_screenshots and before submission. No writes.
| Name | Required | Description | Default |
|---|---|---|---|
| app_dir | Yes | ||
| bundle_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It discloses the read-only nature ('No writes'), the comparison semantics (exactly the local files, in order, all COMPLETE, matching MD5), and that it lists preview sets. It does not detail auth needs or what happens on mismatch, but the behavioral profile is well described.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core purpose, but the middle clause is densely packed and requires parsing around escaped quotes and parenthetical mappings. Three sentences is appropriate, yet the second sentence is information-dense and not optimally readable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained. The description covers the comparison contract, sequencing, and read-only guarantee for a 2-param tool. Minor gaps: no note on failure behavior when the audit finds mismatches.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for two required params (app_dir, bundle_id), so the description doesn't compensate. However, both parameter names are reasonably self-evident and the description's framing implies app_dir is the project directory whose local screenshots are compared. Baseline 3 given the shallow schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource: 'Read-only screenshot readiness audit' of the ASC screenshot set per locale/display type. It names exact display-type mappings (APP_IPHONE_67, APP_WATCH_ULTRA) and describes what it compares, distinguishing it clearly from deliver_screenshots (which writes) and screenshot_sync.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit sequencing: 'Run after deliver_screenshots and before submission.' This tells the agent precisely where this tool sits in the pipeline relative to a named sibling, leaving no ambiguity about when to invoke it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
deliver_subscription_review_screenshotsA
Upload the App Review screenshot of every subscription from store/review-screenshots/.png (hard paywall with real sandbox prices for the default offering, the offer paywall for offer products). Without it a subscription stays MISSING_METADATA. Idempotent by MD5. Human approval.
| Name | Required | Description | Default |
|---|---|---|---|
| app_dir | Yes | ||
| bundle_id | Yes | ||
| approval_id | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does well: it discloses idempotency ('Idempotent by MD5'), the human-approval gate, the MISSING_METADATA blocking state, and the paywall-vs-offer variant logic. It omits permission/auth details and failure modes, so not a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Dense but front-loaded: the action and source path come first, then the blocking consequence, idempotency and approval note. Every sentence carries information, though the paywall parenthetical is slightly heavy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists so return values need no explanation, and the description covers the path convention, idempotency and blocking state for a mutation tool. The gap is parameter meaning, which is left entirely to the undescribed schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate for all three parameters. It only obliquely touches approval via 'Human approval' (hinting approval_id) and never defines app_dir or bundle_id formats, leaving two required parameters completely unexplained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb (Upload) plus precise resource (App Review screenshot of every subscription) and the exact source path convention. An agent can distinguish it from deliver_screenshots and screenshot_capture without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
States the consequence of not running it ('a subscription stays MISSING_METADATA'), which tells the agent when this step is required. It does not name explicit alternatives or exclusions, but the context is clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
design_export_pngA
Rasterize a Claude Design board with local headless Chrome (not Claude in Chrome). serve_url comes from mcp__claude-design__render_preview and is used once, never stored. Icon: B01-AppIcon, 1024×1024 → design/icon.png. Store layout check: an ST0N board at 440×956, scale=3 → 1320×2868.
| Name | Required | Description | Default |
|---|---|---|---|
| scale | No | ||
| width | Yes | ||
| height | Yes | ||
| out_path | Yes | ||
| serve_url | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries full behavioral burden and does disclose meaningful traits: rasterization uses local headless Chrome rather than Claude in Chrome, and serve_url is single-use and never persisted. However, it says nothing about failure modes, whether out_path is overwritten, permissions, or rendering side effects, leaving notable gaps for a file-writing tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the core action and mechanism, then supplies concrete examples. Four compact sentences with no filler, though the mixed prose-plus-example formatting is slightly uneven rather than tight.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Output schema exists so return values need no explanation, and the description covers the confusing inputs and the rendering mechanism. But with five parameters at 0% schema coverage and no annotations, an agent still lacks documented semantics for width/height and any guidance on output overwrite or error behavior.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, and it does for the two most ambiguous parameters: serve_url's origin/lifetime and scale's multiplicative behavior (440×956 at scale=3 → 1320×2868). out_path is illustrated via 'design/icon.png'. Only width/height are left merely implied by the example rather than stated.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Rasterize') and resource ('a Claude Design board') and explicitly disambiguates the mechanism from a similarly-named one ('with local headless Chrome (not Claude in Chrome)'). An agent can tell this apart from siblings like screenshot_capture or design_generate without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Concrete usage context is given: serve_url originates from mcp__claude-design__render_preview, implying this tool runs after a render step, and two example invocations (app icon, store layout check) show intended scenarios. No explicit when-not-to-use or named alternative tool is provided, so this is not a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
design_generateA
Prepare the app's Claude Design project (the factory's ONLY design source) from the design brief, app.spec.json and design/screens.json. Writes STRUCTURAL scaffolds to design/scaffold/ (every screen incl. paywall/offer, the B01-AppIcon slot, one ST board per brief screenshot board) and into design/project/ the mascot motion boards, canvas.json, store_layout.json and upload_plan.json (plan.authored = boards Claude must author in Claude Design from the brief). Refuses without valid research/brief or when screens/tokens are not derived from the brief. Returns the claude-design MCP calls; this tool never uploads.
| Name | Required | Description | Default |
|---|---|---|---|
| app_dir | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does well: it discloses side effects (writes scaffolds and project JSON to disk), refusal conditions, and a hard boundary ('this tool never uploads'). It omits idempotency/overwrite behavior for existing files in design/scaffold/ and design/project/, which matters for a file-writing tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action and follows with the concrete outputs, so the agent gets the essentials in the first clause. The long parenthetical chains are dense and slow to parse, but every clause carries real information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be restated, and the description still adds the useful note that the return is claude-design MCP calls rather than an upload. Given the tool's complexity, the description covers inputs, outputs, and refusal conditions adequately, with only the app_dir contract left unaddressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the single parameter app_dir is never explained in the description (no format, whether it is relative/absolute, or what it must contain). For a tool whose behavior depends heavily on directory layout, the description should compensate and does not.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb ('Prepare') plus a sharply bounded resource (the Claude Design project from brief/app.spec.json/design/screens.json), with the exact artifacts enumerated (design/scaffold/, design/project/canvas.json, store_layout.json, upload_plan.json). An agent can distinguish this from siblings like design_screens_skeleton or design_research_collect without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Preconditions are stated clearly ('Refuses without valid research/brief or when screens/tokens are not derived from the brief'), which tells the agent when this tool will succeed. It does not name the sibling alternatives (e.g., design_screens_skeleton, design_research_brief_template) or state which stage precedes it, so routing between near-neighbors is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
design_record_uploadA
Record a finished Claude Design upload: SHA-256 of every file in upload_plan.json → design/project/claude_design.json. open_url = the claude.ai/design link from render_preview (never the serve_url). The design, icon and screenshots gates fail until this receipt exists and matches the files.
| Name | Required | Description | Default |
|---|---|---|---|
| app_dir | Yes | ||
| open_url | Yes | ||
| project_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden and does so well: it states that a receipt file is written, that SHA-256 hashes of every file in upload_plan.json are recorded, and that the open_url must be the claude.ai/design link from render_preview rather than serve_url. It also explains the downstream gate dependency, which is critical operational context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded, using three sentences to cover the action, the open_url constraint, and the gate consequence. The arrow notation and parenthetical are dense but not wasteful.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has an output schema, so return values need not be described, and the description covers the main workflow and side effects. However, for a three-parameter tool with 0% schema coverage, leaving app_dir and project_id entirely unexplained is a meaningful completeness gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for three undocumented parameters. It explains open_url precisely, including the critical source and exclusion, but app_dir and project_id remain without any semantic explanation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The definition states a specific verb and resource: record a finished Claude Design upload by writing a SHA-256 receipt into design/project/claude_design.json. It distinguishes this from nearby workflow steps by naming render_preview, serve_url, and the design/icon/screenshots gates.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives clear context for when the tool applies: after a finished Claude Design upload and before the design, icon, and screenshots gates can pass. It does not explicitly contrast this with sibling tools such as design_upload_status or design_export_png, so it falls short of full alternative guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
design_research_brief_templateC
The design/research/brief.json shape (pre-filled with the reference ids): purpose, analysis, and a direction per section (palette tokens, typography, components, density, illustration, onboarding flow, paywall, screenshots analysis + concept + boards), each citing references and what it does differently.
| Name | Required | Description | Default |
|---|---|---|---|
| app_dir | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It discloses that the brief comes 'pre-filled with the reference ids', which is a small behavioral hint, but says nothing about whether a file is written, where it lands, whether it overwrites, permissions, or side effects. For a writing/templating tool with zero annotation coverage this is a significant gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
It is a single dense sentence, front-loaded with the file target, but the long nested parenthetical of sections makes it a run-on. It is compact but not cleanly structured, and the section list partly overlaps with what the output schema already conveys.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, return values needn't be described, and the section enumeration covers content adequately. What remains missing is the operational context: what the tool does to disk, where the template is written, and its relationship to the surrounding design_research/design_generate workflow.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The single required parameter app_dir has 0% schema description coverage and is not mentioned at all in the description. The agent gets no clarification of what app_dir points to or how it affects output, so the description fails to compensate for the documentation gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description explains the contents of the brief.json shape and lists the sections it covers, which gives real substance. However, it never states the action the tool performs (create a template? write a file? return a scaffold?), so an agent cannot tell what it actually does versus siblings like design_research_collect or design_research_check. The purpose is implied rather than stated.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no indication of when to call this tool, when not to, or how it relates to the many design_* siblings. The agent is left to infer that this is a templating/scaffolding step from the name alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
design_research_checkB
The design_research gate: ≥6 reference apps with icon + screenshots on disk, and a valid brief (every section cited, screenshot concept citing ≥4 competitor sets, palette/type not the template defaults or another factory app's).
| Name | Required | Description | Default |
|---|---|---|---|
| app_dir | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It usefully discloses the concrete thresholds being checked, which is real behavioral context. However, it does not say whether the tool is read-only, what happens on failure (throw vs. boolean), or whether it triggers downstream actions, leaving mutation/side-effect behavior unspecified.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence that packs the acceptance criteria without filler. The nested parenthetical is dense but each clause is load-bearing, so little is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a validation gate with an output schema, the criteria are adequately specified so the agent knows what 'pass' means. Gaps remain: the lone parameter is undocumented and there is no guidance on when in the pipeline to call it or what the failure output signifies.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The single required parameter app_dir has 0% schema description coverage, so the description must compensate, and it never references the parameter at all. The parameter's role is only inferable from the tool's domain context, not from the description text.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific gate/check and enumerates exactly what it validates: ≥6 reference apps with icon + screenshots on disk, plus a valid brief with specific citation and palette/type conditions. An agent can tell this is a design-research validation step, distinguishing it from collect/brief_template siblings. It stops short of an explicit verb like 'validate', relying on 'gate' to convey the action.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit statement of when to invoke this versus the sibling design_research_collect, design_research_brief_template, or other *_check tools. The 'gate' framing hints it runs as a pass/fail checkpoint, but no prerequisites or timing guidance are given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
design_research_collectB
Design research: the category leaders across storefronts (iTunes Search for terms + the genre's
top-grossing/top-free charts, looked up for screenshots/artwork) → downloads their App Store screenshots
and icons to design/research/apps/-/ + references.json. Study material only (never uploaded).
| Name | Required | Description | Default |
|---|---|---|---|
| terms | Yes | ||
| app_dir | Yes | ||
| max_apps | No | ||
| countries | No | ||
| screenshots_per_app | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It usefully discloses the data sources (iTunes Search, top-grossing/top-free charts), the artifacts downloaded (screenshots and icons), the local destination path with references.json, and the important constraint that material is never uploaded. It still omits permission needs, rate limits, overwrite behavior, and error handling for a tool that writes files to disk.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single dense sentence with arrow notation and no filler; every clause adds information about sources, actions, or destination. It is front-loaded with the purpose. The compact arrow structure is readable but slightly less immediately clear than plain prose, so not a full 5.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained. However, for a tool with five parameters, zero schema descriptions, and no annotations, the description leaves major gaps: most parameter semantics are absent, and behavioral details like permission or overwrite behavior are missing. It conveys the overall workflow but is not complete enough for reliable invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across five parameters, so the description must compensate but largely does not. It mentions `terms` as the input to iTunes Search and the output path design/research/apps/<id>-<slug>/, which loosely implies `app_dir`. It says nothing about `max_apps`, `countries`, or `screenshots_per_app`, leaving three of five parameters with no semantic guidance.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: it uses iTunes Search and genre charts to find category leaders and downloads their screenshots/icons, with a clear output location. It is not a tautology and clearly distinguishes itself from generic design tools. However, it does not explicitly differentiate from siblings like design_research_check or design_generate, so a 4 rather than 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage by framing the tool as 'Design research' and specifying that the output is 'Study material only (never uploaded).' This gives the agent a sense of when the output should be used. It does not, however, name alternative tools, specify prerequisites, or state when-not to use this tool versus siblings such as design_research_brief_template or design_research_check.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
design_screens_skeletonA
STRUCTURAL skeleton for design/screens.json: with a design brief, the onboarding follows the brief's onboarding.flow (stubs for kinds the default funnel lacks) and carries brief_sha256; without one, the default honest funnel. Rules stay: no fake stats/reviews, spin wheel, rating or notification prompt. ADAPT EVERY SCREEN to the app and the brief — the look is authored in Claude Design, not here. write=True saves it to /design/screens.json (refuses to replace an existing file unless overwrite).
| Name | Required | Description | Default |
|---|---|---|---|
| write | No | ||
| app_dir | Yes | ||
| overwrite | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does reasonably well: it discloses that write=True persists to <app>/design/screens.json, that it refuses to replace an existing file unless overwrite, and it enumerates content rules (no fake stats/reviews, spin wheel, rating or notification prompts). The default write=false preview behavior is only implied, not stated outright.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core artifact and purpose, and every clause earns its place (behavioral rules, write/overwrite semantics). It is denser and more run-on than ideal, but there is no padding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained. The description covers the conditional content source, the content constraints, and the persistence semantics, which is nearly everything an agent needs for this three-parameter generator; only the non-writing default mode remains under-specified.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It explains write=True (persists) and overwrite (permits replacing an existing file), and the <app>/design/screens.json path implies app_dir is the app directory. All three parameters get at least implicit meaning, but write=false is not explicitly described as an in-memory preview.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and artifact: it builds a STRUCTURAL skeleton for design/screens.json. It distinguishes itself from the authoring step by noting 'the look is authored in Claude Design, not here,' which helps separate it from siblings like design_generate. The purpose is clear, though the sentence structure is convoluted enough that it takes a second read.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It distinguishes the two input regimes (with a brief vs without) and hints that visual authoring happens elsewhere, but it never explicitly states when to call this versus design_generate or the other design_* siblings by name. Usage is implied rather than prescribed.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
design_upload_statusC
Is the local Claude Design project uploaded and unchanged since (the check every design gate runs)?
| Name | Required | Description | Default |
|---|---|---|---|
| app_dir | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It discloses the semantic it checks (uploaded and unchanged state) but nothing about how 'unchanged' is determined, permission/auth needs, or failure behavior; for a gate check whose result gates downstream steps this is thin.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no padding, and the parenthetical adds genuinely useful gate context without bloat. The only cost is the broken 'unchanged since' construction, which slightly undermines the efficient structure.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Output schema and low complexity mean return values and structure need little explanation, but the description omits any mention of the app_dir parameter and leaves the 'since' clause unresolved, so an agent still lacks the minimal information to call and interpret this gate check confidently.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The single required parameter app_dir has 0% schema description coverage and is not mentioned in the description at all. With one undocumented parameter the description should have compensated but does not, so despite the parameter being plausibly obvious, no semantic meaning is added.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description conveys the resource (local Claude Design project) and the state checked (uploaded and unchanged), and 'the check every design gate runs' hints at its role versus siblings like design_record_upload. However, the phrasing 'unchanged since' is grammatically dangling (since what?) and the intended verb/result is left implicit, so the purpose is understandable but imprecise.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It implies usage context by saying it is 'the check every design gate runs', which tells an agent it is a gate/precondition check. It gives no explicit when-to-use versus alternatives, no when-not, and doesn't name sibling tools, so guidance remains inferential.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
env_doctorB
Check the local toolchain: xcode, swift, node, asc (App Store Connect CLI), fastlane (app creation only), uv, git, maestro + Java 17+ + maestro-live (end-to-end flows, Viewer).
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are present, so the description must carry the behavioral load. It usefully enumerates the exact surface being probed (xcode, swift, node, asc, fastlane, uv, git, maestro, Java 17+, maestro-live), which tells the agent the coverage scope, but it never says whether the tool only reads or also installs/repairs, nor what a passing vs failing result looks like. The presence of a check-style name plus a defined output schema partially compensates.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with the action and the full checklist; every clause carries information. It is dense but not padded, though the trailing parentheticals slightly interrupt the scan.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-argument diagnostic with an output schema that already documents return values, the description's enumeration of checked toolchain components is close to sufficient. The main gap is the absence of any routing guidance relative to the other diagnostic tools.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, which is the baseline-4 case. The description adds no parameter meaning because there are none to describe.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb+resource ('Check the local toolchain') and enumerates exactly which components are inspected, so an agent knows precisely what this tool reports on. It does not, however, differentiate itself from adjacent diagnostic siblings such as config_doctor or setup_status.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no statement of when to run this tool versus config_doctor, setup_status, or orchestrator_preflight, all of which are diagnostic siblings in the same namespace. The parenthetical notes ('app creation only', 'end-to-end flows, Viewer') describe scope of checks, not usage conditions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
firebase_setupB
Set up Firebase Analytics (MANDATORY for every app): GCP project + iOS app + GoogleService-Info.plist → Resources/. Fully automatic (gcloud authed + firebase-tools). Live event tracking (from the phone).
| Name | Required | Description | Default |
|---|---|---|---|
| app_dir | Yes | ||
| app_name | Yes | ||
| bundle_id | Yes | ||
| approval_id | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full load. It discloses automatic execution, required tools (gcloud authed, firebase-tools), and side effects (creates GCP project, iOS app, plist), plus live event tracking. But it omits idempotency, reversibility, permission details beyond gcloud, and failure behavior—important for a setup mutation. Score 3.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with zero waste, front-loaded with the mandatory nature and the concrete outputs. Arrow notation efficiently communicates the artifact flow.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description explains the tool's purpose, mandatory nature, and prerequisites, but with 0% parameter schema coverage it leaves the agent without crucial input guidance (what each argument should contain, what approval_id is for). No annotations compound the gap. Adequate only for high-level understanding, not for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% with 4 parameters; the description provides no mapping or meaning for app_dir, bundle_id, app_name, or approval_id. The agent must guess parameter intent entirely, a severe gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Set up Firebase Analytics') and enumerates the concrete artifacts it creates (GCP project, iOS app, GoogleService-Info.plist). However, it does not explicitly distinguish itself from the many sibling setup_* tools (e.g., setup_services, setup_set), so 4 rather than 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Marks the tool as 'MANDATORY for every app' and gives prerequisites ('gcloud authed + firebase-tools'), providing clear context for when to use it. It stops short of naming when-not-to-use or alternatives, so 4 is appropriate.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
github_create_repoB
Open a PRIVATE GitHub repo under the configured GitHub account (github_user, else the active gh login) + push.
dry_run=True (default) returns the exact command without running it; a live run needs human approval.
| Name | Required | Description | Default |
|---|---|---|---|
| app_dir | Yes | ||
| dry_run | No | ||
| repo_name | Yes | ||
| approval_id | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does disclose meaningful traits: the repo is PRIVATE, the account is resolved from `github_user` falling back to the active gh login, dry_run defaults to true and only echoes the command, and a live run requires human approval. It omits failure/idempotency behavior (e.g. what happens if the repo already exists), which keeps it short of a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences with the critical constraint (PRIVATE, account resolution) front-loaded and no filler. The parenthetical is dense but earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists so return values needn't be explained, and the safety story (dry_run + approval) is solid for a mutation tool with no annotations. Gaps remain: no param docs for app_dir/repo_name and no differentiation from github_push.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for 4 parameters, so the description must compensate. It explains dry_run well ('returns the exact command without running it') and implies approval_id via the human-approval gate, but app_dir and repo_name are left entirely undocumented in both schema and description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Open a PRIVATE GitHub repo ... + push') with the visibility constraint front-loaded, so the agent knows exactly what is created. However, it does not differentiate from the sibling github_push, which it partially overlaps with via '+ push'.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The dry_run default and the human-approval gate for live runs give useful workflow context, implying the intended two-step usage. But it never says when to use this versus github_push or github_issues_bootstrap, and no prerequisites (auth/token) are stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
github_issue_createA
Open an issue for postponed work: imperative title, a role label (+ next|later), body with what/why/done-when. Runs gh as the configured GitHub account. dry_run returns the exact command; live: human approval. n/a (nothing runs) when the user turned off the github_issues run option (pass app_dir).
| Name | Required | Description | Default |
|---|---|---|---|
| body | Yes | ||
| repo | Yes | ||
| title | Yes | ||
| labels | Yes | ||
| app_dir | No | ||
| dry_run | No | ||
| approval_id | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does well: it discloses execution via gh as the configured account, that dry_run returns the exact command, that live requires human approval, and that it is a no-op when the github_issues run option is off. It omits error/auth-failure handling, keeping it below a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Dense but front-loaded: the purpose and content conventions come first, followed by execution behavior and gating. Every sentence carries information; the terse 'n/a (nothing runs)' phrasing is slightly cryptic but not wasteful.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with an output schema already present, the description covers the safety-relevant behavior (dry_run, human approval, run-option gating) an agent needs before calling it. Only the undocumented repo and approval_id parameters leave a small gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% across 7 params, so the description must compensate. It adds real meaning for title, labels (role label + next|later), body (what/why/done-when), dry_run and app_dir, and implies approval_id via 'live: human approval', but repo and approval_id are never named explicitly.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Open an issue for postponed work') and immediately distinguishes it from sibling repo/push tools by naming the intended workflow. The 'postponed work' framing and content conventions make the tool's niche clear without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a clear context for when to use it ('for postponed work') and explains the dry_run vs live approval flow, which is exactly the routing an agent needs. It does not explicitly name the closest sibling (github_issues_bootstrap) as an alternative, so it falls short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
github_issues_bootstrapA
Create/refresh the label set (ios, backend, store, lead, founder, next, later, other-project) on /, gh runs as the configured GitHub account. dry_run returns the commands; live: human approval. n/a (nothing runs) when the user turned off the github_issues run option (pass app_dir to read it from the spec).
| Name | Required | Description | Default |
|---|---|---|---|
| repo | Yes | ||
| app_dir | No | ||
| dry_run | No | ||
| approval_id | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and mostly delivers: it discloses the dry_run behavior (returns commands), the live-mode approval gate ("human approval"), the auth model ("gh runs as the configured GitHub account"), and the n/a condition. It omits details on the mutation's destructive potential (what 'refresh' overwrites) and rate limits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the core purpose and packs behavioral detail into a compact space. The dense multi-clause sentence with parentheticals is slightly run-on, but every clause adds information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 4-param mutation tool with no annotations and an output schema present, the description covers the approval gate, dry-run/live distinction, and the n/a case well. It is largely complete, with only the approval_id parameter and overwrite semantics left unexplained.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate. It clarifies dry_run semantics and notes app_dir is used to read the run option from the spec, and implies repo's <owner>/<repo> format, but approval_id is entirely undocumented in both the description and schema, leaving a gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Create/refresh) and resource (the label set) plus the exact target (<owner>/<repo>) and enumerates the labels. It is clearly distinguishable from github_issue_create (issues) and github_create_repo (repo creation) without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides a clear when-not condition ("n/a ... when the user turned off the github_issues run option") and explains the dry_run vs live mode split. It does not explicitly name alternative tools for related operations, so it stops short of full routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
github_pushB
Commit + push changes (regular tracking). Human approval.
| Name | Required | Description | Default |
|---|---|---|---|
| app_dir | Yes | ||
| message | Yes | ||
| approval_id | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full disclosure burden. It does surface one genuinely useful behavioral trait — the operation requires human approval — but omits destructive/reversibility context, branch behavior, and what happens when approval_id is absent or invalid. This is partial coverage for a mutation tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The action is front-loaded and there is no wasted prose, which suits the dimension. But the telegraphic fragments are terse to the point of under-specification, so brevity here reflects omission rather than efficient communication.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a write operation with no annotations, 0% schema coverage, and a thin one-line description, the definition is not complete enough. An output schema exists so return values need not be described, but the missing branch/commit semantics and approval flow leave an agent with real gaps before calling it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for three parameters, so the description must compensate and largely does not. 'Commit' loosely implies a commit message and 'Human approval' loosely ties to approval_id, but app_dir and the required nature of message are never explained, leaving the schema entirely undocumented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States specific verbs and a resource ('Commit + push changes'), which is enough for an agent to recognize the operation without opening the schema. The parenthetical '(regular tracking)' is cryptic and does not clarify scope or how it differs from any other tracking mode, but the core purpose is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Human approval.' implies this tool is gated and that a review step is required before use, and '(regular tracking)' hints at a variant mode. However, no alternative tool is named and there is no explicit when-to-use/when-not-to-use guidance, leaving usage conditions largely implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
growth_build_slideshowsA
Render viral TikTok/Reels slideshows (SlideSmith pattern, free, no scheduling). CLAUDE writes the hooks/slide text; images are generated with the app's AI (fal). slideshows=[{name, slides:[{text, image_prompt? or image?, cta?}]}]. Output 1080x1920, marketing/growth//NN.png — ready to share. A step for AFTER the app is submitted.
| Name | Required | Description | Default |
|---|---|---|---|
| app_dir | Yes | ||
| slideshows | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does meaningful work: it says the operation is free, unscheduled, generates images via the app's AI (fal), and writes 1080x1920 PNGs into marketing/growth/<name>/. It omits whether files are overwritten, auth/credential needs, or rate limits, but provides substantial behavioral context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action and scope, then supporting detail. Compact overall with little waste, though the parentheticals ('SlideSmith pattern, free, no scheduling') and the object-literal snippet interrupt prose flow slightly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists so return-value explanation is not required, yet the description adds useful output-path detail. It is largely complete for a generation tool, with the main gap being the undocumented app_dir parameter.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate. It admirably documents the required slideshows structure (name, slides[text, image_prompt? or image?, cta?]), covering the nested shape the schema leaves as bare additionalProperties:true. However app_dir is never explained, leaving one of two required params undocumented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Render) and resource (viral TikTok/Reels slideshows) with scope notes like 'free, no scheduling'. It is distinguishable from all siblings, none of which concern slideshows. The 'SlideSmith pattern' jargon is unexplained but does not obscure the core purpose.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Clearly positions the tool in the workflow with 'A step for AFTER the app is submitted', and clarifies 'no scheduling'. It gives when-to-use context but names no explicit alternatives or prerequisites (e.g., auth, dependency on submission).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
icon_generateA
OPT-IN ONLY (spend → approval): fal raster DRAFT → design/icon_drafts/ as reference for the Claude Design icon board. Never installs an icon — the shipped icon comes from Claude Design (icon_install).
| Name | Required | Description | Default |
|---|---|---|---|
| extra | No | ||
| app_dir | Yes | ||
| concept | Yes | ||
| allow_external_generator | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does disclose meaningful traits: it costs money and is gated behind approval, it writes to a specific artifact directory, and it produces reference drafts rather than shippable assets. It omits rate-limit/auth details, but the cost and non-destructive-vs-shipping behavior are the operative facts.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two dense sentences, zero filler, with the most consequential constraint (opt-in/cost/approval) front-loaded before the output-destination and sibling-routing details.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained, and the description covers the action, destination, and gating. The remaining gap is parameter meaning: an agent still cannot tell what app_dir, concept, or extra should contain.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across 4 parameters, and the description never explains app_dir, concept, extra, or allow_external_generator. 'fal raster' only loosely hints at the external-generator flag, so the parameters are essentially undocumented in both schema and description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource with output destination: generates a fal raster DRAFT into design/icon_drafts/ for the Claude Design icon board. It also names the sibling icon_install as the tool that ships the real icon, so an agent can distinguish draft generation from installation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit precondition ('OPT-IN ONLY (spend → approval)') and an explicit boundary ('Never installs an icon — the shipped icon comes from Claude Design (icon_install)'), which routes the agent to the correct sibling for shipping. It stops short of describing when inside the pipeline this step should be invoked.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
icon_installB
Install the Claude Design app icon: the 1024×1024 master exported from B01-AppIcon.dc.html (design_export_png → design/icon.png), already uploaded + recorded (design_record_upload) → AppIcon.appiconset (alpha flattened) + .appfactory/verify/icon.json (the icon gate requires it).
| Name | Required | Description | Default |
|---|---|---|---|
| source | No | design/icon.png | |
| app_dir | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does well: it discloses the alpha-flattening behavior, the exact outputs written (AppIcon.appiconset and .appfactory/verify/icon.json), and that a downstream gate depends on it. It omits idempotency/overwrite behavior and any auth or failure-mode notes.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with the purpose first and pipeline details after. Dense with jargon and nested parentheses, but each clause (source, prerequisites, outputs, gate) carries distinct information, so little is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists so return values need not be explained, and the pipeline context and written artifacts are covered. However, the required app_dir is undefined and mutation semantics (overwrite, permissions) are absent, leaving gaps an agent would need to guess around.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, and it only partly does: it effectively documents the 'source' default (design/icon.png, the 1024×1024 master). The required 'app_dir' parameter is never explained, leaving the single mandatory argument undocumented in both schema and description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Install') and resource ('the Claude Design app icon') with the exact source artifact and destination. It implicitly distinguishes itself from siblings by naming the upstream tools (design_export_png, design_record_upload) whose outputs it consumes, though it does not explicitly contrast with icon_generate.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied through the pipeline parentheses ('already uploaded + recorded') and the note that the icon gate requires the verify artifact. There is no explicit 'use when X vs icon_generate' guidance, so the agent must infer ordering and alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
idea_evaluateC
Evidence bundle to rank one idea (any category): autocomplete count, niche score/verdict, two-window newcomer traction (12 + 6 months), leaders with price ladders, ai_needed (yes|optional|no heuristic + claude_assessment for Claude to set), build_complexity (low|medium|high vs the template), review_risk by genre/term, and the first available name from name_candidates. genre None = inferred from the competitors.
| Name | Required | Description | Default |
|---|---|---|---|
| term | Yes | ||
| genre | No | ||
| country | No | us | |
| leaders | No | ||
| name_candidates | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does disclose useful behavior: it explains that genre None is inferred from competitors, that ai_needed includes a heuristic plus a claude_assessment field, and that a name is selected from name_candidates. It still omits safety-relevant behavior such as whether the tool is read-only, what permissions are needed, and whether it has side effects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the core purpose and is a single dense sentence, which suits a complex tool. However, it is a run-on list that spends most of its length enumerating output fields that an existing output schema already covers, reducing efficiency.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given zero schema description coverage, no annotations, and a rich sibling ecosystem, the description is not complete enough for an agent to call the tool correctly. It explains many return-bundle contents but provides no usage routing, weak parameter semantics, and no safety or side-effect context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for five parameters, so the description must compensate. It adds meaning for genre (inferred from competitors) and name_candidates (first available name is selected from them), and mentions term indirectly, but country and leaders are not explained and term is never defined as the input search term.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool produces an evidence bundle to rank one idea and lists the specific evidence components, so the resource and scope are understandable. However, it does not explicitly distinguish this tool from sibling evaluators such as aso_niche_score, aso_fetch_competitors, or idea_harvest.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description does not say when to use this tool versus alternatives, nor does it mention prerequisites or exclusions. It only implies usage by saying it ranks one idea in any category, which is insufficient guidance for selecting it among many sibling tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
idea_harvestA
Phase 0 idea harvest across EVERY App Store category and storefront (not only AI/photo apps).
Sweeps top-grossing + top-free (legacy RSS) per genre × storefront (default us, gb, de, br, tr, jp, fr; genres None = all 24 charted categories), dedupes, enriches via iTunes lookup (release date, ratings, price) and returns chart_proven (top-grossing), rising_newcomers (released ≤ newcomer_months, sorted by ratings/day) and per-genre clusters (open_niche: open|some|closed|unmeasured). Read-only, free, paced (~0.3 s/request; a full default sweep takes ~1.5 min). A failed feed is listed in failed_feeds (ok:False) — never read as "no apps". exclude_terms drops the founder's own apps (name/seller match).
| Name | Required | Description | Default |
|---|---|---|---|
| feeds | No | ||
| limit | No | ||
| genres | No | ||
| countries | No | ||
| max_results | No | ||
| exclude_terms | No | ||
| newcomer_months | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does so richly: read-only, free, paced at ~0.3s/request, ~1.5 min for a full default sweep. It also discloses critical error semantics (failed feeds appear in failed_feeds with ok:False and must never be read as 'no apps') and that exclude_terms drops the founder's own apps by name/seller match.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The passage is dense but front-loaded: the scope statement leads, then mechanics, then outputs, then operational caveats. Every clause carries signal, though the single long block is slightly heavy to parse relative to an ideal chunked structure.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex 7-parameter tool the description covers scope, timing, failure behavior and output groupings; and since an output schema exists, return values need not be enumerated. The remaining gap is the un-described feeds/limit/max_results parameters and how they interact with the default sweep.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate for all 7 params. It meaningfully explains genres (all 24 charted categories), countries (default us/gb/de/br/tr/jp/fr), newcomer_months (released ≤ N months) and exclude_terms, but leaves feeds, limit, and max_results entirely undocumented, so roughly half the surface remains opaque.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource (idea harvest / Phase 0 sweep) and specifies the exact scope: top-grossing + top-free per genre × storefront. It even pre-empts a misreading by clarifying it is NOT limited to AI/photo apps, which distinguishes it from narrower siblings like aso_top_grossing.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The 'Phase 0' framing and the explicit 'across EVERY category... not only AI/photo apps' contrast give clear context for when this broad-harvest tool is appropriate versus narrower alternatives. However, no sibling tool is named explicitly (e.g. aso_top_grossing, aso_fetch_competitors, idea_evaluate), so the routing is inferred rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
legal_checkC
What still blocks publishing the legal pages (placeholders, unrendered blocks, missing languages).
| Name | Required | Description | Default |
|---|---|---|---|
| app_dir | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It implies a read-only diagnostic but never says so, and gives no information about permissions, scope of the scan, or whether it mutates anything. Listing the check categories is essentially return-value content, which the output schema already covers.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single compact clause with no filler, and the core purpose is front-loaded. It is a fragment rather than a structured sentence, but nothing is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Because an output schema exists, the description needn't explain return values, and for a one-parameter tool it covers the basic purpose. However, with no annotations and no usage context, an agent still lacks enough to confidently distinguish this from legal_verify and related checks.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and there is one required parameter (app_dir) that the description never mentions. The parameter name is fairly self-explanatory, but the description does no work to compensate for the missing schema documentation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific diagnostic purpose — identifying what blocks publishing the legal pages — and enumerates the check categories (placeholders, unrendered blocks, missing languages). It is clear without opening the schema, though it does not differentiate itself from the sibling legal_verify, which sounds like an adjacent check on the same resource.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no statement of when to run this versus legal_verify, legal_render, or the broader pipeline_status/orchestrator checks. The agent must infer the trigger condition entirely from the name.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
legal_renderB
Fill the legal sources (store/privacy/*.md, backend/PRIVACY.md) from the spec + config (APP_NAME, CONTROLLER, SUPPORT_EMAIL, dates, trial/plan sentences) and keep/drop the
| Name | Required | Description | Default |
|---|---|---|---|
| values | No | ||
| app_dir | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It usefully discloses the conditional block handling (keep/drop the consent.health HealthKit section) and that placeholders are substituted, but says nothing about overwriting existing files, required permissions, or idempotency.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single dense sentence with the action and target front-loaded, followed by the config detail and the conditional-block caveat. Efficient, though the heavy parenthetical nesting makes it slightly harder to scan.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists so return values need not be described, and the file targets are specified. However, for a file-writing render tool it omits whether existing files are overwritten, whether it creates directories, and the accepted keys for `values`.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate. It explains that `values` fills product placeholders, adding some meaning beyond the schema, but `app_dir` (the required parameter) is entirely undocumented and the shape/keys of `values` remain unexplained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Fill') and resource ('the legal sources') with concrete target paths (store/privacy/*.md, backend/PRIVACY.md). The verb distinguishes it from siblings like legal_check and legal_verify, though it never names them explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No explicit when-to-use, when-not-to-use, or alternatives are given. The phrase 'from the spec + config' hints at the precondition, but nothing routes the agent between this and legal_check/legal_verify beyond inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
legal_verifyB
LIVE read-only: every privacy/terms page per app language answers 200 in that language with the support email and no placeholder.
| Name | Required | Description | Default |
|---|---|---|---|
| app_dir | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, and it does disclose meaningful traits: it is a live network probe, it is read-only, and it checks 200 status, per-language responses, support email presence, and placeholder absence. It stops short of failure/timeout behavior, auth needs, or what counts as a hard failure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence led by the 'LIVE read-only' qualifier, with no filler. Slightly dense/run-on, but every clause contributes a testable assertion.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The output schema covers return values and there is only one parameter, so the description need not explain results. It conveys the core checks, but omits which page types/URLs are enumerated and how failures surface, leaving an agent short of what it needs for confident invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for the single required parameter app_dir, so the description must compensate and does not. It never explains what app_dir points to or how languages are derived from it; 'per app language' only loosely implies it is a directory containing localized pages.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verification resource (privacy/terms pages per app language) and the concrete assertions made (HTTP 200 in the correct language, support email present, no placeholder). 'LIVE read-only' hints at differentiation from static siblings like legal_check/legal_render, but it never names them or the verb 'verify' explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no statement of when to run this versus legal_check, legal_render, or metadata_listing_check, and no prerequisites (e.g., pages must be deployed first). 'LIVE' implies it targets deployed URLs, but that is inference rather than guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
localize_applyB
Merge translations into the in-app String Catalog for the spec's locales.app (translations = {english_key: {locale: value}}; other locales are ignored).
| Name | Required | Description | Default |
|---|---|---|---|
| app_dir | Yes | ||
| translations | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It usefully discloses that it *merges* (not replaces) and that 'other locales are ignored' — a silent-filtering behavior worth surfacing. It stops short of saying what happens to existing keys or conflicting values, so coverage is partial.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence leading with the action and resource, with the format hint tucked into a parenthetical. No filler, though the parenthetical slightly compresses the flow.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained, and the definition captures the key nuance (locale filtering, merge). The remaining gap is the undocumented `app_dir` parameter and the lack of notes on overwrite behavior for a mutation with no annotations.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate. It documents the shape and meaning of `translations` ({english_key: {locale: value}}) beyond the bare type, but says nothing about `app_dir`, whose purpose and expected value remain undocumented in both schema and description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource: 'Merge translations into the in-app String Catalog'. The phrase 'in-app String Catalog' implicitly separates it from the App Store Connect localization siblings (asc_localize_subscription, asc_localize_group), though it never names them explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No explicit when-to-use or when-not guidance is given, and no alternative is referenced. The 'apply'/'merge' wording only weakly implies this is the write-back step after translation, leaving the agent to infer when to call it versus the many asc_localize_* siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
maestro_testA
Run the app's Maestro flows (.maestro/) on the booted simulator through tools/maestro-live:
it starts maestro mcp, opens the Maestro Viewer (http://localhost:7777, or the next free port)
in the browser so the founder watches the run live, and runs each flow with the MCP run tool.
Returns pass/fail per flow (JUnit also in build/maestro/report.xml). tags: comma-separated
include filter, e.g. "smoke". The app must already be installed (build_for_sim + simctl install).
Writes .appfactory/verify/maestro.json: the features gate and testflight_ship require every
smoke flow green. There is no headless option here: the Viewer is mandatory on the Mac.
| Name | Required | Description | Default |
|---|---|---|---|
| tags | No | ||
| device | No | ||
| app_dir | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations present, the description carries the full burden and does so richly: it discloses that it spawns `maestro mcp`, forces a browser Viewer on localhost:7777 (next free port) with no headless option, writes .appfactory/verify/maestro.json, produces JUnit at build/maestro/report.xml, and returns pass/fail per flow. Side effects, environment constraints, and outputs are all made explicit.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The purpose and Viewer behavior are front-loaded, and the sentences are dense with useful facts rather than filler. It is a long single paragraph, though, and details like the exact port number sit alongside higher-value prerequisites without structural separation.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a multi-step, side-effecting test runner the description covers prerequisites, mandatory Viewer, artifact paths, and gating consequences; an output schema exists so return values need not be restated (it summarizes them anyway). The notable gap is the undocumented `device` parameter and ambiguity about what `app_dir` should point to.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It explains `tags` well ('comma-separated include filter, e.g. "smoke"'), but says nothing about `app_dir` (only inferable) and nothing about `device`, leaving two of three parameters semantically undefined.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource ('Run the app's Maestro flows on the booted simulator') plus the exact mechanism (tools/maestro-live, MCP `run` tool), which clearly separates it from generic build/test siblings like build_test. It never names a sibling to route away from, so it falls just short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives a concrete precondition ('The app must already be installed (build_for_sim + simctl install)') and explains downstream context (the features gate and testflight_ship require every `smoke` flow green), so an agent knows when this step is required. It does not explicitly state when another tool should be used instead, so no exclusion guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
mascot_assetsB
Import approved poses + blink variants into Resources/Assets.xcassets/Mascot/{,-blink}.imageset (namespaced → Image("Mascot/idle")); states without a pose borrow a fallback pose.
| Name | Required | Description | Default |
|---|---|---|---|
| states | No | ||
| app_dir | Yes | ||
| source_dir | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, and it does add real behavioral detail: the exact output path layout, the {,-blink} naming, namespacing to Image("Mascot/idle"), and fallback-pose borrowing for states lacking a pose. It still omits whether existing assets are overwritten, auth/approval mechanics, or failure behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single dense sentence front-loads the destination path, with the namespacing and fallback behavior appended as clarifiers. No wasted prose, though the brace/path notation is heavy for one line.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be described. What is present (target layout, namespacing, fallback) is useful, but with 3 params at 0% schema coverage, no annotations, and no approval mechanics explained, the definition is only partially complete for a write-to-project tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across 3 parameters, so the description must compensate and does not. 'states' is only indirectly implied via the fallback-pose sentence, while app_dir and source_dir are never explained. The one bit of path syntax given is the only parameter-adjacent meaning.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (import) and a concrete resource and destination (poses + blink variants into Resources/Assets.xcassets/Mascot). The fallback-pose note adds a scope distinction. However it does not clearly differentiate from the sibling mascot_blink, which appears related.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No explicit when-to-use, prerequisites, or alternatives are given. 'Approved poses' hints at a precondition but never states who/what approves them or when not to run this tool. The agent is left to infer usage.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
mascot_blinkA
Closed-eye copies of the APPROVED pose PNGs (-blink.png next to each): iris blobs in the upper 55 %,
inpainted with the surrounding color, closed-lid arcs. Default input: /design/mascot/.png.
Needs the optional extra: uv sync --extra mascot. Never rig from a separate parts sheet.
| Name | Required | Description | Default |
|---|---|---|---|
| paths | No | ||
| states | No | ||
| app_dir | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, and it does reasonably well: it discloses the output artifact naming/location, the exact transformation applied to the source art, and a hard dependency requirement. It does not say whether existing files are overwritten, what happens if a pose is missing, or any failure behavior, which keeps it short of a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Compact and front-loaded with the produced artifact first, followed by construction details, then defaults and prerequisites. Some clauses (the inpainting detail, the trailing warning) are dense and slightly run-on, but nearly every sentence carries information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be described. However, for a file-generating tool with zero annotations and 0% parameter schema coverage, the description leaves meaningful gaps: the required `app_dir` and optional `paths` are unexplained and overwrite behavior is unstated.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for all 3 parameters, so the description must compensate. It documents the default input path pattern (`<app>/design/mascot/<state>.png`) and implies the `states` parameter via `<state>-blink.png`, but `paths` and the required `app_dir` are never explained, leaving gaps in the partial compensation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states concretely what it produces: closed-eye PNG variants (`<state>-blink.png`) of approved pose images, with the visual construction spelled out (iris blobs in upper 55%, inpainting, closed-lid arcs). It is clear what the tool does, though the action verb is implied rather than stated and there is no differentiation from the sibling `mascot_assets`.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives a clear when-not ('Never rig from a separate parts sheet') and a prerequisite environment note ('Needs the optional extra: `uv sync --extra mascot`'), which is useful context. But there is no explicit statement of when to run this versus `mascot_assets` or the other mascot/design tools, so selection guidance is only partial.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
metadata_checkB
Validate apple-metadata.md (fields + character limits). No writes.
| Name | Required | Description | Default |
|---|---|---|---|
| source_md | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are supplied, so the description carries the full burden; it does disclose the safety-relevant trait 'No writes', which tells the agent this is a read-only check. Beyond that it says nothing about failure behavior, whether the target file must already exist, or whether the check is blocking.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two clipped clauses, front-loaded with the verb and target, with the safety note last. Nothing is padded and nothing is repeated from the schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return-value explanation is not needed, and the tool has only one parameter. However, the meaning of source_md and the relationship to aso_validate_metadata remain unresolved, so an agent cannot call it confidently.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the single parameter source_md has no description. The prose mentions apple-metadata.md but never clarifies whether source_md is a path, raw markdown content, or a directory, leaving the only required argument ambiguous.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (validate) and resource (apple-metadata.md) with a scoping parenthetical naming the checked concerns (fields + character limits). It does not differentiate itself from close siblings like aso_validate_metadata or metadata_listing_check, so an agent must infer which validation tool applies.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance, no prerequisites, and no named alternative. With a sibling named aso_validate_metadata that also validates metadata, the absence of routing guidance is a real gap.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
metadata_exportB
apple-metadata.md → fastlane/metadata//*.txt (validated).
If app_dir is given the ASO gate applies: export is refused until ASO is complete. primary/secondary_category = appCategories id (e.g. PHOTO_AND_VIDEO, PRODUCTIVITY) → written to the root *.txt files; deliver_metadata pushes them to ASC. Photo/portrait app default: PHOTO_AND_VIDEO+PRODUCTIVITY.
| Name | Required | Description | Default |
|---|---|---|---|
| app_dir | No | ||
| dest_dir | Yes | ||
| source_md | Yes | ||
| privacy_url | No | ||
| support_url | No | ||
| primary_category | No | ||
| secondary_category | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It usefully discloses that output is validated and that a gate can refuse the export, but it omits other consequential traits such as whether existing files are overwritten, required permissions, or failure modes.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Dense but front-loaded: the core transformation leads, followed by the gate rule and category semantics. Telegram-style abbreviations and arrows are efficient, though slightly terse for an unfamiliar reader.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists so return values need not be described, and the most consequential parameters (app_dir, categories) are covered. Still, for a 7-parameter tool with 0% schema coverage, the missing privacy_url/support_url semantics and overwrite/permission behavior leave gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% and there are 7 parameters, so the description must fill the gap. It explains app_dir (gate), primary/secondary_category (appCategories ids), and implies source_md/dest_dir via the transformation line, but privacy_url and support_url are left completely undocumented in both schema and description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a concrete input→output transformation (apple-metadata.md → fastlane/metadata/<locale>/*.txt) that an agent can act on, and differentiates itself from deliver_metadata by noting that the latter pushes to ASC. However, the actual verb ('export'/'write') is implied rather than stated, so it is clear but not perfectly crisp.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives one conditional usage rule (when app_dir is supplied the ASO gate blocks export until ASO completes) and points at deliver_metadata for pushing to ASC. It does not, however, say when to use this over siblings like metadata_check, metadata_render_listing, or localize_apply, leaving the selection context largely inferred.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
metadata_listing_checkC
Validate store/metadata/listing.json: name/subtitle ≤30, keywords 95–100, promo ≤170, no word overlap across fields or cross-indexed storefronts, description ends with the subscription disclosure + Terms + Privacy links (Supabase legal, ?lang=), IAP display copy.
| Name | Required | Description | Default |
|---|---|---|---|
| app_dir | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It lists validation criteria but does not state whether the operation is read-only, what happens on failure, permission requirements, or side effects. The checks are purpose-level detail, not behavioral disclosure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single dense sentence that front-loads the target file and constraint list. It earns its length by covering several specific validation rules, though the run-on list could be slightly more scannable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained. The description thoroughly covers validation rules, but with no annotations and no schema parameter descriptions, it omits parameter meaning and usage context, leaving it minimally adequate for a one-parameter validation tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and there is one parameter, app_dir. The description never mentions app_dir or explains how it maps to the 'store/metadata/listing.json' path, so it fails to compensate for the missing schema documentation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Validate') and resource ('store/metadata/listing.json') and lists the exact constraints checked, so the agent knows this is a metadata listing validator. It does not explicitly distinguish itself from sibling tools like metadata_check or aso_validate_metadata, which is the only gap.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description says what is validated but gives no guidance on when to call this tool versus alternatives, no prerequisites, and no conditions for use. An agent must infer usage from the name alone, which is common among the many validation siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
metadata_render_listingC
Sync listing.json to the spec (store locales, IAP copy slots) and fill the legal URLs.
| Name | Required | Description | Default |
|---|---|---|---|
| app_dir | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. "Sync" and "fill" imply a write to listing.json, but the description says nothing about what gets overwritten, whether existing values are preserved, idempotency, permissions, or failure behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single dense sentence that front-loads the core action and parenthetically enumerates the store locales/IAP copy slots scope. It is appropriately short and free of filler, though the parenthetical is cryptic rather than clarifying.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values needn't be explained, but with no annotations, an undocumented required parameter, and no usage context, the definition is thin for an operation that mutates repository files. An agent could not reliably know when or how to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There is a single parameter, app_dir, with 0% schema description coverage, and the description never mentions it or clarifies its expected form. With one opaque param and no compensating text, the description leaves the schema to speak for itself when the schema doesn't.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a verb ("Sync") and resource ("listing.json to the spec"), plus the sub-content it touches (store locales, IAP copy slots, legal URLs). However, the meaning of "listing.json" and "spec" is jargon-dependent, and it does nothing to distinguish itself from siblings like metadata_check, metadata_listing_check, metadata_export, or legal_render.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no statement of when to run this versus metadata_check, metadata_export, metadata_listing_check, or legal_render, nor any prerequisite or ordering guidance. The agent is left to infer usage entirely from the name.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
onboarding_planA
Return guidance for onboarding step-count selection (does not generate code).
Usage: call this tool to ask the user how many onboarding steps they want, show the returned guidance, let them choose. Then during scaffold, the OnboardingStep array in Views.swift is written by hand with the chosen count and app-specific content (following the per-screen Claude Design boards) — this tool just clarifies the decision point and doesn't modify files.
STANDARD: onboarding is a quiz-style Q&A with ~5 personalization questions (OnboardingStep.kind=.question), so new apps match that count (~5), not "at least one".
| Name | Required | Description | Default |
|---|---|---|---|
| step_count | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It clearly disclaims code generation and file modification, stating it only clarifies the decision point. This is strong transparency for a read-only guidance tool, though it could add details like idempotency or whether it consults external state.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with a clear one-sentence purpose, followed by usage and a standard note. Every section adds value, though the three-part structure is slightly longer than necessary for a one-parameter tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's simplicity, an output schema exists so return values need not be explained. The description covers purpose, usage, non-destructive behavior, and a recommended value. The only minor gap is the incomplete explanation of the step_count parameter and its relationship to the default.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It discusses step-count selection and provides a typical value (~5), which adds meaning beyond the bare integer default of 12. However, it never explicitly maps the step_count parameter or its expected range, leaving some ambiguity about the input's exact role.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: it returns guidance for onboarding step-count selection. It also adds a negative scope (does not generate code, does not modify files), which helps distinguish it from scaffolding tools. However, it does not explicitly name sibling tools or contrast alternatives, so it falls short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit when-to-use instructions: call this tool to ask the user how many onboarding steps they want, show the returned guidance, and let them choose. It also explains the downstream workflow and provides a standard (~5 steps) to guide the choice. Usage is unambiguous.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
orchestrator_needs_humanB
Self-correction exhausted: write NEEDS_HUMAN.md and return its path.
| Name | Required | Description | Default |
|---|---|---|---|
| stage | Yes | ||
| reason | Yes | ||
| app_dir | Yes | ||
| how_to_resolve | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries full behavioral burden. It usefully discloses the side effect (writes a NEEDS_HUMAN.md file) and that it returns a path, but omits whether the run halts, permission/precondition requirements, or what app_dir must contain. It adds some value but is thin for a no-annotation mutation tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single tight sentence with the trigger front-loaded before the action, so nothing is wasted. It is arguably over-sparse given the undocumented parameters, but as a sentence it is well-structured.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be described, and the description covers the core action. But with 4 required params at 0% coverage and no annotations, the definition leaves too much unspecified for an escalation/termination tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across 4 required parameters (app_dir, stage, reason, how_to_resolve). The description supplies no meaning, format, or examples for any of them, so an agent must guess what 'stage' or 'how_to_resolve' should contain.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and artifact ('write NEEDS_HUMAN.md and return its path'), plus a trigger condition ('Self-correction exhausted'). It is clearly distinguishable from pipeline/orchestrator siblings by its escalation framing, though it does not name or contrast any sibling explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'Self-correction exhausted' implies when this tool fires, giving an implicit usage condition. However, it never states when NOT to use it or names the alternative (e.g., orchestrator_record_attempt / orchestrator_next_action) that should be preferred while self-correction is still viable.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
orchestrator_next_actionA
Next action for the driver loop: stage + subagent role + retry budget + instructions. done=True means the pipeline is finished (submit needs human approval). Refuses until the run options are confirmed unless skip_options_check=true; setup_required comes first.
| Name | Required | Description | Default |
|---|---|---|---|
| app_dir | Yes | ||
| skip_options_check | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does well: it discloses the gating/refusal behavior, the meaning of done=True, and that submit requires human approval. It offers no detail on permissions beyond the options check, but the disclosed refusal semantics are exactly the kind of behavioral trait that matters here.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences, front-loaded with the return payload before the gating caveats. No padding, and the most important behavior (refusal condition) is present without burying the purpose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return-value explanation is not required, and the description covers purpose, the done sentinel, and the options-check precondition. For a two-parameter orchestration tool this is largely complete, with only the app_dir meaning left unaddressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It explains the non-obvious parameter (skip_options_check bypasses the options-confirmed gate), which is valuable, but app_dir is never described. Partial compensation warrants a mid score.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states concretely what the tool returns for a 'driver loop': a stage, subagent role, retry budget, and instructions. This is a specific verb+resource outcome, not a restatement of the name. It does not, however, differentiate itself from adjacent siblings like pipeline_next or orchestrator_preflight.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives genuine preconditions — it refuses until run options are confirmed unless skip_options_check=true, and notes setup_required comes first — which implies a correct call order. It stops short of saying when to prefer this over pipeline_next or orchestrator_preflight, so usage is implied rather than explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
orchestrator_preflightB
Pre-flight for an autonomous run: run options confirmed + config keys + fastlane session freshness + caffeinate command. Ready if blockers is empty. Returns setup_required first when AppFactory is not set up yet.
| Name | Required | Description | Default |
|---|---|---|---|
| app_dir | No | ||
| skip_options_check | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It does disclose return semantics (blockers gate, setup_required precedence) and implies the tool inspects external state including a 'caffeinate command', hinting at a possible side effect. However, it never states whether invoking it mutates anything, starts a process, or requires auth, which matters for a tool with no safety annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two dense sentences with the outcome condition ('Ready if blockers is empty') front-loaded in the second. No filler, though the slashed check list is somewhat telegraphic.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return-format detail is not required, and the description still adds the gate logic. For a zero-annotation preflight tool with two undocumented parameters, though, it omits prerequisites and side-effect behavior that an agent would need before calling it in a pipeline.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% and neither of the two parameters is mentioned in the description. 'app_dir' is entirely absent, and 'skip_options_check' is only obliquely implied by the phrase 'run options confirmed' with no indication that it can be bypassed. The description does not compensate for the documentation gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific function ('pre-flight for an autonomous run') and enumerates the concrete checks it performs: run options, config keys, fastlane session freshness, caffeinate command. An agent can tell this is a readiness gate rather than a setup mutator. It does not explicitly differentiate from the nearby setup_status sibling, which keeps it from a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives a decision rule ('Ready if blockers is empty') and an ordering rule ('Returns setup_required first when AppFactory is not set up yet'), which implies when the caller should proceed. But it never says when to call this versus setup_status, setup_services, or orchestrator_next_action, nor whether it must be run before the main run.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
orchestrator_record_attemptB
Count a failed attempt of a stage (after a gate failure). Returns: {stage, attempts, should_retry}.
| Name | Required | Description | Default |
|---|---|---|---|
| stage | Yes | ||
| app_dir | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It discloses the trigger condition (gate failure) and that it increments an attempt counter returning should_retry, which is useful behavioral context, but says nothing about side effects, error behavior, or permissions. It also restates return values already covered by the output schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences with zero waste; the trigger and return contract are front-loaded. Nothing redundant beyond the return-value line that overlaps the output schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Since an output schema exists, return values needn't be explained. But for an orchestration state-mutation tool with two fully undocumented parameters and no annotations, the description does not say enough about preconditions (e.g., whether preflight must run first) or the effect on the orchestration state.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, but it only obliquely references 'stage' and never explains app_dir or the expected format/nature of either parameter. The two required parameters remain undocumented in both places.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource: counting a failed attempt of a stage, with a trigger condition (after a gate failure). This is clear enough to understand the tool's function, though it doesn't distinguish itself from the many sibling orchestrator_* tools (orchestrator_preflight, orchestrator_next_action, orchestrator_needs_human).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides one contextual trigger, 'after a gate failure,' which implies when to call it. However, it names no alternatives and gives no when-not guidance or ordering relative to sibling orchestrator tools, leaving the agent to infer its place in the loop.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
pipeline_markC
Mark a stage (done/in_progress/pending).
| Name | Required | Description | Default |
|---|---|---|---|
| stage | Yes | ||
| status | No | done | |
| app_dir | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It implies a state mutation but does not disclose whether marking persists to disk, whether it is idempotent, whether it overwrites prior status, or what authorization/context (app_dir) it requires. The only added detail is the status vocabulary.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The single front-loaded sentence is tight and wastes nothing. It is arguably over-terse for a three-parameter mutation tool, but as a structural matter it is well-formed and earns its words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no annotations, 0% schema coverage, and an output schema present, the description should at minimum explain the stage namespace and app_dir. It omits both, leaving the agent unable to supply valid parameters without guessing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for three undocumented parameters. It does define the allowed values for 'status' (done/in_progress/pending), which the schema lacks, but 'stage' (which stage identifiers are valid?) and 'app_dir' (format, required-ness) remain entirely unexplained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Mark') and resource ('a stage'), and it enumerates the valid status values. However, it never clarifies what a 'stage' is in this pipeline, and it does nothing to distinguish itself from close siblings like pipeline_status, pipeline_next, or pipeline_validate.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to call this versus the sibling pipeline_* tools, nor any stated prerequisites. It implies marking but never says under what conditions an agent should mark a stage versus check or advance it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
pipeline_nextA
Return the next mandatory step + its instructions (driver). Refuses until the run options are confirmed (run_options → run_options_save) unless skip_options_check=true. Returns setup_required first when AppFactory is not set up yet.
| Name | Required | Description | Default |
|---|---|---|---|
| app_dir | Yes | ||
| skip_options_check | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral burden and does it well: it discloses the refusal/gating behavior, the skip_options_check escape hatch, and the ordered setup_required return. It does not state whether the call is read-only or has side effects, leaving a modest gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences with the core behavior (return the next mandatory step) front-loaded, followed by the gating conditions. Every sentence adds a distinct constraint with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return-value explanation is correctly omitted; the description instead covers prerequisites and ordering behavior, which is the right focus. Only the required app_dir parameter is left implicit.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate. It usefully explains skip_options_check (skipping the options check) but says nothing about app_dir, which is the required parameter. Partial compensation for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: returns the next mandatory step plus its instructions (driver). This is clearly distinct from pipeline_status and pipeline_mark, though it does not explicitly name a sibling to differentiate from. An agent can tell what it produces.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear preconditions: it refuses until run options are confirmed via run_options/run_options_save, and returns setup_required first when AppFactory is not set up. This tells the agent when it will and won't function, though it stops short of explicitly routing to alternative tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
pipeline_statusC
App pipeline status: which stages are done/pending, and what comes next.
| Name | Required | Description | Default |
|---|---|---|---|
| app_dir | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It is a read-oriented tool (status reporting), but it does not disclose whether it requires authentication, whether it has side effects, rate limits, or what the output format is. An output schema exists, so return values are covered, but the description adds little behavioral context beyond the basic purpose.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single concise sentence that front-loads the main purpose. It is appropriately sized and wastes no words, though it could be slightly more informative without becoming verbose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given a pipeline status tool with no annotations, one undocumented parameter, and an output schema, the description is incomplete. It does not explain what the different stages are, what 'what comes next' entails, or how to interpret the status. For a tool that likely returns structured pipeline data, more context about the stages and the meaning of pending/done would be valuable.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has one required parameter `app_dir` with 0% description coverage. The description does not mention the parameter at all, so it provides no semantics for it. With one parameter and no schema description, the description should ideally explain what `app_dir` represents. However, the baseline for a single parameter is 3 when the schema is minimal but the parameter is self-evident from name; still, some explanation would help.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific purpose: reporting app pipeline status including which stages are done/pending and what comes next. However, it does not differentiate itself from sibling tools like `pipeline_next`, `setup_status`, or `orchestrator_next_action`, which appear to have overlapping status-reporting roles. Without distinguishing its unique scope, an agent may struggle to choose it over alternatives.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance, no prerequisites, and no mention of alternatives. The description implies a general status query but does not specify in which contexts it should be preferred over `pipeline_next`, `setup_status`, or other pipeline-related tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
pipeline_validateA
Run a stage's enforced gate WITHOUT modifying the manifest.
For orchestrator/subagent self-checks: does the gate pass before marking it 'done'?
| Name | Required | Description | Default |
|---|---|---|---|
| stage | Yes | ||
| app_dir | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It usefully discloses that the tool does NOT modify the manifest (a read-only/validation effect), which is the most important trait here. However, it does not describe failure behavior, whether it errors or returns a status, or any permission requirements.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, front-loaded with the core action and its key constraint, then the motivating use case. Every sentence earns its place with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained, and the description covers purpose and usage adequately. But with no annotations and zero parameter documentation, an agent lacks guidance on what 'stage' and 'app_dir' must contain to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% and there are two required parameters (app_dir, stage). The description only obliquely references 'a stage,' leaving 'stage' format/uniqueness and 'app_dir' entirely undefined in both schema and description. It fails to compensate for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource ('Run a stage's enforced gate') and explicitly carves out the non-mutating scope ('WITHOUT modifying the manifest'). This distinguishes it from mutation siblings like pipeline_mark, though the term 'enforced gate' is project jargon an agent must infer.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives clear usage context: 'For orchestrator/subagent self-checks: does the gate pass before marking it done?' This tells the agent when to call it (pre-marking self-check) and implicitly positions pipeline_mark as the alternative. No explicit exclusions or named alternative tool, so not a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
playbookA
Return the AppFactory run playbook (markdown). Read it before driving the pipeline.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It discloses the return format (markdown) and that this is instructional rather than mutating, but never states that it is side-effect free/idempotent or what 'driving the pipeline' entails. Low risk for a zero-param doc fetcher, but disclosure is thin.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, no filler, with the action front-loaded and the guidance immediately after. Every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists so return values need not be explained, and a zero-param read-only doc fetch is simple. The description could still say what the playbook covers or how it relates to run_options/orchestrator_next_action, but it is adequate as written.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so there is nothing for the description to disambiguate; baseline 4 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource: return the AppFactory run playbook in markdown. That distinguishes it from the pipeline/orchestrator state tools, though it does not explicitly contrast itself with adjacent guidance siblings like run_options or orchestrator_preflight.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Read it before driving the pipeline' gives explicit timing for when to call it. There is no statement of when not to use it or which sibling supersedes it for guidance, so it stops short of full routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
preview_briefC
Write marketing/preview/BRIEF.md — the HyperFrames brief (destination: app-store-preview) for this
app's App Preview, from app.spec.json (size, fps, duration, poster, locales and shares), and create
marketing/preview/recordings//. Then record the real app (Scripts/sim_store_prep.sh, then
xcrun simctl io <udid> recordVideo) and author the video with the hyperframes skill: real
footage only, motion graphics frame and highlight it.
| Name | Required | Description | Default |
|---|---|---|---|
| app_dir | Yes | ||
| message | No | ||
| overwrite | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does disclose meaningful behavior: a specific output path, directory creation, running Scripts/sim_store_prep.sh, invoking xcrun simctl recordVideo, and the 'real footage only' constraint. It omits idempotency/overwrite behavior, whether the recording step blocks or times out, and what happens if the brief already exists.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the primary artifact (BRIEF.md path) before the follow-on steps, which is good, but it is a single dense run-on sentence mixing file paths, shell commands, and skill invocations. Sized plausibly for the complexity yet hard to scan.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Output schema exists so return values need not be described, but for a multi-step tool with no annotations and 0% parameter coverage the description still leaves gaps: no permission/prerequisite context, no failure behavior, and no guidance on how the record and author steps interact with the schema defaults.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and none of the three parameters (app_dir, message, overwrite) are named or explained in the description. 'This app's App Preview' loosely implies app_dir, but message and overwrite are entirely undefined, leaving the agent dependent on an undocumented schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Names a concrete verb and artifacts: writes marketing/preview/BRIEF.md, creates recordings/<locale>/, records the app, and authors video via the hyperframes skill. This clearly separates it from siblings like preview_check or preview_upload. The only weakness is that it bundles multiple distinct actions into one tool, so the single 'what it does' is a pipeline rather than one operation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No statement of when to use this versus siblings such as preview_check, preview_review_sheets, or design_record_upload, and no prerequisites or exclusions. The pipeline position ('then record... and author...') is implied but never framed as guidance for an agent choosing between tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
preview_checkC
Offline readiness of the App Preview set: BRIEF, real recordings per source locale, one preview per spec.preview locale (or a documented share), every file 886x1920 / <=30 fps / H.264 / 15–30 s / <=500 MB / stereo AAC or silent (ffprobe), plus the self-review status and deterministic rendering.
| Name | Required | Description | Default |
|---|---|---|---|
| app_dir | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, and it does disclose real behavioral substance: the precise validation rules (dimensions, fps, codec, duration, size, audio via ffprobe), self-review status, and deterministic rendering. It still omits pass/fail semantics, whether it mutates state, and what a failure looks like. Reasonable disclosure of what is checked, but incomplete on operational behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose is front-loaded, but the description is one dense run-on sentence packing seven+ technical constraints into a single colon-list. It is information-dense but hard to parse, and some of the specification detail arguably belongs in structured fields rather than prose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained, and the description thoroughly covers what is validated. What remains missing for a 1-param check tool is any explanation of the input and of the operating context in which it should be invoked.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The single parameter app_dir has 0% schema description coverage, and the description never explains it. Given low coverage, the description is expected to compensate, and it does not — the agent cannot learn what app_dir should point to from either source.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action (checking offline readiness) on a specific resource (the App Preview set) and enumerates the exact conditions checked. An agent can tell it validates preview videos, distinguishing it informally from preview_brief, preview_review_*, and preview_upload. However, it never explicitly contrasts itself with those siblings, keeping it short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance, no prerequisites, and no alternatives named. The only hint of usage context is the word 'Offline' and 'readiness', which the agent must infer. Nothing tells the agent when to run this check versus preview_review_log or preview_brief.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
preview_review_logA
Log one self-review round of a final preview (file relative to the app): scores 1–10 for hook, readability, motion, variety, brand, music; while any is under 8, the 3 worst problems with timestamps [{t, issue}]. Fix, re-render, re-review until every score is 8+ on the exact final file (MD5-matched) — preview_upload and the gate refuse otherwise.
| Name | Required | Description | Default |
|---|---|---|---|
| file | Yes | ||
| scores | Yes | ||
| app_dir | Yes | ||
| problems | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does disclose meaningful behavior: the 1-10 scoring scale per dimension, the conditional problem report, and the hard gate that refuses upload unless all scores are 8+ on the MD5-matched final file. It stops short of describing retry limits or what a rejected log returns, but the workflow constraints are unusually well surfaced.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
It is a single dense block that front-loads the core logging action before adding workflow and gating detail, with no filler sentences. The run-on structure and abrupt parentheticals reduce readability slightly but every clause carries information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained, and the description covers the nested scores/problems structure plus the gate conditions an agent needs to act correctly. Given the tool's complexity and zero schema description coverage, it is nearly complete, with only app_dir semantics and failure behavior left implicit.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate, and it does: 'file relative to the app' explains the file path convention, the scores map each key to hook/readability/motion/variety/brand/music with a 1-10 range, and problems is defined as the 3 worst issues with [{t, issue}] timestamps. Only app_dir is left implicit, keeping it below a 5.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a concrete verb and resource (log one self-review round of a final preview) and enumerates the score dimensions, so the agent knows exactly what is being recorded. It also names the related sibling preview_upload, aiding differentiation. The phrasing is somewhat dense and buries the logging action inside a workflow narrative, keeping it short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It conveys the workflow context ('Fix, re-render, re-review until every score is 8+') and hints at the gating relationship with preview_upload, which implies when the tool belongs in the cycle. However, it does not explicitly contrast with the other review siblings (preview_check, preview_review_sheets, preview_brief) or state when not to use it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
preview_review_sheetsA
Self-review material for every final preview (ffmpeg): a contact sheet (fps=2, 270 px, 6x5), a phone-size sheet (fps=1, 360 px, 5x3) and a 12-frame strip around the fastest transition, in marketing/preview/review//. Open them and score with preview_review_log.
| Name | Required | Description | Default |
|---|---|---|---|
| app_dir | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does reasonably well: it discloses that it generates files with specific ffmpeg parameters (fps, px, grid dimensions) and writes to marketing/preview/review/<locale>/. It does not mention prerequisites (a final preview must exist) or whether it overwrites existing sheets, leaving some behavioral gaps.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The core purpose is front-loaded and the verbose parenthetical specs (fps=2, 270 px, 6x5) earn their place by defining the artifacts. It is dense and slightly run-on, but no sentence is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists so return values need not be explained, and the description adequately covers the generated artifacts and their location. It is short of a 5 only because prerequisites and the unexplained app_dir leave an agent with minor unanswered questions.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the sole required parameter app_dir is undocumented in both the schema and the description. The mention of an output path does not explain what app_dir is or how it is used, so the description fails to compensate for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names the concrete artifacts produced (contact sheet, phone-size sheet, 12-frame transition strip), the toolchain (ffmpeg), and the output directory. It is clearly distinguishable from preview_brief/preview_check/preview_upload, though no explicit verb heads the sentence, so it reads as a noun phrase rather than a crisp 'Generate X' statement.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It routes the agent to a downstream step ('Open them and score with preview_review_log'), which is useful sequencing, but it gives no explicit when-to-use or when-not-to-use versus the sibling preview_check/preview_brief tools. Usage is implied rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
preview_uploadA
Upload fastlane/app_previews// to the editable version's IPHONE_67 preview sets through the asc CLI (shares from spec.preview.shared resolved, poster time code from the spec, MD5-idempotent). dry_run=True (default) sends nothing; a live upload needs human approval; refuses while preview_check has problems.
| Name | Required | Description | Default |
|---|---|---|---|
| app_dir | Yes | ||
| dry_run | No | ||
| locales | No | ||
| replace | No | ||
| approval_id | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations to lean on, the description carries the burden well: it discloses the dry-run default, the human-approval requirement for live writes, the refusal condition, and MD5-idempotency (safe re-runs). It stops short of explaining what `replace` destroys or why approval_id is needed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two dense sentences, front-loaded with the action and target path, then constrained by mode/approval/refusal clauses. Dense jargon (spec.preview.shared, MD5-idempotent) is unglossed but every clause adds operational value.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists so return values need not be described, and the mutation's safety profile (dry-run, approval gate, refusal, idempotency) is fully covered without annotations. The main remaining gap is the undocumented `replace` parameter.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across 5 parameters, so the description must compensate. It meaningfully documents dry_run (default true) and implies locales and approval_id via the <locale> path and human-approval requirement, but never explains `replace` or the approval_id semantics.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Upload) plus a precise resource path (fastlane/app_previews/<locale>/ → the editable version's IPHONE_67 preview sets) and the mechanism (asc CLI). An agent can distinguish this from sibling preview tools without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Clearly sets preconditions and modes: dry_run=True sends nothing, a live upload requires human approval, and it refuses while preview_check has problems. It routes through preview_check as a gate, though it does not explicitly name a competing upload alternative.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
pricing_unit_economicsB
Unit economics from the MEASURED AI cost per call (latest backend/eval/results for spec.ai.model, or explicit cost_photo/cost_text) × usage profiles up to the daily caps, per product; writes the generated blocks of store/pricing.md (decisions, unit economics, ASC/RC layout, local prices, anchor).
| Name | Required | Description | Default |
|---|---|---|---|
| app_dir | Yes | ||
| apple_cut | No | ||
| cost_text | No | ||
| cost_photo | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden, and it does disclose the side effect (writes generated blocks into store/pricing.md) and the data source (latest backend/eval/results). It does not say whether existing pricing.md content is overwritten or merged, nor any permission or prerequisite needs, so coverage is partial.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
It is a single dense run-on sentence packed with parentheticals and internal jargon (spec.ai.model, ASC/RC layout), so the core action is front-loaded but the parse cost is high. Information density is good; readability and structure are weak.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be described, and the write target is named. But with zero schema coverage and no annotations, an agent still lacks the meaning of apple_cut, the overwrite semantics of store/pricing.md, and any guidance on choosing this over aso_unit_economics.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and 4 parameters exist, so the description must compensate. It explains the role of cost_photo/cost_text as explicit cost overrides versus the measured default, which is genuinely useful, but apple_cut and app_dir are never explained in the text.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific computation (unit economics from measured AI cost per call times usage profiles) and a specific output artifact (generated blocks of store/pricing.md), so the agent knows what it does. However, it never distinguishes itself from the sibling aso_unit_economics, which sounds like an overlapping concern, so it falls short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use, when-not-to-use, or alternative guidance is given. The clause 'or explicit cost_photo/cost_text' hints at input selection but not at task selection, and the presence of aso_unit_economics makes the missing routing guidance a real gap.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
revenuecat_setupA
Idempotently set up the RC v2 project: entitlement(premium)+4 sub attach+SDK key+secret.
A human creates the RC project + links ASC + provides the v2 key; the rest is automatic. Preflight returns NEEDS_HUMAN if there is no project. env_suffix for apps sharing one Supabase (e.g. _MYAPP). LIVE RevenueCat writes: human approval.
| Name | Required | Description | Default |
|---|---|---|---|
| v2_key | Yes | ||
| bundle_id | Yes | ||
| env_suffix | No | ||
| approval_id | No | ||
| set_secrets | No | ||
| supabase_ref | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does disclose real behavior: idempotency, the human-created project requirement, NEEDS_HUMAN preflight signaling, and human approval for LIVE writes. It omits failure modes, permission specifics, and rate/limit behavior, but the disclosed behaviors are substantive and non-obvious.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the core action and keeps total length tight, with each fragment carrying information. The telegraphic, fragment-heavy style is dense and slightly choppy but wastes nothing.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be described. However, for a complex 6-param setup tool with 0% schema coverage and no annotations, the description leaves key parameter meanings and some behavioral details unaddressed, making it only partially complete for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across 6 params, so the description must compensate. It explains env_suffix meaningfully (with an example, '_MYAPP', for apps sharing one Supabase) and implies v2_key comes from the human, but bundle_id, supabase_ref, approval_id, and set_secrets are left entirely undocumented, so most parameters remain opaque.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Idempotently set up the RC v2 project') and enumerates the concrete artifacts it creates (entitlement premium, 4 sub attach, SDK key, secret). An agent can distinguish this from setup_status/setup_services, though it does not explicitly name a sibling alternative.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Clearly states prerequisites ('A human creates the RC project + links ASC + provides the v2 key') and the precondition behavior ('Preflight returns NEEDS_HUMAN if there is no project'), plus 'LIVE RevenueCat writes: human approval.' This tells the agent when the tool can proceed versus when human action is needed, though it names no explicit alternative tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
run_optionsA
Every optional part of a run (services, monetization, trial, paywalls, onboarding quiz, mascot, languages, store locales, screenshots, preview video, CPPs, analytics, ratings, HealthKit, Sign in with Apple, GitHub issues, e2e smoke tests…) with its question, kind, choices, default and current value. Ask the user each question one at a time (or as a compact checklist if the agent supports multi-select), then call run_options_save. Call this BEFORE any other work when the user says run.
| Name | Required | Description | Default |
|---|---|---|---|
| app_dir | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations exist, so the description carries the full burden. It discloses the return content and the interactive pattern (one question at a time, or checklist), implicitly signaling a non-mutating read since persistence lives in run_options_save. However, it never states permissions/auth needs, whether the read is idempotent, or any failure modes, so it is only partially complete.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The content is split into a payload sentence and an action sentence, with the actionable instruction placed last but clearly marked. The parenthetical enumeration is long but bounded by an ellipsis, so it stays efficient overall.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so explaining return values is optional, yet the description still summarizes them usefully. The workflow (ask then save) is complete for an orchestrator step; the only real hole is the undocumentated app_dir parameter.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There is one parameter (app_dir) with 0% schema description coverage, and the description never mentions it or explains when to supply it versus relying on the default. With coverage this low, the description should have compensated but does not.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names the resource (every optional part of a run) and enumerates the returned fields (question, kind, choices, default, current value), so an agent can tell it is a question-bank reader rather than a writer. The verb is implied rather than stated, and it is not explicitly contrasted against siblings, keeping it out of the 5 band.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It states a clear trigger ('when the user says run'), an ordering constraint ('Call this BEFORE any other work'), a workflow step ('Ask the user each question one at a time'), and names the follow-up sibling run_options_save. That is explicit when-to-use plus a named next tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
run_options_saveA
Save the user's answers ({option id: value}, every id from run_options). Validates types, choices and dependencies; services go to config.toml, the rest to app.spec.json (app_dir) or to a pending file app_scaffold applies. Records options_confirmed.
| Name | Required | Description | Default |
|---|---|---|---|
| answers | Yes | ||
| app_dir | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does meaningfully disclose side effects: it validates types, choices and dependencies, routes services to config.toml and the rest to app.spec.json or a pending file, and records options_confirmed. It still omits error behavior, idempotency, and permission requirements, so it stops short of a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The core action is front-loaded in the first clause and the remaining clauses are information-dense rather than filler. It is a single run-on sentence with semicolons, which slightly hurts readability but wastes little space.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained. For a mutation tool with no annotations, the description covers payload format, validation, and write destinations, leaving only error/retry semantics unaddressed, which is a minor gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, and it mostly does: it clarifies that 'answers' is an {option id: value} map keyed by ids from run_options, and that 'app_dir' governs where non-service results land (app.spec.json). It does not fully spell out the app_dir null/default behavior, so it is not a 5.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Save the user's answers') and pins the expected payload format ('{option id: value}, every id from run_options'). Referencing run_options lets an agent distinguish this from the sibling that produces the ids in the first place.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The mention of 'every id from run_options' implies this is the follow-up to run_options, but there is no explicit when-to-use/when-not statement, no exclusions, and no alternatives named. Usage is implied rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
screenshot_apply_layoutC
Merge the Claude Design store layout (design/project/store_layout.json, the ST0N boards) into marketing/screenshots/config.json so the compositor renders the approved layout for every locale.
| Name | Required | Description | Default |
|---|---|---|---|
| app_dir | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It implies a mutation of marketing/screenshots/config.json but never says whether the merge overwrites existing config, is idempotent, requires prior design generation, or what happens on conflict. The file paths add context but not the behavioral traits a mutation tool needs.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with the action verb first and no filler. It is appropriately sized, though the parenthetical file-path detail is slightly dense.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists so return values needn't be described, and the tool has only one parameter, keeping complexity low. However, the mutation semantics, prerequisites, and the meaning of app_dir are all missing, so the definition is only minimally complete for an agent to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the sole required parameter (app_dir) is never mentioned in the description. The description does not compensate with any hint about what app_dir should contain (e.g., repo root vs app subdirectory), leaving the parameter fully undocumented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ("Merge") and resource (the Claude Design store layout / ST0N boards into marketing/screenshots/config.json), which is more concrete than the sibling screenshot_* tools that capture or build. It does not, however, explicitly distinguish itself from close siblings like screenshot_sync or screenshot_build_all, so an agent must infer the boundary.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains the outcome ("so the compositor renders the approved layout for every locale") but gives no when-to-use/when-not-to-use guidance and never names an alternative among the many screenshot_* siblings. An agent has no explicit routing signal for choosing this over screenshot_sync or screenshot_build_all.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
screenshot_brandA
Branded App Store screenshots via the node compositor, rendering the Claude Design store layout (run screenshot_apply_layout first; screenshot_build_all does both).
| Name | Required | Description | Default |
|---|---|---|---|
| marketing_dir | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries full burden. It discloses the rendering mechanism and a hard prerequisite, which is meaningful, but says nothing about whether existing files are overwritten, what permissions or environment the node compositor needs, or failure modes when the layout step was skipped.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One sentence, front-loaded with the outcome, followed by a parenthetical that carries the mechanism, prerequisite, and alternative with zero filler. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The output schema covers return values and the prerequisite chain is stated, so the tool is callable in broad strokes. However, for a tool whose only parameter is completely undocumented, the definition leaves a real gap about what marketing_dir must point at.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for the single parameter, and the description never mentions marketing_dir — not its expected path, structure, or contents. The one piece of input documentation an agent needs is absent from both schema and description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Branded App Store screenshots') plus the rendering mechanism ('node compositor', 'Claude Design store layout'). It explicitly names two siblings (screenshot_apply_layout, screenshot_build_all) and how they relate, so an agent can distinguish it without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit prerequisite ordering ('run screenshot_apply_layout first') and names the shortcut alternative ('screenshot_build_all does both'), which is exactly the routing information an agent needs. It stops short of stating when NOT to use this tool or what happens if the prerequisite is skipped.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
screenshot_build_allB
MANDATORY turnkey screenshot step (UI-test style, all apps): generate a sample image → build+install → localized captions from the catalog → capture 32 languages × 6 rich screens (onboarding/create/ result/gallery/paywall/settings) → brand → sync to fastlane/screenshots. sample_prompt: a sample prompt for what the app produces (fills Result/Gallery with real output).
| Name | Required | Description | Default |
|---|---|---|---|
| udid | Yes | ||
| scheme | Yes | ||
| app_dir | Yes | ||
| locales | No | ||
| project | Yes | ||
| bundle_id | Yes | ||
| sample_prompt | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does substantial work: it discloses a multi-stage pipeline, that it builds and installs, the scale (32 languages × 6 screens), and the output location (fastlane/screenshots). It still omits permissions/prereqs (e.g. a booted simulator matching udid) and whether files are overwritten, keeping it short of a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the purpose and mandatory framing, with the pipeline conveyed efficiently via arrow notation. Dense but each element describes a real stage; only sample_prompt's trailing gloss is slightly tacked-on.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return-value explanation is rightly omitted, and the pipeline is well described. But for a complex 7-param tool with zero annotation coverage and zero schema-description coverage, leaving six parameters undocumented is a meaningful completeness gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across 7 params, so the description must compensate, but it explains only sample_prompt ('a sample prompt for what the app produces') and indirectly hints at locales via '32 languages'. The other five parameters (app_dir, project, scheme, bundle_id, udid) get no meaning at all, so most semantics are absent.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Uses a specific verb+resource and enumerates the full pipeline (generate sample → build+install → localized captions → capture 32 languages × 6 screens → brand → sync to fastlane/screenshots), so the scope is unmistakable. However, it never explicitly names the sibling component tools (screenshot_capture, screenshot_brand, screenshot_sync) it aggregates, so the agent must infer this is the turnkey aggregate.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'MANDATORY turnkey ... step (UI-test style, all apps)' signals this is the always-use screenshot step, implying when to invoke it. But there is no explicit when-not guidance and no routing to or away from the granular sibling tools, leaving the choice entirely to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
screenshot_captureB
Raw capture from the booted simulator → marketing/raw//.png.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | ||
| udid | Yes | ||
| locale | Yes | ||
| marketing_dir | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, and it does add real signal: the capture is raw (unprocessed) and lands at a fixed path, implying a filesystem write. It says nothing about overwrite behavior for an existing file, error modes when no booted device exists, or whether the operation is safe to repeat, which is meaningful missing context on a write-to-disk tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single short sentence that front-loads the action and ends with the resulting path; there is no filler. It is arguably too terse for four required parameters, but nothing in it is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained, but for a four-required-parameter tool with 0% schema coverage and zero annotations the description leaves too much unstated: what marketing_dir and udid mean, how the file is named on collision, and when to prefer this over the other screenshot tools. It is not sufficient on its own to invoke correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across four required parameters, so the description must compensate. The path template maps <locale> and <name> to filename segments and implies marketing_dir is the root of that tree, but udid is only obliquely referenced through 'booted simulator' and never tied to a specific argument.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a concrete action (raw capture) and its source (the booted simulator) plus the deterministic output path, so an agent can tell it captures a simulator screen to disk. It does not distinguish itself from near-neighbors like build_screenshot, screenshot_brand, or screenshot_generate_sample, leaving that differentiation to the name alone.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'booted simulator' implies a precondition, but there is no explicit when-to-use guidance, no exclusion for the branding/layout screenshot siblings, and no statement of what happens if the simulator is not booted. The reader must infer the entire usage context from the name and the path template.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
screenshot_generate_sampleC
Generate a sample image with fal.ai → Resources/sample_headshot.jpg (enriches Result/Gallery).
| Name | Required | Description | Default |
|---|---|---|---|
| prompt | Yes | ||
| app_dir | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It discloses that generation goes through the external fal.ai service (useful), but says nothing about auth requirements, cost/latency, failure modes, or whether it overwrites an existing file at the stated path.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single sentence with action front-loaded and zero padding, which is good. However, the arrow/telegraphic notation is cryptic and squeezes three distinct facts (provider, path, benefit) into a fragment, trading clarity for brevity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained, and the destination path is stated. But with no annotations and both parameters undocumented, the definition leaves key operational context (auth, overwrite behavior, required inputs) for an agent to guess.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for two required parameters. The description implies a prompt drives generation but never explains the format of 'prompt' or what 'app_dir' scopes, so it fails to compensate for the schema gap that the rubric requires when coverage is under 50%.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Generate'), resource ('sample image'), external provider ('fal.ai'), and concrete output path ('Resources/sample_headshot.jpg'), plus a stated benefit ('enriches Result/Gallery'). It is distinguishable from siblings like screenshot_capture or screenshot_brand by the word 'sample', though it never explicitly contrasts with them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The parenthetical '(enriches Result/Gallery)' hints at why one might call it, but there is no explicit when-to-use, no prerequisites, and no reference to any alternative among the many screenshot_* tools. The agent must infer the trigger condition.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
screenshot_onboarding_heroesA
LEGACY, OPT-IN ONLY: fal hero images per onboarding step (onb_step0..N). Onboarding visuals are designed in Claude Design (design/screens.json boards); use this only for a legacy hero-image onboarding.
| Name | Required | Description | Default |
|---|---|---|---|
| count | No | ||
| app_dir | Yes | ||
| concept | Yes | ||
| allow_external_generator | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, and it does disclose the legacy/opt-in nature of the tool. However, it says nothing about the external 'fal' generator being invoked, whether it makes network/costly calls, required credentials, or side effects of writing hero assets. The allow_external_generator parameter hints at external generation but this is never explained.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the 'LEGACY, OPT-IN ONLY' warning so the most decision-relevant caveat comes first. Efficient, with only mild redundancy between the opening label and the closing 'legacy' clause.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be described. Given that, the routing purpose is covered, but with zero annotations and 0% parameter coverage the description leaves behavioral and parameter gaps an agent would need to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across 4 parameters, so the description must compensate and largely does not. The 'onb_step0..N' phrasing loosely maps to the count default of 11 steps, but app_dir, concept, and allow_external_generator are left entirely undocumented in both schema and description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific resource (hero images per onboarding step, onb_step0..N) and implies generation via 'fal hero images'. It distinguishes itself from the design_* sibling tools by noting that onboarding visuals normally come from Claude Design. The verb is somewhat implicit ('fal hero images' rather than 'generate'), keeping it just below a crisp 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly sets scope with 'LEGACY, OPT-IN ONLY' and 'use this only for a legacy hero-image onboarding', while naming the preferred alternative (onboarding visuals are designed in Claude Design via design/screens.json boards). When-to-use, when-not-to-use, and the alternative are all stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
screenshot_syncD
branded → fastlane/screenshots//.
| Name | Required | Description | Default |
|---|---|---|---|
| marketing_dir | Yes | ||
| fastlane_screenshots_dir | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It does not disclose whether syncing overwrites existing files, requires specific permissions, handles missing locales, or what the mutation effects are.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is very short, but its brevity reflects under-specification rather than clean conciseness. It is a sentence fragment that does not front-load a clear action or resource.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a sync tool with no annotations, 0% schema description coverage, and two required undocumented parameters, the description is far too incomplete. Even with an output schema, an agent lacks enough context to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and both required parameters are undocumented. The description does not name or explain marketing_dir or fastlane_screenshots_dir, nor how they map to the path fragment shown.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives only a cryptic fragment: 'branded → fastlane/screenshots/<locale>/.' It implies movement of branded screenshots into fastlane's locale directory, but does not state a clear verb or fully distinguish the tool from siblings like screenshot_brand or deliver_screenshots.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this tool, when not to use it, or which sibling tools are alternatives. The agent must infer entirely from the tool name and schema.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
setup_approvalsA
Set human approvals for live writes: 'required' (recommended). 'off' is refused here; only the human can switch it off, on the setup_credentials page.
| Name | Required | Description | Default |
|---|---|---|---|
| mode | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the burden, and it does disclose a real behavioral constraint: 'off' cannot be set here because switching it off requires a human on a separate page. That is meaningful beyond a simple setter description, though it does not describe permissions, persistence, or what the tool returns.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences, front-loaded with the purpose and the recommended value, then the restriction. No filler and no repetition of the tool name.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, the description need not explain return values, and it covers the key behavioral constraint. The remaining hole is the incomplete enumeration of valid 'mode' values, which an agent needs in order to call the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
One parameter ('mode') at 0% schema coverage, so the description must compensate. It surfaces two accepted values ('required', 'off') and flags the recommended one, but never states that these are the only valid values, nor whether the 'off' value is rejected with an error here – leaving genuine ambiguity about the accepted input space.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource: configuring human approval mode for live writes. It is not easily confused with setup_status or setup_credentials since it names the exact setting and its recommended value. However it does not name a sibling tool as the alternative, only the sibling page (setup_credentials).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear context ('required' is recommended) and an explicit restriction ('off' is refused here), pointing to where the alternative action lives (setup_credentials page). What remains implicit is when in the onboarding flow an agent should call this versus setup_set/setup_services.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
setup_credentialsA
Open a local browser page (127.0.0.1, random port, one-time token) where the USER types the credentials of the given (default: all enabled) services. Secrets never reach you. Returns at once with the url; tell the user to fill the form and Save, then call setup_status().
| Name | Required | Description | Default |
|---|---|---|---|
| services | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does so well: it discloses the local host/port, the one-time token, the security model ('Secrets never reach you'), and the async return behavior ('Returns at once with the url'). These are exactly the behavioral traits an agent needs before calling.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the action, parentheticals carry mechanism details without bloat, and the final sentence gives the next step. Dense but every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Output schema exists so return values need no explanation, and the description still notes it returns the URL. For a security-sensitive, multi-step setup flow it covers the model and the handoff to setup_status; only the services-parameter detail is thin.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate. It documents the sole parameter's default semantics ('default: all enabled'), which is the key behavior, but adds no format details for the services list itself. Baseline 3 given the single param and partial compensation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific action and mechanism: opens a local browser page for the USER to type service credentials. The purpose is unambiguous. It does not explicitly differentiate itself from nearby siblings like setup_services or setup_set, but the credential-entry framing is distinctive enough.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Clear context for use and an explicit follow-up ('tell the user to fill the form and Save, then call setup_status()'). It names the next step in the workflow but does not say when this tool is the wrong choice relative to setup_services/setup_set.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
setup_servicesB
Turn services on or off (research is always on). Ask the user which ones they want first. Returns setup_status.
| Name | Required | Description | Default |
|---|---|---|---|
| enable | No | ||
| disable | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It discloses two useful behavioral facts: 'research is always on' (an immovable constraint) and the requirement to consult the user first. It does not say what happens to unspecified services, whether toggling is reversible, or what permissions are needed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences with the action front-loaded and no filler. Minor waste in 'Returns setup_status', which restates the declared output schema, but the size is well-matched to the tool's simplicity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained, and the tool has only two optional parameters. The critical missing piece is the set of valid service names — nothing in the schema, annotations, or description tells the agent what values are legal.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate and largely does not. It conveys the on/off polarity that maps to 'enable'/'disable', but never names the parameters nor enumerates the valid service identifiers accepted by those arrays, leaving the agent unable to construct a correct call.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The verb phrase 'Turn services on or off' states an action on a resource, and the parenthetical '(research is always on)' adds a scope detail. But 'services' is never enumerated, and the tool sits among a dense cluster of similarly named setup_* siblings (setup_status, setup_set, setup_approvals, setup_credentials) with no differentiation from them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Ask the user which ones they want first' is an actionable interaction instruction that tells the agent how to sequence the call. However, there is no when-to-use vs. alternatives guidance, no mention of prerequisites, and no indication of when this tool should not be used.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
setup_setA
Set ONE non-secret config key (e.g. asc_key_id, team_id, support_email). Secrets are refused: use setup_credentials so they never pass through the chat. An empty value clears the key.
| Name | Required | Description | Default |
|---|---|---|---|
| key | Yes | ||
| value | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are supplied, so the description carries the burden and does so reasonably: it discloses the secrets-refusal policy and the clear-on-empty-value behavior, both non-obvious traits. It does not state whether an existing key is overwritten, where the value persists, or whether auth is needed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, zero filler, with the core constraint ('ONE non-secret config key') and the critical exclusion (secrets) front-loaded. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need no explanation, and the description covers the constraining behaviors an agent needs before calling. The main residual gap is the relationship to the sibling config_set and confirmation of overwrite behavior.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% and neither parameter is documented in the schema, so the description must compensate. It does: it gives example key values and defines the special semantics of 'value' (empty string clears the key) and 'key' (one at a time, non-secret only). It could be firmer about overwrite semantics for non-empty values.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Set ONE non-secret config key') with scope ('ONE'), and enumerates concrete key examples (asc_key_id, team_id, support_email). An agent immediately knows this mutates a single config entry rather than being a generic setup tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly tells the agent what NOT to use this for (secrets → setup_credentials, 'so they never pass through the chat') and what an empty value does. However it gives no guidance relative to the sibling 'config_set', which appears to serve a very similar purpose, leaving routing between the two ambiguous.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
setup_statusA
Setup state per service: enabled, keys set/missing (never values), missing tools with install commands,
approvals mode, and next steps. Call this first; nothing here needs a config file.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description must carry the full behavioral burden. It discloses that secret values are never shown and that no config file is required, but it does not explicitly state that the tool is read-only or non-mutating, nor does it describe permissions or side effects. This is a notable gap without annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences: the first front-loads the output fields, the second gives the imperative usage note. No filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so the description need not explain return values in full, and it still summarizes the key outputs. For a zero-parameter tool it covers enough to invoke correctly, though the lack of an explicit read-only statement is a minor gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Zero parameters, so the schema baseline is 4. The description does not need to add parameter details, and it appropriately stays focused on output and usage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States the resource ('setup state per service') and enumerates the exact fields it reports (enabled, keys set/missing, missing tools with install commands, approvals mode, next steps), which differentiates it from siblings like setup_services and config_doctor. It lacks an explicit verb such as 'Get', but the scope is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides a direct usage instruction: 'Call this first' and clarifies that no config file is needed. It does not name alternative tools or specify when not to use it, so it stops short of the highest level.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
signing_create_profileC
Create an IOS_APP_STORE provisioning profile + write it to the standard locations.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | AppFactory AppStore | |
| cert_id | Yes | ||
| bundle_id | Yes | ||
| approval_id | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It does disclose one useful trait - it writes the profile to standard disk locations - but says nothing about overwrite/conflict behavior, required Apple credentials, or whether it mutates an existing profile.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single efficient sentence with the action front-loaded and no filler. It is appropriately sized, though the 'plus write it to standard locations' clause could be integrated more precisely.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with no annotations and 0% parameter coverage, the description omits prerequisites, side-effect details, and any routing relative to its sibling. An output schema exists so return values need not be explained, but the input and behavioral gaps leave the definition under-specified.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across four parameters (name, cert_id, bundle_id, approval_id), and the description adds no parameter meaning whatsoever. With zero coverage and no compensation in the text, an agent cannot know what cert_id or approval_id expect.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb (Create) and resource (IOS_APP_STORE provisioning profile), plus the side effect of writing it to standard locations. It is clear on its own, but it never distinguishes itself from the related sibling signing_setup_distribution, which likely also deals with provisioning profiles.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this versus signing_setup_distribution, nor any stated prerequisites (e.g. that a certificate and bundle id must already exist). The agent must infer usage entirely from the name.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
signing_setup_distributionC
Create a distribution cert + install it into a temporary keychain (WWDR included). {cert_id, identity, keychain}.
| Name | Required | Description | Default |
|---|---|---|---|
| approval_id | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It discloses that the keychain is temporary and that WWDR is included, but omits authentication requirements, side effects, cleanup behavior, and whether the certificate is persisted.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The first sentence is concise and front-loaded with the core action. The trailing brace list is cryptic and does not clearly earn its place, weakening structure.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained, but the description still omits usage context and approval_id semantics. With no annotations and 0% schema description coverage, the definition is not complete enough for reliable invocation in a multi-step signing workflow.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The sole input parameter, approval_id, is not explained in the description, and schema coverage is 0%. The brace list {cert_id, identity, keychain} does not map to the input schema and may confuse rather than clarify what the caller must provide.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action: create a distribution certificate and install it into a temporary keychain with WWDR included. It is clear what the tool does, but it does not explicitly distinguish itself from the sibling signing_create_profile or other setup tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance about when to use this tool versus alternatives like signing_create_profile or setup_credentials. The agent must infer usage entirely from the tool name and description.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
storekit_generateA
Write Resources/Configuration.storekit from the app's app.spec.json (group, levels, free-trial intro offers, trial-less offer product, en_US). Deterministic; run after any product change.
| Name | Required | Description | Default |
|---|---|---|---|
| app_dir | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations exist, so the description carries the burden. It usefully discloses that the operation is deterministic, derived from app.spec.json, and what content is emitted. However, it never states that the write overwrites any existing Configuration.storekit, nor anything about path resolution or failure behavior, which matters for a file-producing tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tightly packed sentences: the primary action and output path come first, followed by the determinism/timing note. No filler, no repetition of the tool name.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, return values need not be described, and the description covers source, destination, and generated contents. The remaining gap is overwrite/side-effect semantics on an existing StoreKit file, which is not addressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the single required parameter app_dir is documented nowhere. The description says the file comes from 'the app's app.spec.json' but never explains that app_dir locates that spec or how paths are resolved, so it fails to compensate for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Write) plus the exact output artifact (Resources/Configuration.storekit) and the input source (app.spec.json). It also enumerates what the file contains (group, levels, intro offers, offer product, en_US), so it is clearly distinguishable from the similarly named sibling storekit_parity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit timing guidance: 'run after any product change.' That is a clear context for invocation, but it never names an alternative tool or states when not to run it, so it stops short of full routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
storekit_parityA
Compare a Configuration.storekit with app.spec.json; returns every mismatch (price, period, level, intro offer, missing/extra product, group, locale). ok=true means in parity.
| Name | Required | Description | Default |
|---|---|---|---|
| app_dir | Yes | ||
| storekit_path | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It does disclose meaningful behavioral context: it returns every mismatch and that 'ok=true means in parity', which is useful output semantics beyond structured fields. However, it never states it is a read-only/non-mutating comparison, nor any auth or side-effect profile, leaving the safety burden partly unmet.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with the verb and both operands first, followed by the enumerated mismatch outputs and the parity flag. No filler; every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values needn't be explained, and the description appropriately avoids that. For a simple two-parameter read-only comparison, purpose and output semantics are covered; only the parameter meanings (app_dir, storekit_path) are left unaddressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for 2 parameters. The description references the files being compared (StoreKit config vs app.spec.json), which loosely maps to storekit_path, but the required 'app_dir' parameter and the optional/default-null nature of storekit_path are never explained. It does not compensate for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Compare'), the two resources ('Configuration.storekit' and 'app.spec.json'), and enumerates the exact mismatch categories checked. This clearly differentiates it from the sibling storekit_generate (which produces rather than validates StoreKit config).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by the description (run to verify parity between a StoreKit config and the app spec), but there is no explicit when-to-use/when-not guidance and no mention of adjacent tools like config_doctor, metadata_check, or storekit_generate. The agent must infer it fits a pre-submission validation step.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
store_setupA
Idempotent App Store Connect + RevenueCat setup from app.spec.json + listing.json.
mode: plan (offline) | check (live reads, simulated writes; reports CONFLICT) | apply (live writes; human
approval via appfactory approve <id>).
target: capabilities | asc | rc | all. Order: App ID capabilities FIRST, then ASC (grace period,
group + localizations, products + localizations, availability before prices, equalized USA prices,
price overrides, per-territory intro offers for trial products and none for offer products, age
rating, review contact, app info/version localizations + legal URLs, SKU check, Server
Notifications V2 URL), then RevenueCat (project, app, products, entitlement, offerings/packages,
targeting-rule placements; MCP plan when REST refuses). rc_apple_notification_url: the RevenueCat
dashboard's Apple Server-to-Server URL (stored in app outputs).
| Name | Required | Description | Default |
|---|---|---|---|
| mode | No | plan | |
| target | No | all | |
| app_dir | Yes | ||
| approval_id | No | ||
| rc_project_id | No | ||
| rc_apple_notification_url | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does so well: it declares idempotency, distinguishes simulated vs. live writes, discloses that check reports CONFLICT, and that apply requires human approval via 'appfactory approve <id>'. It also documents execution ordering dependencies (capabilities first, availability before prices) that materially affect safe invocation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The leading sentence is well front-loaded, but the second paragraph is a single dense run-on parenthetical enumerating every ASC/RC step, which hurts parseability. Much of that ordering detail is valuable but could be trimmed or structured as a list.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a high-complexity orchestrator with an output schema (so return values need not be explained), the description covers modes, targets, ordering, and the approval gate adequately. The main gaps are the undefined app_dir and rc_project_id semantics, which leave minor ambiguity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate for all 6 parameters. It well documents mode, target, and rc_apple_notification_url (including where to obtain it), and implies approval_id via the approve flow. But app_dir and rc_project_id are left undefined in both schema and description, so compensation is only partial.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and scope ('Idempotent App Store Connect + RevenueCat setup') plus the input artifacts (app.spec.json + listing.json). An agent can tell this is the full-store orchestration tool versus the granular asc_*/revenuecat_setup siblings. It does not explicitly name the alternatives it supersedes, keeping it just below a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The mode enumeration (plan offline / check live-reads-with-simulated-writes / apply live-writes) with the human-approval gate is effective when-to-use guidance, and target=capabilities|asc|rc|all routes scope. There is no explicit statement of when to choose this over sibling tools like setup_services or revenuecat_setup, so it stops short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
supabase_create_projectA
Create a new Supabase project (LIVE — provisions resources). Human approval.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | ||
| org_id | Yes | ||
| region | Yes | ||
| db_pass | Yes | ||
| approval_id | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, and it does disclose the two most critical traits: the operation is LIVE and provisions real resources (a mutating side effect with likely cost), and it requires human approval. It still omits cost implications, whether the action is reversible, and how approval_id ties into the call.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short clauses, front-loaded with the core action and immediately followed by the critical LIVE/provisioning and approval warnings. Every fragment earns its place with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained, and the description flags the live/provisioning nature. However, for a mutation tool with zero annotation coverage and 0% parameter documentation, it leaves notable gaps around cost, prerequisites, and parameter expectations.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the description gives no meaning for any of the 5 parameters (name, org_id, region, db_pass, approval_id); the schema only declares types. The description fails to compensate, leaving formats and constraints entirely unexplained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ("Create a new Supabase project") that is immediately distinguishable from the read-only siblings supabase_list_projects, supabase_list_orgs, and supabase_get_keys. An agent can tell what this tool does without inspecting the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
"Human approval" implies a gating prerequisite for invoking it and implies this is the provisioning step that must precede supabase_run_sql/supabase_set_secret, but no explicit when-to-use vs. alternative guidance is given. Usage is implied rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
supabase_get_keysB
Get the project's url + anon key (anon→app). The service_role key is never returned to the agent; it is saved to ~/.appfactory/supabase/.json (0600) for server-side use.
| Name | Required | Description | Default |
|---|---|---|---|
| ref | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It discloses that the service_role key is never returned and is saved to a file with 0600 permissions for server-side use, which is valuable behavioral context. It does not cover authentication or rate limits, but covers the key safety aspect.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences, front-loaded with the primary return value. No wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Output schema exists, so return format need not be explained. The description covers the main behavior but omits the meaning of the required 'ref' parameter and any usage context, leaving gaps for an agent with no annotations.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% for the single required parameter 'ref', and the description does not define it. The file path ~/.appfactory/supabase/<ref>.json hints that ref is a project reference, but this is not explicit.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Get) and resource (project's url + anon key), and clarifies what is not returned. It does not explicitly name a sibling alternative, but the resource is distinct enough from siblings like supabase_set_secret.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance or alternatives are provided. It states a security fact about the service_role key but does not say when the agent should call this tool versus supabase_list_projects or supabase_set_secret.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
supabase_list_orgsA
List Supabase organizations (org_id for project creation).
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. 'List' implies a safe read, but the description never confirms read-only behavior, auth requirements, or scoping. For a zero-parameter listing tool the exposure is low, so this is a modest but real gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler; the purpose is the first thing read and nothing is repeated from the schema or title.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With zero parameters and an output schema present, return values need not be described, and the description covers the essential dispatch intent. Only the absence of any auth or read-only confirmation keeps it short of full completeness.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes no parameters, so the baseline is 4. The description usefully attaches meaning to the returned identifier ('org_id for project creation'), which is the only variable an agent actually needs to understand.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('List Supabase organizations'), which separates it from the sibling 'supabase_list_projects'. The parenthetical clarifies the payload it produces (org_id), though it does not explicitly name the sibling it differs from.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The parenthetical '(org_id for project creation)' implies the primary use case — resolving an org_id before calling supabase_create_project — but it never states when to use this versus supabase_list_projects or any prerequisite/ordering constraint explicitly.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
supabase_list_projectsB
List Supabase projects (LIVE).
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. 'List' adequately signals a read-only, side-effect-free operation over zero parameters, and '(LIVE)' hints at live-API data rather than cached state, but nothing is said about authentication needs, scope of the listing, or rate limits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single short sentence with no filler, and the resource is front-loaded. It is arguably too terse, since the parenthetical '(LIVE)' is the only qualifier and its meaning is left unexplained.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is trivial (no parameters, boolean output schema presumably documenting the project list), so little description is needed and an output schema exists to cover the return shape. The only real gap is that '(LIVE)' is never unpacked, leaving the agent to guess what distinguishes live results.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so the baseline is 4 and there is no argument semantics the description could meaningfully add. The empty schema and 'additionalProperties: false' leave nothing ambiguous for an agent to misread.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (List) and resource (Supabase projects), so the agent immediately knows what the tool returns. The '(LIVE)' qualifier signals real remote projects rather than a local config, but it never explicitly distinguishes this from siblings like supabase_list_orgs or supabase_create_project.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use, when-not-to-use, or alternative guidance. An agent must infer from the name alone that this is the discovery entry point before supabase_get_keys or supabase_run_sql, and nothing in the description confirms ordering or prerequisites.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
supabase_run_sqlB
Run SQL on the project (schema/migration) — LIVE write, human approval. Destructive statements (DROP, TRUNCATE, ALTER … DROP, GRANT/REVOKE on auth, DELETE/UPDATE without WHERE) always need approval.
| Name | Required | Description | Default |
|---|---|---|---|
| ref | Yes | ||
| sql | Yes | ||
| approval_id | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Without annotations, the description carries the full burden and does well: it discloses it's a 'LIVE write' requiring 'human approval', and enumerates destructive statement patterns (DROP, TRUNCATE, ALTER...DROP, GRANT/REVOKE on auth, DELETE/UPDATE without WHERE) that need approval. This is valuable behavioral context. However, it doesn't clarify what happens with non-destructive statements (auto-approved?) or the approval_id flow.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the core action and the critical safety caveat. Efficient, though the destructive-statement parenthetical is dense and could be slightly better organized.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values needn't be explained. The description covers the key risk profile (write, approval needed for destructive ops), which is the most important context for this tool. But with 0% param coverage and no annotation safety profile, the missing parameter documentation and approval_id mechanics leave gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for all 3 parameters. It mentions no parameter semantics at all — not 'ref', not 'sql', not how 'approval_id' relates to the approval flow it describes. The approval concept is mentioned but never tied to the parameter that carries it.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource: 'Run SQL on the project (schema/migration)'. The purpose is clear, but the parenthetical '(schema/migration)' is the only hint at intended use case, and there's no explicit differentiation from any sibling. It's clear but not sharply distinguished from other setup/config tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage context ('schema/migration') and names approval requirements for destructive statements, but it gives no explicit when-to-use guidance or named alternatives. For a high-risk live SQL tool, this is merely implied usage rather than clear guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
supabase_set_secretB
Write an edge function secret (FAL_KEY etc.) — server-side, never enters the app. Human approval.
| Name | Required | Description | Default |
|---|---|---|---|
| ref | Yes | ||
| name | Yes | ||
| value | Yes | ||
| approval_id | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations, so the description must carry behavioral disclosure. It usefully states the secret is stored "server-side, never enters the app" and that "Human approval" is required, but omits whether an existing secret is overwritten, permission requirements, and error behavior for a mutation tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single tightly-written sentence with the action and its distinguishing behavior front-loaded. No filler, though the em-dash clause compresses critical info (approval requirement) into a fragment.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be described. However, for a write tool with no annotations and 0% param coverage, gaps around overwrite semantics and the approval_id workflow leave the agent under-informed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for 4 parameters, so the description must compensate. It only indirectly signals that name takes a key like FAL_KEY and that approval_id relates to human approval; ref, value, and the approval_id contract are left undocumented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ("Write") and resource ("edge function secret") with an example (FAL_KEY), which lets an agent distinguish it from read-oriented siblings like supabase_get_keys. Sibling differentiation is only implicit, but the write action is unmistakable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No explicit when-to-use guidance or alternatives (e.g., vs supabase_get_keys or other setup_credentials tools) are named. "Human approval" hints at a workflow prerequisite but doesn't explain when the agent should reach for this tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
team_briefC
Render docs/TEAM.md, docs/team/{ios,backend,store}.md, docs/onboarding-plan.md (from design/screens.json), store/aso-research.md and docs/CHECKLIST.md from the spec. sessions: {lead, ios, backend, store} names.
| Name | Required | Description | Default |
|---|---|---|---|
| repo | No | ||
| app_dir | Yes | ||
| sessions | No | ||
| overwrite | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are present, so the description bears the full behavioral burden. It says it renders files but omits whether it overwrites existing files, how conflicts are handled, what permissions are needed, or whether it is destructive—important for a file-writing tool with an overwrite parameter.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
It is a single dense sentence with no filler, and the rendering action and artifact list are front-loaded. The parenthetical source note and terse sessions notation are compact, though the long comma list is less scannable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be described. However, for a tool that writes multiple files with no annotations and 0% schema coverage, the description leaves critical input semantics and mutation behavior unspecified.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and there are four parameters. The description only partially explains sessions ('{lead, ios, backend, store} names') and says nothing about the required app_dir, repo, or overwrite parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States the specific verb 'Render' and enumerates the output files and source spec, so the agent knows it produces team/checklist/onboarding artifacts. It does not explicitly differentiate itself from sibling renderers like onboarding_plan or backend_render, so it stops short of a top score.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives no when-to-use, when-not-to-use, or alternative-tool guidance. The agent must infer from the artifact list that this is used when team/checklist docs need to be generated from a spec.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
testflight_shipA
End-to-end TestFlight: distribution signing → archive → App Store profiles (the app and every
extension in project.yml, exact bundle-id match, created right before export so Xcode's profile
sweep cannot delete them) → manual-signing export → asc builds upload. Refuses unless the
Maestro smoke flows passed (maestro_test; n/a when the e2e_smoke run option is off). No system keychain password
(temporary keychain). Never submits to App Store review — it only uploads a TestFlight build.
| Name | Required | Description | Default |
|---|---|---|---|
| scheme | Yes | ||
| app_dir | Yes | ||
| project | Yes | ||
| bundle_id | Yes | ||
| approval_id | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does substantial work: it discloses the refusal gate on Maestro results, the temporary-keychain/no-system-password behavior, profile-sweep timing rationale, and the explicit non-submission scope. It omits failure/partial-rollback behavior and side effects like created profiles, which keeps it short of a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The pipeline is front-loaded and compressed into arrow-linked clauses with no filler. It is dense and information-bearing, though the parenthetical about profile sweeps adds minor cognitive load that could be trimmed.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values needn't be re-explained, and the pipeline, refusal gate, and security posture are covered. The unexplained `approval_id` and unlabeled required params leave a modest gap for a 5-parameter orchestration tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across 5 parameters, and the description explains only fragments (project.yml, exact bundle-id matching, `scheme` implied by archive). The required `app_dir` and `project` are not clarified, and `approval_id` — likely a human-approval gate — is entirely unexplained despite being non-obvious. The description does little to compensate for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a precise verb chain (signing → archive → profiles → export → `asc builds upload`) for a specific resource (a TestFlight build), and its end-to-end scope clearly distinguishes it from single-step siblings like build_archive, build_export_ipa, and signing_setup_distribution.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives a concrete precondition (refuses unless Maestro smoke flows passed, referencing maestro_test and the e2e_smoke run option) and a clear negative boundary (never submits to App Store review). It doesn't name an explicit alternative tool for the manual 'step-by-step' path, but the when-to-use context is strong.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
125 tool updates
v0.1.1- First observed
ai_configure - First observed
ai_deploy_proxy - First observed
animation_fetch_recolor - First observed
app_inject_config - First observed
app_scaffold - First observed
app_sync_spec - First observed
asc_add_subscription_group_localization - First observed
asc_append_subscription_disclosure - First observed
asc_create_app - First observed
asc_create_bundle_id - First observed
asc_create_subscription - First observed
asc_create_subscription_group - First observed
asc_ensure_subscription_prices - First observed
asc_finalize_submission_requirements - First observed
asc_finalize_subscription - First observed
asc_get_app - First observed
asc_list_apps - First observed
asc_localize_group - First observed
asc_localize_subscription - First observed
asc_sbp_check - First observed
asc_submit_for_review - First observed
asc_token_check - First observed
aso_check_name - First observed
aso_competitor_iap - First observed
aso_complete - First observed
aso_fetch_competitors - First observed
aso_find_available_name - First observed
aso_niche_score - First observed
aso_run - First observed
aso_scaffold_outputs - First observed
aso_search_hints - First observed
aso_top_grossing - First observed
aso_unit_economics - First observed
aso_validate_metadata - First observed
backend_deploy - First observed
backend_render - First observed
build_archive - First observed
build_boot_sim - First observed
build_export_ipa - First observed
build_for_sim - First observed
build_list_simulators - First observed
build_screenshot - First observed
build_test - First observed
build_xcode_version - First observed
config_doctor - First observed
config_set - First observed
cpp_build_all - First observed
deliver_metadata - First observed
deliver_screenshots - First observed
deliver_screenshots_audit - First observed
deliver_subscription_review_screenshots - First observed
design_export_png - First observed
design_generate - First observed
design_record_upload - First observed
design_research_brief_template - First observed
design_research_check - First observed
design_research_collect - First observed
design_screens_skeleton - First observed
design_upload_status - First observed
env_doctor - First observed
firebase_setup - First observed
github_create_repo - First observed
github_issue_create - First observed
github_issues_bootstrap - First observed
github_push - First observed
growth_build_slideshows - First observed
icon_generate - First observed
icon_install - First observed
idea_evaluate - First observed
idea_harvest - First observed
legal_check - First observed
legal_render - First observed
legal_verify - First observed
localize_apply - First observed
maestro_test - First observed
mascot_assets - First observed
mascot_blink - First observed
metadata_check - First observed
metadata_export - First observed
metadata_listing_check - First observed
metadata_render_listing - First observed
onboarding_plan - First observed
orchestrator_needs_human - First observed
orchestrator_next_action - First observed
orchestrator_preflight - First observed
orchestrator_record_attempt - First observed
pipeline_mark - First observed
pipeline_next - First observed
pipeline_status - First observed
pipeline_validate - First observed
playbook - First observed
preview_brief - First observed
preview_check - First observed
preview_review_log - First observed
preview_review_sheets - First observed
preview_upload - First observed
pricing_unit_economics - First observed
revenuecat_setup - First observed
run_options - First observed
run_options_save - First observed
screenshot_apply_layout - First observed
screenshot_brand - First observed
screenshot_build_all - First observed
screenshot_capture - First observed
screenshot_generate_sample - First observed
screenshot_onboarding_heroes - First observed
screenshot_sync - First observed
setup_approvals - First observed
setup_credentials - First observed
setup_services - First observed
setup_set - First observed
setup_status - First observed
signing_create_profile - First observed
signing_setup_distribution - First observed
store_setup - First observed
storekit_generate - First observed
storekit_parity - First observed
supabase_create_project - First observed
supabase_get_keys - First observed
supabase_list_orgs - First observed
supabase_list_projects - First observed
supabase_run_sql - First observed
supabase_set_secret - First observed
team_brief - First observed
testflight_ship
TDQS
Scored across 125 tools
The server covers many pipeline stages with explicit markers (MANDATORY, OPT-IN, LEGACY, LIVE) and cross-references between descriptions, which helps distinguish most tools. However, several clusters overlap: pipeline_next vs orchestrator_next_action, pipeline_status vs orchestrator_preflight, setup_set vs config_set, metadata_check vs metadata_listing_check vs aso_validate_metadata, and deliver_* vs *_sync variants, creating genuine ambiguity for an agent.
Almost all tool names are snake_case with a domain prefix plus action (asc_*, aso_*, build_*, supabase_*, screenshot_*, design_*, etc.), which is highly consistent within groups. Minor deviations exist: some top-level tools lack a clear domain prefix (deliver_metadata, run_options, playbook) and a few are noun phrases (pricing_unit_economics, team_brief), but there is no camelCase or chaotic mixing.
125 tools is an extreme mismatch for a single server. Even a complex app-factory pipeline likely could consolidate many single-step tools (especially the numerous ESC/design/screenshot/preview stages) into fewer, more purposeful operations.
The surface is extensive, covering idea/ASO research, scaffolding, design, backend, signing, TestFlight, metadata, screenshots, previews, legal, GitHub, and App Store submission. Some gaps remain (no update/delete for several managed resources, limited post-submission monitoring, minimal GitHub PR support), but for the stated pre-submission factory pipeline it is largely complete.
Related MCP Connectors
Build and publish full-stack apps from your coding agent: models, rules, pages, auth, per-app MCP.
MCP server for AI agents to plan, verify, and deploy Cloudflare-native apps.
Official MCP server for subfeed.app — the cloud for agents. 15+ tools for AI agents to register, build, and deploy other agents. Zero human required. Start here: subfeed.app/skill.md
- app-managerOAuthapp.lance
App Store Connect operator for AI agents: icons, TestFlight builds, listings, IAP, rejection fixes.
Related MCP Servers
- AlicenseNot gradedqualityCmaintenanceMCP server that equips AI agents with dev workflow tools including GitHub project management, conventional commits, visual regression testing, Jira/Confluence integration, and a persistent memory knowledge graph.11 npmMIT
- AlicenseNot gradedqualityCmaintenanceMCP server that enables AI agents to run a deterministic orchestration loop with decomposition, subagent execution, and review feedback across multiple LLM backends.61MIT
- FlicenseAqualityCmaintenanceMCP server that enables AI assistants to run multi-step agent pipelines (e.g., Issue Analyst → Code Writer → Test Runner → PR Opener) from conversations, with support for Devin, shell, Python, and HTTP agents.7-
- AlicenseAqualityAmaintenanceMCP server that enables a coordinator AI agent to spawn, control, and supervise local coding agents with interactive gating for high-risk operations.1018 npmMIT