Skip to main content
Glama

Server Details

Design, save, and run outcome-aligned AI workflows and verifiers, with reliable image output.

Ownership verified
Status
Healthy
Uptime
59.5% over 43 days
OAuth
Works in Glama
Last Tested
Transport
Streamable HTTP · MCP 2025-11-25
URL
Repository
Goodeye-Labs/goodeye-cli
GitHub Stars
4

TDQS

A4/5.0

Scored across 80 tools

Disambiguation5/5

Each tool targets a distinct resource and action. Even closely related tools like get_skill_file vs get_skill_files are clearly separated (single vs batch fetch), and optimize_skill vs optimize_description vs audit_skill vs teach_skill have well-defined different purposes. No two tools appear to do the same thing.

Naming Consistency5/5

All 80 tools follow a consistent verb_noun snake_case pattern, e.g., accept_invitation, deploy_verifier, list_teams. There is no mixing of conventions, and the verbs (get/list/create/update/delete/archive/revoke/transfer) are predictable across resource types.

Tool Count2/5

80 tools is far beyond the typical 3-15 range and constitutes an extreme mismatch for an MCP server. Even though the platform covers a wide domain, this number places a heavy cognitive burden on the LLM to select the correct tool, and splitting into separate servers per concern (skills, templates, billing, etc.) would be more usable.

Completeness5/5

The tool surface provides comprehensive coverage of the platform's domain: full lifecycle (create/read/update/delete/archive/transfer/share) for skills, templates, verifiers, image generators, teams, invites, billing, and API keys. Search, listing, and lookup operations exist where needed, with no dead ends or obvious missing operations.

Available Tools

80 tools
accept_invitationAInspect

Accept a pending invitation addressed to the caller.

The invitation must be in pending state and not expired. Accepting a team_membership invitation creates the member row; accepting an ownership invitation transfers ownership immediately. Returns the result payload from the underlying action.

ParametersJSON Schema
NameRequiredDescriptionDefault
invitation_idYesThe invitation to act on, identified by its UUID.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.1/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations providing behavioral hints, the description fully discloses the effects: creating a member row for team_membership or transferring ownership for ownership invitations, and returning the result payload. This exceeds the typical burden for transparency.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise, with three sentences that front-load the main action and efficiently cover preconditions, behavioral details, and return value. No unnecessary words or redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's simplicity (one parameter, output schema exists), the description covers purpose, preconditions, effects, and return. It lacks info on error cases or permissions but is largely complete, warranting a score of 4.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema covers the single parameter with 100% description, including its purpose as a UUID. The description adds no additional semantic value beyond what the schema already provides, so the baseline score of 3 is correct.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool accepts a pending invitation, specifying the two types of invitations (team_membership and ownership) and their outcomes. While the purpose is distinct from siblings like cancel_invitation or decline_invitation, it does not explicitly differentiate itself, so a score of 4 is appropriate.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear context for usage: the invitation must be pending and not expired. However, it does not mention when to use alternatives (e.g., decline_invitation), which prevents a higher score. It achieves a 4 for clear usage context without exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

add_team_memberAInspect

Add a user to a team. Owner only. User accepts UUID, email, or handle.

ParametersJSON Schema
NameRequiredDescriptionDefault
team_identifierYesThe team to act on, identified by its UUID or handle.
user_identifierYesThe target user, identified by their UUID, email, or handle.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate a write operation (readOnlyHint=false). The description adds value by specifying the 'Owner only' constraint, which is critical for correct invocation. It does not detail potential side effects like duplicate handling, but the owner requirement is a meaningful behavioral disclosure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two concise sentences, front-loaded with the core action. Every sentence provides distinct information: the action, the owner constraint, and identifier formats. No filler or redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

While the output schema is present (covering return values), the description lacks important context such as error conditions (e.g., if user already a team member), relationship to invitation flow, or required user existence. For a simple tool, this is adequate but not comprehensive.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and already describes both parameters well. The description only reiterates that user_identifier accepts UUID, email, or handle, which is already in the schema. No additional meaning is added beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description clearly states the verb 'Add' and resource 'user to a team', making the action unambiguous. It also specifies the owner-only restriction and accepted identifier formats, differentiating it from sibling tools like remove_team_member and list_team_members.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage context ('Owner only') but does not explicitly state when to use versus alternatives like invitations or other team management tools. No mention of prerequisites like user existence or team membership status.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

archive_skillA
Idempotent
Inspect

Archive a skill you own. Idempotent.

Sets archived_at and keeps the slug occupied so no other skill can claim the same name for this owner.

ParametersJSON Schema
NameRequiredDescriptionDefault
skill_idYesThe skill to act on, identified by its UUID, slug, or name.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already idempotentHint=true, destructiveHint=false. The description adds valuable context: it sets 'archived_at' and keeps the slug occupied. No contradictions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with action and idempotency, no wasted words. Perfectly concise.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With only one required parameter and an output schema present, the description covers ownership, idempotency, and behavioral details adequately. No gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Input schema has 100% coverage with a clear description for skill_id. The description does not add further parameter-level meaning beyond what the schema provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action 'Archive' and the resource 'skill you own', and distinguishes it from siblings like 'delete_skill' by noting it sets 'archived_at' and keeps the slug occupied. The ownership condition is explicit.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description indicates usage on owned skills and mentions idempotency, providing clear context. However, it does not explicitly state when not to use or mention alternatives like 'unarchive_skill'.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

archive_templateA
Idempotent
Inspect

Archive a template by UUID or @handle/slug.

Owner only. Archiving keeps the slug occupied and hides the template from public listing. Forks pinned at this template keep working; lookup_fork_lineage surfaces parent_template_archived_at. Idempotent on already-archived templates.

ParametersJSON Schema
NameRequiredDescriptionDefault
template_idYesThe template to act on, identified by its UUID or @handle/slug.
archive_reasonNoOptional note recording why the template was archived.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate idempotentHint=true and destructiveHint=false. The description adds rich behavioral context: owner-only access, slug occupation, hiding from public listing, fork persistence, metadata update on lineage, and idempotency for already-archived templates. No contradictions with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise, with four focused sentences. It front-loads the main action and identifier format, then lists key behavioral effects and conditions. No wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the simple parameter set (2 params) and existence of an output schema, the description covers the main action, ownership, idempotency, effects on slugs and forks, and visibility. It sufficiently explains the tool's behavior for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so baseline is 3. The description does not add new meaning beyond what the schema provides for 'template_id' (UUID or @handle/slug) and 'archive_reason' (optional note). It mentions 'Owner only' but that's not tied to a specific parameter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool archives a template by UUID or @handle/slug. It distinguishes from siblings like 'delete_template' (archiving keeps slug occupied and hides from public, but forks keep working) and 'unarchive_template' (reverse action). The verb 'Archive' and resource 'template' are specific.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description specifies 'Owner only' as a usage constraint, and explains effects (slug occupied, hidden from public, forks continue). It implicitly contrasts with deletion and unarchiving, but does not explicitly name alternatives like 'delete_template' or 'unarchive_template'. This is clear but lacks explicit exclusionary guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

audit_skillA
Read-only
Inspect

Review an existing skill against best practices (or a local skill file not yet hosted on Goodeye) and fix what it flags.

Pack-returner: response is {skill_md, references} plus skill_id when one was given. The agent runs the audit locally with the user: it assesses the skill body and every directing sibling file against a documented best-practice rubric, runs the platform quality verifier via run_verifier, produces a priority-ranked report, and applies only the fixes the user approves via save_skill(source='audit'), editing a local copy first when one exists.

skill_id is optional. With it, audits that hosted skill and requires view access. Without it, audits a local skill file the user points at and recommends saving it to Goodeye.

ParametersJSON Schema
NameRequiredDescriptionDefault
skill_idNoSkill UUID, slug, or name to audit. Omit to audit a local skill file the user points at (the report then recommends saving it).

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.8/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description states it applies fixes via save_skill, which is a write operation, but annotations declare readOnlyHint=true. This direct contradiction severely misleads the agent about the tool's side effects.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is thorough and front-loaded with the core purpose, but slightly verbose in detailing the audit process. Every sentence is informative, but some redundancy exists.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity and presence of an output schema, the description provides a complete picture of inputs, process, and outputs. However, the contradiction with annotations undermines trust and completeness.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 100% schema coverage, the description adds significant value by explaining the two modes, the optional nature of skill_id, and the interactive local file audit, far exceeding the schema's brief description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool reviews a skill against best practices and fixes flags, with specific verbs 'Review' and 'fix'. It distinguishes between auditing a hosted skill (via skill_id) and a local file, differentiating it from sibling tools like get_skill or save_skill.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explains two usage modes (with or without skill_id) and implies when to use each. It references related tools (run_verifier, save_skill) but does not explicitly list when not to use or compare to alternatives like optimize_skill.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

cancel_invitationA
Idempotent
Inspect

Cancel a pending invitation you sent (proposer only).

Sets resolution="cancelled". The recipient can no longer accept or decline after this. Returns {invitation_id, resolution}.

ParametersJSON Schema
NameRequiredDescriptionDefault
invitation_idYesThe invitation to act on, identified by its UUID.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description discloses the state change (resolution='cancelled') and the effect on the recipient, beyond what annotations provide. It confirms the mutation (readOnlyHint=false) and non-destructive (destructiveHint=false) nature, with idempotency implied but not elaborated. No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely concise with three short sentences, each adding crucial information: purpose, effect, and return value. No redundant words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the simple tool with one parameter and an output schema, the description is mostly complete. It covers purpose, constraint, outcome, and return format. However, it could mention error scenarios (e.g., if not proposer or invitation already processed).

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description does not add significant meaning beyond the schema's description of invitation_id as a UUID, but it confirms the parameter's role in the cancellation action.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb 'cancel' and the resource 'pending invitation you sent', with a scope constraint 'proposer only'. It effectively distinguishes this tool from its siblings like accept_invitation and decline_invitation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description specifies the user must be the proposer and explains the effect (recipient cannot accept/decline). However, it does not explicitly state when not to use it or mention alternative tools, though the sibling context provides some guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

cancel_subscriptionA
Idempotent
Inspect

Cancel your Pro subscription at the end of the current billing period.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description adds context beyond the annotations by specifying that cancellation occurs at the end of the billing period, not immediately. This aligns with the non-destructive hint and idempotent hint, providing useful behavioral transparency.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence with no wasted words, clearly communicating the action and timing. It is appropriately concise and front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no parameters and the existence of an output schema, the description explains the core action and timing. It could mention the post-cancellation state (e.g., downgrade to free), but overall it is fairly complete for a simple tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There are no parameters, and schema coverage is 100%, so the description does not need to add parameter information. The baseline for no parameters is 4.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool cancels a Pro subscription at the end of the current billing period, which is a specific verb+resource combination and distinguishes it from siblings like upgrade_to_pro and create_billing_portal_session.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies the tool should be used when the user wants to cancel their Pro subscription at period end, but it does not explicitly state when not to use it or provide alternative tools for other cancellation scenarios (e.g., immediate cancellation).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

check_skill_safetyAInspect

Run both system safety verifiers against a saved skill version.

Resolves a UUID or owner-scoped slug. Visibility mirrors get_skill: owner or any active grant; cross-user pointers surface as not found. Defaults to the latest version. Bills two metered verifier runs against the caller. The combined status is one of clean, flagged, blocked, or error; per-side block and advisory carry the verifier id, version, run id, verdict, and reasoning.

ParametersJSON Schema
NameRequiredDescriptionDefault
versionNoPin to a specific version number; omit to use the latest version.
skill_idYesThe skill to act on, identified by its UUID, slug, or name.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description adds value beyond annotations by disclosing billing ('Bills two metered verifier runs') and enumerating possible statuses ('clean', 'flagged', 'blocked', or 'error') with per-side details. No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise (~80 words) and well-structured: first sentence states purpose, then details resolution, visibility, billing, and output. No extraneous content.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a safety check tool, the description covers input resolution, visibility rules, billing impact, and output format (status with per-verifier details). Completeness is adequate given output schema exists.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 100% schema coverage, description adds context: explains 'skill_id' resolution (UUID/slug) and that 'version' defaults to latest. This provides meaning beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description starts with a specific verb+resource+scope: 'Run both system safety verifiers against a saved skill version.' It clearly distinguishes from siblings like 'run_verifier' by specifying it runs two system verifiers, not a single one.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides context on resolution (UUID/slug) and visibility (owner or grant), and mentions defaulting to latest version. However, it does not explicitly state when NOT to use this tool or compare to alternatives like 'run_verifier'.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

check_template_safetyAInspect

Run both system safety verifiers against a published template version.

Resolves a UUID, owner-scoped slug, or @handle/slug. Defaults to the latest live version. Bills two metered verifier runs against the caller. The combined status is one of clean, flagged, blocked, or error; per-side block and advisory carry the verifier id, version, run id, verdict, and reasoning. Anonymous public access (per-IP-hash grant) is available on the REST surface; the MCP transport always authenticates.

ParametersJSON Schema
NameRequiredDescriptionDefault
versionNoPin to a specific version number; omit to use the latest version.
template_idYesThe template to act on, identified by its UUID or @handle/slug.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations (readOnlyHint=false, destructiveHint=false) indicate it's not read-only or destructive. The description adds that it 'Bills two metered verifier runs against the caller', disclosing a cost implication not captured in annotations. It also describes the output structure without contradiction.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two paragraphs, concise and front-loaded with the main purpose. It includes necessary details without being verbose. Could be slightly more structured, but overall efficient.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity (two verifiers, billing, multiple ID formats, output fields) and presence of an output schema, the description adequately covers the tool's behavior and effects. It mentions billing and status outcomes, which is sufficient for the agent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% with descriptions for both parameters. The description adds value by explaining that template_id can be a UUID, slug, or @handle/slug, and reinforces the version parameter's behavior. This goes beyond the schema descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool runs both system safety verifiers on a published template version. It specifies the verb 'run', the resource 'template version', and the scope 'both system safety verifiers'. This distinguishes it from siblings like 'check_workflow_safety' and 'run_verifier'.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explains how to specify the template (UUID, slug, @handle/slug) and defaults to the latest live version. It mentions billing and different access patterns (REST vs MCP). While it doesn't explicitly state when not to use, it provides sufficient context for usage.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

claim_handleA
Idempotent
Inspect

Claim a handle for the authenticated caller.

Validates against the reserved list, rejects collisions, and stamps users.handle + users.handle_claimed_at. Idempotent when the caller re-submits their already-claimed handle.

ParametersJSON Schema
NameRequiredDescriptionDefault
handleYesThe handle to claim for your account.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond annotations (idempotentHint=true, etc.), the description adds validation against reserved list, collision rejection, and field stamping. It also clarifies idempotency behavior for re-submission, providing useful context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Description is three concise sentences with purpose first, then behavior, then idempotency. Every sentence adds value; no unnecessary words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple one-parameter tool with an output schema, the description covers purpose, validation, idempotency, and effects. Given the annotations and schema details, it feels complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema covers 100% with a clear description for the handle parameter. The description adds that handle is validated against a reserved list, but does not add significant new meaning beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description clearly states 'Claim a handle for the authenticated caller', specifying the verb, resource, and actor. This distinguishes it from sibling tools like rename_handle or accept_invitation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage for first-time handle claiming but lacks explicit guidance on when not to use it or mention of alternatives like rename_handle. Contextual information is implied rather than stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

configure_auto_top_upAInspect

Turn on automatic credit top-ups and set (or update) their terms.

Requires a default payment method already on file (a manual credit purchase saves one). Errors clearly when no default payment method is on file, or when the requested amount, threshold, or monthly cap falls outside the allowed range. Re-running this also clears any previously failed automatic top-up state, acting as a retry.

ParametersJSON Schema
NameRequiredDescriptionDefault
amount_usdYesAmount to add, in whole US dollars, each time your balance drops below the threshold.
threshold_usdNoBalance, in whole US dollars, that triggers an automatic top-up. Defaults to the top-up amount when omitted.
monthly_cap_usdNoMaximum total, in whole US dollars, automatic top-ups can spend in a calendar month. Defaults to four times the top-up amount when omitted.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description adds behavioral context beyond annotations, such as error behavior ('Errors clearly when no default payment method is on file, or when the requested amount... falls outside the allowed range') and the fact that re-running clears previous failed state. Annotations are all false and are consistent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise at three sentences, with the main action front-loaded. Every sentence adds value (prerequisite, error handling, retry behavior) without redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (3 parameters with defaults, prerequisite, error cases) and the presence of an output schema, the description is complete. It covers necessary context for the agent to use the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description mentions the parameter types (amount, threshold, monthly cap) but does not add new meaning beyond what the schema provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the main action: 'Turn on automatic credit top-ups and set (or update) their terms.' It uses a specific verb ('turn on and set/update') and resource ('automatic credit top-ups'), and distinguishes from siblings like disable_auto_top_up and purchase_credits.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides usage guidance by stating a prerequisite: 'Requires a default payment method already on file (a manual credit purchase saves one).' It also explains that re-running clears failed state, acting as a retry. While it doesn't explicitly say when not to use it, the context of sibling tools implies its purpose.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_billing_portal_sessionAInspect

Get a link to the Stripe billing portal to manage your subscription and payment.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description does not disclose any behavioral traits beyond the annotations. It mentions 'get a link' but does not specify whether it creates a session (write) or just retrieves a URL. Annotations have readOnlyHint=false, which is not contradicted, but no extra context on auth, rate limits, or side effects.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Single sentence that is clear and to the point. No wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with no parameters and a simple purpose, the description is sufficiently complete. It explains what the tool does and what it returns (a link). Output schema is available for return type details.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

No parameters exist, and the input schema is empty with 100% coverage. The description does not need to add parameter info; baseline of 4 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool 'gets a link' to the 'Stripe billing portal' for managing subscription and payment. It is specific and distinguishes from siblings like cancel_subscription.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool versus alternatives. Sibling tools include cancel_subscription and upgrade_to_pro, but the description does not clarify when to choose this one.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_teamAInspect

Create a team owned by the caller. Handle is immutable post-creation.

Fails with handle_not_claimed if the caller still holds a provisional user handle. Team handles share the handles namespace with user handles.

ParametersJSON Schema
NameRequiredDescriptionDefault
handleYesHandle for the new team, drawn from the shared handle namespace. Immutable after creation.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Discloses handle immutability and namespace sharing, and failure condition. However, with no behavioral hints in annotations (all false), description could elaborate more on side effects or permissions required. It adds some value beyond annotations but is not exhaustive.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two concise sentences that front-load the primary action and add key constraints. No wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter creation tool with an output schema, the description covers the essential failure condition and handle behavior. It feels complete for the tool's complexity, though it could briefly mention what the response contains (but output schema likely covers that).

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% with a detailed description of the handle parameter. The description repeats the immutability and namespace info already present in the schema, adding no new semantic information, so baseline of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states it creates a team owned by the caller, with specific verb and resource. It is unambiguous and distinct from sibling tools like delete_team or list_teams.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides failure condition (handle_not_claimed if provisional handle held) and notes handle namespace sharing, offering context for when to use. Lacks explicit alternatives but given uniqueness of the tool, guidance is clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

decline_invitationA
Idempotent
Inspect

Decline a pending invitation addressed to the caller.

Sets resolution="declined" on the invitation. The proposer may re-invite after this. Idempotent when called on an already-resolved invitation returns InvitationAlreadyResolved.

ParametersJSON Schema
NameRequiredDescriptionDefault
invitation_idYesThe invitation to act on, identified by its UUID.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already provide idempotentHint=true and destructiveHint=false. The description adds detail: sets resolution='declined', mentions idempotency with an error (InvitationAlreadyResolved), and notes re-invite possibility. No contradictions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise with two effective sentences. It front-loads the main action and adds necessary behavioral details without superfluous text.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the simplicity of the tool (single parameter, clear output from idempotency and error behavior), the description covers all necessary context. The presence of an output schema means return values are handled separately.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The only parameter, invitation_id, is fully described in the schema (100% coverage). The tool description does not add extra semantic information beyond what the schema provides. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action: 'Decline a pending invitation addressed to the caller.' It differentiates from siblings like accept_invitation and cancel_invitation by its specific verb and target.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies when to use (declining a pending invitation) and provides context: 'The proposer may re-invite after this.' However, it does not explicitly state when not to use it or compare with alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_imageA
Destructive
Inspect

Hard-delete a hosted image you own by id.

Permanently removes the image and its stored bytes (when no other image shares the same content). There is no recovery path. Raises image_not_found (404) when the image does not exist or belongs to another user (existence masking: the two cases are not distinguished).

Returns {image_id, deleted: true}.

ParametersJSON Schema
NameRequiredDescriptionDefault
image_idYesThe hosted image to act on, identified by its UUID.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate destructiveHint=true. The description adds critical details: permanent removal, condition on shared bytes, no recovery path, and 404 with existence masking. No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four sentences, each providing necessary information: action, effect, error case, return value. No redundancy. Front-loaded with the core action.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the simplicity of the tool (1 param, no output schema shown but description gives return JSON), the description is fully sufficient. It covers behavior, errors, and result.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Only one parameter (image_id) with schema description already explaining it. The description references 'by id' but doesn't add new semantic context. Schema coverage is 100%, so baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states it hard-deletes a hosted image by id. It distinguishes itself from sibling tools like `delete_image_generator` by specifying the resource. The verb 'delete' and resource 'image' are explicit.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It mentions ownership requirement and permanent deletion, implying when not to use (if you don't own the image or want recovery). It also notes existence masking, which is important context. However, it does not explicitly list alternatives for similar actions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_image_generatorA
Destructive
Inspect

Permanently and immediately erase an image generator you own (scope=user only).

generator_id accepts a UUID string.

This is permanent: the generator, all its versions (provider config, model, parameters), all run records, and all anonymous run records are removed from the live system at once. There is no recovery path. Use revoke_image_generator if you want to deactivate the generator without erasing it, keeping the audit trail intact.

Encrypted backups age out within the platform's standard retention window (up to three months), so the data is not instantly erased from all systems everywhere, but it is no longer accessible through any product surface after this call.

Serving gate: if any live published template version carries a snapshot that references this generator, deletion is refused with a Conflict. Unpublish the relevant template version(s) first, then call this operation.

Platform-managed (scope=system) generators and another user's generators always surface as NotFound. No confirmation token is required.

Returns {generator_id, name, deleted: true}.

ParametersJSON Schema
NameRequiredDescriptionDefault
generator_idYesThe image generator to act on, identified by its UUID.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the destructiveHint annotation, the description details the full extent of destruction: all versions, run records, anonymous run records are removed, no recovery path. It also covers edge cases like platform-managed generators surfacing as NotFound and no confirmation token needed. These behaviors are not apparent from annotations alone.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-organized: begins with the primary action, then covers permanence, alternatives, constraints, edge cases, and return value. While detailed, every sentence adds unique information, with no superfluous content. It is slightly lengthy but efficient for the complexity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description comprehensively covers the tool's context: scope, irreversibility, alternative, conflict conditions, edge cases for system generators, return format, and backup retention policy. Given the tool's destructive nature and with an output schema indicated, the description is fully complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The single parameter generator_id is described in the schema as a UUID, and the description adds that it accepts a UUID string, which is not explicitly in the schema format. It also clarifies the scope constraint (own generator) beyond the schema. Since schema coverage is 100% and the description adds practical context, a score of 4 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool 'permanently and immediately erase an image generator you own' with specific verb 'erase' and resource 'image generator'. It distinguishes from sibling tool 'revoke_image_generator', which is explicitly mentioned as an alternative.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit guidance on when to use vs. revoke: 'Use revoke_image_generator if you want to deactivate the generator without erasing it'. Also mentions prerequisite about unpublishing template versions that reference the generator, providing clear conditions for successful invocation.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_skillA
Destructive
Inspect

Permanently and immediately erase a skill you own.

This is permanent: the skill, all its versions, all attached files, and all access grants are removed from the live system at once. There is no recovery path. Use archive_skill if you want a reversible alternative.

Encrypted backups age out within the platform's standard retention window (up to three months), so the data is not instantly erased from all systems everywhere, but it is no longer accessible through any product surface after this call.

Owner only. Works on both live and archived skills.

ParametersJSON Schema
NameRequiredDescriptionDefault
skill_idYesThe skill to act on, identified by its UUID, slug, or name.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already set destructiveHint=true, but description adds critical details: immediate permanent removal of skill, versions, files, access grants; no recovery; encrypted backups age out within 90 days; no longer accessible via product surfaces. No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Description is concise, well-structured, and front-loaded with the key action. Every sentence adds value, with clear sections for consequence, alternative, and additional notes. No unnecessary words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's single parameter, existing annotations, and output schema, the description provides complete context: behavior, prerequisites (owner-only), scope (skill, versions, files, grants), irreversibility, and data retention policy. No gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Only one parameter (skill_id) with full schema description coverage (100%). The description does not add new parameter semantics beyond what schema provides, so baseline score of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states it erases a skill permanently, using specific verb 'erase' and resource 'skill'. It distinguishes from sibling 'archive_skill' by naming it as a reversible alternative, providing clear purpose differentiation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly states when to use (permanent deletion) and when not (use archive_skill for reversible action). Also notes owner-only requirement and applicability to both live and archived skills, offering comprehensive usage guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_skill_versionA
Destructive
Inspect

Permanently and immediately erase a single non-current skill version.

This is permanent: the version row, all its attached files, and any content that is no longer referenced by any surviving version or published template version are removed from the live system at once. There is no recovery path. Use archive_skill if you want a reversible alternative for the whole skill.

Encrypted backups age out within the platform's standard retention window (up to three months), so the data is not instantly erased from all systems everywhere, but it is no longer accessible through any product surface after this call.

The current (live) version cannot be erased with this call. Use delete_skill to permanently remove the entire skill including its current version.

Version numbers remain monotonic with a gap where the erased version was. Surviving versions are not renumbered.

Owner only. Pointing at another user's skill raises NotFound.

ParametersJSON Schema
NameRequiredDescriptionDefault
versionYesThe non-current version number to permanently erase.
skill_idYesThe skill to act on, identified by its UUID, slug, or name.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Description adds extensive behavioral context beyond annotations: permanence, no recovery, erasure of version row and attached files, backup retention window, version numbering gaps, and restriction against erasing current version. No contradiction with destructiveHint: true annotation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Description is well-structured: starts with the core action, then lists consequences, alternatives, and constraints. Each sentence is informative and necessary, with no fluff.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the presence of an output schema and full input schema with annotations, the description covers all essential aspects: action, side effects, prerequisites, and alternatives. It is complete and self-contained.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% with descriptions, but the tool description adds extra meaning: clarifies that version must be 'non-current' and that skill_id can be UUID, slug, or name. This provides useful context beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb ('erase') and resource ('single non-current skill version'), and distinguishes from sibling tools like archive_skill and delete_skill by specifying what each alternative does.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly tells when to use this tool (for permanent erasure of a non-current version) and when not to (use archive_skill for reversible action, delete_skill for entire skill including current version). Also mentions owner-only access.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_teamA
Destructive
Inspect

Delete a team you own by UUID or handle. Releases the handle for reuse.

ParametersJSON Schema
NameRequiredDescriptionDefault
team_identifierYesThe team to act on, identified by its UUID or handle.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, so the destructive nature is known. The description adds the behavioral detail that the handle is released for reuse, which is not deducible from annotations. However, it does not mention other consequences like deletion of team data or member dissociation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely concise with two sentences. Every word serves a purpose: stating the action, identifier methods, and side effect. No redundant or vague phrasing.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity (destructive action, one parameter, output schema exists), the description covers the core purpose and side effect but lacks details on member handling, reversibility, or required permissions beyond ownership. The output schema may compensate, but the description could be more thorough for a deletion tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema provides 100% coverage for the single parameter, describing it as 'UUID or handle'. The description adds the crucial semantic that the team must be owned by the user, which is not in the schema. This helps the agent select the correct identifier.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('Delete a team you own') and the resource ('team'). It specifies that deletion can be done by UUID or handle, and notes the side effect of releasing the handle. This differentiates it from sibling tools like 'remove_team_member' or 'transfer_team_ownership'.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description indicates that the team must be owned by the user, providing a prerequisite. However, it does not explicitly state when not to use the tool (e.g., if the goal is to transfer ownership instead) or mention alternatives among siblings. Additional context about team member impact would enhance guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_templateA
Destructive
Inspect

Permanently and immediately erase a template you own.

This is permanent: the template, all its versions, all attached files, and all version verification records are removed from the live system at once. There is no recovery path. Use archive_template if you want a reversible alternative.

Encrypted backups age out within the platform's standard retention window (up to three months), so the data is not instantly erased from all systems everywhere, but it is no longer accessible through any product surface after this call.

Serving gate: if the template is not archived AND has at least one published version, deletion is refused. Unpublish the relevant version(s) or archive the template first, then call this operation.

Fork severing: any skill forked from this template keeps its own content copy. Only the fork's parent pointer is set to null. The fork remains fully usable; lookup_fork_lineage will report its source as permanently deleted.

Owner only. Pointing at another user's template raises NotFound. Works on both live (serving-gate permitting) and archived templates.

ParametersJSON Schema
NameRequiredDescriptionDefault
template_idYesThe template to act on, identified by its UUID or @handle/slug.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description goes far beyond the annotations by disclosing permanent deletion, absence of recovery, encrypted backup retention window, serving-gate refusal, fork parent pointer behavior, and ownership restrictions. This is model behavior transparency for a destructive op.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Long but every sentence earns its place: permanence, recovery, alternative, serving gate, fork effects, ownership. Front-loaded with the most critical fact (permanent deletion), then consequences. No filler despite length.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Coverage is exceptional for a destructive mutation: permanence, scope of deletion, backup retention, serving-gate refusal conditions, fork behavior, owner restriction, and compatibility with archived templates. With no output schema and destructive implications, all needed operational context is present.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Parameter coverage in the schema is 100%: template_id is fully described as UUID or handle. The description adds ownership and NotFound behavior but no new parameter-specific semantics, so baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('erase'), the resource (template), and the key constraint (you own). It clearly distinguishes from archive_template by emphasizing permanence, and the deletion scope (template, versions, attached files, verification records) is precise.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly tells the agent when to use the alternative (archive_template for reversible removal), when deletion is refused (serving gate for live templates with published versions), and who may call it (owner only). This is complete routing guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_template_versionA
Destructive
Inspect

Permanently and immediately erase a single template version.

This is permanent: the version row, all its attached files, all version verification records, and any content no longer referenced by any surviving version or skill version are removed from the live system at once. There is no recovery path. Use archive_template if you want a reversible alternative for the whole template. Encrypted backups age out within the platform's standard retention window (up to three months), so the data is not instantly erased from all systems everywhere, but it is no longer accessible through any product surface after this call.

Serving gate: the version must first be unpublished (unpublish_template_version) before it can be erased. A still-published version cannot be permanently deleted because it is currently being served to readers. Unpublish it first, then call this operation. To remove the entire template including all its versions, use delete_template instead.

Version numbers remain monotonic with a gap where the erased version was. Surviving versions are not renumbered.

Owner only. Pointing at another user's template raises NotFound.

ParametersJSON Schema
NameRequiredDescriptionDefault
versionYesThe version number to permanently erase.
template_idYesThe template to act on, identified by its UUID or @handle/slug.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate destructiveHint=true, but the description adds critical behavioral details: the permanent erasure, no recovery, the gap left in version numbers, the serving gate requirement, and Owner-only access with NotFound for other users. This goes well beyond the annotation's simple flag.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is detailed yet concise (5 parts), front-loaded with the critical permanence warning, and every sentence provides utility. It uses bullet-like separate paragraphs for clarity without being verbose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive, irreversible operation, the description covers all necessary context: the exact effect, prerequisites, alternatives, consequences to version numbering, and access control. The presence of an output schema also helps complete the picture. Nothing is missing for an agent to use it safely and correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema fully documents both parameters (100% coverage), so the baseline is 3. The description does not add significant new semantics; it only implies the version is the target of erasure bind. That is sufficient given the schema's clarity.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool erases a single template version permanently and list the exact scope. It also distinguishes itself by name from siblings like delete_template and delete_skill_version, making it easy to differentiate.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides explicit when-to-use guidance: only after unpublishing the version, and it names alternatives like archive_template and delete_template. This is a textbook example of clear usage direction.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_verifierA
Destructive
Inspect

Permanently and immediately erase a verifier you own (scope=user only).

verifier_id accepts a UUID string.

This is permanent: the verifier, all its versions (criterion, calibration examples, input contracts), all run records, and all access grants are removed from the live system at once. There is no recovery path. Use revoke_verifier if you want to deactivate the verifier without erasing it, keeping the audit trail intact.

Encrypted backups age out within the platform's standard retention window (up to three months), so the data is not instantly erased from all systems everywhere, but it is no longer accessible through any product surface after this call.

Serving gate: if any live published template version carries a snapshot that references this verifier, deletion is refused with a Conflict. Unpublish the relevant template version(s) first, then call this operation.

Platform-managed (scope=system) verifiers and another user's verifiers always surface as NotFound. No confirmation token is required.

Returns {verifier_id, name, deleted: true}.

ParametersJSON Schema
NameRequiredDescriptionDefault
verifier_idYesThe verifier to act on, identified by its UUID or, for a verifier you own, its name.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Disclosures beyond annotations: lists exactly what gets permanently deleted (verifier, versions, run records, access grants), states no recovery path, mentions backup retention window, and notes no confirmation token needed. Does not contradict annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Well-structured with front-loaded action statement, then detailed bullet points. Every sentence adds value without redundancy. Efficient use of text.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the destructive nature, the description covers all critical aspects: what is deleted, recovery, retention, constraints (serving gate), scope, and return value. No gaps remain for a safe and informed invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema already describes verifier_id as UUID or name. Description adds context: only works for user-scoped verifiers, and reaffirms the return value includes deleted: true. Schema coverage is 100%, so baseline is 3; the additional context raises it to 4.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description clearly states the tool's purpose: 'Permanently and immediately erase a verifier you own'. It distinguishes from sibling tools like revoke_verifier and get_verifier, and specifies scope constraints.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides explicit when-to-use and when-not-to-use guidance: mentions revoke_verifier as alternative for deactivation without erasure, details serving gate conflict with published templates, and clarifies platform-managed/other's verifiers return NotFound.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

deploy_image_generatorAInspect

Deploy a reusable image generator that skills reference to produce images from a chosen model: creates it or appends a version.

An image generator is a named, versioned configuration that routes image generation calls to a specific model. Generators are private and owner-scoped. Skills reference them by UUID or uuid@version. You cannot deploy a new generator whose name matches an active platform scope=system generator (those are tier-level configs that are run-only and not listed or fetched).

Versioning: the first deploy with a given name creates the generator at version 1. Re-deploying the same name appends a new version and requires expected_version_token from the latest known version (returned by deploy/list/get). A new generator must omit the token; an existing one without a token returns Conflict.

Deploy-time validation: the model is checked against the pricing layer. A model that does not resolve to a known image endpoint with an authoritative price is rejected before any row is written.

Returns: {generator_id, name, description, current_version, version, version_token, status, scope, provider, model, generation_contract, config_hash, created_at}. Persist version_token for the next re-deploy.

ParametersJSON Schema
NameRequiredDescriptionDefault
payloadYesPayload for ``deploy_image_generator``: create a new image generator or append a new version to an existing one (owner-scoped, name-uniqueness within owner).

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description goes well beyond the sparse annotations (readOnly=false, idempotent=false, destructive=false) by disclosing meaningful behavior: the model is validated against the pricing layer before creation, name collisions with system-scope generators are rejected, and the operation is non-idempotent because re-deploying appends a new version. This gives an agent an accurate mental model of side effects and failure modes.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is organized into three clear blocks: purpose, versioning/scoping rules, and validation behavior. It is slightly long, but every sentence earns its place—there is no fluff or repetition of the schema. The most decision-relevant facts (create vs. append, system-scope collision, token requirement) are front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the moderate complexity (nested payload, optimistic concurrency, scope constraints), the description covers the essential operational context: what it returns (version, version_token, config_hash, etc.), why it can fail (platform scope collision, model validation), and how it interacts with sibling calls (list/get for tokens). It doesn't spell out the exact return schema, but the description references the return fields adequately for an agent to proceed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema already gives each parameter a descriptionhare and even adds a note explaining version_token semantics and deploy-time validation of model. It adds value on top of the schema by explaining exactly when expected_version_token is required (new vs. existing generator) and what a 'model must resolve to an authoritative price' means in practice. Minor gap: the description does not elaborate on generation_contract values beyond the schema's one-line summary, but the schema carries that load.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb ('Deploy') and resource ('image generator'), then immediately states the core behavior: generators are created or versioned. It also answers 'what is this for' by explaining that skills reference generators to route image-generation calls to a model. This is unmistakably distinct from sibling tools like revoke_image_generator or list_image_generators.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Clear operational rules are given: first deploy creates version 1, re-deploys append a version and require expected_version_token, and the token is obtained from deploy/list/get calls. There is also a concrete exclusion rule ('cannot deploy a new generator whose name matches an active platform-scope generator'). An agent knows exactly when and how to invoke this tool versus alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

deploy_verifierAInspect

Deploy a reusable semantic verifier (an LLM-judge check skills reference by id) that scores agent output against a criterion: creates it or appends a version.

A semantic verifier is a single judgment ("does this output satisfy this criterion?") evaluated by an LLM judge. Verifiers are private and owner-scoped. Skills reference them by verifier_id (or verifier_id@version). You cannot deploy a new verifier whose name matches an active platform scope=system verifier: those definitions are server-owned, never listed or fetched, and only executable through run_verifier.

Versioning: the first deploy with a given name creates the verifier at version 1. Re-deploying the same name appends a new version and requires expected_version_token from the latest known version (returned by deploy/list/get). A new verifier must omit the token; an existing one without a token returns Conflict.

Input contracts:

  • text: input_fields required, media_url rejected.

  • text_image: input_fields plus media_url required at run time.

  • image: input_fields empty, only media_url at run time.

Few-shot examples (3 to 10 typical) calibrate the judge; each example must match the contract (text-only inputs, text+image, or image-only).

Returns: {verifier_id, name, current_version, version, version_token, status, input_contract, config_hash}. Persist version_token for the next re-deploy.

ParametersJSON Schema
NameRequiredDescriptionDefault
payloadYesPayload for ``deploy_verifier``: create a new verifier or append a new version to an existing one (owner-scoped, name-uniqueness within owner).

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With annotations only providing readOnlyHint=false and no destructive hint, the description goes beyond by disclosing that verifiers are private and owner-scoped, that deploying a new verifier returns a conflict when no token is provided, and that the call returns a version token to persist. It also notes platform 'scope=system' verifiers are server-owned and cannot be deployed. A minor gap: it does not explicitly mention whether re-deploying is destructive to prior versions, though versioning implies append-only.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is information-dense and front-loaded with the core purpose ric a clear one-sentence introduction dove. It uses paragraphs and a short list for few-shot examples, which aids scanning. However, it is quite long and repeats some schema details (input contract rules) that might have been left to the schema. Every sentence carries useful information, but it could be tightened slightly without loss.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For an operation with versioning, optimistic concurrency, contract-specific input rules, and platform-scope restrictions, this description covers nearly everything an agent needs: the return payload fields, the token persistence requirement, the conflict behavior, and the input contract constraints. The only minor omission is an explicit note about whether re-deploying changes existing skill references, but that is not essential for invoking the tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is complete (100% – all fields have descriptions), so the baseline is 3. The description adds value by explaining the version_token's role in re-deploys ('Persist version_token for the next re-deploy'), clarifying the input_contract enum semantics (text, text_image, image), and noting the few-shot examples' contract requirements. It does not restate parameter details already in the schema, which is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb+resource: 'Deploy a semantic verifier (an LLM-judge check skills reference by id)'. It explains what the verifier does, distinguishes it from 'run_verifier' (which it names), and clarifies that skill references them by id. This clearly identifies the tool's role and separates it from sibling tools like get_verifier or delete_verifier.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly names the alternative 'run_verifier' for executing verifiers, states that verifiers are owner-scoped, and explains when you cannot deploy (matching an active server-owned platform verifier). It also details the re-deploy path with expected_version_token, making it clear when to deploy vs. update, and notes that a new verifier must omit the token.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

deprecate_template_versionA
Idempotent
Inspect

Mark one template version as deprecated by UUID or @handle/slug.

Owner only. Soft signal: the version stays fetchable; fork_template and lookup_fork_lineage surface the deprecation. message is last-write-wins; deprecated_at anchors on first call.

ParametersJSON Schema
NameRequiredDescriptionDefault
messageYesMessage shown to anyone fetching or forking this version, explaining the deprecation.
versionYesThe version number to deprecate.
template_idYesThe template to act on, identified by its UUID or @handle/slug.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Describes idempotency ('last-write-wins'), non-destructive nature ('soft signal, version stays fetchable'), and deprecation timing ('deprecated_at anchors on first call'). Fully aligns with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four concise sentences, each adding essential information. No fluff. Front-loaded with main action.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Explains return behavior (version stays fetchable), effect on other tools (fork_template, lookup_fork_lineage), and idempotent handling. Complete for a mutation tool with output schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema has 100% coverage. Description adds value by noting 'message is last-write-wins', which is beyond schema. Minor extra context.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Clearly states 'Mark one template version as deprecated' with specific resources. Distinct from siblings like delete_template_version.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly mentions 'Owner only' and describes soft-signal behavior versus alternatives. Lacks explicit when-not-to-use but context is clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

design_skillB
Read-only
Inspect

Create a new skill from scratch: a guided session that designs the skill and its verifiers. The agent should follow it locally and call save_skill() to persist the result.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description reveals that the tool is a 'guided session' and does not persist results, requiring save_skill. However, it contradicts the readOnlyHint annotation by claiming to 'Create,' which is a mutation. This inconsistency severely damages transparency.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences, front-loading the purpose and providing actionable guidance. Every sentence earns its place with no redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers the purpose and sequence, but the contradiction with annotations creates confusion. It does not explain the nature of the 'guided session' or output format, though an output schema exists. Overall, competent but marred by the annotation inconsistency.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With no parameters and 100% schema coverage, the description adds value by explaining the tool's process (guided session) and its outcome (design needing save_skill). It appropriately compensates for the lack of parameter details by clarifying behavior.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states 'Create a new skill from scratch,' but the readOnlyHint annotation indicates the tool is read-only, creating a contradiction. It distinguishes from save_skill by specifying the need to call it afterward, but the core purpose is undermined by the annotation inconsistency.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explains when to use this tool (designing a new skill) and provides a clear follow-up action (call save_skill()). It implicitly distinguishes from save_skill by noting the persistence step is separate, but does not explicitly mention when not to use it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

disable_auto_top_upAInspect

Turn off automatic credit top-ups.

Clears any stored failure state; leaves the previously configured amount, threshold, and monthly cap in place for a later re-enable.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Describes clearing failure state and preserving configuration, going beyond annotations. Could mention idempotency, but adds valuable context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, no waste, front-loaded with main action. Efficient and clear.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple tool with no parameters and an output schema, description covers functionality and side effects fully.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

No parameters, schema coverage 100%, baseline 4. No need for additional parameter info.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description explicitly states the tool turns off automatic credit top-ups and distinguishes from siblings like configure_auto_top_up by noting it preserves configuration.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Clear context for use (disable auto top-up) but lacks explicit when-not-to-use or alternatives; implied but not stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

fork_templateAInspect

Fork a public template into a private skill owned by the caller.

Authentication is required. Returns the new skill's id and lineage metadata; anonymous callers get AuthRequired.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameNoOptional name for the new private skill; defaults to the template's slug.
versionNoPin to a specific version number; omit to use the latest version.
identifierYesThe template to fetch, identified by its UUID, @handle/slug, or @handle/slug@vN to pin a specific published version.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate this is a mutating, non-idempotent operation. The description adds valuable behavioral context: authentication is required, anonymous callers get AuthRequired, and the return value includes the new skill's id and lineage metadata. This goes beyond the annotations and helps the agent anticipate side effects and failure modes.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences with no fluff. The core action is front-loaded, and the authentication note and return value are stated compactly. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers the action, the ownership outcome, authentication requirements, and the return value. With an output schema present and 100% parameter coverage, this is nearly complete. It could mention what happens if the template is not public or if the caller already owns a skill with the same name, but those are edge cases.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all three parameters. The description adds context about the identifier format (UUID, @handle/slug, @handle/slug@vN) and the default behavior for name and version, but this is largely redundant with the schema. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Fork') and resource ('a public template into a private skill owned by the caller'), clearly distinguishing this from sibling tools like get_template, archive_template, or publish_template_version. It also specifies the ownership outcome, which is the core purpose.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies when to use it: when you want to create a private skill from a public template. It doesn't explicitly name alternatives or exclusions, but the verb 'Fork' and the target 'private skill owned by the caller' make the use case clear. It also notes authentication is required, which is a usage prerequisite.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

generate_imageAInspect

Generate one or more images from a text prompt, billed to the caller's credits.

Requires authentication. Anonymous image generation is available only via the REST API (POST /v1/image-generators/{id}/runs); the MCP transport always authenticates.

Resolution order for the generator (highest priority first):

  1. A deployed generator ref (uuid@version or bare UUID): pins the deployed version config.

  2. The model control path (authenticated one-off, ephemeral). Not usable from published templates.

  3. A tier generator ref (system:<tier>): resolves to the tier's current best model (auto-upgrade). Available tiers: system:image-standard (default), system:image-premium, system:image-edit (image-to-image, requires reference_image_url).

  4. Default: system:image-standard when no generator or model is given.

generator and model are mutually exclusive.

For image_to_image generators, reference_image_url is required and must be a public HTTP or HTTPS URL. For text_to_image generators, providing reference_image_url is rejected.

Billing: spend is deducted from the caller's monthly credit balance. BudgetExhausted (402) and AccountSuspended (403) propagate if the balance is zero or the account is suspended.

visibility sets the access level of the hosted copy of each image: public (default) returns a link that opens in any browser; private returns a link only you can open and forward to people you choose, while the plain URL stays locked.

Returns: {run_id, model_tier_or_model, image_url, image_urls, width, height, num_images, cost_usd, duration_ms, status, created_at, error_code, error_message, hosted_images}. hosted_images carries the durable Goodeye-hosted copy of each image with its url (the browser-viewable link) and visibility. The prompt is never stored; only its hash is persisted on the run row.

ParametersJSON Schema
NameRequiredDescriptionDefault
payloadYesPayload for ``generate_image``: produce one or more images from a prompt, billed to the caller's credit balance. Supply ``generator`` to use a named system tier (e.g. ``system:image-standard``) or a deployed generator by ``uuid@version``. Supply ``model`` for an authenticated one-off using a concrete model identifier. ``generator`` and ``model`` are mutually exclusive; omitting both resolves to the default tier.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations are all false (not read-only, not idempotent), so the description carries the disclosure burden and delivers richly: it reveals MCP-transport always authenticates (anonymous only via REST), billing deducted from monthly credit balance, BudgetExhausted (402) and AccountSuspended (403) propagation, the public/private visibility behavior, and that the prompt is never stored (only its hash persisted). This far exceeds what the annotations express and adds genuine safety/cost context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core purpose and well structured into a numbered resolution list and labeled sections (authentication, billing, visibility, returns). Every sentence carries information, and the length is justified by the tool's complexity (11 nested params). Minor redundancy: the visibility explanation is duplicated almost verbatim in the schema description, which could be trimmed.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a high-complexity tool with a nested payload and an output schema (which covers the return shape), the description is remarkably complete: it covers authentication, billing and error propagation, the generator resolution order, mutual exclusivity, per-type reference_image_url rules, and prompt-privacy. Nothing an agent needs to call it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so each of the 11 payload fields is already documented, giving a baseline of 3. The description adds value on top: the generator resolution priority (deployed > model > tier > default), the generator/model mutual exclusivity rule, and the billing-per-image semantics for num_images. The visibility text is somewhat redundant with the schema, but the resolution-order context is net-new and material.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb+resource: 'Generate one or more images from a text prompt, billed to the caller's credits.' This clearly distinguishes the tool from siblings like get_image, list_images, upload_image, and delete_image. The scope (prompt→image, billing) is unambiguous and does not restate the name.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Extremely explicit guidance: a numbered resolution order for the generator (deployed ref > model > tier > default), the mutual exclusivity of generator and model, and the requirement/rejection rules for reference_image_url by generator type (image_to_image requires it, text_to_image rejects it). An agent is told exactly how to select among the parameter paths, with no inference required.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_auto_top_upA
Read-only
Inspect

Check your automatic credit top-up configuration and this month's spend toward it.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, so the description does not need to emphasize safety. It adds context about retrieving monthly spend, but does not disclose any additional behavioral traits like response format or rate limits.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, concise sentence that front-loads the main action ('Check your automatic credit top-up configuration') and then adds the spend detail. No unnecessary words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no parameters and an output schema existing, the description is adequate. It tells the agent exactly what information is returned. A perfect score would require mentioning that it is read-only, but annotations already cover that.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With no parameters, the schema coverage is 100%. The description adds meaning by explaining what the tool retrieves, which is the top-up configuration and spend. Baseline for 0 parameters is 4.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses the verb 'check' and specifies the resource 'automatic credit top-up configuration' along with 'this month's spend'. This clearly distinguishes it from sibling tools like 'configure_auto_top_up' and 'disable_auto_top_up', which are write operations.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies the tool is for viewing current settings and usage. It does not explicitly state when to use it versus alternatives, but the read-only nature is evident from the name and annotations, providing clear context with no exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_imageA
Read-only
Inspect

Fetch a hosted image by id. Caller must own it.

Returns the full image record including the serving URL. Raises image_not_found (404) when the image does not exist or belongs to another user (existence masking: the two cases are not distinguished).

ParametersJSON Schema
NameRequiredDescriptionDefault
image_idYesThe hosted image to act on, identified by its UUID.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Adds return content ('full image record including serving URL') and error behavior with existence masking, going beyond readOnlyHint annotation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three concise sentences: purpose, return value, error handling. No redundancy, front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With output schema existing, description sufficiently covers purpose, ownership, return content, and error behavior. Complete for a simple get tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Only one parameter with 100% schema coverage. Description adds no additional semantic detail beyond what schema provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Clearly states 'Fetch a hosted image by id' with specific verb and resource. Ownership requirement distinguishes it from tools like list_images, update_image, delete_image.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit ownership constraint guides usage, but no mention of when not to use or alternatives like list_images or get_image_generator.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_image_generatorA
Read-only
Inspect

Inspect one image generator you own (its model and full config) at head or a pinned version.

generator_id accepts a UUID string. Platform system:... tier aliases and system generator UUIDs are not returned here (NotFound): system generators are run-only and their internal config never surfaces through list, get, deploy, or revoke.

Defaults to the current version; pass version to pin. Returns the full deploy-time payload (provider, model, generation_contract, default_params) plus config_hash (SHA-256 over the config) so callers can detect drift across versions. Requires ownership; a cross-user or revoked generator surfaces as NotFound.

ParametersJSON Schema
NameRequiredDescriptionDefault
versionNoPin to a specific version number; omit to use the latest version.
generator_idYesThe image generator to act on, identified by its UUID.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With annotations providing readOnlyHint, the description adds value by detailing the return payload, version pinning, and error conditions. This is exemplary transparency.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely concise: three sentences that cover purpose, constraints, and return value, all without unnecessary words. It is front-loaded with the core action and resource.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema available, the description adequately covers the key aspects: what it returns, how parameters work, and error conditions. No gaps are apparent. The tool has low complexity (2 params, read-only), and the description is sufficient.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema already describes both parameters with 100% coverage, so baseline is 3. The description adds semantic value by clarifying that generator_id should be a user-owned UUID (not system), and that version pinning affects the output. This extra context elevates it above baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('Inspect'), the resource ('one image generator you own'), and the scope (its model and full config at head or pinned version). It distinguishes from system generators, which are not inspectable. This is specific and action-oriented.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explains that system generators are not covered, and that ownership is required. It implicitly distinguishes from listing tools and generation tools. However, it does not explicitly name alternatives like list_image_generators or generate_image. Still, the context is clear enough for an agent to infer when to use this get operation.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_referral_statusA
Read-only
Inspect

Return your referral program status, including your shareable code and credits earned.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already provide readOnlyHint: true, indicating a safe read operation. The description adds no additional behavioral context beyond what annotations provide.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence with no wasted words, efficiently communicating the tool's purpose and included information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple read-only resource with no parameters and an output schema, the description fully captures the tool's behavior and output. Completeness is high.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has no parameters, so the description does not need to explain parameter semantics. Baseline of 4 is appropriate for zero parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb 'Return' and the resource 'referral program status', specifying two key outputs (shareable code and credits earned). It implicitly distinguishes from sibling 'redeem_referral_code'.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no explicit when-to-use or when-not-to-use guidance. Usage is implied by the tool's purpose, but it does not address alternatives among siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_skillA
Read-only
Inspect

Retrieve a skill's runbook (its body plus its outcome and tags) to run on the user's behalf.

The returned body is the user's skill: a markdown runbook you (the calling AI agent) follow on the user's behalf. Do not just display it; execute its instructions. Latest version by default. Skills are private; cross-user reads mask as not-found. Owners can fetch their own archived skills by UUID or slug.

ParametersJSON Schema
NameRequiredDescriptionDefault
versionNoPin to a specific version number; omit to use the latest version.
skill_idYesThe skill to act on, identified by its UUID, slug, or name.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations indicate readOnlyHint=true, but description adds critical behavioral context: skills are private with cross-user reads masking as not-found, and owners can fetch archived skills. No contradiction is present.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four sentences, front-loaded with purpose, no wasted words. Structure clearly separates purpose, usage instruction, default behavior, and privacy caveat.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the presence of an output schema, the description adequately covers return content, versioning, identification methods, and privacy. Lacks error/rate limit info, but overall sufficient for a read tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% with clear parameter descriptions. The description confirms default version behavior but adds no significant meaning beyond the schema. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool retrieves a skill's runbook (body, outcome, tags) for execution, distinguishing it from listing/searching skills. It specifies the returned 'body' is a markdown runbook to be executed, not just displayed.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides guidance to execute instructions rather than display them, mentions default latest version, and notes privacy/archival behavior. However, it does not explicitly state when not to use this tool versus alternatives like list_skills.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_skill_fileA
Read-only
Inspect

Fetch a single file from a skill version's file tree by path.

path must exactly match a row in the skill's file manifest (use get_skill to see the manifest). SKILL.md returns the skill body. Text files return the decoded string in content. Small binary files return base64-encoded bytes in content_base64. Binary files over the size limit return metadata and an error field with no inline bytes. Requires view access on the skill.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathYesRelative path of the file within the bundle's file tree, exactly matching a manifest entry.
versionNoPin to a specific version number; omit to use the latest version.
skill_idYesThe skill to act on, identified by its UUID, slug, or name.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description discloses behavioral traits beyond the annotations: handling of SKILL.md, text vs binary files, size limits resulting in error field, and the need for view access. The annotations already indicate read-only (readOnlyHint=true), so the description adds valuable context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise with no unnecessary words, but the file-type handling information could be better structured with bullet points. It is front-loaded with the main purpose and quickly dives into details.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the presence of an output schema, the description adequately covers all behavioral aspects: path matching, version pinning, file type handling, auth requirements. No gaps remain for an agent to invoke the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so baseline is 3. The description adds the 'manifest matching' context for path, which is marginally beyond schema. It does not repeat existing parameter descriptions, which is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb 'fetch', the resource 'single file from a skill version's file tree', and the method 'by path'. It distinguishes this tool from siblings like get_skill (which returns the manifest) and get_skill_files (which likely lists files).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explains that path must exactly match a manifest entry, suggests using get_skill to inspect the manifest, and specifies how different file types are handled. It also mentions the required view access. However, it does not explicitly contrast with get_skill_files or state when not to use this tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_skill_filesA
Read-only
Inspect

Fetch multiple files from a skill version's file tree in a single call.

paths is a list of relative paths, each of which must exactly match a row in the skill's file manifest (use get_skill to see the manifest). Envelopes are returned in lexicographic path order.

Each envelope carries the same fields as get_skill_file. Unknown paths appear as per-path error: "not_found" envelopes rather than failing the whole request. Once the aggregate inline budget is exhausted, remaining inline-eligible files are returned as error: "batch_response_cap_exceeded" references with no content. Every requested path appears exactly once in the response.

Requires view access on the skill. Returns {"files": [...]}.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathsYesList of relative paths to fetch in one call.
versionNoPin to a specific version number; omit to use the latest version.
skill_idYesThe skill to act on, identified by its UUID, slug, or name.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Details error handling for unknown paths and budget exhaustion, ordering, and guarantee every path appears once. Annotations confirm read-only, and description adds significant behavioral context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Well-structured, front-loaded with purpose, each sentence adds value without redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Comprehensive coverage of return format, error modes, ordering, and budget limits; output schema exists but description still clarifies structure.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema covers 100% of parameters with descriptions; description adds practical context (e.g., paths must match manifest entries, version omitted uses latest, skill_id accepts UUID/slug/name).

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states it fetches multiple files from a skill version's file tree, distinguishing from siblings like get_skill and get_skill_file.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explains inputs (paths must match manifest) and suggests using get_skill to see the manifest, but could be more explicit about when to use this vs. get_skill_file.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_templateA
Read-only
Inspect

Retrieve a public template's runbook to run on the user's behalf, by UUID, @handle/slug, or @handle/slug@vN.

The returned body is a skill you (the calling AI agent) execute on the user's behalf as a runbook, not just display. Non-owner reads carry an unverified-template safety banner. Latest live version by default. Anyone can read.

ParametersJSON Schema
NameRequiredDescriptionDefault
versionNoPin to a specific version number; omit to use the latest version.
identifierYesThe template to fetch, identified by its UUID, @handle/slug, or @handle/slug@vN to pin a specific published version.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the readOnlyHint annotation, the description adds valuable behavioral context: the returned body is an executable skill, non-owner reads carry an unverified-template safety banner, and the latest live version is fetched by default. These details inform the agent about execution semantics and safety expectations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and front-loaded with the core function, then adds essential behavioral details in short, focused sentences. No filler is present, and each sentence contributes meaningful information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description is complete for a read-only retrieval tool: it covers identifier formats, version behavior, access permissions, safety banner behavior, and the runbook execution semantics. The presence of an output schema means return-value details do not need to be in the description.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema already documents both parameters fully, including accepted identifier forms and version pinning, so the description does not need to compensate. The description adds minor useful context like 'latest live version by default' and 'anyone can read,' but it largely restates schema information.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'Retrieve a public template's runbook' and clarifies the purpose: 'to run on the user's behalf.' It also distinguishes itself by noting the returned body is a skill to execute, not merely display, which differentiates it from related tools like get_template_file.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear context for when to use the tool: to fetch a public template's runbook for the user. It explains access ('Anyone can read') and default behavior ('Latest live version by default'), but it does not explicitly name alternatives or state when not to use it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_template_fileA
Read-only
Inspect

Fetch a single file from a template version's file tree by path.

identifier accepts UUID, @handle/slug, or @handle/slug@vN. path must match a row in the template's file manifest (use get_template to see the manifest). SKILL.md returns the template body. Text files return the decoded string in content. Small binary files return base64-encoded bytes in content_base64. Binary files over the size limit return metadata and an error field with no inline bytes.

Non-owner responses carry the safety banner and safety_verification_status. Anonymous callers may fetch files from live (published) template versions only; the liveness check is evaluated per read.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathYesRelative path of the file within the bundle's file tree, exactly matching a manifest entry.
versionNoPin to a specific version number; omit to use the latest version.
identifierYesThe template to fetch, identified by its UUID, @handle/slug, or @handle/slug@vN to pin a specific published version.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, but description adds valuable behavioral details: handling of text vs binary files, size limits (metadata+error), identifier formats, and safety banner for non-owners. No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Description is front-loaded with the main action and follows with necessary details. Six sentences each add information, but could be slightly more terse. Overall well-structured.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the output schema exists (context signals), description covers file behavior (text, binary, over-limit), identifier options, and access restrictions. Complete for the tool's complexity with 3 parameters.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, baseline 3. Description adds meaning beyond schema: explains identifier formats (UUID, @handle/slug, @vN), that path must match manifest, and version parameter's purpose. Adds significant value.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states 'Fetch a single file from a template version's file tree by path', using specific verb and resource. It distinguishes from sibling tools like get_template (which returns the template manifest) and get_workflow_file (different resource).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly tells the agent to use get_template to see the manifest for path validation, and notes restrictions for non-owners and anonymous callers, providing clear context for when to use this tool vs alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_usageA
Read-only
Inspect

Check how much of your monthly credit grant you have used, with the remaining balance.

Also includes your automatic credit top-up configuration and this month's spend toward it, null if you have never configured one.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=true and openWorldHint=false. The description adds value by noting that top-up configuration data (including this month's spend) is returned, and that it null if never configured. This provides behavioral context beyond annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with main purpose. Every sentence adds value without redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a parameterless tool with an output schema, the description sufficiently explains the return values (balance, top-up config). No gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There are no parameters, so schema coverage is 100%. Per guidelines, baseline is 4. The description explains what the tool returns, which is sufficient.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description clearly states the tool checks monthly credit grant usage and remaining balance, and includes top-up configuration details. This specific verb+resource pairs well with the tool name and distinguishes it from siblings like get_auto_top_up.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Description indicates when to use (check credit usage and top-up info). It does not explicitly state when not to use or mention alternatives, but the tool's zero parameters and focused purpose make usage clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_verifierA
Read-only
Inspect

Read a verifier's full definition (criterion, calibration examples, judge config); user-scoped verifiers only.

verifier_id accepts a verifier UUID string or an accessible user-scope name. Any caller who can reach the verifier can read it: the owner, and skill grantees at any role (a view/exec grantee can read, not only run). Platform system:... aliases and system verifier UUIDs are never returned (NotFound): system rows are run-only and their internal config never surfaces through list, get, deploy, or revoke.

Defaults to the current version; pass version to pin. Returns the full deploy-time payload (criterion, input_contract, input_fields, few_shot_examples, judge_model_config, reasoning_field_description) plus config_hash (canonical-JSON SHA-256 over the config) so callers can detect drift across versions. A verifier you have no access to (and any revoked one) surfaces as NotFound. Platform-managed verifiers are run-only and never returned here.

ParametersJSON Schema
NameRequiredDescriptionDefault
versionNoPin to a specific version number; omit to use the latest version.
verifier_idYesThe verifier to act on, identified by its UUID or, for a verifier you own, its name.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, but the description goes well beyond this by disclosing the exact payload returned (criterion, input_contract, etc.), the NotFound behavior for system/revoked/inaccessible verifiers, the config_hash for drift detection, and the access model for grantees. No contradiction with annotations; the read-only nature is consistently reinforced.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is longer than average but every sentence earns its place: it front-loads the core purpose and scope, then methodically covers access rules, system exclusion, version defaults, and return payload. No redundancy or filler; it is an efficient, well-organized explanation.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is complex enough (2 params, versioning, access control, specific return fields) that the description covers all necessary aspects: what is returned, when NotFound occurs, how version pinning works, and what caller permissions are required. The presence of an output schema reduces the burden, and this description exceeds that burden by preemptively explaining the key fields.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Although schema coverage is 100%, the description adds substantial semantic depth: verifier_id is explained as accepting a UUID or an accessible user-scope name, and version is described as defaulting to the latest with explicit behavior when pinned. This goes far beyond the schema's terse descriptions and materially helps the agent pick and fill parameters correctly.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a precise verb and resource ('Read a verifier's full definition') and immediately scopes it ('user-scoped verifiers only'), distinguishing it from run_verifier and list_verifiers. It also specifies what is excluded (platform/system verifiers), making its purpose unmistakable among siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

While it does not explicitly name 'use list_verifiers instead for listing', it provides clear when-to-use context: reading a specific verifier's full definition, with explicit exclusions (system verifiers never returned) and access rules (owner and grantees can read). The version pinning guidance and the behavior for inaccessible/revoked verifiers further clarify when this tool is appropriate.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

grant_skillAInspect

Grant skill access to a user or team by UUID, email, or handle.

By default the grantee sees the skill version current at share time and any later versions, but not older ones. Set include_history to true to give the grantee the full version history.

ParametersJSON Schema
NameRequiredDescriptionDefault
roleYesAccess level to grant: 'view', 'edit', or 'admin'.
skill_idYesThe skill to act on, identified by its UUID, slug, or name.
include_historyNoWhen true, the grantee sees the full version history; when false (default), only the version current at share time and later.
grantee_email_or_at_team_handleYesThe grantee: a user's email, or a team's @handle.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations are minimal (non-read-only, non-destructive). The description adds valuable context about version history handling, which is not covered by annotations. No contradictions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, no wasted words. The key information is front-loaded and easy to parse.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the presence of an output schema (not shown but declared), all parameters are well-documented in schema, and the description covers the behavioral nuance of version history. Complete for this tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so baseline is 3. The description adds meaning by explaining the default and effect of include_history, and the flexible identifier formats for the grantee.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: granting skill access to a user or team by identifier. It distinguishes from siblings like revoke_skill_grant and list_skill_grants by focusing on granting access.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides clear context on the include_history parameter and default behavior. Does not explicitly state when not to use this tool or compare with siblings, but the use case is straightforward.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

leave_shared_skillA
Idempotent
Inspect

Remove your direct grant on a shared skill.

ParametersJSON Schema
NameRequiredDescriptionDefault
skill_idYesThe skill to act on, identified by its UUID, slug, or name.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, destructiveHint=false, and idempotentHint=true. The description adds that it removes 'your direct grant', which provides context beyond the structured fields, but does not contradict annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single sentence that is front-loaded and contains no extraneous information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple tool with one parameter and an output schema, the description provides adequate context. It could briefly mention that a 'shared skill' is a skill granted within a team, but the current text is sufficient.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema covers the single parameter with a description of accepted identifiers. The tool description adds no additional meaning beyond the action.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Remove your direct grant on a shared skill' uses a specific verb and resource, clearly distinguishing the action from siblings like grant_skill and revoke_skill_grant.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No explicit guidance on when to use this tool versus alternatives such as revoke_skill_grant. The implied use case is clear but not contrasted with similar operations.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_api_keysA
Read-only
Inspect

List your live (non-revoked) API keys.

Returns {"items": [...], "next_cursor": str | null}. Page until next_cursor is null. Never returns hashes or plaintext.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMaximum number of items to return in one page (1 to 200).
cursorNoPagination cursor from a previous response's next_cursor; omit to start from the first page.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the readOnlyHint annotation, the description adds critical security context: 'Never returns hashes or plaintext.' This discloses important behavioral traits about what the tool does not expose, which is valuable for an API key listing tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is three sentences long, front-loaded with the purpose, then describes the return format, and finally gives pagination instructions. Every sentence is essential, and there is no waste.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has two well-documented parameters, an output schema, and is read-only, the description covers all necessary aspects: what it does, how to paginate, and what it does not return. It is complete for a simple listing tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema already provides 100% coverage with descriptions for 'limit' and 'cursor'. The tool description does not add additional semantic meaning beyond what the schema provides, so a baseline score of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb 'List', the resource 'your live (non-revoked) API keys', and distinguishes from siblings like mint_api_key or revoke_api_key. It is specific and unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explains pagination behavior (page until next_cursor is null) and the return format. While it doesn't explicitly state when not to use it, the context of sibling tools makes it clear this is for viewing existing keys, not for creating or revoking.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_image_generatorsA
Read-only
Inspect

List active (non-revoked) image generators owned by the caller.

Platform-managed system generators (the standard, premium, and image-to-image tiers) are excluded; those are run-only and not listed. Cursor-based pagination mirrors list_verifiers.

Returns {items: [{generator_id, name, description, current_version, version_token, status, scope, created_at, updated_at}], next_cursor: str | null}.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMaximum number of items to return in one page (1 to 200).
cursorNoPagination cursor from a previous response's next_cursor; omit to start from the first page.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare read-only behavior. The description adds detail: lists only non-revoked generators, excludes platform-managed ones, and describes pagination behavior mirroring list_verifiers, all beyond the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences covering purpose, exclusions, pagination, and return format. Front-loaded with main purpose, no unnecessary words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a list tool with simple pagination, the description fully covers what the tool does, what it returns (including field list), and its limitations (exclusions). No gaps given the output schema is provided inline.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% with clear parameter descriptions. The description adds value by referencing list_verifiers pagination behavior, providing extra context for cursor usage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action 'List' and the resource 'active (non-revoked) image generators owned by the caller'. It specifies exclusions (platform-managed) and scope, making it distinct from sibling list tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides clear context that only user-owned generators are listed and platform-managed ones are excluded, informing when not to use this tool. However, no explicit alternatives are named for other use cases.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_imagesA
Read-only
Inspect

List hosted images owned by the caller, with optional filters.

source filters by upload origin: "upload" for directly uploaded images, "generated" for images created via the image generation tools. Omit to return all sources.

visibility filters by access level: "public" or "private". Omit to return both.

Pagination: pass next_cursor from a previous response as cursor to retrieve the next page. Returns {items: [...], next_cursor: str | null}.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMaximum number of items to return in one page (1 to 200).
cursorNoPagination cursor from a previous response's next_cursor; omit to start from the first page.
sourceNoFilter by origin: 'upload' for directly uploaded images, 'generated' for images created by the image tools; omit for all.
visibilityNoFilter by access level: 'public' or 'private'; omit for both.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description goes beyond the readOnlyHint annotation by detailing pagination behavior and return format ('Returns {items: [...], next_cursor: str | null}'). This provides critical behavioral context that helps an agent understand how to iterate results, with no contradictions to annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise at approximately 6 lines, front-loaded with the main purpose, and uses well-structured bullet points for filters and pagination. Every sentence adds value without redundancy, making it easy for an agent to parse.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the presence of an output schema and the description's coverage of filters, pagination, and return format, the description is fully complete for this tool. It addresses all necessary aspects without relying solely on structured fields.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description adds value by elaborating on the source filter (e.g., explicitly mentioning 'generated' images) and explaining pagination cursor usage, which is not fully captured in the schema descriptions. This enhancement justifies a 4.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states 'List hosted images owned by the caller, with optional filters,' specifying a verb (list) and resource (hosted images). This distinguishes it from sibling tools like get_image (single image) and delete_image, making the purpose immediately clear.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies use for retrieving multiple images with optional filters but does not explicitly contrast with alternatives like get_image for individual images or upload_image. The usage context is clear but lacks explicit when-not or sibling differentiation.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_invitationsA
Read-only
Inspect

List invitations visible to the caller.

filter="received" returns invitations where you are the recipient. filter="sent" returns invitations you created. filter="all" combines both. state="pending" excludes expired and resolved rows. Cursor-paginated. Returns {items: [...], next_cursor: str | null}.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMaximum number of items to return in one page (1 to 200).
stateNoWhich states to include: 'pending' only, or 'all' (includes expired and resolved invitations).pending
cursorNoPagination cursor from a previous response's next_cursor; omit to start from the first page.
filterNoWhich invitations to return: 'received' (addressed to you), 'sent' (you created), or 'all'.received

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, and the description adds transparency by explaining cursor pagination, return format, and the effect of filter and state parameters.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise and well-structured: first sentence states purpose, then parameter explanations with examples, then pagination and return format. No wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the parameter count, schema richness, and presence of output schema, the description fully covers behavior: listing with filters, state, and cursor pagination. No gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, but the description adds significant value by providing examples and clarifying parameter behavior (e.g., 'filter="received" returns invitations where you are the recipient').

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states 'List invitations visible to the caller' and elaborates on filters and pagination, distinguishing it from siblings like accept_invitation or cancel_invitation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage for viewing invitations with various filters, but does not explicitly state when to use this tool versus alternatives or provide exclusion criteria.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_skill_grantsA
Read-only
Inspect

List direct user and team grants on a skill. Admin or owner only.

ParametersJSON Schema
NameRequiredDescriptionDefault
skill_idYesThe skill to act on, identified by its UUID, slug, or name.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true. The description adds value by specifying that it lists only direct grants (user and team) and that admin/owner access is required, going beyond the annotation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence, front-loaded with the action ('List'), and contains no unnecessary words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With one well-described parameter, full schema coverage, and an output schema, the description provides all necessary context for the simple list operation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% with a clear parameter description. The description doesn't add new parameter info but reinforces the resource type. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action 'List', the resource 'grants on a skill', and specifies scope 'direct user and team grants', distinguishing it from sibling tools like grant_skill and revoke_skill_grant.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It specifies 'Admin or owner only', providing clear access control guidance. While it doesn't mention alternatives, the context is sufficient for when to use the tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_skillsA
Read-only
Inspect

List skills you own or that are shared with you.

Returns {"items": [...], "next_cursor": str | null}. Page until next_cursor is null. Each item carries an "archived_at" field (ISO-8601 string for an archived skill, null for a live one). Archived skills are excluded unless include_archived is set, in which case your own archived skills are included.

ParametersJSON Schema
NameRequiredDescriptionDefault
tagNoRestrict to skills carrying this tag.
limitNoMaximum number of items to return in one page (1 to 200).
cursorNoPagination cursor from a previous response's next_cursor; omit to start from the first page.
filterNoWhich skills to return: 'mine' (you own), 'shared-with-me' (granted to you), or 'all'.all
searchNoCase-insensitive substring filter over name, description, and tags.
include_archivedNoInclude your own archived skills in the result.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description clarifies that archived skills are excluded by default and only included if include_archived is true for owned skills. This adds detail beyond the readOnlyHint annotation. It also explains the response format and pagination cursor. No contradictions with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is brief and front-loaded with the main purpose, then follows with necessary details about pagination and archived skills. Every sentence adds value without redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the presence of an output schema and six parameters (none required), the description covers pagination, archived inclusion, and filter scope. It omits some edge cases (e.g., behavior when no skills exist), but overall sufficiently completes the specification.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the schema already describes all parameters. The description adds minimal extra meaning for parameters (e.g., it doesn't elaborate on the 'filter' enum beyond what's in schema). The response structure explanation is relevant but not parameter-specific.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool lists skills that are owned or shared, distinct from search_skills and list_skill_grants among siblings. It specifies the resource ('skills') and scope ('you own or shared'), making the purpose precise.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description includes pagination instructions and mentions filtering via parameters, but does not explicitly differentiate when to use list_skills versus alternatives like search_skills or list_skill_grants. Usage context is implied but not formally guided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_team_membersA
Read-only
Inspect

List members of a team by UUID or handle. Owner appears via a synthetic row.

ParametersJSON Schema
NameRequiredDescriptionDefault
team_identifierYesThe team to act on, identified by its UUID or handle.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=true, confirming safe read behavior. The description adds value beyond annotations by disclosing that the owner appears via a synthetic row, which is a non-obvious behavioral trait. No contradictions detected.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description consists of two short, clear sentences. It is front-loaded with the core purpose and includes an important behavioral note without any wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the simple tool (one parameter, read-only, with output schema), the description is complete. It covers the action, identification method, and a notable result behavior. The presence of an output schema means return values need not be detailed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema description coverage is 100% for the single parameter team_identifier. The description mentions 'by UUID or handle', which mirrors the schema's own description. No new meaning is added beyond the schema, so baseline score of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb 'List', the resource 'members of a team', and the identification method 'by UUID or handle'. It also adds a specific detail about the owner appearing via a synthetic row. Among siblings like add_team_member and list_teams, this tool's function is distinct.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description does not provide explicit guidance on when to use this tool versus alternatives, such as listing teams or adding members. It lacks context about prerequisites or conditions that might warrant using this tool over others. While the purpose is clear, usage scenarios are left implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_teamsA
Read-only
Inspect

List teams visible to the caller (owned + member of).

ParametersJSON Schema
NameRequiredDescriptionDefault
filterNoWhich teams to return: 'mine' (you own), 'member' (you belong to but do not own), or 'all'.all

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=false, so the description doesn't need to restate that. It adds value by clarifying the scope ('visible to the caller') and that teams include both owned and member teams. No contradictions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, efficient sentence that front-loads the verb and resource. Every word is meaningful, with no wasted language. It is appropriately sized for the tool's simplicity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the low complexity, full schema coverage, annotation support, and presence of an output schema (which documents return values), the description provides sufficient context. It covers what, scope, and behavior, meeting completeness requirements.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, with the filter parameter fully described in the schema. The tool description does not add any additional parameter semantics beyond the schema, which meets the baseline. No extra context is needed.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'List teams visible to the caller (owned + member of)' clearly states the action (list) and resource (teams), and specifies the scope of visible teams. This distinguishes it from sibling tools like 'list_team_members' (which lists members of a specific team) and other list tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description does not explicitly state when to use this tool versus alternatives. However, the name and context (siblings) imply it's the go-to for listing teams. No usage exclusions or alternative recommendations are provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_templatesA
Read-only
Inspect

List public templates.

Anyone can read. filter='mine' restricts to the caller's templates. Archived templates are excluded unless include_archived is set, and even then only the caller's own archived templates are returned.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMaximum number of items to return in one page (1 to 200).
cursorNoPagination cursor from a previous response's next_cursor; omit to start from the first page.
filterNoWhich templates to return: 'mine' (you own) or 'all' public templates.all
searchNoCase-insensitive substring filter over name, description, and tags.
include_archivedNoWhen true, the caller's own archived templates are included alongside live ones. Archived templates belonging to other owners are never surfaced. Anonymous callers and non-owners always see only live templates.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already mark readOnlyHint=true, so description adds value by specifying that archived templates are excluded unless include_archived is set, and only own archived are returned. No contradictions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two concise sentences, front-loaded with main purpose. Every sentence adds value; no fluff.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Covers key aspects: visibility, filtering, archived behavior. With complete schema, annotations, and output schema, description is sufficient for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, but description adds meaningful context for filter and include_archived parameters beyond the schema. Explains behavior of 'mine' filter and archived inclusion.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Clearly states 'List public templates' with a specific verb and resource. Distinguishes from sibling 'search_templates' by focusing on basic listing with filters.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides clear contexts: 'Anyone can read', 'filter='mine' restricts to own templates', and archived behavior. Lacks explicit exclusion of when to use search_templates instead, but still guides usage well.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_verifiersA
Read-only
Inspect

List active (non-revoked) verifiers visible to the caller.

Returns {items: [{verifier_id, name, description, current_version, status, version_token, created_at, updated_at, role, source_workflow_id}], next_cursor: str | null}. Includes owned verifiers plus skill-derived grants. Platform-managed scope=system verifiers never appear. Anonymous callers get an empty list. Use get_verifier to read the full deploy-time config of a specific version; any grantee who can reach the verifier can read it.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMaximum number of items to return in one page (1 to 200).
cursorNoPagination cursor from a previous response's next_cursor; omit to start from the first page.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=trueebb. The description adds meaningful behavioral detail beyond that: the inclusion/exclusion rules (owned plus grants, never scope=system), the anonymous-caller empty-list behavior, and the visibility constraint. These go beyond what the read-only annotation conveys.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four dense sentences, each carrying distinct information: what is listed, what is returned, what is excluded, and which sibling to use for full config. No filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the low parameter count, existing input schema, readOnly annotation, and output schema signal, the description covers visibility rules, exclusions, pagination, and the related get_verifier tool. Nothing an agent needs to choose or invoke this tool is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% for both parameters (limit and cursor), so the description is not obligated to re-explain them. It mentions next_cursor coupling but does not add new semantic value beyond the schema definitions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: "List active (non-revoked) verifiers visible to the caller." It clearly distinguishes this from get_verifier by specifying that list_verifiers returns the visible set while get_verifier reads full deploy-time config of a specific version. The scope (owned + skill-derived, excluding platform-managed system verifiers) is explicit.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly directs users to get_verifier when they need the full deploy-time config, giving an actionable alternative and the condition that selects it. It also provides important context: anonymous callers get an empty list and platform-managed verifiers are never included.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

lookup_fork_lineageA
Read-only
Inspect

Check whether a forked skill's upstream template has new versions or was unpublished (returns parent_template_id, parent_template_version, upstream_latest_version, is_upstream_unpublished).

ParametersJSON Schema
NameRequiredDescriptionDefault
skill_idYesThe skill to act on, identified by its UUID, slug, or name.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, so the agent knows this is a safe read operation. The description adds value by disclosing the specific behavioral scope: it checks upstream version status and unpublish state, and it returns a defined set of fields. It doesn't mention edge cases like what happens if the skill isn't a fork, but the read-only nature is well covered.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, compact sentence that front-loads the core purpose and lists the return fields in parentheses. Every word earns its place; no filler or repetition of the tool name.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only lookup tool with one fully documented parameter and an output schema, the description is nearly complete. It explains what the tool checks and what it returns. The only minor gap is not stating behavior for non-forked skills or error conditions, but the output schema and read-only annotation cover most of what an agent needs.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already fully documents the single parameter (skill_id). The description doesn't add parameter-level detail beyond what the schema provides, but it does clarify the purpose of the parameter in context (acting on a forked skill). Baseline 3 is appropriate since the schema carries the load.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Check whether') and resource ('a forked skill's upstream template'), and clearly distinguishes its purpose from siblings like get_skill or get_template by focusing on fork lineage and upstream version status. It also lists the exact return fields, making the tool's scope unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies when to use this tool: when you need to know if a forked skill's upstream has new versions or was unpublished. It doesn't explicitly name alternatives or exclusions, but the context of fork lineage is clear enough to guide an agent away from generic get_skill/get_template calls. A brief mention of when not to use it would push this to 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

mint_api_keyAInspect

Mint a new good_live_ API key. The secret is returned ONCE.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYesA label to identify this API key in later listings.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Despite having annotations that are all false (no hints), the description explicitly warns that the secret is returned only once, a critical behavioral trait beyond what annotations provide.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely concise with a single sentence that front-loads the key action and includes a critical note about the one-time return, making it efficient and to the point.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple creation tool with an output schema, the description is minimal. It covers the essential behavioral note but lacks broader context such as security implications, prerequisites, or how it fits into the overall key lifecycle.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema already describes the 'name' parameter fully (100% coverage), and the description adds no extra meaning beyond what the schema provides, meeting the baseline of 3.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('Mint') and resource ('good_live_ API key'), and distinguishes from siblings like list_api_keys and revoke_api_key by its unique creation function.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies the tool is used to create a new API key but does not provide explicit guidance on when to use versus alternatives like revoke_api_key, nor does it mention prerequisites or conditions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

optimize_descriptionA
Read-only
Inspect

Tune an existing skill's trigger: sharpen its description so it fires at the right times.

Pack-returner: response is {skill_md, references, skill_id, max_iterations}. The agent runs the loop locally with the user to tune the skill's description (the text that decides when the skill fires) for trigger accuracy, then persists the winner via save_skill(source='description_optimization') only after explicit user approval. Only the description changes; body, outcome, tags, and sibling files carry forward unchanged. Requires edit access on the skill.

ParametersJSON Schema
NameRequiredDescriptionDefault
skill_idYesThe skill to act on, identified by its UUID, slug, or name.
max_iterationsNoOptimization loop budget (1 to 1000). Defaults to 10 when omitted.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description discloses that only the description changes, while body, outcome, tags, and sibling files carry forward unchanged. It also states the requirement for edit access and explicit user approval before persisting. The annotations declare readOnlyHint=true, which aligns with the description's emphasis on local tuning and approval-gated persistence. The description adds meaningful behavioral context beyond the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and front-loaded with the core purpose. Every sentence adds value: the purpose, the workflow, the scope of changes, and the access requirement. No filler or redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description is complete for a tool with 2 parameters, full schema coverage, and an output schema. It explains the workflow, the constraints, the approval requirement, and the persistence path. An agent has enough context to select and invoke the tool correctly without needing additional details.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the schema already documents both parameters. The description adds context by explaining the optimization loop budget ('max_iterations') and the persistence condition, which helps the agent understand the parameter's role in the workflow. It doesn't add syntax details, but the schema already covers those.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Tune'), a specific resource ('an existing skill's trigger'), and the exact mechanism ('sharpen its description so it fires at the right times'). It also distinguishes itself from the sibling 'optimize_skill' by focusing on description/trigger accuracy rather than general optimization.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly says when to use this tool: to tune a skill's description for trigger accuracy. It also names the persistence path ('save_skill(source='description_optimization')') and the requirement of explicit user approval, which clarifies the workflow and when it should be invoked.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

optimize_skillA
Read-only
Inspect

Automatically improve an existing skill: runs an optimization loop to lift it against its outcome.

Pack-returner: response is {skill_md, references, skill_id, max_iterations}. The agent runs the optimization loop locally with the user, drives Researcher / Editor / Runner subagents over a locked scenario set, and persists the winner via save_skill(source='optimization') only after explicit user approval. Requires edit access on the skill.

ParametersJSON Schema
NameRequiredDescriptionDefault
skill_idYesThe skill to act on, identified by its UUID, slug, or name.
max_iterationsNoOptimization loop budget (1 to 1000). Defaults to 20 when omitted.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.6/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Description describes a mutating operation (persisting via save_skill after user approval), but annotation has readOnlyHint=true, which is a direct contradiction.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Multiple sentences but each adds value; front-loaded with purpose and includes necessary technical details like return format. Could be slightly tighter.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Covers process flow, subagents, user approval, edit requirement, and output structure. Has output schema available, so return values are documented.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, and description adds little beyond schema; mentions max_iterations default but no new semantic detail.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Clear verb+resource: 'automatically improve an existing skill' distinguishes it from siblings like save_skill or teach_skill.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

States requirement of edit access and explains the optimization loop process, but doesn't explicitly mention when not to use or compare with alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

publish_template_versionAInspect

Publish the latest version of your skill as a new public template version.

First publish creates the template (slug reused from the skill). Each call appends a monotonic version. Requires a claimed handle. Bundle a 'demo/README.md' writeup with images referenced by relative path inside 'demo/' to render a visual demo on the public template page.

Publishing makes the template's verifier definitions public: anyone who views the template (including anonymous readers) can read each verifier's criterion and calibration examples. Publishing is hard-blocked if a verifier definition contains a secret or credential, and flagged if it appears to contain private data. When the template references verifiers, the response includes a 'verifier_exposure_notice' restating this.

ParametersJSON Schema
NameRequiredDescriptionDefault
skill_idYesThe skill to act on, identified by its UUID, slug, or name.
release_notesNoOptional notes describing what changed in this published version.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The annotation hints are all false/neutralressive, so the description carries the behavioral disclosure burden. It explains side effects clearly: publishing makes verifier definitions public, a hard block occurs on secrets/credentials, private data triggers a warning flag, and the response includes a verifier_exposure_notice when verifiers are referenced. This is far beyond the annotations and highly relevant for an agent deciding to publish.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Well structured, front-loaded with core purpose, then behavior; no redundant sentences. Each clause carries unique info: versioning, prereq, demo doc, privacy side effects, blocking rule, response field.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Covers the non-obvious aspects an agent must know: template creation on first publish, monotonic versioning, claimed-handle requirement, verifier exposure/privacy implications, blocked vs flagged states, and the response notice field. Nothing critical appears missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Both parameters are already fully described in the input schema (100% coverage), so the description is not required to re-explain them. It adds a context-level prerequisite (claimed handle) but does not deepen the parameter definitions themselves. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and object: 'Publish the latest version of your skill as a new public template version.' It clarifies the key semantic point that the first publish creates the template and each subsequent call appends a version, which fully disambiguates the tool's purpose and scope.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear context: when to publish, that a claimed handle is required, and what happens on first vs. subsequent publishes. It doesn't name alternatives like unpublish_template_version, but the unique publish action is unambiguous, and the prerequisite and sequence guidance is enough for an agent to decide when to call it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

purchase_creditsAInspect

Buy a one-time credit top-up, charging a card on file or falling back to checkout.

Charges the default payment method on file when one exists and returns the new balance. Otherwise returns a secure hosted checkout link so the purchase can complete interactively. Errors clearly when self-service billing is not enabled on this deployment, or when the requested amount falls outside the allowed range.

ParametersJSON Schema
NameRequiredDescriptionDefault
amount_usdYesAmount of credits to buy, in whole US dollars.
idempotency_keyNoOptional caller-supplied key that makes a retried purchase request safe to resend without buying credits twice.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations indicate readOnlyHint false (mutation), idempotentHint false, destructiveHint false. The description adds context: charges default payment method or returns checkout link, and returns new balance. It also notes error conditions. The description aligns with annotations and provides useful behavioral details.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is succinct, using 4-5 sentences with no unnecessary fluff. It effectively front-loads the core action and then details the two possible scenarios and error conditions.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has an output schema, the description explains the two possible outcomes (balance or checkout link) and errors. This covers the essential context for an agent to invoke the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema has 2 parameters, both with descriptions (schema coverage 100%). The description does not add new information about parameters beyond what the schema provides, so baseline score of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states it buys a one-time credit top-up, distinguishing it from sibling tools like configure_auto_top_up or cancel_subscription. The verb 'buy' and resource 'credit top-up' are specific and unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explains when to use (one-time purchase) and describes two paths: charging card on file or falling back to checkout link. It also mentions error conditions (self-service billing disabled, amount out of range). However, it does not explicitly state when not to use or suggest alternatives like configure_auto_top_up for recurring purchases.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

redeem_referral_codeAInspect

Redeem a referral code to claim your one-time new-user bonus credits.

ParametersJSON Schema
NameRequiredDescriptionDefault
codeYesThe referral code to redeem for one-time bonus credits.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations indicate mutation (readOnlyHint=false) and non-destructive (destructiveHint=false). The description adds 'one-time' constraint, which is a key behavioral trait beyond annotations. However, it does not detail error handling or side effects.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Single sentence, no wasted words, front-loads the action and outcome. Perfectly concise for the tool's simplicity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the low complexity (1 parameter, output schema exists), the description covers the main behavioral constraint ('one-time') and purpose. Minor lack of explicit prerequisites or error cases, but adequate for a simple redeem action.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and the schema already describes the parameter 'code' as 'The referral code to redeem for one-time bonus credits.' The description does not add new semantic information beyond reinforcing the purpose, so baseline score applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('Redeem'), the resource ('referral code'), and the outcome ('claim your one-time new-user bonus credits'). It distinguishes from sibling 'get_referral_status' which checks status rather than redeeming.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Implies usage for new users with a referral code to claim bonus credits, but does not explicitly state when not to use it (e.g., if already redeemed) or mention the alternative 'get_referral_status' for checking eligibility.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

remove_team_memberBInspect

Remove a team member by UUID, email, or handle.

ParametersJSON Schema
NameRequiredDescriptionDefault
team_identifierYesThe team to act on, identified by its UUID or handle.
user_identifierYesThe target user, identified by their UUID, email, or handle.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description only says 'Remove' but does not elaborate on side effects (e.g., impact on shared workflows, permissions, or content). Annotations set destructiveHint to false, which might mislead about the actual effects of removal. No mention of reversibility or required permissions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Single sentence, front-loaded with the verb, no redundancy. Every word contributes to understanding the tool's core function.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple tool with two parameters and an output schema (not shown), the description provides the essential information. However, it omits behavioral context like idempotency (idempotentHint false) and potential failure modes, which could be important for an AI agent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and both parameters have clear descriptions matching the tool's description. The description reinforces the accepted identifier formats but adds no new meaning beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('Remove'), the resource ('team member'), and the allowed identification methods (UUID, email, or handle). It distinguishes from the sibling tool 'add_team_member' by specifying removal rather than addition.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool vs alternatives like 'cancel_invitation' for pending invites or 'leave_shared_workflow'. The description does not mention prerequisites (e.g., ownership/admin status) or scenarios where removal is appropriate.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

rename_handleAInspect

Rename your already-claimed handle.

Enforces 1 rename per rolling 90 days plus 3 per UTC calendar year. Self-reclaim of your own released handle within its 90-day reservation window is free and does not consume a rename slot.

ParametersJSON Schema
NameRequiredDescriptionDefault
handleYesThe new handle to rename to.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations provide little behavioral info (all false), so the description carries the burden. It discloses important constraints like rate limits and the free reclaim exception, which are critical for the agent to understand side effects. However, it does not mention error behavior or permissions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely concise with two sentences that front-load the main action. Every sentence adds essential information with no redundancy or filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's simplicity (one parameter, basic operation) and existence of an output schema, the description covers purpose and key constraints. It lacks details on return values or error conditions, but those are partially addressed by the output schema and annotations.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema covers 100% of the parameter 'handle' with a clear description. The tool description adds no additional semantic detail about the parameter beyond the schema, so it meets the baseline of 3.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states 'Rename your already-claimed handle,' providing a specific verb and resource. It distinguishes itself from sibling tools like 'claim_handle' by focusing on renaming an existing handle, not claiming a new one.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly mentions rename limits (1 per 90 days + 3 per year) and the special case of free self-reclaim within the 90-day reservation window. While it doesn't explicitly state when not to use or name alternatives, these details guide appropriate usage.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

revoke_api_keyB
Idempotent
Inspect

Revoke an API key you own. Idempotent.

ParametersJSON Schema
NameRequiredDescriptionDefault
key_idYesThe API key to act on, identified by its id.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already provide idempotentHint=true and destructiveHint=false. The description only repeats 'Idempotent' and adds 'you own' for authorization, but no additional behavioral traits beyond annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, front-loaded with verb, no wasted words. Highly concise.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present and simple action, the description covers ownership and idempotency. Could mention effect on existing tokens, but adequate for the tool's simplicity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and the parameter description is clear. The tool description does not add extra meaning to the parameter beyond schema, meeting baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('Revoke') and the resource ('API key you own'), distinguishing it from sibling tools like mint_api_key and list_api_keys. It is specific and not a tautology.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool versus alternatives (e.g., revoke_image_generator, delete_*). The idempotency note is useful but does not provide context on selection.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

revoke_image_generatorA
Idempotent
Inspect

Revoke a user-owned image generator by UUID.

Platform scope=system generators cannot be revoked (NotFound).

Sets status="revoked" and revoked_at. Revoked generators disappear from list/get/generate for you (subsequent calls surface as NotFound). Existing generation run rows are kept for audit. There is no un-revoke; deploy a fresh generator under a new name to replace one. Returns {generator_id, name, revoked: true}.

ParametersJSON Schema
NameRequiredDescriptionDefault
generator_idYesThe image generator to act on, identified by its UUID.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations provide idempotentHint=true and destructiveHint=false, and the description adds rich context: status change, field updates, visibility effects, audit retention, and irreversibility. This goes well beyond annotations, fully disclosing behavioral traits.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise, front-loaded with the purpose, and uses bullet-point style sentences that each add essential information. No redundant or excessive wording.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's simplicity (one parameter, clear output), the description covers all necessary aspects: action, conditions, side effects, irreversibility, and return format. Output schema exists but the description still sufficiently explains the result.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the parameter is straightforward (generator_id). The description repeats that it's a UUID but adds no new semantic meaning beyond what the schema already provides. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('Revoke') and resource ('user-owned image generator'), with explicit verb+resource. It distinguishes from siblings by specifying that system-scope generators cannot be revoked and by detailing the post-revocation behavior that differs from deletion tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explains when to use (user-owned generators) and when not (system generators cause NotFound). It also advises that revocation is irreversible and suggests deploying a fresh generator as an alternative, providing good guidance without explicitly comparing to delete_image_generator.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

revoke_skill_grantA
Idempotent
Inspect

Revoke a direct skill grant for a user or team identifier.

ParametersJSON Schema
NameRequiredDescriptionDefault
skill_idYesThe skill to act on, identified by its UUID, slug, or name.
grantee_email_or_at_team_handleYesThe grantee whose access to revoke: a user's email, or a team's @handle.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already provide idempotentHint=true and destructiveHint=false. Description adds no extra behavioral context beyond the verb 'revoke', which implies a mutation. Does not contradict annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, clear sentence that fully conveys the purpose. No redundant or missing words. Front-loads the core action.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the high schema coverage, annotations, and presence of an output schema, the description adequately covers the operation. Slightly lacking in specifying the effect or prerequisites, but acceptable.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Both parameters are fully described in the input schema (100% coverage). The description adds no additional meaning beyond the schema definitions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action (revoke) and the resource (direct skill grant) with target (user or team identifier). It distinguishes from sibling tools like 'grant_skill' and 'list_skill_grants'.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool versus alternatives like 'grant_skill' or when not to use it. No context about prerequisites or error states.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

revoke_verifierA
Idempotent
Inspect

Revoke a verifier you own (scope=user only).

verifier_id accepts UUID string or caller-owned name. Platform scope=system verifiers cannot be revoked (NotFound).

Sets status="revoked" and revoked_at. Revoked verifiers disappear from list/get/run for you (subsequent calls surface as NotFound). Existing verifier_runs rows are kept for audit. There is no un-revoke; deploy a fresh verifier under a new name to replace one. Returns {verifier_id, name, revoked: true}.

ParametersJSON Schema
NameRequiredDescriptionDefault
verifier_idYesThe verifier to act on, identified by its UUID or, for a verifier you own, its name.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description discloses all relevant behavioral traits: it sets status and revoked_at, causes revoked verifiers to disappear from list/get/run, existing runs are kept for audit, and cannot be reversed. Annotations (idempotentHint=true) are consistent because once revoked, re-revoking has no further effect.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-structured and concise, front-loading the primary purpose and then adding important details. A minor improvement could be combining the first two sentences, but it is already efficient.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the single parameter, the existence of an output schema (implied by the return description), and the rich behavioral details, the description is fully complete. No further information is needed for correct use.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% with a description for verifier_id. The description adds value by clarifying that it accepts UUID or caller-owned name, which is not in the schema. This aids correct invocation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action (revoke a verifier you own) and distinguishes it from sibling tools like delete_verifier by explaining that revoked verifiers can be replaced but not un-revoked. The verb+resource combination is specific and unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly states when the tool is applicable (user-owned verifiers) and when it is not (scope=system verifiers, which cause a NotFound). It also provides clear guidance on what to do instead of un-revoking (deploy a fresh verifier).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

run_verifierAInspect

Run a verifier on agent output to get a pass/fail judgment with reasoning.

verifier_id resolves to an active verifier: UUID string, accessible user-scope name (owned or one unambiguous grant), or platform alias system:<name> for platform-managed judges (no criterion/config leakage via list/get).

inputs keys must match the version's input_fields exactly (no missing or extra). media_url is required for text_image and image contracts, forbidden for text. Caller-shape errors return a 400 with no row written; judge runtime errors persist a row with status="error" and an error_code of runtime_error, verifier_unavailable, or timeout.

Provenance fields (skill_id, skill_version, skill_ref, run_id) are optional and stamped onto the row. skill_id is access-checked: a skill the caller cannot see surfaces as NotFound.

Returns: {verifier_run_id, verifier_id, version, status, passed, reasoning, duration_ms, created_at} on success, plus error_code and error_message on error. passed is null on errors. On a successful run that you attributed to a skill you own, the result also carries skill_id.

ParametersJSON Schema
NameRequiredDescriptionDefault
inputsYesField values for the verifier, keyed by the version's input field names (exact match, no missing or extra keys).
run_idNoOptional caller-supplied run identifier to correlate this run, for provenance.
versionNoPin to a specific version number; omit to use the latest version.
skill_idNoOptional id of the skill this run is associated with, for provenance; a skill you cannot see is rejected.
media_urlNoPublic image URL to judge; required for image and text+image verifiers, rejected for text-only ones.
skill_refNoOptional free-form skill label (slug or name) to stamp on the run, for provenance.
verifier_idYesThe verifier to act on, identified by its UUID or, for a verifier you own, its name.
skill_versionNoOptional skill version number to stamp on the run, for provenance.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description goes well beyond the annotations (readOnlyHint=false, destructiveHint=false, etc.) by disclosing side effects: successful runs persist a row, judge runtime errors persist a row with status='error', caller-shape errors return 400 with no row written, and provenance fields are stamped. It also explains access-checking behavior for skill_id. This is rich behavioral context that annotations alone cannot convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but well-organized: it front-loads the core purpose, then covers identifier resolution, input constraints, error semantics, provenance, and return shape. Every sentence earns its place. It loses one point because it is quite long and could be slightly tightened, but the density is justified given the tool's complexity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (8 params, nested inputs object, output schema, error states), the description is remarkably complete. It covers identifier resolution, input validation, media rules, error behavior, provenance, and return values. The output schema exists, so return values are documented, but the description still summarizes them helpfully. Nothing an agent needs to call this correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the schema already documents all 8 parameters. The description adds value by explaining the verifier_id resolution rules (UUID, user-scope name, or system: alias), the exact-match requirement for inputs keys, and the media_url contract rules. It doesn't repeat every parameter but adds semantic context that the schema lacks. A 4 is appropriate because the description meaningfully supplements the schema without being redundant.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Run a verifier on agent output to get a pass/fail judgment with reasoning.' This clearly distinguishes the tool from siblings like get_verifier, list_verifiers, deploy_verifier, and revoke_verifier. The purpose is unambiguous and action-oriented.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides explicit guidance on when to use this tool and how to avoid errors: it explains verifier_id resolution, input key matching requirements, media_url rules per contract type, and error behavior. It also implicitly distinguishes from get_verifier (which retrieves config) and deploy_verifier (which manages lifecycle). This is comprehensive usage guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

save_skillAInspect

Save a skill to your private registry so you can run, share, or improve it later (name, description, body, outcome, tags).

Addressing: omit skill_id to save in your own namespace, which updates your own skill of this name or creates a new one. Saving by name never edits a skill someone else shared with you. To edit a skill you own or have edit access to, pass its skill_id (and a matching name). If a skill of this name is shared with you at edit access and you save by name alone, the reference is ambiguous and the call asks you to address it by id.

Optional source ('manual' | 'teach' | 'optimization' | 'description_optimization' | 'audit') attributes the version to a manual save, a teach_skill session, a skill optimization session, a description optimization session, or an audit_skill session; omit to leave NULL.

Returns skill_id and version. Skills are private to the caller; share via publish_template_version.

files is a full snapshot: any path you do not include is deleted. Only pass files when you hold the complete tree. To change specific files in a skill that already exists, use update_skill_files, which keeps the paths you do not name. When a save does drop paths, the response lists them under deleted, so check that key and restore anything you did not mean to remove.

ParametersJSON Schema
NameRequiredDescriptionDefault
payloadYesPublish-skill payload aligned with the Claude Code skills convention. Required top-level fields mirror a Claude Code ``SKILL.md`` frontmatter: ``name`` (kebab-case handle) and ``description`` (one-sentence summary). ``outcome`` and ``tags`` are top-level discovery facets surfaced by ``list_skills``; everything else lives in the ``body`` and is opaque to the registry. Skills are always private to the caller. Public sharing is the ``templates`` surface (``publish_template_version`` / ``fork_template``); there is no ``visibility`` knob on a skill row. Unknown fields are silently dropped (``extra="ignore"``). Fields such as ``manifest``, ``visibility``, and ``kpi`` do not cause validation errors and are simply ignored. The server accepts both ``slug`` and ``name`` as input for the skill handle. Responses return both ``name`` and ``slug`` fields with identical values.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Adds significant behavioral context beyond annotations: skills are private (shared via publish_template_version), files is a full snapshot (deleting omitted paths), response includes deleted files for recheck, and saving by name never edits shared skills. No contradiction with annotations; clarifies conditional destructiveness not captured by destructiveHint.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Well-structured with clear paragraphs: first sets purpose, then addressing behavior, source, return/sharing, and files. Every sentence adds value. Slightly lengthy but justified by complexity; could potentially combine some points without losing clarity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (nested schema, file snapshot semantics, sharing nuances) and presence of output schema, the description is highly complete. Covers edge cases (shared skill ambiguity, concurrency token), relationship with siblings, and notes what's not included (visibility knob). Fulfills all informational needs.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, but description adds extra meaning: explains relationship between skill_id and name, files snapshot versus null preservation, source attribution, expected_version_token concurrency, and the fact that unknown fields are silently dropped. Also notes server acceptance of both slug and name.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool saves a skill to a private registry for later run, share, or improvement. It highlights key fields (name, description, body, outcome, tags) and distinguishes between create and update via skill_id, with explicit mentions of sibling tools like update_skill_files.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides explicit when/when-not guidance: omit skill_id to save in own namespace (create or update own skill); pass skill_id to edit a specific skill; warns about ambiguous name sharing for edit-access skills. Also explains files snapshot behavior versus incremental update using update_skill_files, and advises to check the deleted key in response.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_skillsA
Read-only
Inspect

LLM-ranked natural-language search over skills visible to you.

This does not perform lexical query prefiltering. Use list_skills with search=... for deterministic metadata filtering.

ParametersJSON Schema
NameRequiredDescriptionDefault
payloadYesNatural-language search request over skills visible to the caller.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=true, and the description adds 'LLM-ranked' to indicate AI-based ranking, which goes beyond annotations. No contradictions. Slightly more detail about visibility could be included, but it's well covered.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two concise sentences: the first defines the tool's primary function, the second clarifies its limitations and alternative. No extraneous words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's moderate complexity, the rich schema, and annotations (readOnlyHint=true), the description is largely complete. The presence of an output schema compensates for lack of return value details. Could mention that results are ranked and what 'visible to you' means precisely, but still adequate.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 100% schema description coverage, the schema already documents all parameters (tag, limit, query, filter) thoroughly. The description adds only a high-level phrase ('Natural-language search request over skills visible to the caller') without new parameter-specific details. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states 'LLM-ranked natural-language search over skills visible to you,' establishing a specific verb (search), resource (skills), and method (natural language, LLM-ranked). It distinguishes itself from sibling tools like list_skills by mentioning lexical prefiltering.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly says 'This does not perform lexical query prefiltering. Use list_skills with search=... for deterministic metadata filtering.' This provides clear when-to-use and when-not-to-use guidance with an alternative tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_templatesA
Read-only
Inspect

LLM-ranked natural-language search over public or owned templates.

This does not perform lexical query prefiltering. Use list_templates with search=... for deterministic metadata filtering.

ParametersJSON Schema
NameRequiredDescriptionDefault
payloadYesNatural-language search request over public or owned templates.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, so no mutation expected. Description adds value by specifying 'LLM-ranked' and 'natural-language search', and clarifies that it does not do lexical prefiltering. No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two concise sentences. First sentence states purpose, second provides usage guidelines (when not to use and alternative). No wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity of a search tool with an output schema, the description is complete: it explains the search method, scope (public/owned), and limitation (no lexical prefiltering). With output schema present, no need to describe return values.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% with detailed parameter descriptions (query, limit, filter). The tool description does not add additional parameter semantics beyond what schema provides, so baseline score is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: 'LLM-ranked natural-language search over public or owned templates.' It uses a specific verb (search) and resource (templates), and distinguishes from sibling 'list_templates' by implying semantic vs metadata search.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly states when to use this tool vs alternative: 'This does not perform lexical query prefiltering. Use list_templates with search=... for deterministic metadata filtering.' Provides clear exclusion and alternative.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

teach_skillA
Read-only
Inspect

Improve an existing skill by hand: teach it from examples and corrections you provide.

The agent follows the pack locally with the user and persists changes via save_skill(source='teach'). Requires edit access on the skill.

ParametersJSON Schema
NameRequiredDescriptionDefault
skill_idYesThe skill to act on, identified by its UUID, slug, or name.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.6/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description indicates mutation ('improve', 'persists changes') while the annotations declare readOnlyHint: true, which is a contradiction. This severely undermines transparency.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is three sentences, front-loaded with the key action, and every sentence adds essential information without fluff.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the simple single-parameter tool, the description covers purpose, permissions, and persistence. However, the annotation contradiction detracts from overall completeness.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 100% schema coverage and only one parameter, the description adds no extra semantic meaning beyond the schema's existing description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb 'improve' and the resource 'existing skill', and distinguishes it from siblings like 'save_skill' by specifying the teaching process.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description mentions the prerequisite of edit access and the persistence mechanism via save_skill, but does not explicitly contrast with alternative modification tools like optimize_skill.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

transfer_skill_ownershipAInspect

Transfer skill ownership to another user. Owner only.

ParametersJSON Schema
NameRequiredDescriptionDefault
skill_idYesThe skill to act on, identified by its UUID, slug, or name.
new_owner_user_id_or_emailYesThe user to transfer ownership to, by UUID or email.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.5/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations provide behavioral context, and the description only states 'transfer ownership' without disclosing behavioral traits such as role changes, acceptance requirements, or reversibility.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, front-loaded with the core action and constraint, no redundant information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple mutation tool with good schema coverage and output schema, the description is adequate but lacks context about consequences or post-transfer state.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% for both parameters; the description adds no extra meaning beyond the schema, warranting a baseline score of 3.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb 'transfer', the resource 'skill ownership', and the constraint 'Owner only', distinguishing it from sibling tools like transfer_team_ownership.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage is restricted to the current owner but does not provide explicit guidance on when to use this tool versus alternatives like transfer_template_ownership.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

transfer_team_ownershipBInspect

Transfer ownership to a user by UUID, email, or handle.

ParametersJSON Schema
NameRequiredDescriptionDefault
team_identifierYesThe team to act on, identified by its UUID or handle.
new_owner_user_identifierYesThe user to transfer ownership to, by UUID, email, or handle.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate non-read-only (readOnlyHint=false) and non-destructive (destructiveHint=false). The description adds that ownership is transferred, but does not disclose additional behavioral traits such as permission requirements, effect on current owner, or reversibility. No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence that is efficient and front-loaded. It is concise but could be more structured (e.g., separating team and user identifiers). No wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the presence of an output schema, the description does not need to explain return values. However, for a mutation tool, it lacks context on side effects (e.g., current owner loses access) or required permissions. It is minimally adequate.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so baseline is 3. The description repeats that identifiers can be by UUID, email, or handle, which matches the schema. It adds no new semantic information beyond what the schema already provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action (transfer ownership) and the object (a team), with specific methods (by UUID, email, or handle). It distinguishes from sibling tools like transfer_template_ownership and transfer_workflow_ownership, but doesn't explicitly mention 'team' in the description text, relying on the tool name.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool versus alternatives like transfer_template_ownership or transfer_workflow_ownership. There is no mention of prerequisites (e.g., being an admin) or when not to use it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

transfer_template_ownershipAInspect

Transfer template ownership by UUID or @handle/slug. Owner only.

Team targets are rejected. Existing publishing-handle stamps on prior versions are immutable per VISION 14.

ParametersJSON Schema
NameRequiredDescriptionDefault
template_idYesThe template to act on, identified by its UUID or @handle/slug.
new_owner_user_id_or_emailYesThe user to transfer ownership to, by UUID or email.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations are all false but the description adds important behavioral traits: ownership requirement, rejection of team targets, and immutability of handle stamps on prior versions per VISION 14. This goes beyond what annotations provide.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is very concise, using three short sentences that each add unique value. It is front-loaded with the core action and includes key constraints without unnecessary detail.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given that an output schema exists (not shown), the description covers all essential aspects: action, identification, ownership restriction, invalid targets, and a behavioral note about immutability. It is complete for a transfer operation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already fully describes both parameters. The description restates identification methods (UUID, @handle/slug, email) but does not add new meaning beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action (transfer template ownership) and identifies the resource (template). It specifies identification by UUID or @handle/slug and notes 'Owner only', which distinguishes it from sibling transfer tools like transfer_workflow_ownership and transfer_team_ownership.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly states 'Owner only' and 'Team targets are rejected', providing clear context on when to use and what inputs are invalid. While it doesn't name alternatives, the sibling list implies other transfer tools exist for different cases.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

unarchive_skillA
Idempotent
Inspect

Restore a skill you previously archived. Idempotent.

Reverses archive_skill. Clears archived_at and re-derives verifier grants that were removed during archiving.

ParametersJSON Schema
NameRequiredDescriptionDefault
skill_idYesThe skill to act on, identified by its UUID, slug, or name.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the annotations (idempotentHint=true), the description adds specific behavioral details: it clears archived_at and re-derives verifier grants removed during archiving. No contradictions with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise with three sentences: purpose, idempotency, and detailed effect. No redundant information. Well-structured and front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers key behavioral aspects and has an output schema. Minor gap: no mention of error conditions or state prerequisites (e.g., skill must be archived). Otherwise complete for a simple restoration tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema already fully describes the single parameter (skill_id) with a clear description. The description adds no additional semantic information beyond what is in the schema, so baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb 'Restore' and the resource 'a skill you previously archived'. It explicitly distinguishes itself from sibling tools like archive_skill and delete_skill by indicating it reverses the archiving operation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description states it reverses archive_skill, which implies usage context. However, it does not explicitly specify when not to use or mention prerequisites like ensuring the skill is archived, leaving some ambiguity.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

unarchive_templateA
Idempotent
Inspect

Restore an archived template by UUID or @handle/slug.

Owner only. Idempotent on a live template. Cannot raise Conflict because the archived row already holds its slug slot in the full unique constraint.

ParametersJSON Schema
NameRequiredDescriptionDefault
template_idYesThe template to act on, identified by its UUID or @handle/slug.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description adds value beyond annotations by explaining why no conflict can occur ('archived row already holds its slug slot'), which aligns with idempotentHint=true and destructiveHint=false. No contradiction.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences with no wasted words, front-loaded with the core purpose, and each subsequent sentence adds necessary behavioral context.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers purpose, ownership, idempotency, and conflict behavior. It could mention the effect of restoration on template state (e.g., becomes active), but overall sufficient given output schema exists.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema already covers 100% of the parameter description (template_id as UUID or @handle/slug). The description repeats this without additional meaning, so baseline of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description explicitly states 'Restore an archived template' which is a specific verb and resource, and distinguishes from siblings like 'archive_template' and 'delete_template' by focusing on the unarchive action.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description clearly states 'Owner only' and 'Idempotent on a live template', providing context for appropriate use. It does not explicitly mention when not to use or alternatives, but the ownership constraint is key.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

unpublish_template_versionA
Idempotent
Inspect

Soft-unpublish a single template version. Owner-only.

Existing forks pinned to this version keep working. The catalog hides the template if no live version remains.

ParametersJSON Schema
NameRequiredDescriptionDefault
versionYesThe version number to unpublish.
template_idYesThe template to act on, identified by its UUID or @handle/slug.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Describes behavioral traits: 'soft-unpublish' (non-destructive mutation), existing forks pinned to this version keep working, catalog hides template if no live version remains. Consistent with annotations (idempotentHint=true, destructiveHint=false). Adds detail beyond annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Extremely concise: three short sentences. Front-loaded with core purpose in first sentence, then behavioral details. Every sentence adds value without redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Covers key aspects: purpose, owner-only restriction, side effects on forks and catalog visibility. With output schema present, return value explanation is not needed. Could mention idempotency or error cases but not necessary for this simple operation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and already provides descriptions for both parameters (version number, template_id as UUID or handle). Description adds no additional parameter meaning beyond the schema, so baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Clearly states 'Soft-unpublish a single template version' with specific verb 'soft-unpublish' and resource 'template version'. Distinguishes from siblings like publish_template_version, deprecate_template_version, and delete_template_version. Also notes owner-only restriction.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides clear context: behavior regarding forks and catalog hiding when no live version remains. Implies use case for removing from catalog without breaking forks. However, lacks explicit comparison to alternatives like delete_template_version or deprecate_template_version.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update_imageAInspect

Update visibility, TTL, or the private view link of a hosted image.

Caller must own it.

visibility flips the access level: "public" or "private". Omit to leave unchanged.

ttl_seconds and permanent are mutually exclusive. permanent clears the expiry so the image lives indefinitely. ttl_seconds sets a new expiry relative to now (positive integer). Omitting both leaves the current expiry unchanged.

rotate_view_secret issues a fresh private view link and invalidates every link shared earlier for this image, so use it to un-share a private image. For a private image the returned url is the current view link.

Returns the updated image record. Raises image_not_found (404) when the image is absent, expired, or owned by another user.

ParametersJSON Schema
NameRequiredDescriptionDefault
image_idYesThe hosted image to act on, identified by its UUID.
permanentNoClear the expiry so the image lives indefinitely; mutually exclusive with ttl_seconds.
visibilityNoNew access level: 'public' or 'private'; omit to leave unchanged.
ttl_secondsNoNew expiry in seconds from now; mutually exclusive with permanent.
rotate_view_secretNoIssue a fresh private view link and revoke every link shared earlier for this image. Use to un-share a private image.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Discloses all behavioral traits: visibility flips access level, ttl vs permanent effects, rotate_view_secret invalidates old links, and return of updated record. No contradiction with annotations (all false).

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Well-organized with bullet points and front-loaded purpose. Each sentence serves a purpose, though could be slightly more concise for very high score.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the presence of an output schema (not shown but indicated), the description covers usage, parameters, errors, and ownership fully, leaving no gaps for an agent to make mistakes.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, but description adds critical context like mutual exclusivity and the consequence of rotate_view_secret, going beyond schema descriptions. Slight deduction because the schema already covers basics well.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description explicitly states 'Update visibility, TTL, or the private view link of a hosted image', identifying specific verb and resource. It clearly distinguishes from siblings like delete_image, get_image, and upload_image.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides clear when-to-use conditions: 'Caller must own it.' and explains mutual exclusivity of ttl_seconds and permanent. Describes error scenarios (404 for absent/expired/owned by another).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update_skill_filesAInspect

Change specific files in a skill that already exists, including its runbook (SKILL.md). Paths you do not name are kept unchanged, so you never need to read back or resend the rest of the tree. Use this for any edit to a skill whose files you did not just author in full. To create a skill, change its name, description, or tags, or replace the entire tree at once, use save_skill.

Send each file's complete new content, not a diff. A file's executable flag and purpose label keep their current values unless you set them. The skill's description, outcome, tags, and verifier references carry forward untouched.

This writes the next version of the skill. Requires edit access. Returns skill_id, version, the new version_token, the sorted changed and deleted paths, and carried_forward, the count of files left untouched.

ParametersJSON Schema
NameRequiredDescriptionDefault
filesNoThe files to write, one entry per path, each carrying that file's complete new content. A path may appear only once. Use the path 'SKILL.md' to rewrite the runbook itself.
sourceNoAttributes the new version to how it was produced ('manual', 'teach', 'optimization', 'description_optimization', or 'audit'); omit to leave it unset.
skill_idYesThe skill to act on, identified by its UUID, slug, or name.
delete_pathsNoPaths to remove from the skill's file tree. Each must already be in the tree, and 'SKILL.md' cannot be removed.
expected_version_tokenYesConcurrency token for the version you last read, returned by get_skill and by every write. A stale token is rejected so you never overwrite a change someone else made.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Discloses that it writes new version, requires edit access, returns version details, and that executable/purpose default to carry-forward. Doesn't contradict annotations (readOnlyHint=false, destructiveHint=false).

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Concise, front-loaded, every sentence adds value, no fluff.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given output schema exists and annotations, description covers auth, return values, and key behavioral traits; no gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% with detailed descriptions; the description adds minor behavioral context (complete content, no diff) but does not significantly enhance parameter meaning beyond schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Specific verb 'Change specific files' + resource 'skill that already exists', distinguishes from 'save_skill' for creation/rename/replace.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly says when to use (existing skill, not just authored) vs when to use save_skill; also describes that untouched paths are preserved.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

upgrade_to_proAInspect

Start a Pro subscription and return a secure Stripe checkout link.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations indicate mutation (readOnlyHint=false) and non-destructive (destructiveHint=false). The description adds the Stripe checkout link detail but doesn't disclose that the subscription is in pending state until payment, or that this does not immediately grant Pro access.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Single sentence of 12 words, front-loaded with action and result. Every word adds value. No wasted text.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with zero parameters and an output schema (existing), the description covers the core action and return value. It is adequate for a simple checkout initiation, though could mention payment flow expectations.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

No parameters in schema (100% coverage trivially). With 0 parameters, baseline is 4. Description adds no param info, which is acceptable as none exist.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('Start a Pro subscription') and the outcome ('return a secure Stripe checkout link'). It uses a specific verb and resource, and distinguishes itself from sibling tools like cancel_subscription or create_billing_portal_session.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool versus alternatives. For example, it does not mention when to use create_billing_portal_session for existing subscriptions or prerequisites like requiring a team.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

upload_imageAInspect

Upload an image and return the hosted image record.

data must be the image bytes encoded as standard base64 (RFC 4648). Accepted image formats are PNG, JPEG, WebP, and GIF.

visibility controls who can access the served URL: "public" makes it accessible to anyone with the link; "private" (default) requires the owner's credentials. Accepted values: "public", "private".

ttl_seconds sets an expiry relative to now (positive integer). Omit to create a permanent image.

Returns: {id, token, url, visibility, expires_at, size_bytes, content_type}.

ParametersJSON Schema
NameRequiredDescriptionDefault
dataYesRaw image bytes encoded as standard base64 (RFC 4648). Accepted formats: PNG, JPEG, WebP, GIF.
visibilityNoWho can access the served URL: 'public' for anyone with the link, or 'private' (default) for owner-only.private
ttl_secondsNoExpiry in seconds from now; omit for a permanent image.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Discloses behavioral details beyond annotations: requires base64 encoding, accepted formats, visibility control, TTL expiry, and return fields. Annotations only indicate non-readonly, non-destructive, non-idempotent; description fills in specifics.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Concise and well-structured with bullet points for key details. Every sentence earns its place without redundancy. Efficiently communicates all necessary information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given output schema presence, description explains return fields (id, token, url, etc.), covers input requirements, and provides complete behavioral context. No gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, and the description adds meaningful explanations for each parameter: data format, visibility acceptable values, TTL semantics. Adds value beyond the schema alone.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Clearly states 'Upload an image and return the hosted image record.' Specifies verb (upload) and resource (image), differentiating it from siblings like get_image, delete_image, and update_image.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides detailed usage guidance on data format, accepted image types, visibility options, and TTL. Does not explicitly mention when not to use or alternatives, but the context is clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 2 tool updates
    • Changedsave_skill3 fields changed
      • changedInput schema / properties / payload / properties / files / description
        Previous value: -"Sibling files bundled with the skill. Include a 'demo/README.md' writeup with images referenced by relative path inside 'demo/' to render a visual demo on the public template page once the skill is published as a template."New value: +"Sibling files bundled with the skill. Include a 'demo/README.md' writeup with images referenced by relative path inside 'demo/' to render a visual demo on the public template page once the skill is published as a template. A supplied list is a full snapshot: any path you do not include is deleted. Omitting the field entirely preserves the existing tree unchanged; to change specific files without resending the rest, use `update_skill_files` instead."
      • changedInput schema / properties / payload / properties / outcome / description
        Previous value: -"The business outcome this skill is meant to drive, surfaced as a discovery facet."New value: +"The business outcome this skill is meant to drive, surfaced as a discovery facet. When omitted on an update, the existing outcome is preserved; send null to clear it."
      • changedInput schema / properties / payload / properties / tags / description
        Previous value: -"Discovery tags surfaced when listing and searching skills."New value: +"Discovery tags surfaced when listing and searching skills. When omitted on an update, the existing tags are preserved; send an empty list to clear them."
    • Addedupdate_skill_files
  2. 46 tool updates
    • Addedarchive_skill
    • Removedarchive_workflow
    • Addedaudit_skill
    • Removedaudit_workflow
    • Addedcheck_skill_safety
    • Removedcheck_workflow_safety
    • Addeddelete_skill
    • Addeddelete_skill_version
    • Removeddelete_workflow
    • Removeddelete_workflow_version
    • Addeddesign_skill
    • Removeddesign_workflow
    • Changedfork_template1 field changed
      • changedInput schema / properties / name / description
        Previous value: -"Optional name for the new private workflow; defaults to the template's slug."New value: +"Optional name for the new private skill; defaults to the template's slug."
    • Changedgenerate_image6 fields changed
      • addedInput schema / properties / payload / properties / skill_id
        Added value: +{
        +  "anyOf": [
        +    {
        +      "type": "string"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null,
        +  "description": "Optional skill UUID for provenance. Access-checked: a skill the caller cannot see surfaces as NotFound rather than being stamped onto the run."
        +}
      • addedInput schema / properties / payload / properties / skill_ref
        Added value: +{
        +  "anyOf": [
        +    {
        +      "maxLength": 256,
        +      "type": "string"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null,
        +  "description": "Free-form skill label (slug, name) for human-readable provenance."
        +}
      • addedInput schema / properties / payload / properties / skill_version
        Added value: +{
        +  "anyOf": [
        +    {
        +      "type": "integer"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null,
        +  "description": "Skill version invoking this run, paired with ``skill_id``."
        +}
      • removedInput schema / properties / payload / properties / workflow_id
        Removed value: -{
        -  "anyOf": [
        -    {
        -      "type": "string"
        -    },
        -    {
        -      "type": "null"
        -    }
        -  ],
        -  "default": null,
        -  "description": "Optional workflow UUID for provenance. Access-checked: a workflow the caller cannot see surfaces as NotFound rather than being stamped onto the run."
        -}
      • removedInput schema / properties / payload / properties / workflow_ref
        Removed value: -{
        -  "anyOf": [
        -    {
        -      "maxLength": 256,
        -      "type": "string"
        -    },
        -    {
        -      "type": "null"
        -    }
        -  ],
        -  "default": null,
        -  "description": "Free-form workflow label (slug, name) for human-readable provenance."
        -}
      • removedInput schema / properties / payload / properties / workflow_version
        Removed value: -{
        -  "anyOf": [
        -    {
        -      "type": "integer"
        -    },
        -    {
        -      "type": "null"
        -    }
        -  ],
        -  "default": null,
        -  "description": "Workflow version invoking this run, paired with ``workflow_id``."
        -}
    • Addedget_skill
    • Addedget_skill_file
    • Addedget_skill_files
    • Removedget_workflow
    • Removedget_workflow_file
    • Removedget_workflow_files
    • Addedgrant_skill
    • Removedgrant_workflow
    • Addedleave_shared_skill
    • Removedleave_shared_workflow
    • Addedlist_skill_grants
    • Addedlist_skills
    • Removedlist_workflow_grants
    • Removedlist_workflows
    • Changedlookup_fork_lineage3 fields changed
      • addedInput schema / properties / skill_id
        Added value: +{
        +  "description": "The skill to act on, identified by its UUID, slug, or name.",
        +  "type": "string"
        +}
      • removedInput schema / properties / workflow_id
        Removed value: -{
        -  "description": "The workflow to act on, identified by its UUID, slug, or name.",
        -  "type": "string"
        -}
      • changedInput schema / required
        Previous value: -[
        -  "workflow_id"
        -]New value: +[
        +  "skill_id"
        +]
    • Changedoptimize_description3 fields changed
      • addedInput schema / properties / skill_id
        Added value: +{
        +  "description": "The skill to act on, identified by its UUID, slug, or name.",
        +  "type": "string"
        +}
      • removedInput schema / properties / workflow_id
        Removed value: -{
        -  "description": "The workflow to act on, identified by its UUID, slug, or name.",
        -  "type": "string"
        -}
      • changedInput schema / required
        Previous value: -[
        -  "workflow_id"
        -]New value: +[
        +  "skill_id"
        +]
    • Addedoptimize_skill
    • Removedoptimize_workflow
    • Changedpublish_template_version3 fields changed
      • addedInput schema / properties / skill_id
        Added value: +{
        +  "description": "The skill to act on, identified by its UUID, slug, or name.",
        +  "type": "string"
        +}
      • removedInput schema / properties / workflow_id
        Removed value: -{
        -  "description": "The workflow to act on, identified by its UUID, slug, or name.",
        -  "type": "string"
        -}
      • changedInput schema / required
        Previous value: -[
        -  "workflow_id"
        -]New value: +[
        +  "skill_id"
        +]
    • Addedrevoke_skill_grant
    • Removedrevoke_workflow_grant
    • Changedrun_verifier6 fields changed
      • addedInput schema / properties / skill_id
        Added value: +{
        +  "anyOf": [
        +    {
        +      "type": "string"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null,
        +  "description": "Optional id of the skill this run is associated with, for provenance; a skill you cannot see is rejected."
        +}
      • addedInput schema / properties / skill_ref
        Added value: +{
        +  "anyOf": [
        +    {
        +      "type": "string"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null,
        +  "description": "Optional free-form skill label (slug or name) to stamp on the run, for provenance."
        +}
      • addedInput schema / properties / skill_version
        Added value: +{
        +  "anyOf": [
        +    {
        +      "type": "integer"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null,
        +  "description": "Optional skill version number to stamp on the run, for provenance."
        +}
      • removedInput schema / properties / workflow_id
        Removed value: -{
        -  "anyOf": [
        -    {
        -      "type": "string"
        -    },
        -    {
        -      "type": "null"
        -    }
        -  ],
        -  "default": null,
        -  "description": "Optional id of the workflow this run is associated with, for provenance; a workflow you cannot see is rejected."
        -}
      • removedInput schema / properties / workflow_ref
        Removed value: -{
        -  "anyOf": [
        -    {
        -      "type": "string"
        -    },
        -    {
        -      "type": "null"
        -    }
        -  ],
        -  "default": null,
        -  "description": "Optional free-form workflow label (slug or name) to stamp on the run, for provenance."
        -}
      • removedInput schema / properties / workflow_version
        Removed value: -{
        -  "anyOf": [
        -    {
        -      "type": "integer"
        -    },
        -    {
        -      "type": "null"
        -    }
        -  ],
        -  "default": null,
        -  "description": "Optional workflow version number to stamp on the run, for provenance."
        -}
    • Addedsave_skill
    • Removedsave_workflow
    • Addedsearch_skills
    • Removedsearch_workflows
    • Addedteach_skill
    • Removedteach_workflow
    • Addedtransfer_skill_ownership
    • Removedtransfer_workflow_ownership
    • Addedunarchive_skill
    • Removedunarchive_workflow
  3. 4 tool updates
    • Addedconfigure_auto_top_up
    • Addeddisable_auto_top_up
    • Addedget_auto_top_up
    • Addedpurchase_credits
  4. 3 tool updates
    • Addedcancel_subscription
    • Addedcreate_billing_portal_session
    • Addedupgrade_to_pro
  5. 6 tool updates
    • Changedgenerate_image1 field changed
      • addedInput schema / properties / payload / properties / visibility
        Added value: +{
        +  "default": "public",
        +  "description": "Access level for the hosted copy of each generated image. ``public`` (default) returns a link that opens in any browser. ``private`` returns a link only you can open and forward to people you choose, while the plain URL stays locked. Anonymous generations are always public.",
        +  "type": "string"
        +}
    • Changedlist_invitations2 fields changed
      • changedInput schema / properties / filter / description
        Previous value: -"'received' | 'sent' | 'all'"New value: +"Which invitations to return: 'received' (addressed to you), 'sent' (you created), or 'all'."
      • changedInput schema / properties / state / description
        Previous value: -"'pending' | 'all'"New value: +"Which states to include: 'pending' only, or 'all' (includes expired and resolved invitations)."
    • Changedlist_teams1 field changed
      • changedInput schema / properties / filter / description
        Previous value: -"'mine' | 'member' | 'all'"New value: +"Which teams to return: 'mine' (you own), 'member' (you belong to but do not own), or 'all'."
    • Changedlist_templates1 field changed
      • changedInput schema / properties / filter / description
        Previous value: -"'mine' | 'all'"New value: +"Which templates to return: 'mine' (you own) or 'all' public templates."
    • Changedlist_workflows1 field changed
      • changedInput schema / properties / filter / description
        Previous value: -"'mine' | 'shared-with-me' | 'all'"New value: +"Which workflows to return: 'mine' (you own), 'shared-with-me' (granted to you), or 'all'."
    • Changedupdate_image1 field changed
      • addedInput schema / properties / rotate_view_secret
        Added value: +{
        +  "default": false,
        +  "description": "Issue a fresh private view link and revoke every link shared earlier for this image. Use to un-share a private image.",
        +  "type": "boolean"
        +}
  6. 72 tool updates
    • First observedaccept_invitation
    • First observedadd_team_member
    • First observedarchive_template
    • First observedarchive_workflow
    • First observedaudit_workflow
    • First observedcancel_invitation
    • First observedcheck_template_safety
    • First observedcheck_workflow_safety
    • First observedclaim_handle
    • First observedcreate_team
    • First observeddecline_invitation
    • First observeddelete_image
    • First observeddelete_image_generator
    • First observeddelete_team
    • First observeddelete_template
    • First observeddelete_template_version
    • First observeddelete_verifier
    • First observeddelete_workflow
    • First observeddelete_workflow_version
    • First observeddeploy_image_generator
    • First observeddeploy_verifier
    • First observeddeprecate_template_version
    • First observeddesign_workflow
    • First observedfork_template
    • First observedgenerate_image
    • First observedget_image
    • First observedget_image_generator
    • First observedget_referral_status
    • First observedget_template
    • First observedget_template_file
    • First observedget_usage
    • First observedget_verifier
    • First observedget_workflow
    • First observedget_workflow_file
    • First observedget_workflow_files
    • First observedgrant_workflow
    • First observedleave_shared_workflow
    • First observedlist_api_keys
    • First observedlist_image_generators
    • First observedlist_images
    • First observedlist_invitations
    • First observedlist_team_members
    • First observedlist_teams
    • First observedlist_templates
    • First observedlist_verifiers
    • First observedlist_workflow_grants
    • First observedlist_workflows
    • First observedlookup_fork_lineage
    • First observedmint_api_key
    • First observedoptimize_description
    • First observedoptimize_workflow
    • First observedpublish_template_version
    • First observedredeem_referral_code
    • First observedremove_team_member
    • First observedrename_handle
    • First observedrevoke_api_key
    • First observedrevoke_image_generator
    • First observedrevoke_verifier
    • First observedrevoke_workflow_grant
    • First observedrun_verifier
    • First observedsave_workflow
    • First observedsearch_templates
    • First observedsearch_workflows
    • First observedteach_workflow
    • First observedtransfer_team_ownership
    • First observedtransfer_template_ownership
    • First observedtransfer_workflow_ownership
    • First observedunarchive_template
    • First observedunarchive_workflow
    • First observedunpublish_template_version
    • First observedupdate_image
    • First observedupload_image

Related MCP Connectors

Related MCP Servers

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.