Unsora
Server Details
Generate AI images, video, avatars, music and voiceovers, clip videos, and schedule social posts.
- Status
- Healthy
- OAuth
- Works in Glama
- Last Tested
- Transport
- Streamable HTTP · MCP 2025-11-25
- URL
TDQS
Scored across 51 tools
Tools are mostly distinct despite the large surface: create_image vs create_image_and_schedule, compose_post vs create_post, and the wait_for_* vs *_status pairs have clear descriptions that explain when to use each. Some overlap remains among video-generation tools and voice-listing tools, but detailed guidance reduces misselection.
The 51 tools overwhelmingly follow a snake_case verb_noun pattern (create_video, list_posts, get_price, wait_for_image), but the three admin skill tools use kebab-case (get-skill, update-skill, upload-skill). This is a minor separator inconsistency, otherwise names are predictable and readable.
51 tools is far above the typical 3–15 range and likely overwhelming for an agent, even though the server covers a broad AI media generation and social scheduling domain. Many separate status/wait tools for each media type contribute to the high count.
Core lifecycle coverage is strong: posts have create/get/list/update/delete/publish/retry/analytics; media generation, uploads, voices, credits and subscription are covered. Minor gaps include no list_skills/delete-skill for the admin catalog and no way to remove already-published content from platforms, but agents can work around these.
Available Tools
51 toolsaudio_statusAudio StatusCRead-onlyIdempotentInspect
Single-shot generation status check (used by the preview UI).
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | ||
| tool | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint/idempotentHint/openWorldHint=false, so the safety profile is covered structurally. The description contributes one genuinely useful behavioral fact — "single-shot" implies it does not poll/wait, unlike the wait_for_* siblings — but says nothing about the returned state values or whether a job may be pending.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single short sentence, front-loaded and free of filler, but its brevity comes at the cost of substance rather than through efficient coverage of the essentials.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 2-param, no-output-schema tool with 0% schema coverage and no annotation detail on return values, the description needs to carry more weight than one parenthetical. It leaves the purpose of `tool`, the expected state output, and the distinction from sibling status/wait tools unaddressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for both required parameters. The description never mentions `tool` or `id`, so it does not compensate. Worse, the mismatch between the "audio" framing and the broad `tool` enum leaves the meaning of the `tool` parameter unclarified.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
It names a verb-and-resource pair ("generation status check"), which is clearer than a bare "status", but it is prefixed with "audio" while the required `tool` enum spans image, video, music, avatar, etc. The result is a name/description that under-describes the actual scope, so an agent cannot be confident what it queries.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
"(used by the preview UI)" hints this is an internal-facing helper, but there is no guidance on when an agent should call it versus the many siblings (video_status, image_status, clipping_status, wait_for_*) or what distinguishes those tools from this one. No when-not conditions are given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
change_voiceChange VoiceAInspect
Re-voice existing speech into one of the user's cloned voices (Voice Changer, speech-to-speech): keeps the timing and delivery of the source, swaps the voice. The source can be audio or video, up to 5 minutes. Costs about 10 credits per minute. voiceId must be a cloned voice id from list_voices (create one with create_voice_clone). Poll with wait_for_voiceover. Renders a live audio player in app-capable hosts.
| Name | Required | Description | Default |
|---|---|---|---|
| voiceId | Yes | Cloned voice id (list_voices → clones). | |
| sourceUrl | Yes | Public URL of the speech audio or video to convert. | |
| outputFormat | No | Default mp3. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden and does so well: it discloses cost (~10 credits/minute), a hard 5-minute source limit, accepted input types, the async/polling nature of the operation, and the output behavior ('renders a live audio player'). These are exactly the traits an agent needs before invoking and typically missing from structured fields.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single dense paragraph, front-loaded with the core action and transformation, then constraints, cost, id sourcing, polling, and output. Every sentence adds a distinct, non-redundant fact with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 3-param async mutation tool with no annotations and no output schema, the description covers purpose, input constraints, cost, id source, polling path, and output rendering. An agent has everything needed to call it correctly and know what to expect.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all three parameters including the cloned-voice constraint on voiceId and video acceptance for sourceUrl. The description largely reinforces the schema (cloned voice id, audio-or-video source) rather than adding new syntax or format detail, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Re-voice existing speech into one of the user's cloned voices', with 'Voice Changer, speech-to-speech' as a gloss) and clarifies the transformation ('keeps the timing and delivery of the source, swaps the voice'). This clearly separates it from sibling create_voiceover (text-to-speech) and create_voice_clone (voice creation), which an agent can distinguish without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear usage context: source must be audio or video up to 5 minutes, voiceId must come from list_voices, and results should be polled with wait_for_voiceover. It routes the agent to the right supporting tools, though it never states an explicit 'do not use this for X' exclusion (e.g., versus create_voiceover for text input).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
clipping_statusClip StatusBRead-onlyIdempotentInspect
Single-shot AI clipping job status + clips (used by the preview UI).
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | ||
| tool | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent and closed-world behavior, so the safety profile is covered. The description still adds useful traits beyond that: it is single-shot (no waiting/polling) and returns clips in addition to status, which matters since there is no output schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler. It is efficient, though borderline terse given how much remains unexplained about the return content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description should carry return-shape burden; it only notes '+ clips'. It omits the possible status values, error/failure states, and how to interpret an in-progress job, leaving the agent to infer polling semantics.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Both parameters (id, tool) have 0% schema description coverage and the description adds no meaning for either. The 'tool' enum is self-evident from the schema, but 'id' (job id) is never clarified, so the description fails to compensate for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific resource and action: AI clipping job status plus clips for a given job id. 'Single-shot' implicitly distinguishes it from the polling-style sibling wait_for_clipping, though no sibling is named explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The parenthetical '(used by the preview UI)' describes an internal consumer rather than telling an agent when to call this instead of wait_for_clipping or create_clipping. No when-to-use or when-not-to-use guidance is provided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
compose_postCompose PostAInspect
Open the interactive post composer: the user picks accounts, edits the caption and media, sets per-platform options (YouTube title/visibility/category/tags, TikTok privacy/comments/duet/stitch/branded content/AI label, Pinterest board/title/link, Google Business call-to-action button, video cover) and posts now, schedules or saves a draft — the composer creates the post itself. Use it whenever the user wants to create or schedule a post and the host can show apps; prefill everything the user already gave (caption, media URLs, accounts, time). After it opens, do NOT call create_post yourself; tell the user to finish in the composer. In hosts without app support, gather the details in chat and call create_post instead.
| Name | Required | Description | Default |
|---|---|---|---|
| title | No | YouTube video title / Pinterest pin title. | |
| caption | No | ||
| coverUrl | No | Video cover image URL. | |
| mediaType | No | video = one video, images = one or more images (slideshow), text = no media. | |
| mediaUrls | No | Media to post: one video URL, or image URLs in order. | |
| accountIds | No | Accounts to preselect (ids from get_accounts). | |
| scheduled_at | No | ISO datetime to preselect for scheduling. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does well: it discloses that this is an interactive UI invocation, that the composer performs the actual creation, and that the host must support apps or the workflow falls back to create_post. It does not cover auth requirements, what the call returns (e.g., whether the created post is returned), or how long the composer session lives.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The first sentence is a long enumeration but it is front-loaded with the core action and the 'composer creates the post itself' clarification. The subsequent usage sentences all earn their place; only the long per-platform list is arguably more detail than an agent needs.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-required-param UI-opening tool with no output schema and no annotations, the description covers the important behaviors: host requirement, prefill expectation, and the do-not-double-create rule. It leaves the return value of the call unspecified, which is the main remaining gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is already 86% and the schema documents title, mediaType, mediaUrls, accountIds, and scheduled_at clearly. The description restates some of this and adds per-platform option names (YouTube visibility, TikTok privacy/duet/stitch, Pinterest board, etc.) that are not actual parameters, which risks confusing the agent about what can be passed. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource (open the interactive post composer) and immediately clarifies the crucial distinction that the composer, not this tool directly, creates the post. It is unmistakably separable from the sibling create_post, list_posts, and update_post.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives explicit when-to-use ('whenever the user wants to create or schedule a post and the host can show apps'), a prefill instruction, an explicit when-not ('do NOT call create_post yourself after it opens'), and names the alternative path (call create_post in hosts without app support). This is a complete routing rule.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_avatar_videoCreate Talking AvatarAInspect
Turn a portrait into a talking-head video (AI Avatar Maker, SkyReels V3): the person in imageUrl speaks a transcript (voiced with voiceId) or lip-syncs a supplied audioUrl. Costs 10 credits. Use list_voices to pick a voice — MiniMax presets, Eleven v3 voices, or the user's cloned voices. Poll with wait_for_video. Renders a live preview in app-capable hosts.
| Name | Required | Description | Default |
|---|---|---|---|
| prompt | No | Optional guidance for expression, gestures or setting. | |
| emotion | No | Delivery emotion. Default neutral. | |
| voiceId | No | Voice for the transcript — an id from list_voices. Default Friendly_Person. | |
| audioUrl | No | Speech audio to lip-sync instead of a transcript. | |
| imageUrl | Yes | Portrait of the person or character, face clearly visible. | |
| resolution | No | Default 720p. | |
| transcript | No | What the avatar says (max 2,000 characters). Required unless audioUrl is set. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the disclosure burden and does well: it names the 10-credit cost, discloses the asynchronous job nature requiring wait_for_video polling, and mentions live preview behavior in app-capable hosts. It stops short of saying what happens on failure or whether credits are refunded.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four tight sentences, front-loaded with the core action and mode selection before cost, routing and polling hints. No filler or restatement of the title.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Covers the essentials for a 7-parameter async generation tool with no output schema: cost, polling, voice sourcing, and dual input modes. Missing only edge-case behavior such as failure/refund handling and validation limits already visible in the schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3; the description adds value beyond it by explaining where voiceId comes from (MiniMax presets, Eleven v3, cloned voices via list_voices) and by framing voiceId/transcript versus audioUrl as alternative input paths.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Turn a portrait into a talking-head video') and names the underlying engines (AI Avatar Maker, SkyReels V3). The two operating modes (transcript+voiceId vs audioUrl lip-sync) make it clearly distinct from siblings like create_video, create_influencer or create_motion_control.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly routes the agent to list_voices for voice selection and to wait_for_video for polling, and explains the condition selecting each of the two audio modes. It does not state when-not-to-use it (e.g. versus create_video for non-avatar footage), so no exclusions are covered.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_clippingCreate ClipsAInspect
Turn a long video into short clips (AI clipping). Use this for any request to clip, cut, chop, trim or repurpose a long video, podcast, stream, webinar or interview into shorts, Reels, TikToks or YouTube Shorts, or to find and extract specific moments or highlights (set query to describe the moments in natural language). Provide a public videoUrl. DO NOT call this tool until the user has explicitly chosen: aspect ratio, captions on/off (and style if on), number of clips, and clip length (not needed when query is set). This tool spends the user's credits — if any of these choices is missing from the conversation, you MUST stop and ask the user for all missing ones first. Never assume defaults, never infer unstated preferences. In app-capable hosts this renders a live preview that polls progress and displays the clips by itself — when the preview is rendered, do NOT call wait_for_clipping or clipping_status afterwards; summarize the job settings and end your turn. Only poll with wait_for_clipping in hosts without the preview.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| query | No | Find Moments mode: natural-language description of the moments to extract (e.g. 'funny reactions', 'product demos', 'goal moments and key plays'). When set, targetDuration is ignored. Omit to auto-detect the most viral clips. | |
| ratio | No | Output aspect ratio. AI reframe keeps the main subject centered. 9:16 for TikTok/Reels/Shorts, 1:1 or 4:5 for feed posts, 16:9 for YouTube. Omit or use 'original' to keep the source ratio. | |
| videoUrl | Yes | YouTube link or direct video file URL (e.g. .mp4). Page links from other platforms (TikTok, Instagram, Vimeo) are not supported. | |
| sourceLang | No | ||
| targetLang | No | ||
| captionStyle | No | Caption look, named by appearance: colors are the accent on the spoken word (e.g. classic-yellow = white text, yellow active word). glow-* add a soft glow, static-* don't animate, gaming-* are bold streamer styles, white-card/black-box put text on a card. Default: classic-yellow. | |
| enableCaption | No | Burn animated captions into the clips. Auto-enabled when captionStyle is set. | |
| targetDuration | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full behavioral burden and does well: it discloses that the call spends the user's credits, requires a public videoUrl, renders a live preview in capable hosts, and that polling must be skipped there. It does not cover failure behavior, credit cost magnitude, or processing time, so it stops just short of a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with purpose, then prerequisites, then host-conditional polling guidance — a sensible order. It is long for a description and repeats itself ('never assume defaults, never infer unstated preferences' restates the preceding sentence), but almost every clause carries routing or constraint value.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a credit-spending, 9-parameter tool with no annotations and no output schema, the description covers the critical unknowns: consent gating, input URL constraints, preview rendering, and the polling alternative. The gaps are secondary parameter definitions (sourceLang/targetLang/limit) and any statement of what the call returns.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 56%, so the description must compensate, and it partly does: it frames the user-facing choices (aspect ratio, captions/style, number of clips, clip length) and explains that clip length is irrelevant when query is set. However sourceLang, targetLang, and limit remain unexplained in both places, and the ratio/captionStyle enum semantics come from the schema rather than the description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Opens with a concrete verb+resource ('Turn a long video into short clips') and enumerates the request types it covers (clip, cut, chop, trim, repurpose, find moments). It also names the siblings it coordinates with (wait_for_clipping, clipping_status), so an agent can separate it from the status/polling tools without opening schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit when-to-use (any clipping/repurposing/highlight request), explicit gating (do not call until the user chose aspect ratio, captions, clip count, clip length), and explicit host-conditional behavior for polling. This is close to a complete decision procedure, including what to do instead in preview-capable vs preview-less hosts.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_imageCreate ImageAInspect
Generate an image for a social post, thumbnail or ad creative. Models include GPT Image 2.5 Flare / Sunburst, GPT Image 2, Nano Banana Pro / 2 / 2 Lite and Seedream 5.0 Pro — see list_models for each model's inputs. Attaching images in inputs edits / uses them as references. Priced live (get_price). Returns generation.id — poll with wait_for_image. Renders a live preview in app-capable hosts.
| Name | Required | Description | Default |
|---|---|---|---|
| mode | No | Mode key from list_models. Omit to pick it from the attached media. | |
| model | No | Model key from list_models (category "image"). Default: nano-banana-2. | |
| inputs | No | WaveSpeed input fields for the chosen model and mode, exactly as list_models shows them (e.g. aspect_ratio, resolution, duration, generate_audio, image, last_image, reference_images). Media fields take URLs; files hosted elsewhere are imported into the user's library automatically. Omitted fields use the model's defaults. | |
| prompt | Yes | What to generate | |
| maxCredits | No | Refuse to start if the live price is above this. Pass the credits get_price returned. | |
| idempotencyKey | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does well: it discloses the async lifecycle (returns generation.id, poll with wait_for_image), live pricing and credit refusal via maxCredits, automatic import of externally hosted media, and live preview in app-capable hosts. It omits auth requirements, rate limits, and idempotency behavior, keeping it below a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose is front-loaded in the first sentence, followed by tight, information-dense clauses covering models, inputs, pricing, and polling with no filler. The model enumeration is slightly long, but every sentence carries actionable content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 6-parameter, nested-input tool with no output schema and no annotations, the description covers the full call lifecycle: model selection, input construction, cost guardrails, return value, and follow-up polling. Only permission/prerequisite context is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 83%, so the baseline is 3, and the description adds genuine meaning beyond it: attaching `images` inside `inputs` means edit/reference, `mode`/`model` keys come from list_models, and omitted fields fall back to model defaults. This enriches the nested `inputs` object that the schema only partly documents.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Generate an image') and names concrete use cases (social post, thumbnail, ad creative). It is clear what the tool does, but it does not explicitly differentiate itself from core siblings like create_thumbnail or create_image_and_schedule, which overlap with the stated use cases.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit routing for the workflow: see list_models for per-model inputs, get_price for live pricing, and wait_for_image to poll. It also explains that attaching `images` in inputs switches to an edit/reference flow. It stops short of stating when NOT to use it versus the scheduling/thumbnail variants.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_image_and_scheduleCreate Image & Schedule PostBInspect
Generate an image for a social post and schedule it in one step: create image → wait → schedule slideshow post to accounts. If any target account is Instagram, prefer aspectRatio 4:5 (or 1:1); other ratios are padded to fit. In app-capable hosts the result renders as an interactive panel the user can see and act on — do NOT repeat its contents as a list or table in your reply; add only what the panel doesn't say.
| Name | Required | Description | Default |
|---|---|---|---|
| model | No | Image model key from list_models. Default: nano-banana-2. | |
| prompt | Yes | ||
| caption | Yes | ||
| accountIds | Yes | ||
| resolution | No | ||
| aspectRatio | No | ||
| scheduled_at | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does disclose real behavioral traits: the internal generate→wait→schedule sequence, automatic padding when ratios don't match, and the host-dependent interactive-panel rendering with an explicit instruction not to duplicate its contents. It still omits auth/permission needs, what happens on partial failure, and scheduling constraints.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the tool's purpose, then the pipeline, then the conditional guidance, then the host-behavior note — a sensible ordering with little filler. The final sentence is dense but each clause carries distinct instructions.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex combined tool with 7 params, no output schema, and no annotations, the description covers the operation flow and rendering behavior well but leaves key invocations underspecified. An agent still lacks semantics for most inputs and any failure-mode or dependency information.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 14% (just 'model'), so the description must compensate for six largely undocumented parameters. It only adds meaning for aspectRatio (4:5/1:1 preference and padding behavior); prompt, caption, accountIds, resolution, and scheduled_at are left with no format, constraint, or intent explanation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb pair and resource: 'Generate an image for a social post and schedule it in one step,' and even spells out the internal pipeline (create image → wait → schedule slideshow post to accounts). This clearly distinguishes it from atomic siblings like create_image and create_post without naming them, though it never explicitly contrasts itself against those alternatives.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by 'in one step,' suggesting it is the combined-operation choice when a user wants both generation and scheduling, but there is no explicit when-to-use vs. when-to-decompose guidance. The only concrete conditional rule ('If any target account is Instagram, prefer aspectRatio 4:5') is parameter-shaped rather than routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_influencerCreate AI InfluencerBInspect
Generate an AI influencer portrait for a social account or UGC-style post. Renders a live preview grid in app-capable hosts; poll each id with wait_for_image otherwise.
| Name | Required | Description | Default |
|---|---|---|---|
| age | No | Subject age in years (18–70). | |
| count | No | ||
| prompt | Yes | ||
| styleMode | No | Look / lighting preset. | |
| aspectRatio | No | ||
| cameraAngle | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It usefully discloses the async/preview-grid behavior and the need to poll ids, but omits credit cost, whether the prompt is free-form or safety-filtered, and what happens on multi-count requests.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, zero filler, with the core action front-loaded and the async handling second. Appropriately sized.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 6-parameter generation tool with no annotations and no output schema, the description covers only the async retrieval pattern. Enum parameters like styleMode, aspectRatio, and cameraAngle and the meaning of count are left entirely to a sparse schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 33% (only age and styleMode have descriptions). The description mentions no parameters at all — nothing about count, aspectRatio, cameraAngle, or the prompt's expected content — so it does not compensate for the coverage gap on a 6-parameter tool.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Generate an AI influencer portrait') and scopes it to social accounts or UGC-style posts. It is distinguishable from create_image by the influencer/UGC framing, though it never explicitly contrasts with that close sibling.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives conditional guidance on retrieval ('Renders a live preview grid in app-capable hosts; poll each id with wait_for_image otherwise'), which is genuinely useful. However it says nothing about when to pick this over create_image, create_avatar_video, or other generation siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_motion_controlMotion ControlAInspect
Make a character copy the motion of a reference video: the character in inputs.image performs the movement, dance or acting from inputs.video. Models: Kling 3.0 Pro / Standard, Kling 2.6 Pro, Wan 2.2 Animate, DreamActor v2 (see list_models). Priced live by the motion video's length — call get_price first. Poll with wait_for_video. Renders a live preview in app-capable hosts.
| Name | Required | Description | Default |
|---|---|---|---|
| mode | No | Mode key from list_models. Omit to pick it from the attached media. | |
| model | No | Model key from list_models (category "motion-control"). Default: kling-mc-3.0-pro. | |
| inputs | No | WaveSpeed input fields for the chosen model and mode, exactly as list_models shows them (e.g. aspect_ratio, resolution, duration, generate_audio, image, last_image, reference_images). Media fields take URLs; files hosted elsewhere are imported into the user's library automatically. Omitted fields use the model's defaults. | |
| prompt | No | Optional guidance | |
| maxCredits | No | Refuse to start if the live price is above this. Pass the credits get_price returned. | |
| idempotencyKey | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and delivers key traits beyond the schema: async execution requiring wait_for_video polling, live credit pricing tied to motion-video length, and live preview rendering in app-capable hosts. It omits auth requirements and failure/refund behavior, keeping it short of a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One tight paragraph, front-loaded with the core action, then models, pricing, polling, and preview in descending priority. Every sentence carries information an agent needs to invoke this correctly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 6-param async creation tool with no output schema, the description covers the essentials: what it produces, how to choose models, how it is priced, and how to await results. Only minor gaps remain (no output/return shape detail, no failure handling), but polling guidance compensates for the missing output schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 83% (baseline 3), and the description adds genuine meaning by explaining the role split of the media inputs (image = performing character, video = motion reference) and pointing to list_models for mode/model keys and the WaveSpeed input fields. That extra role context lifts it above the schema baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Make a character copy the motion of a reference video') and pins the exact input slots (inputs.image = character, inputs.video = motion source). This clearly distinguishes it from siblings like create_video and create_avatar_video without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a clear operational sequence: check models via list_models, price via get_price first, then poll with wait_for_video. It stops short of an explicit when-not-to-use this over create_video/create_avatar_video, but the routing points to companion tools are strong and specific.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_movie_materialCreate Movie MaterialAInspect
Generate pre-production reference images for AI films and ads (Movie Materials Generator, GPT Image 2): character face references, full-body references, turnaround sheets, location references, video first frames, style mood boards, and 2x4 / 1x4 storyboards. Feed the results to create_video as reference images for consistent characters and locations. Costs 3 / 4 / 6 credits at 1k / 2k / 4k. Poll with wait_for_image. Renders a live preview in app-capable hosts.
| Name | Required | Description | Default |
|---|---|---|---|
| mode | Yes | face = character face reference; wide-body = full body with outfit and pose; sheet = multi-angle turnaround; location = scene/environment; first-frame = opening frame of a shot; style-collage = mood board; multishot-2x4 = 8-panel storyboard; multishot-1x4 = 4-panel strip. | |
| ratio | No | Default 3:4 for face and wide-body, 16:9 otherwise. | |
| params | No | Style controls, each 'auto' by default. All modes: cinematography (auto, cinematic, anime, realistic, cartoon, fantasy). face: age (auto, child, teen, adult, elderly), gender (auto, male, female, neutral). first-frame: camera-angle (auto, eye-level, low-angle, high-angle, dutch-angle, birds-eye, worms-eye, over-shoulder). style-collage: color-palette (auto, warm, cool, muted, vibrant, monochrome, pastel, neon), lighting-mood (auto, natural, golden-hour, blue-hour, neon-lit, studio, dramatic, low-key, high-key), era-vibe (auto, modern, retro-70s, 80s-synth, 90s-grunge, noir, victorian, futuristic, analog-film). | |
| prompt | Yes | Describe the character, location, frame or shot sequence. | |
| resolution | No | Default 2k. | |
| referenceImageUrls | No | Reference images — face, outfit, location or style inspiration, or earlier character/location materials. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does well: it discloses exact credit costs (3/4/6 at 1k/2k/4k), reveals the operation is asynchronous by instructing 'Poll with wait_for_image', and notes a live preview in app-capable hosts. It omits the return format (image URL vs. ID) and any failure/rate-limit behavior, so it stops short of full transparency.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose and output list are front-loaded in the first sentence, then short declarative sentences cover credit costs, polling, and rendering. Every sentence carries information, though the opening enumerates seven output types and is somewhat dense. No filler or restated boilerplate.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 6-parameter tool with a nested object, 3 enums, no output schema, and no annotations, the description covers cost, async polling, downstream usage, and output types. The main gap is that with no output schema it should more explicitly state what the call returns (e.g., an image reference/URL to poll for), but it is otherwise complete enough to invoke correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, including a detailed enum key for mode and a rich nested params object, so the schema already does the heavy lifting and the baseline is 3. The description's listing of output types loosely maps to the mode values but adds no syntax or constraints (e.g., resolution/ratio interplay, referenceImageUrls count) beyond what the schema states.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Generate) and resource (pre-production reference images for AI films/ads) and enumerates exactly what it produces: faces, full-body refs, turnarounds, locations, first frames, mood boards, and storyboards. It distinguishes itself from siblings like create_image and create_video by scoping to pre-production reference material. An agent can tell what this does without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly names the downstream consumer ('Feed the results to create_video as reference images') and the polling tool ('Poll with wait_for_image'), giving clear workflow context. However, it never states when to use this over create_image or create_thumbnail, nor any exclusion conditions, so the routing guidance against the closest sibling is implied rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_musicCreate MusicAInspect
Generate a music bed or song for a social video, Reel or ad (Mureka AI, song or instrumental BGM). Song models (auto, mureka-9, mureka-8, mureka-o2, mureka-7.6) require lyrics; mureka-7.5 generates instrumental BGM and treats lyrics as optional. The prompt sets genre, mood, tempo, and vocal style (e.g. 'r&b, slow, passionate, male vocal'). Returns generation.id — poll with wait_for_music. Renders a live audio player in app-capable hosts.
| Name | Required | Description | Default |
|---|---|---|---|
| model | No | ||
| lyrics | No | Song lyrics (max 3000 chars). Required for song models; optional for mureka-7.5 BGM. Section labels like [Verse] and [Chorus] are supported. | |
| prompt | Yes | Style prompt — genre, mood, tempo, vocal style. | |
| output_format | No | ||
| idempotencyKey | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does well: it discloses the async pattern (returns generation.id, poll with wait_for_music) and host-dependent rendering behavior. It omits cost/credit implications and failure/retry behavior, keeping it short of a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four dense sentences, all front-loaded with the core action and model rules before the return-value note. No filler, though the model enumeration is verbose enough to make scanning slightly heavy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 5-param async generation tool with no output schema, the description covers the action, model constraints, prompt semantics, and polling path. Only the two minor params (output_format, idempotencyKey) and cost behavior are unaddressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 40%, but the description compensates substantially: it explains what the prompt controls (genre, mood, tempo, vocal style with an example), and the lyrics/model coupling rule. output_format and idempotencyKey remain undocumented in both places.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Opens with a specific verb+resource: 'Generate a music bed or song for a social video, Reel or ad', and names the provider (Mureka AI). An agent can distinguish this from create_voiceover or create_video immediately.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explains model-dependent usage (song models need lyrics; mureka-7.5 treats lyrics as optional) and points to wait_for_music for polling, which is clear operational context. It does not explicitly state when to prefer this over sibling audio tools, so it falls short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_postSchedule PostAInspect
Create, schedule, queue, publish or cross-post content to connected social accounts. Use this for any request to post now, schedule for later, publish to multiple platforms at once, or add something to the posting queue. Supported platforms: YouTube, TikTok, Instagram, Facebook, LinkedIn, Pinterest, Threads, Bluesky, X and Google Business Profile. When the user wants to set the post up themselves or per-platform options are still open (TikTok privacy, YouTube title, Pinterest board), prefer compose_post, which opens an editable composer in app-capable hosts. Without scheduled_at the post is saved as a DRAFT — set publishNow: true to post immediately. Requires paid plan. Images are fitted to each platform automatically at publish (format, size, Instagram's 4:5–1.91:1 ratio by padding, consistent carousels); 4:5 images look best on Instagram, 9:16 suits TikTok and Reels. Posts that break a rule Unsora can't fix (video length/size, image count, caption length) are rejected with an issues list per platform. YouTube only takes video and uses title (video title). Pinterest needs media (1 image, 2–5 images or a video). X takes text, 1 image, 2–4 images or 1 video (up to 140 s, 512 MB) with a caption of at most 280 characters. Google Business Profile takes text or exactly 1 JPG/PNG image (max 5 MB), no video, summary max 1500 characters, with an optional call-to-action button (google_business). In app-capable hosts the result renders as an interactive panel the user can see and act on — do NOT repeat its contents as a list or table in your reply; add only what the panel doesn't say.
| Name | Required | Description | Default |
|---|---|---|---|
| title | No | Title applied to every account (YouTube video title). Ignored by platforms without titles. | |
| tiktok | No | TikTok settings (applies to tiktok accounts). | |
| caption | Yes | ||
| youtube | No | YouTube upload settings (applies to google accounts). | |
| coverUrl | No | Video cover image URL (Instagram Reel cover, YouTube thumbnail, Pinterest video cover). | |
| mediaUrl | No | ||
| timezone | No | IANA timezone the user scheduled in, e.g. Europe/London (display only). | |
| No | Instagram settings (applies to instagram accounts). | ||
| mediaType | No | ||
| mediaUrls | No | ||
| No | Pinterest pin settings (applies to pinterest accounts). | ||
| accountIds | Yes | ||
| publishNow | No | Publish immediately after creating. Ignored when scheduled_at is set. | |
| external_id | No | ||
| scheduled_at | No | ISO datetime at least 2 minutes ahead. Omit for a draft or publishNow. | |
| google_business | No | Google Business Profile post settings (applies to google_business accounts). | |
| accountOverrides | No | Per-account title/caption overrides (id must also be in accountIds). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full behavioral burden and does so well: it discloses draft behavior, publishNow semantics, paid-plan requirement, automatic image fitting, rejection with per-platform issues list, and app-panel rendering instructions. These are exactly the traits an agent needs before invoking a write operation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long but front-loads purpose, usage, and the compose_post alternative before diving into platform rules. Most sentences carry operational detail, though the platform list and some image-ratio guidance could be tightened; still, it is structured for scanning.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 17-parameter, multi-platform tool with no output schema and no annotations, the description covers core invocation behavior, platform constraints, error handling, and UI rendering. It still omits how to source accountIds (get_accounts), alternatives like update_post/publish_post for existing drafts, and constraints for several supported platforms.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 65%, so the description must add value, and it does: it gives platform-specific constraints for YouTube title, Pinterest media, X character/media limits, and Google Business Profile limits. It does not explain every parameter (e.g., external_id, accountOverrides, mediaUrl vs mediaUrls, timezone), leaving some gaps.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific set of verbs (create, schedule, queue, publish, cross-post) and the resource (content to connected social accounts), lists supported platforms, and explicitly names compose_post as the alternative for self-serve editing. An agent can distinguish this from siblings like compose_post, update_post, and publish_post.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It states exactly when to use it ('any request to post now, schedule for later, publish to multiple platforms at once, or add something to the posting queue') and when to prefer compose_post ('when the user wants to set the post up themselves or per-platform options are still open'). It also clarifies draft vs immediate behavior with scheduled_at/publishNow.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_thumbnailCreate ThumbnailAInspect
Generate a YouTube thumbnail (16:9) for a video you are publishing. Returns one job per variation. Renders a live preview grid in app-capable hosts; poll each id with wait_for_image otherwise.
| Name | Required | Description | Default |
|---|---|---|---|
| prompt | No | ||
| context | No | ||
| expression | No | ||
| variations | No | ||
| templateImageUrls | No | ||
| referenceImageUrls | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does disclose the critical behavioral trait: this is asynchronous and returns one job per requested variation, with polling required via wait_for_image. It omits cost/credit implications, whether generated thumbnails persist or can be deleted, and what happens when all six optional parameters are omitted.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, no waste, and the most decision-relevant facts (what it makes, and how you get the output) are front-loaded. Nothing repeats the tool name or title.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a six-parameter generation tool with zero required parameters, no annotations and no output schema, the description covers the async return contract but leaves parameter meaning and the zero-required-parameter behavior unexplained. It is minimally viable: enough to invoke the tool and handle its jobs, not enough to invoke it well.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and six parameters (prompt, context, expression, variations, templateImageUrls, referenceImageUrls) have no schema-level documentation. The description only incidentally hints at 'variations' via 'one job per variation' and says nothing about the prompt, context, expression, or the template vs reference image distinction, so it fails to compensate for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource with format detail: 'Generate a YouTube thumbnail (16:9) for a video you are publishing.' That is enough to distinguish it from generic image tools, but it never explicitly contrasts itself with siblings like create_image or create_movie_material, nor does it explain what 'context' and 'expression' contribute to the result.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives real routing guidance: results render as a live preview grid in app-capable hosts, while other hosts must poll each returned id with wait_for_image. That tells the agent how to consume the output in two environments. It does not say when to prefer this over create_image for non-YouTube artwork, so the alternative-selection guidance is incomplete.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_videoCreate VideoAInspect
Generate a short video for a social post, Reel, TikTok or ad. Models include Seedance 2.5 / 2.5 Turbo / 2.0, Veo 3.1 (Fast, Lite), Gemini Omni 1.1 Flash, Kling O3 Pro / 3.0, Wan 3.0, Grok Imagine Video 1.5 and Sora 2 — see list_models for modes and inputs. Video is expensive: call get_price first, tell the user the credits, and pass them as maxCredits. Poll with wait_for_video. Renders a live preview in app-capable hosts.
| Name | Required | Description | Default |
|---|---|---|---|
| mode | No | Mode key from list_models. Omit to pick it from the attached media. | |
| model | No | Model key from list_models (category "video"). Default: seedance-2.5. | |
| inputs | No | WaveSpeed input fields for the chosen model and mode, exactly as list_models shows them (e.g. aspect_ratio, resolution, duration, generate_audio, image, last_image, reference_images). Media fields take URLs; files hosted elsewhere are imported into the user's library automatically. Omitted fields use the model's defaults. | |
| prompt | Yes | What to generate | |
| maxCredits | No | Refuse to start if the live price is above this. Pass the credits get_price returned. | |
| idempotencyKey | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does well: it flags that video is 'expensive' (cost profile), that generation is asynchronous and must be polled with wait_for_video, and that a live preview renders in app-capable hosts. It does not cover failure/refund behavior or auth requirements, so it falls just short of complete.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four sentences, front-loaded with purpose and use cases, then workflow, with zero filler. The inline model enumeration is long but plausibly useful for model selection; it is the only element that mildly dilutes density.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an async, credit-gated generation tool with no output schema and no annotations, the description covers cost gating, polling, model discovery and preview behavior. Minor gaps remain around idempotencyKey usage and what happens if a render fails or is rejected on credits.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 83%, so the schema already documents mode, model, inputs and maxCredits, giving a baseline of 3. The description adds real meaning by tying maxCredits to the get_price workflow ('pass them as maxCredits') and by directing the agent to list_models for valid mode/model/input keys, which the schema only references abstractly.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Generate a short video') plus the concrete use cases (social post, Reel, TikTok, ad), and enumerates the supported models. Combined with the sibling list (create_image, create_avatar_video, create_motion_control), an agent can immediately distinguish this from the other creation tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit prerequisite chain: see list_models for modes/inputs, call get_price first, tell the user the credits, pass them as maxCredits, then poll with wait_for_video. It names three alternative sibling tools and the exact condition under which each is needed, leaving nothing to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_voice_cloneClone VoiceAInspect
Create an instant voice clone from a speech recording (ElevenLabs). Only clone the user's own voice or a voice they have permission to use. Best results: 1–2 minutes of clean speech with no music or background noise. Costs 15 credits; max 10 clones per account. Returns clone.id — use it as voice_id in create_voiceover, create_avatar_video and change_voice. In app-capable hosts the result renders as an interactive panel the user can see and act on — do NOT repeat its contents as a list or table in your reply; add only what the panel doesn't say.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | Name for the cloned voice. | |
| sampleUrl | Yes | Public URL of the voice sample (mp3, wav, m4a, ogg, webm or mp4). Use upload_file first for a file the user provides. | |
| description | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does so well: discloses cost (15 credits), a per-account cap (max 10 clones), the return handle (clone.id), and non-obvious host-rendering behavior ('renders as an interactive panel... do NOT repeat its contents'). These are exactly the traits an agent cannot infer from the schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Dense but front-loaded: purpose, legal/prereq constraint, quality guidance, cost/limits, return value, host-rendering rule. Every sentence earns its place with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, yet the description explains the salient return value (clone.id) and its downstream use, plus the host-panel behavior. For a 3-param mutation tool, nothing an agent needs to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 67% and the schema already documents name and sampleUrl (including formats and the upload_file hint). The description adds input-quality guidance ('1–2 minutes of clean speech with no music or background noise') that the schema does not provide, though it leaves the optional description field unaddressed.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Create an instant voice clone from a speech recording') and names the provider. It is clearly distinguishable from siblings like delete_voice_clone, list_voices and change_voice.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit when-to-use constraints (only the user's own voice or one they have permission for), prerequisite routing ('use upload_file first'), and downstream routing (use clone.id as voice_id in create_voiceover, create_avatar_video, change_voice). This is near-complete routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_voiceoverCreate VoiceoverAInspect
Generate a voiceover for a social video, Reel or ad from a script. Use list_voiceover_voices first and let the user pick an ElevenLabs Eleven v3 voice from the previews; the user's cloned voices and MiniMax presets (see list_voices) work too. Costs: Eleven v3 6 credits per started 1,000 characters; cloned voices 3 credits per started 500 characters; MiniMax presets 1 credit per started 1,000 characters. Supports <#x#> tags between words to pause for x seconds (0.01–99.99). Returns generation.id — poll with wait_for_voiceover. Renders a live audio player in app-capable hosts.
| Name | Required | Description | Default |
|---|---|---|---|
| text | Yes | Script to voice (max 10,000 characters). | |
| speed | No | Speaking speed 0.5–2, default 1. | |
| emotion | No | Delivery emotion — MiniMax preset voices only. Default neutral. | |
| voice_id | Yes | Eleven v3 voice (Aria, Roger, Sarah, Laura, Charlie, George, Callum, River, Liam, Charlotte, Alice, Matilda, Will, Jessica, Eric, Chris, Brian, Daniel, Lily, Bill) — see list_voiceover_voices for previews — or a cloned voice id / MiniMax preset id from list_voices. | |
| stability | No | Eleven v3 only. 0–1, default 0.5. Higher = more consistent delivery, lower = more expressive. | |
| similarity | No | Eleven v3 only. 0–1, default 1. How closely the output sticks to the base voice. | |
| idempotencyKey | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden and largely succeeds: it discloses per-provider credit costs, the pause-tag syntax, that it is async (generation.id + polling), and host rendering behavior. It omits failure modes, rate limits, and whether credits are refunded on failure, keeping it just below the top band.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Roughly five dense sentences, each front-loaded with a distinct payload: purpose, voice-selection workflow, cost matrix, pause-tag syntax, and return/polling contract. No filler or repeated schema content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description proactively supplies the return contract ('generation.id — poll with wait_for_voiceover'), and it also covers cost and text-formatting concerns an agent would otherwise miss. Only the idempotencyKey parameter is left unexplained, a minor gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 86%, so the schema already documents nearly every parameter and a baseline of 3 applies. The description adds genuinely new meaning beyond the schema, notably the pause-tag syntax that governs the 'text' field and the voice-selection provenance (Eleven v3 list vs cloned vs MiniMax presets).
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Generate a voiceover ... from a script') plus the target artifact class (social video, Reel or ad). An agent can distinguish this from create_music or create_voice_clone without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit prerequisite workflow ('Use list_voiceover_voices first and let the user pick') and a downstream step ('poll with wait_for_voiceover'), which is strong context. It never states when NOT to use it or contrasts directly with a conflicting sibling like change_voice, so it falls short of the full 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
delete_generationDelete GenerationADestructiveInspect
Permanently delete one result from the user's library (type + id from list_generations), or an uploaded file (type upload, id from list_uploads). Confirm with the user first — this cannot be undone. In app-capable hosts the result renders as an interactive panel the user can see and act on — do NOT repeat its contents as a list or table in your reply; add only what the panel doesn't say.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | ||
| type | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, so the safety flag is covered, but the description adds real value beyond it: the operation is permanent and cannot be undone, requires user confirmation, and in app-capable hosts renders an interactive panel that should not be echoed back as text. These are non-obvious behavioral traits an agent could not infer from the schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose is front-loaded, followed by the irreversibility/confirmation constraint and then the host-rendering guidance. The final sentence is long but every clause is actionable; nothing is redundant.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 2-parameter, no-output-schema destructive tool, the description covers both operating modes, id sourcing, irreversibility, confirmation, and host-specific output handling. Nothing material for correct invocation is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must carry the load. It explains the origin of 'id' (list_generations / list_uploads) and the special meaning of type='upload', which compensates for the highest-value ambiguity, but it does not explain the 14 other enum values or the type/id pairing relationship.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (permanently delete) and resource (a generation result or uploaded file), and disambiguates the two modes via the type field. An agent can tell it apart from delete_post and delete_voice_clone without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly routes the agent to the source tools for the id (list_generations for results, list_uploads for the 'upload' type) and mandates confirming with the user first. It does not, however, contrast against the sibling delete_post/delete_voice_clone tools, so the exclusion side is implied rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
delete_postDelete PostBInspect
Delete a post. Does not remove already-published content from the platforms. In app-capable hosts the result renders as an interactive panel the user can see and act on — do NOT repeat its contents as a list or table in your reply; add only what the panel doesn't say.
| Name | Required | Description | Default |
|---|---|---|---|
| postId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It usefully discloses that deletion is scoped and does not touch already-published platform content, and it explains the interactive-panel rendering and the instruction not to echo its contents. It omits irreversibility, permission/auth requirements, and the fate of scheduled-but-unpublished posts.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The core action and its scope are front-loaded in the first two sentences. The final clause about panel rendering is somewhat tangential to deletion but is short and actionable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive tool with no annotations and no output schema, the description covers scope and UI rendering but leaves key gaps: whether the operation is irreversible, what permissions it needs, and how scheduled/published posts are affected.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The single parameter postId has 0% schema description coverage and the description adds nothing about it, so the agent must infer it identifies the post to remove. The name is conventional enough that the gap is minor, but no meaning is added beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ("Delete a post") and clarifies scope with "Does not remove already-published content from the platforms." It does not, however, distinguish itself from the many sibling delete tools (delete_generation, delete_voice_clone) or from update_post/list_posts.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this versus update_post, publish_post, or the other delete siblings, and no prerequisites or preconditions (e.g. must the post be unpublished first). The scoping sentence describes behavior, not usage conditions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
delete_voice_cloneDelete Voice CloneADestructiveInspect
Permanently delete one of the user's cloned voices. Confirm with the user first. In app-capable hosts the result renders as an interactive panel the user can see and act on — do NOT repeat its contents as a list or table in your reply; add only what the panel doesn't say.
| Name | Required | Description | Default |
|---|---|---|---|
| cloneId | Yes | Clone id from list_voices. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare destructiveHint=true. The description adds significant behavioral context beyond that: the deletion is permanent, requires user confirmation, and specifies how output renders in app-capable hosts, instructing the agent not to duplicate the panel. This exceeds what annotations alone provide.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, front-loaded with the core action. The confirmation and output-panel instructions are focused and earn their place, though the third sentence is somewhat long. No obviously wasted text.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given a simple, single-parameter destructive tool with annotations and no output schema, the description covers the essentials: permanence, confirmation, and output handling. It omits error behavior and explicit routing to list_voices, but those are minor gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% for the single cloneId parameter, and the schema itself documents it as 'Clone id from list_voices.' The description adds no parameter-level details beyond the schema, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('delete') and resource ('cloned voices') with scope ('one of the user's'). It distinguishes from other delete tools such as delete_generation and delete_post by resource type. However, it does not explicitly route to or differentiate from siblings like list_voices or create_voice_clone.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides a key prerequisite: 'Confirm with the user first.' However, it does not state when to use this tool versus alternatives (e.g., using list_voices to find the cloneId) or when not to delete. Usage is implied by the tool name and purpose.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_accountsConnected AccountsARead-onlyInspect
List connected social accounts for scheduling. Supported platforms: YouTube, TikTok, Instagram, Facebook, LinkedIn, Pinterest, Threads, Bluesky, X and Google Business Profile (one account per business location). Each account has a provider field identifying its platform (YouTube accounts have provider "google"). In app-capable hosts the result renders as an interactive panel the user can see and act on — do NOT repeat its contents as a list or table in your reply; add only what the panel doesn't say.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnlyHint=true, openWorldHint=false), so the description's value is the added behavioral context: supported-platform enumeration, the provider field mapping, and the app-capable-host interactive-panel rendering instruction with an explicit "do NOT repeat" directive. Auth requirements and pagination are unstated, keeping this short of a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with purpose, then platform coverage, then return-field detail, and finally the output-presentation instruction. The platform enumeration is lengthy but informational rather than redundant, so each sentence largely earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, so the description must carry the return-value burden, and it does: it names the platforms, the provider field, and the panel rendering behavior. Minor gaps remain (auth/permission prerequisites, ordering), but for a parameterless list tool this is nearly complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With zero input parameters, the baseline is 4. The description adds useful semantic detail about the returned data shape (provider field, e.g. YouTube accounts use provider "google"), though this concerns outputs rather than inputs.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ("List connected social accounts") plus the scope/purpose ("for scheduling"). This is clearly distinguishable from sibling list tools (list_posts, list_generations, list_voices), which concern posts or media rather than connected accounts.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied by "for scheduling" — the agent can infer it enumerates the accounts available before composing/publishing. There is no explicit when-to-use/when-not statement and no named alternative among the many sibling tools, so it meets only the minimum viable bar.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_creditsCheck CreditsARead-onlyInspect
Get current credit balance. In app-capable hosts the result renders as an interactive panel the user can see and act on — do NOT repeat its contents as a list or table in your reply; add only what the panel doesn't say.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=false. The description adds genuinely new behavioral context: that in app-capable hosts the result renders as an interactive user-facing panel and that the agent should not duplicate it. That is useful output-presentation behavior not derivable from the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action in the first sentence, followed by one longer sentence of rendering guidance. Slightly wordy but every clause earns its place by preventing redundant output.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, but the description explains the return shape that matters (an interactive panel plus the caveat that it may not render in all hosts). For a no-param, read-only tool this is close to complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Zero parameters, so the baseline of 4 applies. There is nothing for the description to clarify and it introduces no parameter-related confusion.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Get current credit balance'), which is clear enough to separate it from get_subscription and get_price. It does not explicitly name or contrast those siblings, so it stops short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives real guidance on how to use the *result* (don't repeat the panel contents), but says nothing about when to call this tool versus get_subscription, get_price, or get_accounts. Usage is only implied by the resource name.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_postGet PostARead-onlyInspect
Fetch one post by id, including per-account publish status and media. In app-capable hosts the result renders as an interactive panel the user can see and act on — do NOT repeat its contents as a list or table in your reply; add only what the panel doesn't say.
| Name | Required | Description | Default |
|---|---|---|---|
| postId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish that this is a safe, non-open-world read (readOnlyHint=true, openWorldHint=false). The description adds useful behavior beyond that: the payload includes per-account publish status and media, and in app-capable hosts the result is rendered as an interactive panel with an explicit instruction not to duplicate it as text. That is meaningful context not derivable from the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the core purpose. The second sentence is dense but each clause earns its place by telling the agent how to handle rendered output; no filler or repetition of the title.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description steps in to name the return contents (publish status, media) and the host-dependent rendering, so an agent knows what to expect. It stops short of covering failure cases such as an unknown or deleted post id, a minor gap for a simple read.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There is a single required parameter (postId) with 0% schema description coverage, so the description must carry the meaning. 'Fetch one post by id' tells the agent the parameter identifies a post, which is adequate but adds no format, source, or retrieval detail beyond the obvious.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource with scope: 'Fetch one post by id'. The singular 'one post' and 'by id' clearly distinguish it from the plural list_posts sibling, and the mention of per-account publish status and media sharpens what is returned.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied rather than stated: an agent can infer this is the single-item counterpart to list_posts, but no alternative or condition is named explicitly. The only real guidance concerns how to present the result in app-capable hosts, not when to choose this tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_post_analyticsPost AnalyticsAInspect
Performance of the user's published posts (Scheduler Analytics): totals for views, likes, comments and shares plus per-platform breakdowns over the last N days. Pass refresh: true to pull fresh metrics from the platforms first (slower). In app-capable hosts the result renders as an interactive panel the user can see and act on — do NOT repeat its contents as a list or table in your reply; add only what the panel doesn't say.
| Name | Required | Description | Default |
|---|---|---|---|
| days | No | Window in days, 7–90. Default 30. | |
| refresh | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden and does disclose real behavioral traits: refresh incurs a latency cost because it hits the platforms first, and results render as an interactive panel whose contents should not be duplicated. It omits auth/permission requirements and any staleness/pagination behavior, so it isn't fully exhaustive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core purpose, followed by the refresh tradeoff, then the rendering instruction — a logical order. The panel-response directive is wordy ('do NOT repeat its contents as a list or table in your reply') but earns its place by preventing redundant output.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and no annotations, the description compensates by enumerating the return shape (totals plus per-platform breakdowns) and covering the refresh semantics and host-dependent rendering. Adequate for an agent to call and interpret it, though auth/permission context and metric freshness defaults are absent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 50% (the refresh boolean has no schema description), so the description must compensate — and it does, explaining that refresh: true pulls fresh platform metrics and is slower. The days parameter adds little beyond the schema's 7–90 range and default, but the uncovered parameter is well handled.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Combines a specific verb (get/retrieve analytics) with a precise resource (the user's published posts) and names the returned payload: totals for views, likes, comments and shares plus per-platform breakdowns. This clearly distinguishes it from siblings like get_post or list_posts, which return post data rather than aggregated performance metrics.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explains the refresh: true tradeoff (pulls fresh metrics from platforms first, slower) and gives an explicit when-to-do-what instruction for reply formatting in app-capable hosts. It stops short of naming alternatives or exclusions, but the operative context for invoking it correctly is present.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_priceGet Generation PriceARead-onlyInspect
Live credit price for one video, image or motion-control generation, quoted by the provider for these exact settings (resolution, duration, audio, attached media length). Free. Use before expensive video jobs, tell the user the cost, and pass the credits as maxCredits to the create tool.
| Name | Required | Description | Default |
|---|---|---|---|
| mode | No | ||
| model | No | Model key from list_models | |
| inputs | No | WaveSpeed input fields for the chosen model and mode, exactly as list_models shows them (e.g. aspect_ratio, resolution, duration, generate_audio, image, last_image, reference_images). Media fields take URLs; files hosted elsewhere are imported into the user's library automatically. Omitted fields use the model's defaults. | |
| prompt | No | ||
| category | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=false, so safety is covered. The description adds behavior annotations don't carry: it is free, it is a live provider quote, and the price is only valid for the exact settings supplied. It still says nothing about latency or failure modes when a model/settings combination is invalid.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences, front-loaded with what it returns, and every clause carries actionable information (settings scope, free, downstream maxCredits use). No filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description correctly explains the return ('live credit price') and what to do with it, and it flags the nested inputs object implicitly via the settings list. For a 5-param nested tool it could say more about required vs optional inputs and invalid-setting behavior, but nothing essential is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 40% across 5 params, but the description names the settings that drive the quote (resolution, duration, audio, attached media length) and explains the output's use as maxCredits. It adds no guidance on mode, prompt, or how category interacts with the other fields, so it only partially compensates for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Live credit price for one video, image or motion-control generation') and pins the scope to 'these exact settings', which separates it from the sibling get_credits (account balance). An agent can tell immediately this is a per-generation cost quote, not a balance check.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear when-to-use ('Use before expensive video jobs') and a concrete downstream workflow ('tell the user the cost, and pass the credits as maxCredits to the create tool'). It doesn't state exclusions or contrast with get_credits, so it falls just short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get-skillGet SkillARead-onlyInspect
ADMIN ONLY — fetch a skill from the Unsora skills catalog by id or slug. Returns the full record (any status) including page content, uploaded files, and attached media. Use before update-skill to see the current state. Non-admin accounts get a 403.
| Name | Required | Description | Default |
|---|---|---|---|
| skill | Yes | Skill id or slug. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only cover readOnlyHint and openWorldHint; the description adds the admin gating, the 403 failure mode, and the fact that the full record is returned regardless of status, including page content, files, and media. That is substantive behavioral context beyond the structured hints.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences with the most consequential constraint (ADMIN ONLY) front-loaded, followed by return contents, then the workflow hint. Every sentence carries distinct information and none is padding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter read with no output schema, the description covers access requirements, lookup key, and return payload composition, which is everything an agent needs to call it correctly. No meaningful gap remains.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the single parameter is already documented as "Skill id or slug." The description's phrase "by id or slug" merely restates the schema, adding no format, prefix, or resolution-order detail, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ("fetch a skill from the Unsora skills catalog by id or slug") and immediately qualifies the scope with ADMIN ONLY. It also distinguishes itself from the sibling update-skill by framing itself as the read-side counterpart, so an agent can route without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit call-site ("Use before update-skill to see the current state") and an explicit exclusion ("Non-admin accounts get a 403"), which tells the agent both when to reach for it and when it will fail. Nothing about selection is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_subscriptionCheck SubscriptionARead-onlyInspect
Get plan and Stripe subscription state (check isActive before scheduling posts). In app-capable hosts the result renders as an interactive panel the user can see and act on — do NOT repeat its contents as a list or table in your reply; add only what the panel doesn't say.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnlyHint=true, openWorldHint=false), but the description adds genuinely non-obvious behavior: in app-capable hosts the result renders as an interactive panel, and the agent must not duplicate its contents in the reply. That host-dependent rendering rule is not inferable from schema or annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, zero filler, front-loaded with the resource and the field the agent should act on. The second sentence is dense but each clause (panel rendering, don't duplicate, add only missing info) carries distinct guidance.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, so the description carries the burden of indicating what comes back; it names plan, Stripe subscription state, and the isActive flag, and explains how the result is surfaced. It could say more about edge cases (e.g., no subscription present), but it is sufficient for a zero-parameter read.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters and schema coverage is 100%, so there is nothing for the description to clarify. Baseline 4 applies for a no-argument tool.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Get plan and Stripe subscription state') plus the key field of interest ('isActive'). It is distinguishable from related siblings like get_credits and get_accounts, though it does not name them explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a concrete triggering condition ('check isActive before scheduling posts'), which tells the agent when this call matters. It stops short of naming an alternative tool or stating when not to call it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
image_statusImage StatusCRead-onlyIdempotentInspect
Single-shot generation status check (used by the preview UI).
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | ||
| tool | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and openWorldHint=false, so the safety profile is covered. The description adds one genuinely useful trait beyond the annotations: 'single-shot' signals a non-blocking, immediate-return check, in contrast to the wait_for_* siblings. It does not disclose what states can be returned or what happens for an unknown/idle id.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single short sentence with no filler, and the core function is front-loaded before the parenthetical qualifier. It is efficiently sized for a two-parameter status tool, though the terseness edges toward under-specification rather than true concision.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and no annotations describing return values, the description should explain what a status result looks like (pending/succeeded/failed) and how it relates to wait_for_image; it does neither. Given two undocumented parameters and 0% schema coverage, an agent lacks enough information to call this confidently versus the dedicated wait tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for both required parameters, so the description carries the full burden and fails it. It never explains that `id` is the generation identifier to check or that `tool` selects among the 13 generation types in the enum, leaving the agent to infer the meaning of a large enum from bare strings.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a verb (status check) and a resource (single-shot generation), which is clearer than a tautology, but it never says which generation type this covers or how it relates to siblings like video_status, audio_status, and clipping_status. The parenthetical '(used by the preview UI)' muddies rather than sharpens the purpose, hinting the tool is UI plumbing rather than an agent-facing query.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use or when-not-to-use guidance is given. The most important routing decision — whether an agent should call this or the sibling wait_for_image (which presumably polls until completion) — is left entirely unaddressed. The 'used by the preview UI' note implies a context of use without stating whether an agent should adopt it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_generationsList GenerationsARead-onlyInspect
Browse the user's past results for one feature, newest first, with status and output URLs: image, video, music, voiceover, influencer, thumbnail, clipping, image_upscale, video_upscale, watermark_removal, motion_control, avatar, movie_material, voice_change. Use it to find something made earlier (e.g. to post or reuse it). Uploaded files are in list_uploads. In app-capable hosts the result renders as an interactive panel the user can see and act on — do NOT repeat its contents as a list or table in your reply; add only what the panel doesn't say.
| Name | Required | Description | Default |
|---|---|---|---|
| page | No | ||
| type | Yes | ||
| limit | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare readOnlyHint=true and openWorldHint=false. The description adds real behaviour: results are ordered newest-first, carry status and output URLs, and render as an interactive panel in app-capable hosts with instructions not to duplicate its contents. It does not cover pagination or how status values should be interpreted, so it falls short of a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core scope before the routing note and the host-rendering instruction; every sentence serves a purpose. The inline list of all 14 type values duplicates the schema enum and slightly bloats the sentence.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and thin annotations, the description does useful work by naming the returned fields (status, output URLs) and the sort order, plus panel-rendering behaviour. The gaps are pagination semantics and status interpretation, which leaves it a step short of complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must carry parameter meaning. It clarifies that 'type' selects one feature (though the enum values are already in the schema) but never explains 'page' or 'limit' — no mention of pagination, defaults, or the 100 cap. Two of three parameters are undocumented in both places.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('browse the user's past results for one feature'), its ordering ('newest first'), and its payload ('status and output URLs'), then enumerates the supported feature types. It also explicitly routes uploaded files to the sibling list_uploads, so an agent can distinguish it from neighbours without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit use case ('find something made earlier — e.g. to post or reuse it') and names the alternative for a related need ('Uploaded files are in list_uploads'). It additionally states how to behave in app-capable hosts, which is directly actionable guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_modelsList Generation ModelsARead-onlyInspect
Every video, image and motion-control model with its modes and exact input fields (WaveSpeed's own names, enums, ranges and defaults). Call this before create_video / create_image / create_motion_control to pick a model and build inputs.
| Name | Required | Description | Default |
|---|---|---|---|
| category | No | Only list one category |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=false, so safety is covered. The description adds useful behavioral context about the return payload (model names, modes, enums, ranges, defaults) and that it is a prerequisite step for the create_* tools, though it says nothing about size of the catalog or whether results are cached/live.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences, zero filler, with the what (full model catalog) front-loaded and the when (before the create_* calls) immediately after.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, yet the description explains what an agent receives (models with modes and exact input fields) and why it matters (building `inputs`). For a read-only one-parameter list tool, nothing needed to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the single `category` enum is fully documented in the schema. The description's 'video, image and motion-control' wording loosely mirrors the enum values but adds no filtering syntax or behavior beyond what the schema already gives; baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific resource (video, image and motion-control models) and what each entry carries (modes, exact input fields, enums, ranges, defaults). It is clearly distinguished from the create_* siblings it feeds, which it names explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit when-to-call rule: 'Call this before create_video / create_image / create_motion_control to pick a model and build `inputs`.' The alternatives and the ordering constraint are both stated, leaving nothing to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_postsList PostsARead-onlyInspect
List the user's posts, newest first: drafts, scheduled, publishing, published and failed. Filter with status (DRAFT, SCHEDULED, PUBLISHING, PUBLISHED, PARTIALLY_PUBLISHED or FAILED). In app-capable hosts the result renders as an interactive panel the user can see and act on — do NOT repeat its contents as a list or table in your reply; add only what the panel doesn't say.
| Name | Required | Description | Default |
|---|---|---|---|
| page | No | ||
| limit | No | ||
| status | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnlyHint=true, openWorldHint=false). The description adds genuine behavioral context beyond that: in app-capable hosts the result renders as an interactive panel and the agent must not duplicate its contents. Missing are pagination/ordering semantics for large result sets.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, each serving a purpose: capability, filtering, and host-rendering behavior. Front-loaded with the core action. Minor redundancy listing statuses twice (prose and parenthetical).
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only list tool with no output schema, the description covers scope, ordering, filtering, and return rendering. The remaining gap is pagination semantics (page/limit), which an agent may need for large post collections.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must carry parameter meaning. It documents the status filter (though the enum already lists these values in the schema) but says nothing about page or limit, leaving two of three parameters unexplained. Baseline 3 for partial compensation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (List) and resource (the user's posts), plus ordering ('newest first') and the covered statuses. An agent can distinguish it from get_post (single-item) and list_generations without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly documents the status-filter use case and enumerates filter values, and adds the important host-conditional rendering instruction. It does not name sibling alternatives (get_post, list_generations) or state when not to use it, so it stops short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_uploadsList UploadsBRead-onlyInspect
List the user's uploaded media assets (from upload_file or the app). In app-capable hosts the result renders as an interactive panel the user can see and act on — do NOT repeat its contents as a list or table in your reply; add only what the panel doesn't say.
| Name | Required | Description | Default |
|---|---|---|---|
| page | No | ||
| limit | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=false, so safety is covered. The description adds meaningful context beyond that: the result may render as an interactive panel in app-capable hosts and the agent should not duplicate it in text. That host-dependent rendering behavior is real information not derivable from the schema or annotations, though pagination and return-shape behavior remain unstated.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with purpose before the presentation rule, with no filler. The second sentence is slightly dense with the em-dash aside, but every clause carries actionable instruction.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, so the description does carry the burden of explaining results — and it does describe the interactive-panel rendering well. However, it omits result volume, pagination behavior, and what fields assets expose, which an agent needs when the panel does not render (non-app hosts).
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% — neither 'page' nor 'limit' carries any explanation in the schema, and the description does not mention pagination, defaults, or the 100 maximum at all. With two undocumented parameters and a tool that can return many assets, the description fails to compensate for the schema gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('List the user's uploaded media assets') and clarifies provenance with '(from upload_file or the app)'. This is clearer than a bare name restatement, but it never contrasts itself with the many other list_* siblings (list_generations, list_posts) that could plausibly return overlapping content, so an agent must infer the boundary.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides a clear directive for how to present results in app-capable hosts, but no guidance on when to choose this tool over list_generations or list_posts, nor any indication of when the tool is unnecessary. The negative instruction ('do NOT repeat its contents') is useful but is about response formatting, not tool selection.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_voiceover_voicesList Voiceover VoicesARead-onlyInspect
List the 20 ElevenLabs Eleven v3 voices available for create_voiceover, each with a description, gender, accent, and a short preview clip URL. In app-capable hosts (e.g. Claude.ai) this tool renders an interactive voice picker with tap-to-play previews — the user already sees every voice, so do NOT repeat the voices as a table, list, or summary in your reply; just tell the user to pick from the picker above. Only in hosts without the UI, share the preview URLs so the user can hear a voice before choosing it.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare a safe read (readOnlyHint=true, openWorldHint=false), but the description adds meaningful behavior beyond that: the fixed count of 20, the exact returned fields, and the host-dependent UI rendering (interactive picker with tap-to-play previews). That rendering behavior is not derivable from any structured field.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, front-loaded with what the tool returns, then the host-dependent response instructions. Every sentence earns its place, though the picker instruction is dense enough that it could be tightened slightly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no parameters and no output schema, the description fully carries the burden: it names the output fields, the count, and the crucial response-handling rule for UI vs non-UI hosts. An agent has everything needed to call it and respond correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema is empty with zero parameters, so there is nothing to document and the baseline of 4 applies. The description correctly implies the tool is a no-argument listing call, though it adds no parameter detail (none exists to add).
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a precise verb+resource with scope ('the 20 ElevenLabs Eleven v3 voices available for create_voiceover') and lists the returned fields, so it is immediately distinguishable from the sibling list_voices and tied to its consumer create_voiceover.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit conditional guidance: in app-capable hosts use the picker and do NOT repeat the voices as a table/list/summary; only in hosts without the UI share preview URLs. This is exactly the when/how-to-present routing an agent needs and leaves nothing to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_voicesList All VoicesARead-onlyInspect
List every voice the user can use: MiniMax presets (presets), ElevenLabs Eleven v3 voices with preview URLs (elevenV3Voices), and the user's own cloned voices (clones — use the clone id). Voices work as voice_id in create_voiceover and create_avatar_video; change_voice needs a cloned voice. Also returns the clone limit (maxClones) and cloneCreditCost. In app-capable hosts the result renders as an interactive panel the user can see and act on — do NOT repeat its contents as a list or table in your reply; add only what the panel doesn't say.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare readOnlyHint and openWorldHint; the description adds substantially more — that ElevenLabs entries carry preview URLs, that clones must be referenced by clone id, that clone limit and credit cost are returned, and that app-capable hosts render an interactive panel whose contents should not be echoed in the reply. That last point is non-obvious host-dependent behavior an agent would otherwise get wrong.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core purpose, then layered with return categories, downstream usage, and a rendering caveat. The final panel instruction is comparatively long but carries non-obvious guidance, so it earns its place; the sentence structure is dense but readable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no input parameters and no output schema, the description carries the full burden of explaining returns, and it does so thoroughly — categories, identifier conventions, credit-related fields, and host rendering. An agent has everything needed to call it and to interpret and present the result correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so the baseline is 4 per the rubric. The description does name the result field keys (presets, elevenV3Voices, clones, maxClones, cloneCreditCost), which helps with return interpretation, but there are no parameter semantics to clarify.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource and goes further by enumerating the three voice categories returned (presets, elevenV3Voices, clones) plus the aggregate fields (maxClones, cloneCreditCost). It does not, however, distinguish itself from the sibling list_voiceover_voices, which an agent could easily confuse with this tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear downstream context: voices returned here work as voice_id in create_voiceover and create_avatar_video, and change_voice specifically requires a cloned voice. It stops short of stating when to call this versus list_voiceover_voices, so no explicit alternative or exclusion is named.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
pinterest_boardsPinterest BoardsBRead-onlyIdempotentInspect
Boards a connected Pinterest account can pin to (used by the composer UI).
| Name | Required | Description | Default |
|---|---|---|---|
| accountId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and openWorldHint=false, so safety and determinism are covered. The description adds one genuinely useful behavioral fact — that boards belong to a connected Pinterest account, implying an account-linking prerequisite — but says nothing about rate limits, pagination, or what happens if the account isn't connected.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler; the resource is stated before the provenance note. It is efficient, though its brevity is partly under-specification rather than pure economy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter read-only list tool with no output schema, the definition is minimally viable but leaves the agent guessing about the return shape — notably whether board IDs usable by a pinning/composing flow are included. The connected-account prerequisite is mentioned, which is the most important missing piece, but the rest is thin.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for the single accountId parameter, so the schema provides no meaning. The description partially compensates by framing it as 'a connected Pinterest account,' but does not clarify whether accountId is an internal account identifier or a Pinterest user handle, nor where to obtain it.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific resource (boards) restricted to a connected Pinterest account, which reads as a retrieval/list operation and clearly distinguishes it from the write-oriented siblings like create_post or publish_post. The verb is implied rather than stated, so it falls just short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The parenthetical '(used by the composer UI)' notes where the data is consumed but gives the agent no when-to-use guidance, no prerequisites (must the account be connected first?), and no comparison to alternatives such as tiktok_creator_info or get_accounts. Context is hinted at, not stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
publish_postPublish Post NowAInspect
Publish a DRAFT or SCHEDULED post to its accounts right now instead of waiting for its schedule. Requires paid plan. Publishing runs in the background: this returns status PUBLISHING straight away. Call get_post after a minute or two to see each account's result; if some accounts fail, use retry_post. In app-capable hosts the panel tracks publishing live, so don't poll get_post there. In app-capable hosts the result renders as an interactive panel the user can see and act on — do NOT repeat its contents as a list or table in your reply; add only what the panel doesn't say.
| Name | Required | Description | Default |
|---|---|---|---|
| postId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and delivers: a paid-plan prerequisite, asynchronous background execution, the immediate PUBLISHING return status, the recommended follow-up delay ('a minute or two'), and host-specific panel behavior including a directive not to duplicate panel contents. This is unusually complete behavioral disclosure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action and scope before the operational details. Each sentence carries distinct information, though the two app-capable-host sentences overlap slightly and could be merged without loss.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description must explain returns — and it does, stating the tool returns PUBLISHING immediately and where to look for per-account results. Prerequisites, async behavior, and error recovery are all covered.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Only one parameter exists and the schema documents none of it, but postId is self-evident from the name. The description adds real meaning by constraining the target to a DRAFT or SCHEDULED post, implying posts in other states are invalid input.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (publish) and resource (a DRAFT or SCHEDULED post) and clarifies the semantics: publish now rather than waiting for the schedule. This distinguishes it cleanly from siblings like create_post, update_post, and retry_post.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says when to use it (publish a draft/scheduled post immediately), when not to act (don't poll get_post in app-capable hosts), and names the alternatives (get_post to check results, retry_post for failures). Routing guidance is fully specified.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
remove_watermarkRemove Subtitles & WatermarksAInspect
Remove burned-in subtitles, captions, watermarks, logos or on-screen text from a video (Subtitle Remover). Costs 10 credits per started 10 seconds of video. Pass a direct video file URL (e.g. .mp4) — page links (YouTube, TikTok) don't work; use upload_file first for a file the user provides. Poll with wait_for_video. Renders a live preview in app-capable hosts.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | Label in the user's library. Defaults to the file name. | |
| videoUrl | Yes | Direct URL of the video file. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations the description carries the full burden, and it discloses key traits: credit cost (10 credits per started 10s), asynchronous polling, an input-type restriction, and live preview rendering. It omits auth/permission requirements and any failure modes, keeping it short of a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose is front-loaded, followed by cost, input constraint, and polling instruction in a compact sequence. Every sentence carries distinct operational information with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 2-param generation tool with no output schema and no annotations, the description covers cost, async workflow, input constraints, a fallback tool, and rendering behavior. Nothing an agent needs to invoke it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the 3 baseline applies, but the description adds real meaning for videoUrl beyond the schema's generic 'Direct URL of the video file': it must be a direct file URL (e.g. .mp4) and page links are invalid. The 'name' parameter is left to the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb ('remove') and a precise set of resources ('burned-in subtitles, captions, watermarks, logos or on-screen text from a video'), with an alternative alias '(Subtitle Remover)'. An agent can distinguish this from siblings like create_video or upscale_video without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives an explicit when-not ('page links (YouTube, TikTok) don't work') and an alternative tool with its condition ('use upload_file first for a file the user provides'), plus the follow-up action ('poll with wait_for_video'). The routing is unambiguous.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
retry_postRetry PostAInspect
Retry a FAILED or PARTIALLY_PUBLISHED post. Re-attempts only the accounts that failed; already-published accounts are never re-posted. Runs in the background: call get_post after a minute or two for the outcome (in app-capable hosts the panel tracks it live). In app-capable hosts the result renders as an interactive panel the user can see and act on — do NOT repeat its contents as a list or table in your reply; add only what the panel doesn't say.
| Name | Required | Description | Default |
|---|---|---|---|
| postId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and delivers: partial-retry semantics (only failed accounts), a no-duplicate guarantee (published accounts are never re-posted), asynchronous background execution, and a polling instruction. These are non-obvious behaviors an agent could not infer from the schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads eligibility first, then behavior, then outcome retrieval, then rendering constraints. Four tight sentences with no filler; every clause carries operational information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, and the description compensates by explaining that results arrive asynchronously and where to read them, plus how the host renders the panel. An agent has enough to invoke it and handle the response correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the single postId parameter is undocumented, but the description's repeated references to 'post' make the referent obvious. It adds no format, ID-type, or sourcing guidance beyond that implicit context.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Retry a ... post') and narrows scope with the exact eligible states (FAILED or PARTIALLY_PUBLISHED). This clearly distinguishes it from siblings like publish_post, create_post, and update_post, which do not imply re-attempting a failed publish.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a precise precondition (only failed/partially published posts) and tells the agent how to obtain the outcome via get_post, which is effectively the follow-up step. It stops short of naming an explicit alternative or stating what to do when the post is not in a retryable state.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
tiktok_creator_infoTikTok Posting OptionsARead-onlyIdempotentInspect
TikTok creator info for a connected account: allowed privacy levels, whether comments/duets/stitches are disabled, max video length (used by the composer UI).
| Name | Required | Description | Default |
|---|---|---|---|
| accountId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and openWorldHint=false, so the safety profile is covered. The description usefully discloses the shape of the returned data (privacy levels, disabled-feature flags, max length), but says nothing about auth requirements for the connected account or how to obtain a valid accountId.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with the key scope ('TikTok creator info for a connected account') first, followed by the enumerated contents. The trailing parenthetical is mildly informational but not wasteful.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read-only, single-parameter lookup with annotations covering safety, the description supplies the important missing piece: what the return payload contains. It is close to complete, with only the accountId provenance unaddressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 1 parameter at 0% schema coverage, the schema documents the accountId field only as an untyped-format string. The description partially compensates by framing it as 'a connected account', but adds no format, source, or lookup guidance beyond that.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: retrieves 'TikTok creator info for a connected account', and enumerates what that info contains (privacy levels, comment/duet/stitch flags, max video length). This is clearly a read/info tool, distinguishable from the creation and posting siblings, though it doesn't explicitly name an alternative.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase '(used by the composer UI)' implies this is fetched ahead of composing a TikTok post, which gives context. However, it never states when an agent should call it versus siblings like compose_post or create_post, nor any prerequisites beyond a 'connected account'.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
update_postUpdate PostAInspect
Edit a DRAFT or SCHEDULED post (published posts cannot be edited). Only the fields you pass change. Passing scheduled_at reschedules; passing scheduled_at: null converts it back to a draft. Passing accounts replaces the full target account list. In app-capable hosts the result renders as an interactive panel the user can see and act on — do NOT repeat its contents as a list or table in your reply; add only what the panel doesn't say.
| Name | Required | Description | Default |
|---|---|---|---|
| postId | Yes | ||
| caption | No | ||
| timezone | No | ||
| accountIds | No | Replaces the post's target accounts. | |
| scheduled_at | No | ISO datetime to (re)schedule, or null to unschedule. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does well: patch semantics ('only the fields you pass change'), the destructive replacement of accounts, the scheduled_at null-to-draft transition, and the interactive panel output with an explicit instruction not to duplicate it. It omits auth/permission and error behavior, but covers the mutation's risky traits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The editability constraint is front-loaded, then field semantics, then output handling, all in compact sentences with no filler. Every clause carries information an agent needs.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with no annotations and no output schema, it covers editability, patch behavior, scheduling transitions, and result rendering. Only auth requirements, error handling, and caption/timezone semantics are absent, which are minor given the tool's simplicity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 40% and caption, timezone, and postId are undocumented in both places. The description adds real meaning for scheduled_at (null converts to draft) and accounts (full replacement), partially compensating, but leaves several parameters unexplained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb (Edit) and resource (DRAFT or SCHEDULED post) and immediately states the scope boundary that published posts cannot be edited. This clearly separates it from siblings like create_post, publish_post, and delete_post.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It states the precondition for use (only DRAFT or SCHEDULED posts) and the behavior of key fields, which tells the agent when this tool applies versus publish_post. It does not explicitly name an alternative tool for the excluded case, so it falls short of a full routing statement.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
update-skillUpdate SkillAInspect
ADMIN ONLY — partially update a skill in the Unsora skills catalog by id or slug. Only the fields you pass change. Can rewrite any landing page content, replace the SKILL.md file, publish/unpublish, and attach showcase media (images/videos) by URL via addMedia. Non-admin accounts get a 403.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | ||
| slug | No | New URL slug. | |
| skill | Yes | Skill id or slug to update. | |
| steps | No | Replaces the whole how-it-works steps list. | |
| status | No | Publish or unpublish the skill. | |
| ctaBody | No | ||
| skillMd | No | Replacement SKILL.md content (stored as a new version). | |
| tagline | No | ||
| addMedia | No | Showcase media to ATTACH to the skill page (appended, nothing is removed). | |
| examples | No | Replaces the whole example prompts list. | |
| highlights | No | Replaces the whole highlights list. | |
| ctaHeadline | No | ||
| description | No | ||
| requirements | No | Replaces the whole requirements list. | |
| installCommand | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden and does disclose the key traits: admin-only auth with a 403 failure mode, the partial-update contract ('only the fields you pass change'), and the mutation surface (rewriting landing content, replacing SKILL.md, publish/unpublish, appending media). It omits reversibility, versioning behavior, and any warning that supplied fields overwrite existing content, which is meaningful for a mutation tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences with the critical 'ADMIN ONLY' constraint front-loaded and no filler. Each clause carries distinct information: eligibility, update semantics, capability surface, and failure mode.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 15-parameter mutation tool with no annotations and no output schema, the description gives an adequate behavioral picture (auth, partial update, capability scope). It does not touch most of the individual content fields, but the schema covers those at 60%, so the agent has enough to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 60%, so the schema already documents many parameters. The description adds meaning for a few (id-or-slug targeting, addMedia attaching media by URL, SKILL.md replacement, status as publish/unpublish) but leaves the many content fields (name, tagline, description, examples, highlights, requirements, installCommand) unexplained, so it does not compensate for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('partially update a skill in the Unsora skills catalog') with an unambiguous scope ('by id or slug'). The 'partially update' framing and the enumerated capabilities clearly separate it from create-only siblings like upload-skill and read-only get-skill.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Establishes a clear precondition ('ADMIN ONLY') and outcome for the wrong caller ('Non-admin accounts get a 403'), which tells an agent exactly who may invoke it. It does not name a sibling alternative (e.g., upload-skill for creation), so it stops short of full when-not guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
upload_fileUpload FileAInspect
Upload media to the user's Unsora library (available to every authenticated user). Pass a public source URL (preferred) or a base64 payload for small files (≤ ~7MB). The file is stored and recorded as an upload asset; the returned url can be used anywhere a media URL is accepted (create_post media, reference images, clipping input). In app-capable hosts the result renders as an interactive panel the user can see and act on — do NOT repeat its contents as a list or table in your reply; add only what the panel doesn't say.
| Name | Required | Description | Default |
|---|---|---|---|
| url | No | Public http(s) URL of the file to import (max 200MB). | |
| base64 | No | Base64 file contents (raw or data: URL). Use only for small files; prefer url. | |
| fileName | No | File name (required with base64, optional with url). | |
| contentType | No | MIME type override, e.g. image/png. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries full behavioral burden and does well: it discloses auth scope ('available to every authenticated user'), persistence behavior ('stored and recorded as an upload asset'), and an important host-dependent UI effect (renders as an interactive panel the user can act on). It omits rate limits, error cases, and whether re-uploads dedupe/overwrite, keeping it below the top.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads purpose, then method, then downstream use, then UI caveat, in a tight sequence with no filler. The final sentence about the interactive panel is longer but earns its place by preventing a redundant list/table response.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 4-param, no-output-schema mutation tool, the description covers storage, auth, method choice, size limits, downstream consumption, and return value ('the returned url can be used anywhere a media URL is accepted'). Only finer details like failure modes and idempotency are missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds value the schema lacks: a ~7MB practical ceiling for base64 (the schema states only a 200MB url cap) and a stated preference for url. It does not clarify contentType or fileName precedence beyond what the schema already says.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Upload media to the user's Unsora library') and clearly distinguishes itself from generation siblings like create_image/create_video by describing an import path rather than a synthesis path. An agent immediately knows this ingests existing media, not generates new media.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit preferred path ('Pass a public source URL (preferred) or a base64 payload for small files') which selects between the two input modes, and names downstream uses (create_post media, reference images, clipping input). It lacks an explicit 'when not to use' exclusion versus e.g. upscale_image or remove_watermark, so it stops short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
upload-skillUpload SkillAInspect
ADMIN ONLY — upload a skill to the Unsora skills catalog and write its full landing page content (tagline, highlights, how-it-works steps, example prompts, requirements, CTA copy, install command). Pass the raw SKILL.md content (frontmatter name/description is parsed automatically; explicit arguments override it). Created as DRAFT unless status is PUBLISHED. Non-admin accounts get a 403.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | Skill name. Defaults to the SKILL.md frontmatter name. | |
| slug | No | URL slug. Defaults to a slugified name. | |
| steps | No | How-it-works steps shown on the page. | |
| status | No | DRAFT (default) or PUBLISHED (immediately live on the landing site). | |
| ctaBody | No | Call-to-action body copy. | |
| skillMd | Yes | Raw SKILL.md content, including the --- frontmatter block with name/description. | |
| tagline | No | Short one-liner shown under the skill name. | |
| examples | No | Example prompts users can try with the skill. | |
| highlights | No | Bullet points of what the skill does. | |
| ctaHeadline | No | Call-to-action headline. | |
| description | No | Defaults to the SKILL.md frontmatter description. | |
| requirements | No | What the user needs before using the skill. | |
| installCommand | No | Command shown on the page to install the skill (e.g. a curl of SKILL.md). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the burden and does well: it discloses the admin-only auth requirement, the 403 failure mode, the DRAFT-vs-PUBLISHED default, and the frontmatter-parsing/override precedence. It stops short of describing behavior on duplicate slugs or existing skills (overwrite vs conflict).
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The critical 'ADMIN ONLY' constraint is front-loaded, and each clause carries information. The parenthetical enumeration of landing page fields is dense, but it maps to real parameters rather than filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 13-parameter mutation tool with no output schema and no annotations, the description covers auth, defaults, and frontmatter handling well. It omits the side-effect/conflict behavior for pre-existing slugs, which would round out completeness.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the structured fields already document every parameter. The description still adds non-schema value by explaining that SKILL.md frontmatter name/description is parsed automatically and explicit arguments override it, clarifying the precedence among the 13 fields.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb and resource (upload a skill to the catalog) and expands the scope to writing full landing page content, listing the concrete copy artifacts. It is clearly distinguishable from the get-skill and update-skill siblings by its 'upload' action.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It states the strong precondition 'ADMIN ONLY', warns that non-admin accounts get a 403, and clarifies the default DRAFT outcome unless status is PUBLISHED. It does not explicitly route the agent to update-skill for existing skills, so the sibling distinction is left implicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
upscale_imageUpscale ImageAInspect
Upscale an image to 2K, 4K or 8K (Image Upscaler). Costs 2 / 3 / 5 credits. Pass a public imageUrl — a generation outputUrl or an upload_file url. Poll with wait_for_image. Renders a live preview in app-capable hosts.
| Name | Required | Description | Default |
|---|---|---|---|
| imageUrl | Yes | Public URL of the image to upscale. | |
| resolution | No | Target resolution. Default 2k. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations present, the description carries the burden well: it discloses the credit cost per resolution (2/3/5), that the operation is asynchronous and must be polled via wait_for_image, and that it renders a live preview in app-capable hosts. It does not cover failure behavior or how long polling typically takes, which is the remaining gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four short sentences, each carrying unique payload (capability, cost, input sourcing, polling), with the core purpose front-loaded. No filler or repetition of the tool name.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and no annotations, the description supplies cost, async/polling behavior, input provenance, and host preview support, which is nearly everything needed to invoke it correctly. Minor omissions (failure modes, expected wait time) keep it from a 5.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds real meaning: imageUrl is constrained to public URLs sourced from generation outputUrl or upload_file, and resolution choices are tied to specific credit costs. That is value beyond the schema's terse field descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (upscale) and resource (image) with an explicit range of outputs (2K/4K/8K) and names the product surface (Image Upscaler). An agent can immediately separate it from the sibling upscale_video without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives the workflow context an agent needs: the input must be a public URL, either a generation outputUrl or an upload_file url, and results are polled with wait_for_image. No explicit when-not-to-use or alternative routing (e.g., vs. create_image) is given, so it stops short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
upscale_videoUpscale VideoAInspect
Upscale a video's resolution and sharpness (Video Upscaler). Models: standard (10 credits), ultra-1080p (6 / 11 / 16 credits for up to 5s / 10s / longer), ultra-4k (22 / 43 / 64 credits). Pass a direct video file URL (e.g. .mp4 — a generation outputUrl or an upload_file url). Poll with wait_for_video. Renders a live preview in app-capable hosts.
| Name | Required | Description | Default |
|---|---|---|---|
| model | No | Default standard. | |
| videoUrl | Yes | Direct URL of the video file to upscale. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does well: it discloses that the operation is asynchronous (poll with wait_for_video), that input must be a direct file URL rather than a page URL, that cost scales with model and duration, and that some hosts render a live preview. It omits failure behavior, typical completion time, and whether credits are consumed on failure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the purpose, then the credit/model table, then the operational steps, ending with the preview note. Every sentence earns its place, though the compressed credit notation ('6 / 11 / 16 credits for up to 5s / 10s / longer') takes a moment to decode.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No annotations and no output schema, so the description must cover the workflow — and it does, naming the polling companion tool and the cost model. The main residual gap is that it never states what the tool returns (e.g. an outputUrl) or how long the job typically takes, which an agent must know to hand off to wait_for_video.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds real meaning: videoUrl must be a direct video file URL with concrete .mp4 / outputUrl / upload_file examples, and the model enum is annotated with credit pricing at different durations. That is substantive semantics beyond the schema's 'Direct URL of the video file to upscale.'
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource — 'Upscale a video's resolution and sharpness' — with the parenthetical '(Video Upscaler)' clarifying the feature identity. It is trivially distinguishable from upscale_image by resource type and from create_video by being an enhancement rather than a generation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives concrete operating guidance: pass a direct video file URL (explicitly a generation outputUrl or upload_file url) and poll with wait_for_video. The per-model credit tables implicitly guide model selection by cost. It does not state exclusions — e.g. max input duration/resolution or when upscaling is inappropriate — so it stops short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
video_statusVideo StatusBRead-onlyIdempotentInspect
Single-shot generation status check (used by the preview UI).
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | ||
| tool | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and openWorldHint=false, covering the safety profile. The description adds the meaningful behavioral point that it is a single-shot (non-blocking, non-polling) check, but says nothing about what status values are returned, which matters for a status tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with a clarifying parenthetical and no filler. Efficient, though arguably too terse for the information an agent needs to call it correctly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, so the description should explain the return shape (status values, terminal states), and it does not. It also leaves the id/tool relationship unstated, making it inadequate for a status-check tool whose entire value is interpreting returned state.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%: 'id' is a bare string with no explanation, and the 13-value 'tool' enum is unaddressed by the description. The word 'generation' lets an agent loosely infer that id identifies a generation, but the description does not compensate for the documentation gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('generation status check') and the qualifier 'single-shot' implies it differs from polling tools like wait_for_video. The name 'video_status' is slightly misleading since the 'tool' enum spans image, music, voiceover, etc., though the description's generic 'generation' wording partially corrects this.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Single-shot' and '(used by the preview UI)' hint at when this is appropriate versus the wait_for_* siblings, but no explicit when/when-not or named alternative is given. Usage must be inferred rather than read.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
wait_for_clippingWait for ClipsARead-onlyInspect
Poll an AI clipping job until COMPLETED or FAILED. Returns clips with output URLs. Do NOT use this when the live preview panel is rendered (it polls by itself) — only in hosts without the preview, or when the user explicitly asks for output URLs.
| Name | Required | Description | Default |
|---|---|---|---|
| clippingId | Yes | ||
| maxAttempts | No | ||
| intervalSeconds | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnlyHint=true, openWorldHint=false), and the description adds genuinely useful non-structured context: this is a blocking poll loop with defined terminal states and a return payload of clips with URLs. It omits polling budget/timeout behavior, which would matter to a caller of a blocking tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, front-loaded with the core action and terminal states, then the return value, then the usage caveat. No filler or repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, so the description compensates by stating what is returned (clips with output URLs), and it covers the blocking-poll semantics and the preview-panel exclusion. The main gap is the unexplained polling parameters, which leaves the caller unable to reason about attempt/interval tuning.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across three parameters, so the burden falls entirely on the description, which explains none of them. clippingId, maxAttempts, and intervalSeconds (with their bounds) are left for the agent to interpret solely from their names.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (poll) and resource (AI clipping job) with the terminal conditions (COMPLETED or FAILED) and the return value (clips with output URLs). It does not name or contrast with the adjacent sibling clipping_status, so an agent must infer the distinction between polling-to-completion and a single status check.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit when-not ('Do NOT use this when the live preview panel is rendered') plus two when-to-use conditions (hosts without preview, or user asks for output URLs). It stops short of naming the alternative tool (clipping_status) to use otherwise, leaving that routing to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
wait_for_imageWait for ImageARead-onlyInspect
Poll an image job until COMPLETED or FAILED: create_image, create_influencer, create_thumbnail, upscale_image and create_movie_material.
| Name | Required | Description | Default |
|---|---|---|---|
| maxAttempts | No | ||
| generationId | Yes | ||
| intervalSeconds | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
readOnlyHint=true already tells the agent this is a safe read, so the description only needs to add behavior beyond that. It discloses the polling loop and terminal states, which is useful, but omits what happens on timeout, how maxAttempts/intervalSeconds bound the wait, or what is returned when the job never completes.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler; the purpose clause leads and the enumeration of applicable creator tools carries real routing information rather than padding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a blocking poll tool with 0% schema coverage and no output schema, the description covers purpose and applicability but leaves the polling controls and the return-on-timeout/return-value behavior undocumented. Adequate for identification, thin for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the schema documents none of the three parameters and the description must compensate. It never mentions generationId (required), maxAttempts, or intervalSeconds, leaving polling limits and the polling target unexplained beyond suggestive parameter names.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Poll), resource (an image job), and termination semantics (until COMPLETED or FAILED). It also enumerates the creator tools whose jobs it applies to, which separates it from the one-shot image_status sibling without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The trailing list (create_image, create_influencer, create_thumbnail, upscale_image, create_movie_material) makes clear which upstream calls produce the jobs this tool waits on, so when-to-use is well covered. It stops short of naming when *not* to use it (e.g., use image_status for a non-blocking check), so it's clear context without explicit exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
wait_for_musicWait for MusicCRead-onlyInspect
Poll music status until COMPLETED or FAILED
| Name | Required | Description | Default |
|---|---|---|---|
| maxAttempts | No | ||
| generationId | Yes | ||
| intervalSeconds | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=false, so safety is covered. The description adds genuinely useful behavior: that this blocks/polls until a terminal state. It omits what happens on timeout or exhausted attempts, so the added context is partial.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single tight sentence with the action and terminal states front-loaded and no filler. It is arguably too terse given three undocumented parameters, but there is zero waste.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, so the description should explain what a completed/failed poll returns and how long the tool waits, but it does neither. Combined with 0% parameter documentation for a 3-parameter blocking tool, an agent lacks enough to call it reliably.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across three parameters, so the description carries the full burden and provides nothing: it never explains maxAttempts, intervalSeconds, or which ID to pass. The polling semantics hint at generationId's role but no format or defaults are given.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Poll) and resource (music status) plus the two terminal conditions (COMPLETED/FAILED), which lets an agent tell it apart from create_music and audio_status. It does not explicitly name a sibling, but the resource noun is distinctive enough.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No statement of when to use this versus audio_status or the other wait_for_* tools, and no prerequisite such as needing a generationId produced by create_music. The polling-until-done intent is only inferrable from the wording.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
wait_for_videoWait for VideoBRead-onlyInspect
Poll a video job until COMPLETED or FAILED: create_video, create_motion_control, create_avatar_video, upscale_video and remove_watermark.
| Name | Required | Description | Default |
|---|---|---|---|
| maxAttempts | No | ||
| generationId | Yes | ||
| intervalSeconds | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish this is a read-only, closed-world operation. The description adds genuine behavioral context beyond them: this tool blocks and loops until a terminal state, rather than returning a single snapshot. It stops short of disclosing timing behavior, defaults, or the cost of long polls.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One front-loaded sentence: verb and behavior come first, and the generator list is a compact clause rather than padding. Slightly list-heavy, but every element is load-bearing.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and 0% parameter coverage, the description should at least say what a completed poll returns or how the wait is bounded. It gives terminal states but nothing about return payload (URL/asset id), timing bounds, or how maxAttempts/intervalSeconds shape the wait.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the three parameters (generationId, maxAttempts, intervalSeconds) are entirely undocumented in both schema and description. The word 'Poll' faintly implies interval/attempt semantics, but there is no statement of what maxAttempts or intervalSeconds mean, their units, or defaults.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Poll') and resource ('a video job') with clear terminal outcomes (COMPLETED or FAILED), and enumerates the five generator tools whose jobs it watches. It does not explicitly contrast itself with the sibling video_status, so an agent must infer the difference between blocking poll and simple status read.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The list of generator tools implies 'call this after create_video/upscale_video/etc.', which is real routing guidance. However, it never says when NOT to use it or names the non-blocking alternatives (video_status, wait_for_clipping, wait_for_image) that a sibling-aware agent would want ruled out.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
wait_for_voiceoverWait for VoiceoverBRead-onlyInspect
Poll a voiceover or change_voice job until COMPLETED or FAILED.
| Name | Required | Description | Default |
|---|---|---|---|
| maxAttempts | No | ||
| generationId | Yes | ||
| intervalSeconds | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=false, so safety is covered. The description adds the terminal states it watches for, but omits key behavioral facts: what happens when maxAttempts is exhausted, total wait bounds, and whether it throws or returns FAILED on timeout.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single tight sentence with the verb and terminal conditions front-loaded. No filler, no redundancy, and every word contributes to understanding the tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and no annotation detail beyond safety hints, the description should say what a successful wait returns (final status, job payload) and what a timeout yields, but it does neither. For a blocking poll tool with three undocumented parameters, this is materially incomplete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the description mentions no parameters at all. It does not explain that generationId targets the job or clarify how maxAttempts and intervalSeconds control polling cadence, leaving the agent to infer everything from bare schema types.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (poll) and resource (a voiceover or change_voice job) plus the terminal conditions it waits for, so the agent knows the core behavior. It does not, however, distinguish itself from the sibling audio_status or explain why a polling wait is needed instead of a single status read.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'until COMPLETED or FAILED' implies this blocks/waits, contrasting implicitly with one-shot status tools like audio_status, but there is no explicit 'use after create_voiceover' or 'do not use if you already have a terminal status' guidance. Usage is inferred rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
51 tool updates
- First observed
audio_status - First observed
change_voice - First observed
clipping_status - First observed
compose_post - First observed
create_avatar_video - First observed
create_clipping - First observed
create_image - First observed
create_image_and_schedule - First observed
create_influencer - First observed
create_motion_control - First observed
create_movie_material - First observed
create_music - First observed
create_post - First observed
create_thumbnail - First observed
create_video - First observed
create_voice_clone - First observed
create_voiceover - First observed
delete_generation - First observed
delete_post - First observed
delete_voice_clone - First observed
get_accounts - First observed
get_credits - First observed
get_post - First observed
get_post_analytics - First observed
get_price - First observed
get_subscription - First observed
get-skill - First observed
image_status - First observed
list_generations - First observed
list_models - First observed
list_posts - First observed
list_uploads - First observed
list_voiceover_voices - First observed
list_voices - First observed
pinterest_boards - First observed
publish_post - First observed
remove_watermark - First observed
retry_post - First observed
tiktok_creator_info - First observed
update_post - First observed
update-skill - First observed
upload_file - First observed
upload-skill - First observed
upscale_image - First observed
upscale_video - First observed
video_status - First observed
wait_for_clipping - First observed
wait_for_image - First observed
wait_for_music - First observed
wait_for_video - First observed
wait_for_voiceover
Related MCP Servers
- AlicenseAqualityCmaintenanceEnables brand visibility monitoring across major AI platforms like ChatGPT, Claude, Gemini, and Perplexity. It allows users to track visibility scores, analyze competitor data, and receive actionable insights to improve AI-generated brand recommendations.1624 npm1MIT
- AlicenseCqualityBmaintenanceCompetitor Monitor AI - MCP server providing AI-powered tools and automation by MEOK AI Labs1114 npmMIT
- AlicenseAqualityCmaintenanceRevnuvo Company Intelligence tells AI agents what changed at a company, with evidence. It observes company websites, technologies, and DNS over time and returns timestamped, confidence-aware changes, signals, and monitoring.9MIT

industrylens-mcpofficial
AlicenseNot gradedqualityBmaintenanceBrowse IndustryLens's published competitive-intelligence reports and head-to-head competitor comparisons from any AI agent — real, source-backed data.MIT
Glama MCP Gateway
Add one secure layer between your agents and this server.