SuperCMO Skills
OfficialIt lets an AI agent plan and run a full marketing production pipeline: generate and edit images, videos and voiceovers, analyze products and competitor ads, and assemble finished campaign assets.
Generate media: create images, video clips, and speech/voiceovers from text prompts or reference media, with model selection, batching, and dry-run previews.
Build marketing assets: product photos, UGC/ad/cartoon videos, storyboards, AI actors, scripts, and voiceovers.
Analyze content: extract structured data from URLs, analyze images/videos, and transcribe audio/video.
Research competitors: pull read-only public data from ad libraries and social platforms (Meta, Instagram, TikTok, Reddit, LinkedIn, etc.).
Post-produce: stitch clips, add narration/music/subtitles, burn captions, and overlay logos/text/end cards.
Manage setup and jobs: check configured keys and readiness, retrieve pending generations, and preview requests without spending.
Provides text-to-speech generation for voiceovers and narration, including listing available voices and audio models and generating spoken audio from scripts.
Enables researching Instagram profiles, posts, comments, and related public data for competitor and audience analysis.
Enables querying the Meta Ad Library for competitor ads and advertiser data across Meta's platforms for market and competitor analysis.
Enables subreddit and trend discovery, posts, and comments research for audience listening and market research.
Enables researching TikTok profiles, posts, comments, transcripts, and hashtag/keyword search for trend and competitor analysis.
Enables researching YouTube videos, transcripts, comments, and related public data for competitor and content analysis.
SuperCMO Skills
Enable AI agents to create marketing videos and images
Open-source skills that empower any AI agent (Claude, Cursor, Codex, Hermes, etc.) to generate end-to-end marketing campaigns - UGC videos, ad videos, product photography and more.
Just provide a product photo and a brief. The agent extracts the product details, selects the optimal AI image and video models, casts AI actors, and edits the generated content into finished, campaign-ready assets. It can produce videos of any length while maintaining perfect actor and product consistency.
The open-source alternative to closed AI marketing agents.
create an ugc video for this tshirt. it should be a get ready with me reel. snappy. very genz. [product-photo]
create a cartoon video ad in stylized 3D style, promoting my tumbler. make it fun and entertaining. [product-photo]
make me an ad video for this energy drink [product-photo]
Contents
Related MCP server: MCP Server for Google Ads Meta Ads GA4
See how it works
A full production studio inside Claude. You provide a product photo and a high level (or detailed) brief.
The AI agent then plans the video → writes the script → generates the AI actor → creates storyboards → generates the video clips → stitches them together, and delivers the final video.
The best part? You stay in charge all along.
Quick start
Paste this to your coding agent (Cursor, Codex, …):
Run `npx --yes github:SupercmoHQ/superCMO-skills --all`. Once it's installed, run the supercmo-setup skill to complete the setup.Prefer to run it yourself:
npx --yes github:SupercmoHQ/superCMO-skills --all # every detected hostRun the supercmo-setup skill or ask your agent to "set me up".
Prefer a preview before spending? Every generation supports dry_run — the exact request and cost, no API call.
Start from a product photo
Starting from your product is faster than describing a video from a blank prompt.
Hand SuperCMO a single product photo and it builds a UGC ad around it - a creator on camera reviewing, unboxing, trying on, or demoing your product:
Upload a product photo and give a short brief
It analyzes the product, picks the format, casts the creator, and writes the script. Then shows you the concept and waits for your approval before spending a cent
Once approved, it generates the storyboards, renders each clip, stitches them together, and hands back the finished video
"Make a 30-second unboxing UGC ad for this product." [attach product photo]What you get back is not a random clip. You get:
You stay in charge - approve the creator, the concept, and the script before anything renders, and change any part along the way.
Consistency built in - maintains the same actor and product across every clip, through the storyboard and anchor references
A finished, scroll-stopping video - delivered at the length you asked for, ready to post
Works inside Claude Code, Cursor, Codex, Hermes - any agent that supports the Agent Skills spec.
Why SuperCMO Skills?
You direct, it produces | SuperCMO doesn't decide what your campaign should say — you do. It handles the production work underneath: scripting to your brief, casting, shooting, editing. It checks in at every step, so you can redirect it mid-run instead of accepting whatever comes out the end. |
No tool hopping | Making one video ad today means jumping between tools: script in one, images in another, voice in a third, video in a fourth, then stitching it in an editor. SuperCMO does it all from a single brief, in one place. |
The right model, picked for you | Every AI image/video model works differently and needs different prompting. Knowing that Veo wants one kind of prompt and Kling another — and which to reach for in the first place — is most of the work. SuperCMO's skills already know. You write the brief; it handles the rest. |
Consistency across every shot | One good clip is easy. Ten clips where the product and the face stay identical is the part that breaks. SuperCMO storyboards before it renders and locks a reference into every generation, so continuity holds across the whole cut. |
Editing included | It doesn't just hand back a raw 10-second clip. It generates the media, splits long clips, trims footage, adds voiceover, and stitches everything into a finished asset. |
Pay per use, no subscription | Most marketing tools charge monthly whether you use them or not. SuperCMO doesn't — bring your own vendor keys (free), or generate on SuperCMO's keys and pay per use. No lock-in. |
Runs where you already work | Inside Claude Code, Cursor, OpenAI Codex, Hermes, Openclaw, or any agent that supports the Agent Skills spec. Your keys and files stay on your machine. |
How the skills work together
You act as the creative director, and your agent uses SuperCMO skills to orchestrate the entire production pipeline.
Here is an example of how the skills chain together behind the scenes to execute a complex brief like: "Make a 1-minute video ad for this product [URL] with a voiceover."
[Brief] "Make a 1-minute video ad for this product [URL] with a voiceover."
│
├── 1. Analyze ──> Scrapes the URL to understand the product and fetch product images.
│
├── 2. Plan ──> Breaks the 60s brief into four 15-second shots.
│ Writes the script and shot list.
│
├── 3. Image ──> Generates master reference images of the product
│ using the best image models.
│
├── 4. Video ──> Generates the 4 video clips. Locks the anchor reference
│ into every shot so the product doesn't shape-shift.
│
├── 5. Audio ──> Adds a voiceover, budgeting the script to match the
│ exact length of the video.
│
└── 6. Stitch ──> Stitches the finished clips together to create one
continuous 1-minute video file.
The Marketing Production Pipeline
SuperCMO installs a full creative agency into your text-based agent. It doesn't just generate media; it orchestrates the specific steps required to make high-converting ads:
Product Analysis - Scrapes your URLs to extract brand guidelines, materials, and features so the AI models don't hallucinate your product.
Ad Copy & Scripting - Breaks the brief into a shot-by-shot ad script tailored for platforms like Meta, TikTok, or YouTube.
Visual Generation - Produces master reference images and locks them as anchors so your product looks consistent in every video frame.
Audio & Voiceover - Generates natural, high-energy voiceovers paced perfectly to your video duration.
Post-Production - Trims footage, syncs audio, and stitches the final MP4.
These are the discrete skills installed into your agent to run that pipeline:
Skill | Description |
Produces UGC videos of any length - like review, unboxing, try-on, tutorial and more. It casts the actor, writes the script, and storyboards multiple clips to lock in actor and product consistency. It lets you direct every step to get a cohesive, ready-to-run ad in minutes. | |
Produces a polished, cinematic product commercial in your brand's voice — a no-actor product showcase or an actor-led story ad, at any length. It storyboards every clip to hold your real product identical across cuts, writes the voiceover and casts the actor when the ad calls for one, then stitches it into one ready-to-run spot you approve step by step before anything expensive renders. | |
Produces a drawn, animated video for your product — cartoon, anime, illustrated or painted, at any length. Your product either stays exactly as photographed or is drawn into the style, and a cast of characters carries the story around it. It settles one art style, writes the story, draws your cast once so every cut is the same cast, then films it clip by clip and lays a voiceover over the top — with your approval before anything expensive renders. | |
Recreates a video ad you admire with your own product. It watches the reference ad, works out the shots, pacing, camera and voiceover that make it work, then rebuilds the same structure around your product — swapping the product, brand and spoken lines for yours while holding the original's craft. You approve the plan before anything expensive renders, and it comes back as one finished clip. | |
Produces a scroll-stopping image ads for your product — one image, or a set in every placement ratio your feed needs. It keeps your real product and your brand identity perfectly consistent across every frame, and shapes the offer and claims you give it into a well designed ad, following industry best practices. | |
Directs a commercial photoshoot around your product, automatically applying the correct lighting, framing, camera angles and props to deliver studio-quality product shots across 10+ formats. | |
Analyzes what your competitors are advertising and shows you what is actually working — the ads that have run longest, the ones they quietly dropped, and what the winners have in common. A vision model watches the image and video ads that carry the signal — the long-runners, the ones just launched, the ones they dropped — the whole way through. So you learn how each one is built: the hook, the pacing, when the product lands, the proof and the offer. You also get the angles nobody in your category has taken. | |
Audits your own advertising the way you would a competitor's — pulling every ad you run from the public library and watching the creative the whole way through. You get the map of what you already cover: which angles, hooks and formats you run, which ads have held longest, what you quietly stopped and how fast, and which winners are still carrying the account with nothing newer built beside them. | |
Decides what your next ads should be. It reads your brand, your product, the ads you already run and what competitors are doing, then writes concepts ranked by what is worth testing — each a full description of one ad, the buyer it targets and the bet it tests, with the proof it rests on and the skill that would build it. Approve the plan and it builds the concepts you pick, one by one, through the specialist skills. | |
Sets your brand up from one link. Give it your website and it works out who you are — your voice, your audience, what you sell — finds who you compete with, and writes both where every future session will find them. | |
Gets you from installed to generating in a few simple steps | |
Takes an ad you already have and remakes it in every other size you need it in. It reads your ad first, works out what has to move and what fills the space the new shape opens up, and shows you that plan before it spends anything — then rebuilds each shape around the original so your product, palette and wording stay recognisably the same rather than being stretched, cropped or letterboxed into a mess. | |
Works out your brand from your website — your colours, fonts and logo, how you sound, who you sell to, what you sell, and the claims you stand behind. Anything your site doesn't actually say is written down as unconfirmed rather than invented, so the skills that plan and make your ads work from your real brand. | |
Analyzes an e-commerce link or photo to map out your product's exact physical mechanics and extract clean reference images, building a highly accurate creative brief without any manual research or extraction. | |
Casts a distinct, demographic-specific AI actor as a reusable identity reference, allowing you to reliably feature the exact same brand face across an entire multi-asset image and video campaign. | |
Converts scripts into audio. Lets you preview voice candidates from top models before generating. It then formats your script's numbers, dates, and URLs so the chosen AI model pronounces them accurately. | |
Routes your brief to the best of 10+ SOTA image models (like Nano Banana, GPT Image), writes detailed prompts as per model specifications and generates multiple variations instantly. | |
Draws out your entire video concept clip-by-clip as static images, enforcing strict character and product continuity so you can approve camera angles and visual flow before committing to expensive video renders. | |
Routes your brief to the best of 10+ SOTA video models (like Veo, Seedance), writes detailed prompts as per model specifications and generates multi-clip stories that hold continuity across cuts, stitched into one file. | |
Finds who a brand competes with: searches the web for the alternatives buyers compare it against, proposes a shortlist with a source per name, and confirms the real ones for other skills to build on. | |
Writes the text that runs alongside your ad — headline, primary text, description and call to action — tuned to the fields and best practices of each platform it runs on: Facebook and Instagram, Google Search, LinkedIn, TikTok, X and Reddit. It works one angle several ways and sizes each field so nothing gets cut off in the feed. | |
Translates your scenes into the highly specific, detailed instructions that different video models require. It times the camera motion and steers around known AI rendering limits to match your exact visual brief. | |
Drafts spoken scripts built around proven opening hooks designed specifically to stop the scroll. It sizes the text to fit your exact runtime and paces words across natural breath points for human delivery. |
Install
Every host loads the same skills + the same local MCP server - only the wiring differs.
Prerequisite: uv on PATH — the MCP
server runs via uvx, which provisions Python and the server for you (nothing to pip install). The
root pyproject.toml / package.json are packaging scaffolding — you don't build anything by hand to
run or contribute to the skills.
npx installer - registers the MCP server via each host's own mechanism (codex/claude mcp add where a CLI exists; a .cursor/mcp.json for Cursor) and places the skills. One command does every detected host:
npx --yes github:SupercmoHQ/superCMO-skills --claude # Claude Code
npx --yes github:SupercmoHQ/superCMO-skills --cursor --project-dir . # Cursor (per project)
npx --yes github:SupercmoHQ/superCMO-skills --codex # Codex
npx --yes github:SupercmoHQ/superCMO-skills --all # every detected hostIf the SuperCMO tools don't show up in your agent afterward, restart it once.
Claude Code plugin - an alternative to npx … --claude (use one, not both): the whole repo installs as one plugin, managed by Claude Code:
/plugin marketplace add SupercmoHQ/superCMO-skills
/plugin install supercmo@superCMO-skillsCodex plugin - an alternative to npx … --codex (use one, not both): installs the skills + MCP server as one plugin, managed by Codex:
codex plugin marketplace add SupercmoHQ/superCMO-skills
codex plugin add supercmo@superCMO-skillsClaude Cowork / Claude desktop - download supercmo-plugin.zip from the
latest release, then
Settings → Plugins → Upload local plugin.
Found us on skills.sh? npx skills add SupercmoHQ/superCMO-skills copies the
skills but not the MCP server they call — so image_generate/video_generate won't exist yet.
Install the server, then reload your agent:
npx --yes github:SupercmoHQ/superCMO-skills --all # or --cursor --project-dir . · --codex · --allSet up a key
Generation needs a key. Two ways — pick one:
Option A · Managed — one command
Generate on SuperCMO's keys, pay per use — no vendor signups, nothing to paste.
npx --yes github:SupercmoHQ/superCMO-skills loginOpens SuperCMO in your browser to sign in and authorize this device; the key is written to
~/.supercmo/.env automatically. Buy credits in the web app when you're ready.
Option B · Bring your own keys — free
Bring your own vendor keys — requests go directly to the model vendor; nothing routes through SuperCMO.
One file, every host. The installer creates ~/.supercmo/.env; the MCP server loads your keys from
there on any host (Claude Code, Cursor, Codex, …). Open it and add a key:
# ~/.supercmo/.env
WAVESPEED_API_KEY=your-key # image + video (start here) — https://wavespeed.ai
FAL_KEY= # image + video (alternative) — https://fal.ai
ELEVENLABS_API_KEY= # voiceover (optional) — https://elevenlabs.io
GEMINI_API_KEY= # image/video analysis (opt.) — https://aistudio.google.com
FIRECRAWL_API_KEY= # url extraction (optional) — https://firecrawl.devOne key to start:
WAVESPEED_API_KEYcovers image + video and is what we recommend — users report better reliability there.FAL_KEYis fully supported too and serves the same models, so use it if that's the account you already have. With both set, WaveSpeed is used;SUPERCMO_MEDIA_PROVIDER=falpicks the other. The rest are optional — add one when you want that capability (voiceover, analysis, extraction).Prefer environment variables? Exporting the key (or putting it in your host's MCP config
envblock) also works and takes precedence —~/.supercmo/.envis just the reliable default that also works for GUI-launched hosts that don't inherit your shell.Check what's set: ask your agent to run the
setup_statustool (host-agnostic, no path needed).
Security & trust
These skills run inside your agent. Exactly what happens:
Open source & inspectable. Every skill, script, and the MCP server lives in this repo under Apache-2.0 — read, diff, or pin before you run it.
Your keys stay local (BYOK). With your own vendor keys, they live in
~/.supercmo/.env(or your host's MCP configenv/ your shell); the server reads them from its process environment and requests go directly to the vendor — nothing routes through SuperCMO. With the managed key, requests go to SuperCMO's proxy (billed to your credits) — you never hand us a vendor key.The MCP server is local + minimal. A stdlib-only Python package (
supercmo-skillson PyPI, source inscripts/supercmo_skills/mcp/), fetched and run on demand viauvx supercmo-skills@<version>— it runs on your machine, launched by your host, and starts only when your host enables the plugin.Dry-run everything. Generation tools support
dry_run- a free preview of the exact request (keys masked), no API call.
Found something off? Open an issue.
Telemetry
SuperCMO sends anonymous, opt-out usage counts from the MCP server so we can see which tools get
used and prioritize. Full details in TELEMETRY.md.
What we send: the tool name, whether it succeeded, how long it took, versions (OS / Python / SuperCMO), and a random install id. Nothing else.
What we NEVER send: your prompts, tool arguments, generated media, file paths, keys, hostname, username, or IP address.
Turn it off (any one):
SUPERCMO_TELEMETRY=false,DO_NOT_TRACK=1, orDISABLE_TELEMETRY=1. It also honors Claude Code'sCLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC.See what would be sent: run with
SUPERCMO_TELEMETRY=log- prints each payload instead of sending.Events go to our own endpoint (
api.getsupercmo.ai), never a third-party analytics host. The random install id is never linked to any account.
Community
Questions, ideas, or something you built? Open an issue or discussion.
Contributing
New skills and improvements welcome - see CONTRIBUTING.md. In short:
Create
skills/<your-skill>/(folder name = thenamein frontmatter), usingskills/generating-images/as the reference layout.Keep
SKILL.mdshort; push detail intoreferences/, deterministic work intoscripts/(stdlib-only where possible, BYO-keys from env,--dry-runon anything that mutates).Validate - CI runs the same on every PR:
python3 scripts/quick_validate.py # structural + strict-YAML frontmatter (blocking)
python3 scripts/listing_gate.py # scripts compile + --dry-run gates (blocking)
python3 scripts/check_shared_client.py # no raw vendor HTTP - the brokering seam (blocking)
python3 scripts/check_catalog_sync.py # provider-key catalog single-sourced (blocking)License
Apache-2.0 - see LICENSE and NOTICE.
The skill files, scripts, and MCP server in this repository are Apache-2.0. The hosted SuperCMO
product (getsupercmo.ai) is a separate service governed by its own terms.
If SuperCMO saved you time, a ⭐ helps others find it.
Built by SuperCMO · Report an issue
Available Tools
22 toolsaudio_generateA
For a user's voiceover request, load the generating-audio skill BEFORE calling this — it picks the right model and voice and prepares the script for reading aloud (this tool does none of that, and calling it raw gives a flat, mispronounced read). Turn written text into spoken audio: voiceovers, narration, ad reads, character lines, or any script read aloud. Pass requests: ONE object per clip (wrap even a single clip — { requests: [ { text } ] }); add more objects (up to 10) to generate DIFFERENT lines in one call — a single approval covers the batch. Each result carries the spoken audio plus a local file path, or a structured error with a hint. This generates speech and nothing else: no sound effects, music, or ambience, no re-voicing an existing recording, and no dubbing a video. If the user asks for one of those, say so plainly rather than substituting a different tool. Every request needs a voice — the voice_id of a row from list_voices. There is no default voice. Models differ in expressiveness, language coverage, speed, price, and per-request character limit — call list_audio_models to compare them. Set dry_run=true to preview the exact requests without generating (no credits spent).Each entry in results is one of three things: finished media; a {status:"pending", ...} job handle to rejoin with job_status; or a failure carrying ok: false and an error. A failed entry is terminal — report its error and never poll or re-submit it. Read every entry rather than the top-level counters alone.
| Name | Required | Description | Default |
|---|---|---|---|
| dry_run | No | If true, return the requests that would be sent (keys masked), make no API call. | |
| requests | Yes | One object per audio clip (wrap even a single clip); add more objects to generate different lines in one call (up to 10). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description fully discloses behavioral traits: there is no default voice, raw calls produce flat/mispronounced audio, dry_run previews without spending credits, results can be finished media, pending jobs, or terminal failures, and failed entries should not be polled. It also clarifies what the tool does NOT do, such as dubbing or sound effects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long but tightly packed with actionable information. It is front-loaded with a critical prerequisite warning, and each subsequent section covers a distinct aspect (purpose, request format, result handling, exclusions, voice/model requirements). It could improve by placing the core purpose statement first, but overall no sentence is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite lacking an output schema, the description fully explains return values (audio, path, structured errors, pending job handles), error handling (terminal failures, rejoin via job_status), and external dependencies (list_voices, list_audio_models). It also covers the batch approval and dry_run semantics, making the tool self-sufficient for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description adds value by explaining the requests array structure (wrap single clip, up to 10), emphasizing the required voice field, and clarifying the dry_run behavior—information that goes beyond the schema's per-property descriptions. The added context is meaningful but not extensive enough to warrant a 5.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool converts written text to spoken audio, enumerating use cases (voiceovers, narration, ad reads, character lines), and explicitly excludes other audio tasks (sound effects, re-voicing, dubbing), which distinguishes it from sibling tools like video_generate and image_generate.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides explicit when-to-use guidance (voiceover requests) and prerequisites (load the generating-audio skill, call list_voices, call list_audio_models). It also gives clear when-not-to-use instructions, naming alternatives and telling the agent to decline non-speech requests rather than substitute another tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
calendar_addA
Schedule a new calendar event that fires once or on a recurrence. Two kinds: 'task' re-invokes the agent with prompt when it fires — write the prompt as a complete, self-contained instruction, since the agent has no memory of this call when it runs. 'post' publishes prepared content through a named tool with no model call in the loop — requires content, channel, and publish_tool. Exactly ONE of at (a one-shot ISO-8601 timestamp, which must be in the future) or rrule (an RFC 5545 recurrence rule) is required — never both, never neither. channel is free text, not a router: it is a hint the tool named in publish_tool reads to decide where to post, so word it the way that tool expects, never a fixed enum. publish_tool must name a tool the caller can actually call right now — a currently-connected tool, never invented or assumed; the event hard-fails at fire time if the named tool is not connected. media is an optional list of local file paths and/or URLs to publish alongside content. timezone is an IANA name (e.g. 'America/New_York') the schedule is interpreted in.
| Name | Required | Description | Default |
|---|---|---|---|
| at | No | One-shot fire time, ISO-8601 (e.g. '2026-09-01T14:30:00+05:30'). Must be in the future. Exactly one of `at` or `rrule` is required — never both, never neither. | |
| kind | Yes | Which kind of event this is. 'task' re-invokes the agent with `prompt` when it fires; 'post' publishes `content` via `publish_tool` with no model call in the loop. Determines which other fields are required — see their descriptions. | |
| media | No | Optional media to publish alongside `content` — a list of local file paths and/or URLs. | |
| rrule | No | Recurrence rule, RFC 5545 (e.g. 'FREQ=WEEKLY;BYDAY=MO,WE,FR;BYHOUR=9'), for an event that fires repeatedly. Exactly one of `at` or `rrule` is required — never both, never neither. | |
| title | Yes | Short human-readable label for the event, shown in calendar listings. | |
| prompt | No | REQUIRED when kind='task'. The instruction the agent is re-invoked with when the event fires — write it as a complete, self-contained instruction, since the agent has no memory of this call at fire time. | |
| channel | No | REQUIRED when kind='post'. Free text naming where to publish (e.g. 'LinkedIn company page', 'Instagram @brand') — a hint the tool named in `publish_tool` reads to decide where to post. This is NOT a routing enum: word it however `publish_tool` expects. | |
| content | No | REQUIRED when kind='post'. The exact prepared text to publish. | |
| timezone | No | IANA timezone name (e.g. 'America/New_York', 'Asia/Kolkata') that `at` or `rrule` is interpreted in. | |
| publish_tool | No | REQUIRED when kind='post'. The exact name of a tool the caller can currently call to publish (e.g. 'linkedin_post'). Must be a real, currently-connected tool — never guessed or invented — since the event hard-fails at fire time if the named tool isn't connected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full behavioral burden and does so thoroughly: it explains that task agents have no memory, post fires with no model call, `channel` is a non-enum hint, and the event hard-fails if `publish_tool` is not connected. These are meaningful execution traits beyond the schema definitions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Every sentence in the description encodes a real constraint or behavioral fact; there is no filler. The core purpose is front-loaded, and the detailed clauses about kind-specific requirements and constraints are organized naturally given the 10-parameter conditional schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex 10-parameter tool with no output schema and no annotations, the description covers parameter requirements, execution behavior, and failure conditions well. It does not describe the success return value or how to later manage the created event, but those are partially covered by sibling tools and are not essential for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Even though schema coverage is 100%, the description adds critical semantics: future-only `at` timestamps, IANA timezone names, free-text channel behavior, self-contained prompts, and the requirement that `publish_tool` be a currently-connected real tool. This goes well beyond the field descriptions in the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening sentence names a specific verb and resource: 'Schedule a new calendar event that fires once or on a recurrence.' It immediately distinguishes this creation tool from sibling tools such as calendar_update, calendar_list, and calendar_remove.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly states when each mode applies: 'task' for agent re-invocation with prompt, 'post' for publishing content, plus the exact mutually-exclusive `at`/`rrule` requirement. It does not explicitly contrast with calendar_update or calendar_remove, but the creation-focused purpose and sibling names make the intended use clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
calendar_listA
List the caller's calendar events — tasks and posts — ordered by when they next fire. Each event includes a computed next_occurrence: the ISO-8601 timestamp of its next fire time, derived from the event's at or rrule so the caller never has to compute recurrence itself. Defaults to upcoming, still-scheduled events; pass status to also see cancelled or historical (done/failed/missed) events. Narrow further with window_start/window_end (a time range), kind (task vs post), or channel (exact match on the free-text channel), and cap how many come back with limit. Use this before calendar_add to check for a clash, or before calendar_update/calendar_remove to find the event's id.
| Name | Required | Description | Default |
|---|---|---|---|
| kind | No | Filter to only this kind of event. Omit to list both. | |
| limit | No | Maximum number of events to return, ordered by next occurrence (soonest first). | |
| status | No | Filter to events in this lifecycle state. Omit to default to 'scheduled' (upcoming) events only. | |
| channel | No | Filter to events whose `channel` matches this free text exactly. Omit to include all channels. | |
| window_end | No | Only include events whose next occurrence is at or before this ISO-8601 timestamp. Omit for no upper bound. | |
| window_start | No | Only include events whose next occurrence is at or after this ISO-8601 timestamp. Omit for no lower bound (defaults to now). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral burden and does well: it explains the computed next_occurrence, defaulting to scheduled events, and how status widens visibility to cancelled/historical events. It does not describe full return fields or error behavior, but the key non-obvious behaviors are disclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the core purpose, then adds derived behavior, filtering options, and sibling usage in a logical progression. Every sentence carries useful information with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite having no output schema and no annotations, the description covers all 6 parameters, explains default behavior, names the computed field, and ties usage to sibling tools. An agent has enough to invoke this tool correctly in most workflows.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Input schema coverage is 100%, so the baseline applies. The description adds useful context around filters and the computed recurrence, but largely restates what the schema already documents for kind, status, channel, window, and limit.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource ('List the caller's calendar events — tasks and posts') and adds ordering and scope. It clearly differentiates this tool from calendar_add, calendar_update, and calendar_remove by naming those siblings directly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly says when to use this tool: before calendar_add to check for clashes and before calendar_update/calendar_remove to find an event's id. It also explains default behavior and how to expand or narrow results, leaving little ambiguity about selection.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
calendar_removeA
Cancel a calendar event by id — it stops firing but is not deleted, the same effect as calendar_update with status='cancelled'. Use calendar_list first if you don't already have the event's id.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | The event id returned by calendar_add or calendar_list, identifying which event to act on. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden, and it does meaningful work: 'it stops firing but is not deleted' tells the agent the operation is not a destructive delete, and equivalence to 'calendar_update with status='cancelled'' clarifies the effective state change. It doesn't mention authorization, idempotency, or return payload, but the core side-effect is stated honestly.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two compact sentences; the first front-loads the action and the key side-effect, the second handles the prerequisite. Every clause provides either behavior or usage information and there is no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter, no-output-schema cancellation tool, the description covers what it does, what happens to the event, the equivalent update call, and how to obtain the id. It doesn't describe result/response format, but the tool is simple enough that this is a minor gap; overall it is complete enough.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The only parameter, id, is fully documented in the schema (100% coverage), including where the id comes from ('returned by calendar_add or calendar_list'). The description repeats the by-id mechanism but adds no semantics beyond the schema, matching the baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description uses a specific verb-resource pair: 'Cancel a calendar event by id' and adds a key distinguishing fact ('same effect as calendar_update with status='cancelled''). This clearly separates it from calendar_update and calendar_list, so an agent can pick the right sibling without opening schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly instructs 'Use calendar_list first if you don't already have the event's id', which is a concrete precondition. It also references calendar_update as the semantic equivalent, though it stops short of giving a condition for choosing one over the other, so it's not a full when-not-to-use guide.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
calendar_updateA
Change fields on an existing calendar event, or re-arm/cancel it. Pass only the fields you want to change — anything omitted is left as it was. kind is fixed at creation and can't be changed here. The same rules as calendar_add apply to whatever you set: exactly one of at/rrule if you're changing the schedule, and content+channel+publish_tool together if the event is a 'post'. channel stays free text read by publish_tool, never a routing enum, and publish_tool must still name a currently-connected tool. Pass status: "cancelled" to cancel the event without deleting it, or status: "scheduled" to re-arm a cancelled one.
| Name | Required | Description | Default |
|---|---|---|---|
| at | No | One-shot fire time, ISO-8601 (e.g. '2026-09-01T14:30:00+05:30'). Must be in the future. Exactly one of `at` or `rrule` is required — never both, never neither. | |
| id | Yes | The event id returned by calendar_add or calendar_list, identifying which event to act on. | |
| media | No | Optional media to publish alongside `content` — a list of local file paths and/or URLs. | |
| rrule | No | Recurrence rule, RFC 5545 (e.g. 'FREQ=WEEKLY;BYDAY=MO,WE,FR;BYHOUR=9'), for an event that fires repeatedly. Exactly one of `at` or `rrule` is required — never both, never neither. | |
| title | No | Short human-readable label for the event, shown in calendar listings. | |
| prompt | No | REQUIRED when kind='task'. The instruction the agent is re-invoked with when the event fires — write it as a complete, self-contained instruction, since the agent has no memory of this call at fire time. | |
| status | No | Set to 'cancelled' to stop the event from firing without deleting it, or 'scheduled' to re-arm a cancelled event. | |
| channel | No | REQUIRED when kind='post'. Free text naming where to publish (e.g. 'LinkedIn company page', 'Instagram @brand') — a hint the tool named in `publish_tool` reads to decide where to post. This is NOT a routing enum: word it however `publish_tool` expects. | |
| content | No | REQUIRED when kind='post'. The exact prepared text to publish. | |
| timezone | No | IANA timezone name (e.g. 'America/New_York', 'Asia/Kolkata') that `at` or `rrule` is interpreted in. | |
| publish_tool | No | REQUIRED when kind='post'. The exact name of a tool the caller can currently call to publish (e.g. 'linkedin_post'). Must be a real, currently-connected tool — never guessed or invented — since the event hard-fails at fire time if the named tool isn't connected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden, and it excels: it discloses partial-update semantics, immutability of `kind`, conditional parameter grouping, the free-text nature of `channel`, the hard-fail risk if `publish_tool` is not connected, and the difference between cancelling and deleting. This is exactly the kind of context an agent needs to avoid destructive or invalid updates.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but every sentence earns its place, and the most important fact—this modifies an existing event—is front-loaded. The later clauses efficiently pack conditional rules, cancellation semantics, and critical constraints on `publish_tool` without repeating the schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an 11-parameter mutation tool with no output schema and no annotations, the description provides strong coverage of the update contract, including partial updates, cancellation/re-arm, and conditional rules. The only minor gap is that it does not describe what the tool returns on success or failure, but that does not materially impede an agent's ability to select and invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds cross-parameter semantics beyond the schema: it states that omitted fields are unchanged, that schedule changes require exactly one of `at`/`rrule`, that post events require `content`+`channel`+`publish_tool` together, and that `kind` cannot be changed. These relationships are not fully captured by individual parameter descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Change fields on an existing calendar event,' which clearly distinguishes this from the sibling create/list/remove tools. It also names the two additional allowed actions—re-arm and cancel—so the agent immediately understands the tool's scope.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description makes clear this is for existing events, not new ones, and refers to calendar_add for shared rules, which orients the agent toward the correct sibling. It lacks an explicit statement like 'use calendar_remove to delete an event,' so alternatives are implied rather than fully enumerated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
caption_videoA
Burn styled, social-style captions into a video from a word-timed transcript — local ffmpeg, no credits. The usual chain is transcribe -> caption_video: run transcribe on the video (or its voiceover) to get word timestamps, then pass those here. Captions are styled and positioned with a font bundled in the package (no system-font dependency); optional karaoke highlights each word as it is spoken. Timestamps are relative to the video's own audio (t=0). Returns the output file path with its duration, resolution, and size, or a structured error with a hint. Requires ffmpeg. Set dry_run=true to preview without rendering.
| Name | Required | Description | Default |
|---|---|---|---|
| style | No | Optional caption styling. | |
| video | Yes | The video to caption — a local file path (e.g. a video_generate `path`) or an http(s) video URL. | |
| output | No | Optional output file path. Omit to write a default filename into the media output directory. | |
| dry_run | No | If true, return the planned output and line count; run no ffmpeg. | |
| transcript | Yes | The words to show, in order — each an object with the word text and its timing in seconds. This is exactly the `words` list transcribe returns. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries full disclosure burden. It reveals key behaviors: local ffmpeg (no credits), bundled font (no system dependency), karaoke optional, timestamp base (t=0 relative to video audio), return format (path, duration, resolution, size, or structured error), and the dry_run option. This is rich and transparent, though it does not cover potential edge cases like file overwrite behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is several sentences but each adds essential information: purpose, usage chain, font handling, karaoke, timestamps, return format, ffmpeg requirement, and dry_run. It is front-loaded with the primary purpose and concise without fluff. A slightly better structure could group related details, but it remains efficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's moderate complexity (5 parameters, nested style object), no output schema, and no annotations, the description covers the entire workflow: prerequisite chain, input specifics, styling behavior, output details, error handling, and a preview mechanism. It is sufficiently complete for an agent to correctly invoke the tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema already has 100% coverage of parameter descriptions, so the baseline is 3. The description adds value by linking the transcript parameter to transcribe's output ('exactly the `words` list transcribe returns') and by clarifying the video input can be a video_generate path or URL. This contextual addition justifies a 4.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb+resource combination: 'Burn styled, social-style captions into a video' and clearly identifies the input as 'word-timed transcript'. It distinguishes from siblings by emphasizing the captioning function and the local ffmpeg/no-credits aspect, which is unique among the sibling tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly provides the usage chain 'transcribe -> caption_video' and explains that the transcript should come from transcribe, giving clear context for when to use this tool. It also mentions dry_run for preview without rendering, though it does not explicitly list when not to use it or suggest alternatives. The guidance is adequate but not exhaustive.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
image_analysisA
Look at one or more images (local file paths or image URLs) and answer a question about each — returns text, not new images. Use to read a product photo (category, materials, on-pack text, distinctive details), to judge whether a shot is product-only or shows a face, or to describe any image's content, layout, or text. Pass requests to read several images in ONE call — they are analyzed in parallel, so a batch costs about the same wall time as its slowest image. Give a specific 'prompt' for a focused answer; omit it for a general description. Set dry_run=true to preview the request without spending.
| Name | Required | Description | Default |
|---|---|---|---|
| image | No | The image to analyze — a local file path or an http(s) image URL. | |
| prompt | No | The question to answer about this image — e.g. 'What product is this, how is it used, how does it open, and what color/material/label details define it?' Omit for a general description. | |
| dry_run | No | If true, return the request that would be sent (key and image masked), make no API call. | |
| requests | No | Analyze several images in one call (1-10). Each entry takes its own `image` and optional `prompt`. Use this instead of one call per image whenever you have more than one to read — they run in parallel. Supply either `requests` or a single `image`, not both. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It clearly states that it returns text, does not generate images, and that batch requests run in parallel ('they are analyzed in parallel, so a batch costs about the same wall time as its slowest image'). It also explains dry_run behavior ('preview the request without spending'). It does not mention rate limits or auth, but for a read-only analysis tool the key behaviors are covered.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is thorough but not bloated. It front-loads the core purpose and then adds details about batching, prompting, and dry_run in a logical order. Each sentence contributes useful information; there is no repetition or filler. It is slightly longer than necessary, but given the complexity of multiple parameters, the length is justified.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with 4 parameters, no output schema, and no annotations, the description is quite complete. It explains when to use batching, how to use dry_run, and how to give prompts. It does not describe the return format, but since there is no output schema, that is acceptable. The only minor gap is that it doesn't mention whether the tool requires any special permissions, but that is negligible for a read-only analysis. Overall, it covers all essential aspects for correct usage.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3. The description adds value by explaining the distinction between single image vs. `requests`, the parallel execution benefit, the purpose of `dry_run`, and how to craft a focused prompt (specific prompt vs. general description). It also gives an example prompt directly in the parameter description. This goes beyond the bare schema definitions, justifying a 4.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('look at'), a specific resource ('images'), and the output type ('returns text, not new images'). It also gives concrete use cases (product photos, face detection, general content description) and explicitly differentiates from image generation by noting it does not create images. This clearly distinguishes it from siblings like image_generate and video_analysis.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides explicit use cases ('Use to read a product photo... to judge whether... to describe any image's content') and explains when to use the `requests` parameter for batching. It implicitly excludes generation by stating 'returns text, not new images,' which helps an agent choose between this and image_generate. However, it does not explicitly name alternatives or state when not to use it, leaving some inference to the agent.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
image_generateA
For a user's image request, load the generating-images skill BEFORE calling this — it picks the right model and builds the prompt (this tool does neither, and calling it raw gives weak, inconsistent results). Generate one or many still images from text prompts, optionally guided by reference images (a product photo, a character, a style or composition to follow). Pass requests: ONE object per image (wrap even a single image — { requests: [ { prompt } ] }). Generate a batch of DIFFERENT images in a SINGLE call by adding more request objects (up to 10), each with its own prompt/model/aspect_ratio/resolution/reference_images; a single approval covers the whole batch. Each result carries a hosted image URL plus a local file path, or a structured error with a hint. Use for graphics, mockups, product/marketing visuals, logos, concept art, or to render a product or character from a supplied reference. Images are polled for you; a heavy image (large model / 4k / big batch) that runs long returns {status:"pending", ...} (a job handle, not an error) — pass that exact handle to job_status to retrieve it, and never re-submit a pending image. Set dry_run=true to preview the exact requests and cost without generating (no credits spent).Each entry in results is one of three things: finished media; a {status:"pending", ...} job handle to rejoin with job_status; or a failure carrying ok: false and an error. A failed entry is terminal — report its error and never poll or re-submit it. Read every entry rather than the top-level counters alone.
| Name | Required | Description | Default |
|---|---|---|---|
| dry_run | No | If true, return the requests that would be sent (keys masked), make no API call. | |
| requests | Yes | One object per image (wrap even a single image); add more objects to batch different images in one call (up to 10). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does so thoroughly. It explains polling behavior, pending job handles, terminal errors, result structures (URL/path/error), reference image constraints, and no silent resolution downgrades—far beyond basic operation details.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long but dense with actionable details; it is front-loaded with the most critical instruction (load skill first). Some redundancy exists (e.g., discussing result entries twice), but every sentence earns its place overall, so it loses only one point.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (batching, references, dry_run, pending, errors) and no output schema, this description is remarkably complete. It covers invocation, result types, failure handling, polling, and cost preview, leaving no obvious gaps for an agent to misuse the tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Although schema coverage is 100%, the description adds substantial meaning: it explains the nested request structure (wrap single image, batch up to 10), guides model selection, elaborates on reference_images use cases, explains dry_run's behavior, and clarifies resolution/aspect ratio error handling—all beyond the schema text.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool generates 'one or many still images from text prompts,' optionally with reference images. It distinguishes itself from siblings focused on video, audio, analysis, and status by specifying image generation, and lists concrete use cases like graphics, mockups, and logos.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit guidance: load the `generating-images` skill first, use `list_image_models` for model selection, use `job_status` for pending jobs, and never re-submit pending or failed entries. It also clarifies batching behavior and when to use dry_run, giving clear when-to vs. when-not-to context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
job_statusA
Retrieve a long-running generation that was submitted earlier but hasn't finished — any result from a generation tool that came back as {status:"pending", ...} (a job handle, not media). Pass the exact pending handle object(s) in jobs; NEVER re-submit a pending job with the tool that created it — that starts (and bills) a new one. Each job comes back one of three ways: finished (a hosted URL plus a local file path); still pending ({status:"pending", ...}), in which case call job_status again with the same handle after a short wait; or failed, carrying ok: false and an error. A failed job is terminal — it will never finish, so report the error and never poll or re-submit that handle. A batch can mix all three, so read every entry in results rather than the top-level counters alone. This works for any kind of pending generation and only rejoins an existing job — it neither starts nor bills a new one.
| Name | Required | Description | Default |
|---|---|---|---|
| jobs | Yes | The pending job handle object(s) to retrieve — each exactly as returned by a prior video_generate / job_status call. Add more than one to retrieve a batch in one call. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are present, so the description carries full responsibility for behavioral disclosure. It thoroughly explains that jobs can finish, remain pending, or fail; that failure is terminal; that batches mix results; and that the tool has no billing side-effects. This is exemplary transparency.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long but every sentence earns its place. It is front-loaded with the core purpose and then systematically covers return states, batch behavior, and costing. Bold formatting for 'finished', 'pending', and 'failed' improves readability. Nothing is redundant or wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and no annotations, the description must fully inform the user about the tool's behavior, return values, and side effects. It does so comprehensively: it explains the three possible outcomes, the structure of results, error handling, polling discipline, and the fact that it neither starts nor bills a new job. This is complete for a polling tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
While the schema already covers the single 'jobs' parameter well (100% coverage), the description adds critical meaning: the handles must be the exact objects returned by a prior generation call, and passing the same handle is required to rejoin the job. It also clarifies that multiple handles can be passed for a batch. This goes beyond the schema's basic type description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: retrieving long-running generation jobs that were submitted earlier and returned a pending job handle. It distinguishes it from generation tools by explicitly saying it only rejoins existing jobs and never starts a new one. The verb 'retrieve' plus the resource 'pending generation job' is specific and unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly instructs when to use this tool (when a generation returns a pending status), and gives a critical exclusion: never re-submit a pending job with the creating tool to avoid duplicate billing. It also tells the user to wait a short time before polling again, and that failed jobs should not be re-polled. This is clear, actionable guidance with explicit alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_audio_modelsA
List the available speech models — for each, its strengths, price, per-request character limit, language coverage, and the audio types it supports — plus the output formats audio_generate accepts. Every model works with every voice, so voices are a separate concern — use list_voices for those. This is the authoritative source for what a model accepts; call it when choosing a model for an open-ended request, or to check a value before setting it. Pass an optional 'query' to filter models by use-case keyword (e.g. 'expressive', 'long-form', 'fast').
| Name | Required | Description | Default |
|---|---|---|---|
| query | No | Optional keyword to filter models by use-case (matches the name, display name, and strengths). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It reveals the return contents in detail and notes that 'Every model works with every voice,' which is a non-obvious behavioral trait. It also labels itself as the 'authoritative source' for model acceptance, setting expectations for reliability. It doesn't explicitly state read-only or error behaviors, but for a list operation, the transparency is strong.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three sentences and front-loaded with the main action and output details. The first sentence is packed with information but remains specific; the second clarifies the relationship with voices; the third gives usage guidance and parameter behavior. No filler words, though the first sentence is dense. It earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has one optional parameter and no output schema, the description thoroughly covers what is returned (strengths, price, character limit, language coverage, audio types, output formats for audio_generate) and when to use it. It also addresses the relationship with voices and the authoritative nature of the data. This is complete enough for an agent to select and invoke the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%: the schema already describes 'query' as 'Optional keyword to filter models by use-case (matches the name, display name, and strengths).' The description adds illustrative examples ('expressive', 'long-form', 'fast') and reinforces that filtering is by use-case keyword. This added value is helpful but not substantial beyond the schema, so the baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool lists available speech models and details what attributes are included (strengths, price, character limit, language coverage, audio types). It distinguishes itself from list_voices by explicitly saying voices are a separate concern, and the resource is specific to audio models, setting it apart from sibling list_video_models and list_image_models.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides explicit guidance on when to use this tool: 'call it when choosing a model for an open-ended request, or to check a value before setting it.' It also names an alternative: 'use list_voices for those' for voice-related queries. This gives clear decision rules for the agent.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_image_modelsA
List the available image-generation models (with strengths, price, the aspect ratios each accepts and how many reference images it takes), plus the valid aspect ratios and resolution tiers that image_generate accepts. Use when you need to choose a model and don't already have one in mind (e.g. an open-ended request), or to check the valid aspect_ratio / resolution values, or how many reference images a model will take, before calling image_generate — most of the time the model is the default or already specified. Pass an optional 'query' to filter models by use-case keyword (e.g. 'text', 'photorealistic', 'fast').
| Name | Required | Description | Default |
|---|---|---|---|
| query | No | Optional keyword to filter models by use-case (matches the name, display name, and strengths). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the burden of describing behavior. It details what data the tool returns (strengths, price, aspect ratios, reference image count, valid aspect ratios/resolution tiers) and the query filter behavior. It does not mention output format or potential errors, but for a read-only listing tool this is adequate context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is tightly written with essential information only: purpose, contents, usage conditions, and parameter example. Despite being dense, every clause adds value, and the structure front-loads the action.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple listing tool with one optional parameter and no output schema, the description is comprehensive: it lists the return contents, the query semantics, and the relationship to image_generate. No further information is needed for an agent to decide when to invoke it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The parameter 'query' is fully described in the schema (coverage 100%), and the description adds concrete examples ('text', 'photorealistic', 'fast') and clarifies the filtering matches use-case keywords. This adds value beyond the schema's description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool lists available image-generation models with details (strengths, price, aspect ratios, reference image count) and additionally the valid aspect ratios and resolution tiers for image_generate. It uses the specific verb 'List' and resource 'image-generation models', distinguishing it from sibling list tools for video/audio models.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly says when to use it: when you need to choose a model and don't already have one in mind, or to check valid aspect_ratio/resolution values before calling image_generate. It also states when not needed ('most of the time the model is the default or already specified'), providing clear exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_research_sourcesA
List the available research sources for social_research — every platform, its endpoints, and each endpoint's required and optional params plus per-call cost. Call this FIRST whenever you need competitor ads, profiles, posts, comments, transcripts, or platform search and don't already know the exact platform + endpoint + params. Pass an optional 'query' to filter by platform, endpoint, or keyword (e.g. 'ads', 'reddit', 'comments').
| Name | Required | Description | Default |
|---|---|---|---|
| query | No | Optional keyword to filter sources by platform, endpoint, or description. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full load. It clearly discloses the output (list of sources with endpoints and costs) and the absence of side effects is implied by its read-only nature, though not explicitly stated. It does not mention potential limitations like pagination or ordering, but for a simple listing tool this is not a significant gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is ~80 words, front-loaded with the primary purpose, then usage context, then parameter detail. Every sentence serves a purpose with zero fluff. It avoids repeating the schema description verbatim and uses structure to guide the agent from 'what' to 'when' to 'how'.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has only one optional parameter, no output schema, and no annotations. The description fully explains what the response contains (platforms, endpoints, params, costs), when to use it, and how to filter. For a discovery tool, this is complete; nothing necessary is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The single parameter 'query' is already described in the schema with its purpose. The description adds specific examples ('ads', 'reddit', 'comments') and clarifies that it filters 'by platform, endpoint, or description', which enhances the schema's generic 'platform, endpoint, or description' and clarifies the exact matching anatomy. This exceeds the baseline for covered parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description explicitly states the tool 'List[s] the available research sources for social_research', naming the exact resource and action. It enumerates what the list contains (platforms, endpoints, required/optional params, per-call cost) and clearly distinguishes it from the sibling 'social_research' tool by positioning it as the prerequisite discovery step.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit when-to-use guidance: 'Call this FIRST whenever you need competitor ads, profiles, posts, comments, transcripts, or platform search and don't already know the exact platform + endpoint + params.' It also provides examples of filter keywords ('ads', 'reddit', 'comments') and implies that social_research should only be called after consultating this list, effectively stating when not to use it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_video_modelsA
List the available video-generation models with, for each, its full schema: modes (text / image / first-last-frame / reference), the aspect ratios, durations and resolutions it accepts, which media it takes (start/end frame and reference image/video/audio with max counts), whether it has native audio, plus strengths and price. This is the authoritative source for a model's exact ranges — call it when choosing a model for an open-ended request, or to check what a model accepts before setting aspect_ratio / duration / resolution / media. Pass an optional 'query' to filter by use-case keyword (e.g. 'cinematic', 'fast', 'audio').
| Name | Required | Description | Default |
|---|---|---|---|
| query | No | Optional keyword to filter models by use-case (matches the name, display name, and strengths). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It discloses the behavior in detail: what the tool returns (full schema, modes, aspect ratios, etc.) and its authoritative nature. It does not explicitly mention side effects, but 'List' implies a read-only operation. The description adds useful context about the tool's role as the source of truth for model capabilities.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-structured and front-loaded: it opens with the primary action and returns in detail, then provides usage context, then explains the parameter. Every sentence carries meaningful information without redundancy or fluff. It is compact despite the rich detail.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple tool with one optional parameter and no output schema, the description is complete. It tells the agent exactly what the tool provides (full details for each model), when to use it (for model selection and parameter validation), and how to filter. No critical information is missing, and it aligns with the tool's simple interface.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema already covers the single 'query' parameter with 100% coverage, so baseline is 3. The description adds examples ('cinematic', 'fast', 'audio') and clarifies that the filter matches name, display name, and strengths, which slightly exceeds the schema's description. It provides practical guidance on how to use the filter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose with a specific verb ('List') and resource ('available video-generation models'), and enumerates the detailed information returned (schema, modes, aspect ratios, durations, resolutions, media, native audio, strengths, price). It is clearly distinguished from sibling tools like list_image_models and list_audio_models by focusing on video models.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly says when to use this tool: 'call it when choosing a model for an open-ended request, or to check what a model accepts before setting aspect_ratio / duration / resolution / media.' It also positions it as the 'authoritative source' for exact ranges. However, it does not explicitly mention when not to use it or point to alternative tools, so it lacks explicit exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_voicesA
Find a voice to speak with, and get the voice_id that audio_generate requires. Returns the voices saved in the active ElevenLabs account — the user's own on their key, or the shared SuperCMO set on a managed key — each with its gender, accent, age, use-case and a preview_url you can hand the user so they hear it before committing. Filter by what the brief actually demands (a stated gender or accent is not negotiable) and keep limit small: offer a few candidates with their previews rather than a long list. A voice missing the attribute you filtered on is kept rather than dropped, because a voice the user cloned themselves often carries no labels at all. If the account holds no voices the result says so — a newly created ElevenLabs account starts empty, and voices must be added in the ElevenLabs dashboard before anything can be spoken.
| Name | Required | Description | Default |
|---|---|---|---|
| age | No | Filter by apparent age of the voice. | |
| limit | No | How many voices to return. Keep it small — a handful of good candidates beats a catalogue. | |
| accent | No | Filter by accent as the provider labels it (e.g. 'american', 'british', 'indian', 'australian'). Free text, since the set grows. | |
| gender | No | Filter by voice gender. Apply whenever the user stated one. | |
| search | No | Free-text match over name, description and labels (e.g. 'warm', 'storyteller'). Passed to the provider. | |
| language | No | Filter by primary language as a short code (e.g. 'en', 'hi', 'es'). | |
| use_case | No | Filter by what the voice is built for (e.g. 'advertisement', 'conversational', 'narrative_story', 'social_media', 'informative_educational'). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does so thoroughly. It discloses account-specific behavior (own vs. shared voices), the inclusion of preview_url, the unusual behavior of keeping voices that lack filtered attributes, and the empty-account case. This goes well beyond basic expectations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is longer than typical but every sentence adds value: purpose, return contents, filtering advice, behavioral quirk, and account edge case. It is front-loaded with the core purpose and remains structured, though it could be tightened slightly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no output schema, the description covers return values (voices with gender, accent, age, use-case, preview_url), account context, filtering behavior, and empty results. It is complete enough for an agent to select and invoke the tool correctly without further clarification.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, providing baseline 3. The description adds meaningful context beyond the schema: it explains that limit should be small, that missing filtered attributes are tolerated, and that preview_url can be handed to the user. This enriches parameter understanding without restating schema fields.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific action and resource: 'Find a voice to speak with, and get the voice_id that audio_generate requires.' It clearly identifies the tool's role in the audio generation workflow and distinguishes it from sibling tools like list_audio_models by focusing on voices rather than models.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear guidance on when to use the tool (when a voice_id is needed for audio_generate), how to filter based on brief requirements, and advises keeping limit small and offering previews. It does not explicitly mention alternatives or exclusions, but the context is strong enough to infer appropriate usage.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
setup_statusA
Check which SuperCMO media-generation keys are configured and which capabilities (image / video / audio) are ready — the setup doctor. Call this FIRST when a user is setting up SuperCMO, asks which keys they need, or a generation failed with 'no_provider_configured'. Returns each vendor key (set/missing, what it enables, where to get it), managed-key state, and per-capability readiness. Set check=true for a FREE key-validity probe where one exists (never a paid generation). Reports only key NAMES and set/missing — never key values.
| Name | Required | Description | Default |
|---|---|---|---|
| check | No | If true, also run a FREE key-validity probe where a provider implements one (no paid call). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries full burden. It discloses that the tool 'Reports only key NAMES and set/missing — never key values', a key security behavior, and that the optional check is 'FREE... never a paid generation'. It also explains the return structure (vendor keys, managed-key state, per-capability readiness), though it does not explicitly state read-only or idempotency.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is moderately sized but every sentence contributes value: purpose, usage, return content, check behavior, and privacy guarantee. It is front-loaded with the core purpose and avoids fluff, though the check behavior is slightly redundant with the schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given its simple parameter set (1 boolean) and no annotations, the description is highly complete: it covers purpose, trigger conditions, return details, cost implications, and data privacy. It lacks specifics on error handling or exact vendor key names, but for a setup status tool this is sufficient.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The single parameter 'check' already has a detailed description in the input schema (100% coverage), so the description adds little beyond repetition. The phrase 'never a paid generation' reinforces the schema's 'no paid call' but does not introduce new semantic meaning.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool checks SuperCMO media-generation key configuration and capability readiness with the specific verb 'check' and resource 'SuperCMO media-generation keys'. It is labeled the 'setup doctor' and explicitly distinguishes its diagnostic role from sibling generation/list tools by instructing to 'Call this FIRST' for setup or configuration issues.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides explicit scenarios for use: 'when a user is setting up SuperCMO, asks which keys they need, or a generation failed with no_provider_configured'. It does not directly name alternatives or exclusion criteria, but the context and 'FIRST' positioning make the intended usage clear relative to siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
transcribeA
Transcribe speech from an audio or video file into text with word-level timestamps. Use it to caption a video (chain transcribe -> caption_video), to read a voiceover back, or to analyse a competitor ad's spoken script. audio is a local file path or an http(s) URL (audio or video). Returns {ok, text, words:[{word, start, end}], duration, language}, or a structured error. Set dry_run=true to preview the request without spending.
| Name | Required | Description | Default |
|---|---|---|---|
| audio | Yes | The audio or video to transcribe — a local file path or an http(s) URL. | |
| dry_run | No | If true, preview the request (key masked); make no API call. | |
| language | No | Optional ISO language-code hint (e.g. 'en'); omit to auto-detect. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations are absent, so the description carries the burden. It discloses the exact return format ({ok, text, words, duration, language}), input types (local path or URL), error shape, and dry_run cost-saving behavior. It could mention auth, file size limits, or processing time, but core behavior is transparent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four dense sentences, front-loaded with the core capability. Every sentence serves a purpose: use cases, input contract, return contract, and cost-saving option. No filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no output schema, it still explains return values thoroughly and covers common invocation concerns. It is complete enough for selection and basic use; missing details like format limitations or authentication are not critical for an agent deciding to invoke it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline 3 applies. The description reinforces the audio param type and dry_run purpose, but adds no meaningful parameter-level detail beyond what the schema already provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Clearly states the verb 'Transcribe speech' and the resource ('audio or video file') and differentiates itself with word-level timestamps. The mention of chaining into caption_video distinguishes this tool from sibling tools like video_analysis and audio_generate.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides concrete use cases: captioning a video, reading a voiceover, analyzing a competitor ad's spoken script, and explicitly suggests the transcribe -> caption_video chain. It lacks explicit when-not-to-use guidance or named alternatives, but the context is strong enough for an agent to decide.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
url_extractionA
Extract structured data from a web page — a product listing (Amazon, Shopify, AliExpress, any store) or any URL — guided by a prompt and/or a JSON schema. Returns the requested fields (e.g. name, brand, price, description, specs) and any gallery image URLs as a compact JSON object, plus page metadata — not the page's full text. Use when you need specific data or image URLs from a page. Set dry_run=true to preview the exact request without spending.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | The page URL to extract from (an http(s) URL). | |
| prompt | No | What to extract, in plain language — e.g. 'product name, brand, price, currency, variant, full description, feature bullets, specs, and all product-gallery image URLs (front/side/back/close-up/packaging); exclude review photos, related products, banners, logos'. | |
| schema | No | Optional JSON Schema describing the exact shape to return. Use for a strict, typed result; omit to let the prompt guide the extraction. | |
| dry_run | No | If true, return the request that would be sent (key masked), make no API call. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Without annotations, the description carries the full burden. It discloses the return type (compact JSON object, not full text), mentions page metadata and gallery image URLs, and explains the dry_run parameter. It also hints at cost via 'without spending.' It could be more explicit about errors or rate limits, but overall it offers strong transparency.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is only three sentences, front-loaded with the core action, and contains no redundant information. Every word contributes to understanding the tool's purpose and usage.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no output schema, the description does an excellent job covering purpose, output details, usage context, and the dry_run option. It gives a complete picture of what to expect from the tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3. The description adds value by explaining that prompt and/or schema guide extraction, and that dry_run previews the request without spending. This goes beyond the schema's basic field descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states that the tool extracts structured data from a web page, with a specific focus on product listings and any URL. This distinguishes it from sibling tools that handle media generation or analysis.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context: 'Use when you need specific data or image URLs from a page.' It does not explicitly mention alternatives or exclusions, but the intended use case is well defined.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
video_analysisA
Watch one or more videos (local file paths or video URLs) and answer a question about each — returns text, not new video. Use to read a clip before generating or matching it, to describe what happens in it, or to transcribe what is said. Pass requests to watch several videos in ONE call — they are analyzed in parallel, so a batch costs about the same wall time as its slowest clip. Analyzes a clip inline, so a very large file may be rejected — trim or link a shorter clip if so. Set dry_run=true to preview the request without spending.
| Name | Required | Description | Default |
|---|---|---|---|
| video | No | The video to analyze — a local file path or an http(s) video URL. | |
| prompt | No | The question to answer about this video — e.g. 'Describe the shots, the camera moves, the pacing, and transcribe what is said.' Omit for a general description. | |
| dry_run | No | If true, return the request that would be sent (key and video masked), make no API call. | |
| requests | No | Analyze several videos in one call (1-10). Each entry takes its own `video` and optional `prompt`. Use this instead of one call per video whenever you have more than one to watch — they run in parallel. Supply either `requests` or a single `video`, not both. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries full burden. It discloses the key behavioral traits: output is text (not video), parallel execution for batches, potential rejection of very large files due to inline analysis, and dry_run for previewing without spending. It also implicitly notes cost by referencing 'spending' in the dry_run context. No contradictions and substantial transparency.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the core purpose, then delivers use cases, batch advice, a caveat, and dry_run in a logical flow. Every sentence carries necessary information without redundancy. It is packed but well-organized, striking a good balance between detail and brevity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with no output schema and no annotations, the description adequately covers what an agent needs to call it: it names the input types, explains the output type ('text'), covers the batch scenario, flags the large-file risk, and explains dry_run. Missing details like error handling or exact response format are not critical given the simplicity of the output promise.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so all parameters are already documented. The description adds practical meaning beyond the schema: for 'requests' it explains parallel run and the mutual exclusion with 'video', for 'prompt' it gives an example and the default behavior if omitted, and for 'dry_run' it clarifies what the preview returns. This adds genuine value, though the baseline of 3 is raised slightly.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Watch') and resource ('one or more videos') and clearly defines its output ('returns text'). It distinguishes from siblings like video_generate by explicitly saying 'not new video', and from image_analysis by the video domain. It also mentions transcription, which overlaps with the transcribe sibling but clarifies it happens as part of video analysis.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides explicit use cases: 'Use to read a clip before generating or matching it, to describe what happens in it, or to transcribe what is said.' It also gives guidance on batching via the 'requests' parameter and explains when to prefer it. It mentions inline analysis and the large-file caveat, and describes dry_run for previewing. Clear situational guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
video_generateA
For a user's video request, load the generating-videos skill BEFORE calling this — it picks the right model and builds the motion prompt (this tool does neither, and calling it raw gives weak, generic clips). Generate one or many short video clips from text prompts, optionally guided by a start (and end) frame or by reference images, videos, or audio. Pass requests: ONE object per clip (wrap even a single clip — { requests: [ { prompt } ] }); add more objects (up to 10) to batch DIFFERENT clips in one call, and repeat an object for variations of one prompt — a single approval covers the batch. Models differ in the aspect ratios, durations, resolutions, and media they accept — call list_video_models to check. Video generation is long-running: each clip is submitted and polled for you. A clip that finishes in time returns a hosted video URL plus a local file path; a clip still generating returns {status:"pending", ...} (a job handle, NOT an error) — pass that exact handle to job_status to retrieve it, and never re-submit a pending clip. Set dry_run=true to preview the exact requests without generating (no credits spent).Each entry in results is one of three things: finished media; a {status:"pending", ...} job handle to rejoin with job_status; or a failure carrying ok: false and an error. A failed entry is terminal — report its error and never poll or re-submit it. Read every entry rather than the top-level counters alone.
| Name | Required | Description | Default |
|---|---|---|---|
| dry_run | No | If true, return the requests that would be sent (keys masked), make no API call. | |
| requests | Yes | One object per clip (wrap even a single clip); add more objects to batch different clips in one call (up to 10). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries full behavioral burden. It discloses long-running async behavior with polling, that a pending result is a job handle not an error, that failures are terminal with ok:false and error, that dry_run spends no credits, that durations snap to valid values, and that start/end frames require each other and cannot be combined with reference_* inputs. This is comprehensive behavioral disclosure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long but dense with essential information—prerequisite skill, batching, model differences, async behavior, result types, and failure handling. Every sentence earns its place, though the run-on structure makes it less scannable than it could be. It is appropriately front-loaded with the most critical warning about the skill prerequisite.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is complex (batching, multiple reference types, long-running jobs, model variations) and has no output schema, yet the description covers all key aspects: purpose, prerequisites, model checking, job handles, failure terminality, dry_run, result entry types, and edge cases like snappng and incompatible inputs. This leaves no significant gap for an agent to use the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Even though schema description coverage is 100%, the description adds crucial semantics not in the schema: wrapping single clips in an array, batching up to 10 different clips, repeating an object for variations, a single approval covering the batch, and dry_run previewing requests without generation. It clarifies how requests should be structured and adds the pending-result meaning, which the schema cannot convey.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states 'Generate one or many short video clips from text prompts' with optional frame/reference guidance, giving a specific verb and resource. It also distinguishes itself from siblings by warning that it does not pick the model or build the motion prompt (that's the generating-videos skill), and by pointing to list_video_models for model selection, seting it apart from video_stitch, job_status, and image_generate.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly instructs to 'load the generating-videos skill BEFORE calling this' because the tool itself 'does neither' model selection nor prompt building, and calling raw gives weak clips. It also directs users to call list_video_models to check model-specific capabilities, and to use job_status for pending jobs, with a warning never to re-submit pending clips or poll failed entries. This gives clear when-to-use and when-not-to-use guidance with named alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
video_overlayA
Stamp a logo, timed text, and a branded end card onto a video — local ffmpeg, no credits. Overlay a logo watermark at a chosen corner, drop in timed text (CTAs, offers, captions you place yourself), and/or append an end-card image as a short closing still. Pass at least one of logo / texts / end_card. Text is rendered with a bundled font (no system-font dependency). Returns the output file path with its duration and resolution, or a structured error. Requires ffmpeg. Set dry_run=true to preview.
| Name | Required | Description | Default |
|---|---|---|---|
| logo | No | Optional logo image (PNG with transparency recommended) — path or URL. | |
| texts | No | Timed text overlays. | |
| video | Yes | The video to decorate — a local file path or an http(s) video URL. | |
| output | No | Optional output file path. Omit to write a default filename into the media output directory. | |
| dry_run | No | If true, return the plan; run no ffmpeg. | |
| end_card | No | Optional end-card image (path or URL) appended as a closing still. | |
| logo_scale | No | Logo width as a fraction of the video width (default 0.15). | |
| logo_position | No | Where the logo sits (default bottom-right). | bottom-right |
| end_card_duration | No | How long the end card holds, in seconds (default 3). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full transparency burden and does well: it discloses local processing, dependency on ffmpeg, bundled-font rendering, output contents (path, duration, resolution), structured errors, and dry-run behavior. It does not detail potential side effects like overwriting files or network timeouts for URLs, but it covers the main behavioral traits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is tight and front-loaded, with the core purpose in the first sentence and supporting constraints in the following sentences. Every sentence earns its place, and there is no filler or repetition of schema content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 9-parameter tool with no output schema, the description adequately covers the main workflow, constraints, prerequisites, and return values. It could be slightly richer on how multiple overlays combine or what happens with URL inputs, but the visible guidance is sufficient for most agent decisions.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3; the description adds meaningful value by grouping parameters into three feature families (logo, timed text, end card), stating the 'at least one of' constraint, and explaining text use cases. It does not redundantly restate schema details.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific action ('Stamp a logo, timed text, and a branded end card onto a video') and clearly identifies the resource and scope. It also differentiates itself from siblings by emphasizing 'local ffmpeg, no credits' and its overlay-specific capabilities versus stitching or captioning tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives concrete usage context: pass at least one of logo/texts/end_card, requires ffmpeg, and dry_run previews without running ffmpeg. It does not explicitly name alternatives or state when not to use the tool, but the 'captions you place yourself' phrase hints at a contrast with automated captioning.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
video_stitchA
Join finished video clips into one file, in the order given, with a hard cut between each and each clip's audio kept — this assembles existing clips, it does not generate new video. Use it to build a video longer than a single model clip: generate the shots with video_generate, then stitch them. Do NOT use it for a single clip, or for a batch of clips meant to stay separate. Three optional layers, each its own parameter: lay a voiceover over the picture (pass narration — ONE take per clip, in clip order, NOT one joined track; each take is aligned to its own clip so nothing drifts), lay a background-music track under the whole thing (pass music), or burn in subtitles from an SRT file (pass subtitles); clips of different sizes are scaled to a common frame. Returns the output file path with its duration, resolution, and size, or a structured error with a hint. Requires ffmpeg on the system. Set dry_run=true to preview the plan without running anything.
| Name | Required | Description | Default |
|---|---|---|---|
| clips | Yes | The clips to join, in play order — local file paths (e.g. the `path` a video_generate result returns) or direct http(s) video URLs. At least two. | |
| music | No | Optional audio file (a local path or a URL) laid under the whole video as background music, mixed below the clips' own audio. | |
| output | No | Optional output file path. Omit to write a default filename into the media output directory. | |
| dry_run | No | If true, return the planned output path and inputs; run no ffmpeg. | |
| narration | No | Optional voiceover — ONE audio take per clip, in the same order and the same count as `clips`. Each take is padded with silence to its clip's length, so take N is heard over clip N with no timecodes to keep in sync. The narration sits at full level and the clips' own audio is ducked beneath it. A take longer than the clip it belongs to is an error naming that clip, never a truncation — shorten the line and re-voice that one take. | |
| subtitles | No | Optional SRT subtitle file (a local path or a URL) burned into the video. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It discloses key behaviors: hard cuts, audio kept, scaling of different-sized clips, narration alignment per clip, music mixing, subtitle burning, dry_run behavior, and ffmpeg requirement. It also mentions error handling for narration takes longer than clips. The only minor gap is not explicitly stating that the operation is non-destructive to input files, but the description is quite thorough.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but well-structured. It front-loads the core purpose, then provides usage guidance, then details optional parameters, then return value and requirements. Every sentence adds value. It's slightly long but justified given the complexity of the narration parameter. The structure is logical and scannable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (6 parameters, optional layers, alignment semantics, error conditions), the description is remarkably complete. It covers all parameters, explains the return value, mentions the ffmpeg dependency, and provides dry_run behavior. No output schema exists, so the description's mention of return fields ('path' with duration, resolution, size) is essential and provided. This is a model description.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description adds significant value beyond the schema: it explains the narration alignment model ('ONE take per clip, in clip order, NOT one joined track; each take is aligned to its own clip so nothing drifts'), the ducking behavior, and the error condition for long takes. It also clarifies the clips parameter accepts paths or URLs. This goes well beyond the schema's basic descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: 'Join finished video clips into one file, in the order given, with a hard cut between each and each clip's audio kept'. It specifies the verb (join), resource (video clips), and key behaviors (order, hard cut, audio kept). It also distinguishes from video_generate by explicitly stating 'this assembles existing clips, it does not generate new video'.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides explicit when-to-use guidance: 'Use it to build a video longer than a single model clip: generate the shots with video_generate, then stitch them.' It also gives clear exclusions: 'Do NOT use it for a single clip, or for a batch of clips meant to stay separate.' This is exemplary usage guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
TDQS
Each tool targets a clearly distinct concern — generation (image/video/audio), model discovery, analysis, research, transcription, and job handling — with descriptions that explicitly separate overlapping cases (e.g., video_stitch vs video_overlay vs caption_video; raw generation vs list_* discovery). No two tools appear redundant, and each description names what it does not do to prevent misselection.
The server follows a consistent snake_case, media-first convention (image_*, video_*, audio_*, list_*) that makes related tools instantly groupable, and the list_* prefix is applied uniformly. Minor deviations: caption_video inverts the media-first order (vs video_stitch/video_overlay), transcribe stands as a bare verb, and noun-style names like url_extraction and image_analysis mix with verb-style ones, though all remain predictable and readable.
18 tools exceeds the ideal 3–15 range, but each maps to a defensible function in the media-generation and marketing-analysis workflow, organized into coherent families (generate/models/analyze/post-process). It's a wide surface, yet every tool has a clear place and earns its presence; none feel like filler or duplication.
The domain is well covered across the full creative pipeline: setup, generation for all three media types, model/voice discovery, analysis, transcription, research, and comprehensive video post-production. Minor gaps exist — no image post-processing/editing, no asset management or cleanup, and no configuration beyond a status check — but the core workflows are fully traversable without dead ends.
Maintenance
Related MCP Connectors
Deterministic visual marketing engine. Your agent plans, renders, and posts on-brand campaigns.
Create, manage, schedule, and publish short-form user-generated content through AI agents.
On-brand ad creative generation: teach it your brand once, generate images and video forever.
On-brand creative studio for AI agents: images, video, audio, and 3D.
Related MCP Servers
- AlicenseNot gradedqualityAmaintenanceOpen-source AI marketing agent toolkit with plan-first workflow, human checkpoints, and BYOK support.MIT
- FlicenseNot gradedqualityCmaintenanceAutomates marketing workflows across Google Ads, Meta Ads, and GA4, enabling content generation, scheduling, and analytics via a multi-agent pipeline.
- AlicenseNot gradedqualityCmaintenanceEnables AI clients to compose and execute marketing campaigns by mapping free-form intent to a deterministic plan of over 1,000 production-tested skills, supporting multiple workflows across paid ads, SEO, content, and more.MIT
- AlicenseNot gradedqualityBmaintenanceEnables AI agents to produce video commercials from product descriptions via a multi-stage pipeline with human-in-the-loop approval gates and explicit spend authorization.4MIT
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/SupercmoHQ/superCMO-skills'
If you have feedback or need assistance with the MCP directory API, please join our Discord server
social_researchA
Pull read-only structured public data from social platforms and ad libraries — competitor ads (Meta/Facebook + Instagram, LinkedIn), profiles, posts, comments, transcripts, hashtag/keyword search, and subreddit / trend discovery. Two steps: call list_research_sources FIRST to see the platforms, their endpoints, and each endpoint's params; then call this with
platform,endpoint, and aparamsobject built from that endpoint's required/optional params. Returns the source's structured JSON indata, or a structured error naming the missing or unknown params. Known endpoints are projected to their readable fields and one media URL per item, with page-level facts carried once inadvertisersrather than repeated on every row, andshapingsays what was dropped;fieldswidens or narrows that. Every response is also written to a file:savedcarries itspath, the run'soutput_dirfor anything built from it, the count and the cursor, so a response can be handed straight to a script without being copied out of the conversation. Where the response is too large to read,saved.inlineis false anddatais omitted — use the file. Use for competitor and market research, audience listening, and trend discovery — this is read-only public data, not posting and not private data. Set dry_run=true to preview the exact request without spending.TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full behavioral burden and does so exceptionally. It discloses read-only nature, file persistence (every response written to a file with path, output_dir, count, cursor), large-response handling (saved.inline false omits data), projection/truncation behavior (advertisers carries page-level facts once, shaping says what was dropped), error structure ('a structured error naming the missing or unknown params'), and the no-spend dry_run behavior. This is rich, accurate behavioral disclosure beyond any structured field.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long but dense, and every sentence earns its place for a tool of this complexity. It is well front-loaded: purpose first, then the two-step workflow, then response/error behavior, file persistence, and finally use cases and dry_run. It is slightly on the verbose side — the precise character counts and the detailed saved-field accounting could arguably be trimmed — but the structure is logical and the ordering is correct, making it a 4 rather than a 5.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description must carry the return-value explanation — and it does, thoroughly: data (JSON), advertisers (page-level facts carried once), shaping (what was dropped), saved (path, output_dir, count, cursor), and saved.inline semantics. Combined with nested-object parameters, zero annotations, and a companion discovery tool, all the gaps an agent would face are covered. Nothing essential for correct invocation is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, setting a baseline of 3. The description adds meaningful value beyond the schema: concrete example values for platform ('meta_ad_library', 'instagram', 'tiktok', 'reddit', 'x', 'linkedin') and endpoint ('company_ads', 'profile', 'posts', 'search'), the size warning for fields="*" (~185,000 characters, a third signed CDN query strings), and the workflow context that params objects are built from list_research_sources output. It doesn't reach 5 only because per-endpoint parameter specifics are intentionally deferred to the companion tool.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Pull read-only structured public data from social platforms and ad libraries', then enumerates the concrete scope (competitor ads, profiles, posts, comments, transcripts, hashtag/keyword search, subreddit/trend discovery). It clearly distinguishes itself from siblings by contrasting against the generation tools (image_generate, video_generate, audio_generate, transcribe) and explicitly routing the discovery task to list_research_sources. An agent can tell exactly what this tool does and what it does not do.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The usage guidance is explicit and actionable: 'call list_research_sources FIRST to see the platforms, their endpoints, and each endpoint's params; then call this with platform, endpoint, and a params object'. It further states when to use it ('Use for competitor and market research, audience listening, and trend discovery') and what it is not for ('not posting and not private data'), plus the dry_run preview option. The companion tool is named and sequenced, leaving nothing to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.