kie-mcp
kie-mcp is an MCP server that gives Claude (or any MCP client) cost-aware access to 60+ image models, 85+ video models, and 20+ audio tools on the kie.ai API, with built-in model intelligence, vertical playbooks, and file-management utilities.
Image generation (60+ models) —
generate_imageacross OpenAI (GPT Image 2.5), Google (Nano Banana 2/Pro, Imagen 4), Flux Kontext/2, Seedream 4.5/5 Lite, Ideogram v3, Qwen, Grok Imagine, Recraft, Topaz; supports text-to-image, image-to-image, transparent backgrounds, text/logo rendering, and upscaling/background removal.Video generation (85+ models) —
generate_videowith Veo 3.1, Seedance 2.x, Kling 3.0, Wan 3.0, Hailuo H3, PixVerse V6, Grok Imagine 1.5, Runway; text-to-video, image-to-video, first/last-frame, multi-shot scripting, native audio, and 4K.Avatar & lip-sync — OmniHuman 1.5, Volcengine Video Lip-Sync, Kling AI Avatar, Infinitalk for talking heads, audio-driven avatars, and re-dubbing existing footage.
Music & sound (Suno) —
generate_music,generate_sfx,generate_sounds(loop/BPM/key), plus extend, cover, add instrumental/vocals, replace section, mashup, personas, style boost, lyrics, MIDI, stems/vocal separation, WAV conversion, music videos, and cover art.Speech (ElevenLabs + Gemini) —
generate_tts,generate_dialogue,generate_gemini_tts(30 voices, 2-speaker dialogue, inline tone tags, style direction), speech-to-text with diarization, and audio isolation.Targeted editing & decomposition —
grok_segment_map(free named-region segmentation) +grok_image_editfor surgical region or whole-image edits, andseedream_layer_decomposeto split any image into independent layers.Video post-processing —
veo_extend,veo_upscale_1080p,veo_upscale_4k, andrunway_extend.Vertical playbooks —
profile_briefreturns professional intake questions, per-deliverable model routing with live costs, prompt formulas, and multi-tool workflows across 10 verticals (architecture, game assets, advertising, product photo, film, brand, web product, editorial, social video, audio branding).Discovery & cost control —
list_modelswith smart filtering (e.g. "lip sync", "cheapest video"), per-model pricing in credits and USD, embeddedresearchverdicts, andcheck_credits.Async & task management —
wait: falseasync mode,check_task,list_tasks,download_result, andlist_raw_assetsto track and retrieve outputs.File handling —
upload_file(URL, base64, or local path with magic-byte-correct naming) to make local media publicly reachable for downstream generations, plus per-calldownload_dircontrol and configurable concurrency/polling budgets.
Provides access to ByteDance's image and video generation models (e.g., Seedream, Seedance) through the kie.ai API.
Provides access to ElevenLabs' text-to-speech, audio isolation, and speech-to-text services through the kie.ai API.
Provides access to Google's image generation models (Nano Banana, Imagen) and video generation models (Veo) through the kie.ai API.
Provides access to Google Gemini's text-to-speech service with style-directed speech and multiple voices through the kie.ai API.
Provides access to Kuaishou's Kling video generation and AI avatar models through the kie.ai API.
Provides access to OpenAI's image generation models (GPT Image, GPT-4o Image) through the kie.ai API.
Provides access to Suno's music generation, sound effects, lyrics, and audio manipulation tools through the kie.ai API.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@kie-mcpGenerate a 10-second video of a cat playing piano in Pixar style"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
kie-mcp
A comprehensive Model Context Protocol server for the kie.ai generation API. Gives Claude (and any MCP client) access to 60+ image models, 85+ video models, and 20+ audio tools with deep model intelligence built in.
Why this exists
Most MCPs are thin API wrappers. This one is different:
Vertical profiles (NEW in 5.0) — per-domain playbooks:
profile_briefreturns the intake questions a professional would ask, model routing per deliverable with live costs, per-model prompt formulas, and multi-tool workflows. Profiles double as MCP prompts (/kie-art:architecturein Claude Code). Verticals (10): architecture & interiors, video game assets, advertising & marketing, web & software product imagery, film & storyboarding, product photography & e-commerce, brand & graphic design, editorial & publishing, short-form social video, and audio branding & music.Deep research embedded — Every major model has a
researchfield with verdicts, prompt techniques, weaknesses, cost-efficiency analysis, and competitor comparisons. Researched by Averiguare, our model intelligence agent.Cost-aware — Every model has pricing in credits and USD. The MCP tells you the cheapest option for your use case.
Smart filtering —
list_models filter="lip sync"orfilter="architecture"orfilter="cheapest video"— searches across capability tags, descriptions, AND research fields.Dual-mode transport — stdio for local Claude Code, HTTP Streamable for remote Cowork/cloud usage.
Related MCP server: kie-ai-mcp-server
What you can do with it
Just ask Claude things like:
"Generate a brand presentation board for a perfume launch" — picks GPT Image 2 (best for text-heavy layouts)
"Make a 10s video of fruit scarecrows defending against crows, Pixar style" — recommends Veo 3.1 or Wan 2.7
"Generate music for a fantasy adventure game" — Suno V5
"Lip-sync this audio to my character image" — Kling AI Avatar or Infinitalk
"Upscale this video to 4K" — Veo 4K upscale or Topaz
"Replace the wall color in this room photo" — Flux Kontext Pro (best for surgical edits)
Model coverage
Image (60+)
OpenAI: GPT Image 2.5 Flare / Sunburst (NEW — 1K-4K, transparent background), GPT Image 2, GPT-4o Image, GPT Image 1.5
Google: Nano Banana 2.1 (NEW, 4 cr), Nano Banana 2 / 2 Lite / Pro / Edit / Original, Imagen 4 (Fast/Standard/Ultra)
Black Forest Labs: Flux Kontext Pro/Max, Flux 2 Pro/Flex
ByteDance: Seedream 3.0 / 4.0 / 4.5 / 5.0 Lite
Alibaba: Wan 2.7 Image / Image Pro
Ideogram: v3, Character, Edit, Remix, Reframe
xAI Grok Imagine Image 2.0 (#2 Arena T2I + edit; free segment map → region-targeted edit chain; whole-image edits of ANY uploaded image)
ByteDance Seedream 5.0 Pro (NEW — T2I/I2I + layer decomposition: split any image into layer files)
Qwen Image 3.0 / 3.0 Pro (NEW — seed, negative prompts, 2K at the 1K price on standard)
Qwen Image 2.1 (NEW — transparent background, mask inpainting, 10-ref compositing)
Others: Qwen/Qwen2, Z-Image, Grok Imagine 1.x, Recraft, Topaz
Video (85+)
Google Veo 3.1: Quality / Fast / Lite (T2V + I2V), Extend, 1080p/4K upscale
Alibaba HappyHorse: 1.1 (NEW — T2V/I2V/R2V with native audio + 7-language lip-sync), 1.0 (T2V/I2V/R2V/Video Edit)
ByteDance Seedance: 2.5 (NEW — 30s single takes, live Aug 2026) / 2.0 / 2.0 Fast / 2.0 Mini / 1.5 Pro
Kuaishou Kling: 3.0 Omni "O3" (NEW — per-shot multi_prompt scripting, 4K, video Transformation), 3.0, 3.0 Turbo, 2.6, V2.5 Turbo, V2.1 Master/Pro/Standard, AI Avatar
Alibaba Wan: 3.0 + 3.0 Prime (NEW — unified prompt-or-media, audio), 2.7 (T2V/I2V/Edit/R2V), 2.6, 2.5, 2.2 Turbo, Animate
Google Gemini Omni: Video + 1.1 Flash (NEW — first→last-frame mode, 360p-4K; image/voice/video/character references)
MiniMax Hailuo: H3 (NEW — 2K + native stereo audio, image+video+audio references, first→last-frame I2V), 2.3 Pro/Standard, 02 Pro/Standard
xAI Grok Imagine: Video 1.5 preview (NEW — I2V with native audio, cheapest audio video), T2V, I2V, Upscale, Extend
Avatar / lip-sync: OmniHuman 1.5 (NEW — audio-driven full-body avatar + free subject-detection + human-identification utilities), Volcengine Video Lip-Sync (NEW — re-dub existing footage), Kling AI Avatar, Infinitalk
PixVerse V6 (NEW): T2V, I2V (viral templates), Transition (first→last morph), Fusion R2V (@ref_name), Extend — budget all-rounder with native audio
Runway: Aleph, Aleph Edit, Extend
Others: ByteDance V1 Pro/Lite, Topaz upscale
Audio (20+)
Suno: Music Gen, Extend, Cover, Add Instrumental/Vocals, Replace Section, Lyrics, Sounds, Sound Effects, MIDI, Music Video, Cover Art, Mashup, Persona, Timestamped Lyrics, Boost Style, Vocal Separation, WAV, Custom Voice cloning (experimental)
ElevenLabs: TTS (Turbo 2.5 + Multilingual V2), Text-to-Dialogue V3, Audio Isolation, Speech-to-Text
Google Gemini TTS (NEW): style-directed speech, 30 voices, 2-speaker dialogue, inline tone tags — ~4.2 cr/min
Utility
File upload (URL or base64)
Veo Extend, 1080p Upscale, 4K Upscale
Runway Extend
Task status, credit check, raw asset listing
Installation
Prerequisites
Node.js 18+
A kie.ai API key from kie.ai/api-key
Setup
git clone https://github.com/YOUR_USERNAME/kie-mcp.git
cd kie-mcp
npm installRun as stdio MCP (Claude Code, Claude Desktop)
Add to your Claude config (~/.claude.json for Claude Code, or your MCP client's equivalent):
{
"mcpServers": {
"kie-art": {
"command": "node",
"args": ["/absolute/path/to/kie-mcp/server.mjs"],
"env": {
"KIE_API_KEY": "your-kie-ai-api-key",
"KIE_PROJECT_ROOT": "/optional/path/for/outputs"
}
}
}
}Or use the Claude Code CLI:
claude mcp add -s user kie-art /usr/bin/env -- KIE_API_KEY=your-key node /path/to/server.mjsRun as HTTP MCP (Cowork, remote clients)
HTTP mode requires a bearer token and listens on 127.0.0.1 by default.
export KIE_MCP_AUTH_TOKEN=$(openssl rand -hex 32)
KIE_API_KEY=your-key node server.mjs --http --port=3100Configure your MCP client with the URL and the token:
{
"mcpServers": {
"kie-art": {
"type": "http",
"url": "http://127.0.0.1:3100/mcp",
"headers": { "Authorization": "Bearer <your KIE_MCP_AUTH_TOKEN>" }
}
}
}To reach it remotely, put it behind a tunnel (ngrok / Cloudflare Tunnel) or a reverse proxy. The server stays on loopback and the tunnel forwards to it. Add the tunnel's hostname to KIE_MCP_ALLOWED_HOSTS. Treat the token like a password: anyone who has it can spend your kie credits and upload media files from the server's disk.
Security note (5.2.2): earlier versions ran HTTP mode on
0.0.0.0with no authentication andAccess-Control-Allow-Origin: *. Any web page open in your browser, or anyone who had the tunnel URL, could call the tools, including reading local files throughupload_file. Upgrade if you use--http.
Environment variables
Variable | Required | Purpose |
| yes | Your kie.ai API key |
| no | Server-wide default for where generated files are saved (default: server cwd; files go to |
| no | Port for HTTP mode (default: 3100) |
| no |
|
| no |
|
| HTTP mode | Bearer token HTTP clients must send ( |
| no | HTTP bind address (default |
| no | Comma-separated extra |
| no | Comma-separated browser origins allowed to call the server (CORS). Requests with any other |
| no | Callback URL sent with Suno generation requests (kie.ai requires the field; results are fetched by polling regardless). Defaults to an inert placeholder — set this only if you want to receive the callbacks yourself |
| no | Max simultaneous task-creation calls (default 4). Excess parallel generations queue inside the server instead of hitting kie.ai's rate limits — parallel tool calls are safe |
| no | Blocking-mode polling budget per tool category, in seconds (defaults: 600 / 900 / 300 / 300). Per-call |
Tools available
generate_image, generate_video, generate_music, generate_sfx,
generate_tts, generate_gemini_tts, generate_dialogue, generate_sounds, generate_lyrics,
generate_persona, generate_mashup, generate_cover_art,
generate_midi, create_music_video,
prepare_voice_clone, create_voice_clone, regenerate_voice_clone,
create_omni_voice, create_omni_character,
extend_music, cover_audio, upload_extend_audio,
add_instrumental, add_vocals, replace_section,
convert_to_wav, separate_vocals, boost_style,
get_timestamped_lyrics, audio_isolation, speech_to_text,
profile_brief,
list_models, check_task, list_tasks, check_credits,
download_result, list_raw_assets, upload_file,
grok_segment_map, grok_image_edit, seedream_layer_decompose,
veo_extend, veo_upscale_1080p, veo_upscale_4k, runway_extendSmart model recommendations
Try these queries in any MCP client:
list_models filter="reasoning" # GPT-4o, Nano Banana, GPT Image 2
list_models filter="lip-sync" # OmniHuman 1.5, Volcengine, Kling Avatar, HappyHorse 1.1
list_models filter="multi-shot" # Kling 3.0/Turbo
list_models filter="cheapest video" # Grok Imagine 1.5, Wan Flash
list_models filter="alibaba" # HappyHorse 1.0/1.1 family
list_models filter="best visual quality" # Veo Quality, Seedance 2.0
list_models filter="text rendering" # Ideogram v3, GPT Image 2
list_models filter="character" # Ideogram Character, Kling AI AvatarArchitecture
server.mjs # Transport, helpers, tool handlers (~2700 lines)
├── createMcpServer() # Factory for stdio + HTTP modes
├── Tool handlers # generate_*, list_*, etc.
└── helpers # polling, recovery, pricing, validation, download
data/ # Pure data, imported (and re-exported) by server.mjs
├── registry-image.mjs # MODEL_REGISTRY — image models (47+)
├── registry-video.mjs # VIDEO_MODEL_REGISTRY — video models (80+)
├── registry-audio.mjs # AUDIO_TOOLS_REGISTRY — audio tool metadata
├── pricing.mjs # PRICING, PRICING_ESTIMATED, PROMPT_CAPS
└── voices.mjs # ELEVENLABS_VOICES catalogThe registries and pricing live in data/*.mjs so model-catalog changes are reviewable diffs instead of edits buried in a 5000-line file; server.mjs imports and re-exports them (tests and downstream keep importing from server.mjs).
Each model entry has:
name,description,capabilities(tags),pricing(credits)aspectRatios,options(with types and defaults)buildBody/buildInput(request builders)research(Averiguare verdicts, prompt techniques, weaknesses, comparisons, sources)
Development
npm run check # node --check server.mjs (syntax)
npm test # offline unit tests for the pure helpers (test/*.test.mjs)
npm run smoke # live end-to-end over MCP stdio — needs KIE_API_KEY
# (spends ~0 credits; uses the free subject-detection model)server.mjs guards its side effects behind a main-module check, so it can be imported by tests (test/unit.test.mjs) to exercise the pure helpers without starting a server. test/harness.mjs is a reusable stdio JSON-RPC client for driving the real server in smoke/integration checks. CI (.github/workflows/ci.yml) runs the syntax check + unit tests on Node 20 and 22 for every push and PR.
Drift watch
kie.ai changes things without notice — advertised prices, model availability, even API shapes. .github/workflows/drift-watch.yml runs scripts/drift-watch.mjs weekly (and on demand) to scan for it: paused/removed slugs, pricing that no longer matches the PRICING table, and new models in kie's catalog. Findings land in a single rolling GitHub issue. Add a KIE_API_KEY repo secret to enable the per-slug liveness probes (0 credits — empty-input validation errors); the pricing and new-model scans need no secret. Run locally with node scripts/drift-watch.mjs.
Releasing
Releases are automated by .github/workflows/release.yml. To cut a release:
Bump the version in
package.json,server.json(both the top-levelversionandpackages[0].version), andserver.mjs(SERVER_INFO+ the/healthhandler), and add a## [X.Y.Z]section toCHANGELOG.md. Merge tomain.Tag and push:
git tag vX.Y.Z && git push origin vX.Y.Z
The workflow verifies the tag matches every in-repo version string, publishes to npm with provenance (NPM_TOKEN repo secret), and creates the GitHub Release using the matching CHANGELOG section as the notes. A tag whose version doesn't match the code fails fast without publishing. workflow_dispatch is an emergency manual publish of the current package.json version.
Credits
Built with the MCP TypeScript SDK
Powered by kie.ai — affordable unified API for 100+ AI models
Model intelligence by Averiguare — "No sabes hasta que averiguas — y averiguo en todas partes."
License
MIT
Available Tools
46 toolsadd_instrumentalC
Add instrumental backing to uploaded vocal audio via Suno.
| Name | Required | Description | Default |
|---|---|---|---|
| tags | No | Style tags for the instrumental | |
| wait | No | Set false to submit and return immediately with the task_id (async mode) — then poll with check_task and fetch with download_result. Recommended for long generations to avoid client-side watchdog timeouts. | |
| model | No | V5 | |
| title | No | ||
| filename | No | ||
| uploadUrl | Yes | URL of vocal audio | |
| download_dir | No | Absolute directory to save the file(s) into (created if missing). Defaults to the server's kie/assets/raw/. Must be absolute — the MCP server's working directory is not the caller's. | |
| negativeTags | No | ||
| max_wait_seconds | No | Override the blocking-mode polling budget in seconds (default: audio 300). Ignored when wait=false. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It does not disclose that by default the call blocks up to ~300s (watchdog timeouts possible), that async polling with check_task/download_result is required for long runs, or any prerequisite about the uploaded vocal audio. These behaviors are buried in schema field descriptions, not stated in the description.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, front-loaded sentence with no waste. Being terse is efficient, though it borders on under-specification given the tool's async complexity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 9-parameter audio-generation tool with an async polling workflow, no annotations, and no output schema, the description is far too thin. It omits the blocking-vs-async behavior, failure modes, and parameter semantics that an agent needs to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 56%, leaving several parameters (title, filename, negativeTags, model) undocumented. The description adds no parameter meaning at all. With over a third of parameters uncovered and no compensating text, this is weak.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Add') and resource ('instrumental backing') with the input ('uploaded vocal audio') and provider ('Suno'). This is clear but does not differentiate it from the close sibling add_vocals, which is essentially the inverse operation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this versus siblings like generate_music, add_vocals, or extend_music. The async workflow is described inside a parameter schema rather than in the description, and the description gives no context selection criteria.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
add_vocalsC
Add AI vocals to uploaded instrumental audio via Suno.
| Name | Required | Description | Default |
|---|---|---|---|
| wait | No | Set false to submit and return immediately with the task_id (async mode) — then poll with check_task and fetch with download_result. Recommended for long generations to avoid client-side watchdog timeouts. | |
| model | No | V5 | |
| style | No | ||
| title | No | ||
| prompt | Yes | Lyrics or vocal description | |
| filename | No | ||
| uploadUrl | Yes | URL of instrumental audio | |
| download_dir | No | Absolute directory to save the file(s) into (created if missing). Defaults to the server's kie/assets/raw/. Must be absolute — the MCP server's working directory is not the caller's. | |
| negativeTags | No | ||
| max_wait_seconds | No | Override the blocking-mode polling budget in seconds (default: audio 300). Ignored when wait=false. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden, and it says nothing about blocking vs async execution, expected generation latency, credit consumption, or what files are produced. The only hints (wait, max_wait_seconds) live in the schema, not the description, and the far more complex async/polling workflow is never surfaced.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single sentence with no filler is front-loaded and easy to scan, but it is under-specified rather than concise — the brevity comes at the cost of any operational context.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 10-parameter generation tool with an async polling mode, no annotations, and no output schema, a one-line description is not sufficient. An agent cannot infer the polling contract, output artifacts, or how the vocal models differ from what the schema alone provides.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 50%, leaving model, style, title, filename, and negativeTags undocumented in both places. The description adds zero parameter meaning (no lyric/format hints for prompt, no model-selection advice), so it fails to compensate for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource combination ('Add AI vocals to uploaded instrumental audio') and names the underlying provider (Suno), which makes it distinguishable from the reverse-direction sibling add_instrumental. It does not explicitly name or route against alternatives, but the purpose is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'uploaded instrumental audio' weakly implies a prerequisite that an instrumental already exists at a URL, but there is no explicit when-to-use guidance, no mention of when to prefer add_instrumental/cover_audio/generate_music, and no note on credit cost or eligibility.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
audio_isolationB
Isolate vocals or audio from background noise using ElevenLabs via kie.ai. Input an audio URL, get clean isolated audio back.
| Name | Required | Description | Default |
|---|---|---|---|
| filename | No | Output filename. Auto-generated if omitted. | |
| audio_url | Yes | Audio URL to process (max 10MB) | |
| download_dir | No | Absolute directory to save the file(s) into (created if missing). Defaults to the server's kie/assets/raw/. Must be absolute — the MCP server's working directory is not the caller's. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations are absent, so the description carries full burden. It does not disclose processing time, supported formats, failure modes, or the max file size (10MB is only in schema). Minimal behavioral insight.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Single 17-word sentence, front-loaded with key verb and resource. No wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple input/output tool without output schema, the description is adequate but lacks context on quality, latency, or supported audio types. Could be more complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3. The description adds no meaning beyond the schema (e.g., no extra context for audio_url, filename, or download_dir).
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('isolate'), resource ('audio'), and output ('clean isolated audio'), with specific reference to ElevenLabs via kie.ai. It distinguishes from siblings like 'separate_vocals' by mentioning 'background noise,' but does not explicitly differentiate.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this tool versus alternatives (e.g., separate_vocals). No when-not or prerequisites mentioned. The description only states the basic input/output flow.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
boost_styleA
NEW — Convert concise style input (e.g. "Pop, Mysterious") into enhanced style description for music generation.
| Name | Required | Description | Default |
|---|---|---|---|
| content | Yes | Concise style description to enhance |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description must disclose behavioral traits. It only states the conversion action without mentioning whether it is read-only, has side effects, requires permissions, or handles errors, leaving key behavioral aspects unspecified.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single concise sentence prefixed by 'NEW —', which front-loads relevance. It contains no fluff, but could be slightly more structured for readability.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no output schema, the description fails to explain the output format or how the enhanced description is used, and does not integrate with sibling tools (e.g., generate_music) despite being a likely preprocessing step.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, but the description adds value by providing an example input ('Pop, Mysterious') and clarifying the transformation purpose, which goes beyond the schema's minimal description 'Concise style description to enhance'.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool converts concise style input into an enhanced style description for music generation, using a specific verb ('Convert') and resource ('style input'), and distinguishes it from generation siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage as a preprocessing step for music generation tools (e.g., generate_music) with the example 'Pop, Mysterious', but does not explicitly state when to use it versus alternatives or provide when-not-to-use guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
check_creditsB
Check remaining kie.ai account credits
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are present, so the description carries the full burden of behavioral disclosure. It only states the action without revealing whether the operation is read-only, whether it consumes credits, or what the response format looks like.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, short sentence that directly states the tool's purpose with no extraneous words. It is optimally concise and front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a parameterless tool, the description covers the essential purpose. However, it lacks any mention of return values or format, which would help the agent interpret the result. Given the simplicity, it is mostly complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has zero parameters with 100% coverage, so no parameter documentation is needed. The description adds no parameter semantics, but this is adequate given the absence of parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses the specific verb 'check' and clearly identifies the resource as 'remaining kie.ai account credits'. Among sibling tools focused on generation and editing, this is the only credit-related tool, making its purpose unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No usage guidelines are provided. The description does not indicate when to use this tool versus alternatives, nor does it mention any prerequisites or context for checking credits.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
check_taskA
Check the status of a kie.ai generation task by taskId
| Name | Required | Description | Default |
|---|---|---|---|
| task_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description must fully disclose behavior. It only states 'check status', omitting whether the tool is read-only, what the response contains, or any side effects. Minimal behavioral context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single sentence that is front-loaded and to the point, with no wasted words. Efficiently conveys the core action.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple tool with one parameter and no output schema, the description is minimally adequate but lacks details on response content, error handling, or behavior under different task states. Enough to understand basic function but not fully complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has one string parameter (task_id) with 0% description coverage. The description mentions 'by taskId', adding context that task_id is the identifier, but does not explain format, origin, or examples. Some value added but insufficient for full compensation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb 'check' and the resource 'status of a kie.ai generation task' with the mechanism 'by taskId'. It distinguishes itself from sibling tools like list_tasks, which lists tasks, and generate_*, which create tasks.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage when you have a taskId and want to query its status, but provides no explicit guidance on when to use this tool versus alternatives like list_tasks (which might also show status) or when not to use it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
convert_to_wavC
Convert a Suno track to lossless WAV format. Downloads to kie/assets/raw/.
| Name | Required | Description | Default |
|---|---|---|---|
| wait | No | Set false to submit and return immediately with the task_id (async mode) — then poll with check_task and fetch with download_result. | |
| taskId | Yes | Task ID of the Suno generation | |
| audioId | Yes | Audio ID from sunoData | |
| filename | No | ||
| download_dir | No | Absolute directory to save the file(s) into (created if missing). Defaults to the server's kie/assets/raw/. Must be absolute — the MCP server's working directory is not the caller's. | |
| max_wait_seconds | No | Override the blocking-mode polling budget in seconds (default: audio 300). Ignored when wait=false. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description must disclose behavioral traits. It mentions the download path but does not disclose safety, destructiveness, or other important behaviors like blocking vs. async mode, rate limits, or permissions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is extremely concise, consisting of just two sentences that get straight to the point. No unnecessary information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite having 6 parameters and no output schema, the description does not explain the tool's behavior in sync vs async mode, return values, or how it interacts with other tools like check_task and download_result. This is insufficient for a tool that requires polling in async mode.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 83%, so the schema already explains parameters well. The description adds no extra semantic value beyond what is in the schema, so baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Convert a Suno track to lossless WAV format') and provides the download location. However, it does not explicitly distinguish from other audio conversion tools among the siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description lacks any guidance on when to use this tool versus alternatives. No mention of prerequisites, context for conversion, or integration with other tools like check_task or download_result.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
cover_audioC
Create an AI cover from uploaded audio — custom vocals, style, and instrumentation via Suno.
| Name | Required | Description | Default |
|---|---|---|---|
| wait | No | Set false to submit and return immediately with the task_id (async mode) — then poll with check_task and fetch with download_result. Recommended for long generations to avoid client-side watchdog timeouts. | |
| model | No | V5 | |
| style | No | ||
| title | No | ||
| prompt | No | Description of desired cover style | |
| filename | No | ||
| uploadUrl | Yes | URL of audio to cover | |
| customMode | No | ||
| vocalGender | No | Vocal gender preference | |
| download_dir | No | Absolute directory to save the file(s) into (created if missing). Defaults to the server's kie/assets/raw/. Must be absolute — the MCP server's working directory is not the caller's. | |
| instrumental | No | ||
| negativeTags | No | Tags to avoid in the cover | |
| max_wait_seconds | No | Override the blocking-mode polling budget in seconds (default: audio 300). Ignored when wait=false. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, and it discloses almost nothing: it does not say the operation is asynchronous/blocking by default, that it consumes credits, that generation can take minutes, or what comes back. The single sentence gives only the creative intent, not the operational behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One compact, front-loaded sentence with no filler. It is efficient, though the extreme brevity is part of the tool's under-specification problem rather than a virtue.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 13-parameter generation tool with no annotations, no output schema, and only 54% parameter coverage, a single sentence is not enough. The async/polling workflow, credit cost, and result shape are all unaddressed at the description level.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 54%, and the description adds nothing beyond loosely echoing three fields (vocals, style, instrumentation). It does not clarify the distinction between style, prompt, and negativeTags, what customMode does, or how model choice affects the cover.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: create an AI cover from uploaded audio, with the field hints (vocals, style, instrumentation) and the backend (Suno). An agent can distinguish it from siblings like generate_music or add_vocals, but the description never names or contrasts those alternatives explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to choose a cover over generate_music, extend_music, or add_vocals, nor any stated prerequisites for the source audio. Usage is only implied by the name and the mention of 'uploaded audio'.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_music_videoB
Generate an MP4 music video visualization from a Suno track. Downloads to kie/assets/raw/.
| Name | Required | Description | Default |
|---|---|---|---|
| wait | No | Set false to submit and return immediately with the task_id (async mode) — then poll with check_task and fetch with download_result. | |
| author | No | Author name for video credits | |
| taskId | Yes | Task ID of the Suno generation | |
| audioId | Yes | Audio ID from sunoData | |
| filename | No | ||
| domainName | No | Domain name for video branding | |
| download_dir | No | Absolute directory to save the file(s) into (created if missing). Defaults to the server's kie/assets/raw/. Must be absolute — the MCP server's working directory is not the caller's. | |
| max_wait_seconds | No | Override the blocking-mode polling budget in seconds (default: audio 300). Ignored when wait=false. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Description discloses the download destination (kie/assets/raw/), but without annotations, it fails to detail other behavioral traits like wait mode, error handling, or rate limits. Some transparency but significant gaps.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, no filler, front-loaded with purpose. Highly concise.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With 8 parameters, 2 required, no output schema, and no annotations, the description is too brief. It does not explain the full workflow, return format, or dependencies on other tools like check_task.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is high (88%), so description adds marginal value beyond the schema. It provides context for download_dir default but does not explain other parameters like author or domainName beyond their schema descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description clearly states it generates an MP4 music video from a Suno track and specifies the download destination. However, it does not differentiate from sibling tools like generate_video, which could be ambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this tool versus alternatives like generate_video, or prerequisites such as needing a Suno generation task. Usage context is only implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_omni_characterB
NEW — Create a reusable visual character for Gemini Omni video generation. Combines image + optional voice. Returns characterId.
| Name | Required | Description | Default |
|---|---|---|---|
| audio_ids | No | Optional voice IDs from create_omni_voice | |
| image_urls | Yes | Exactly 1 image URL (≤20MB) | |
| descriptions | Yes | Character appearance, identity, style, clothing, personality | |
| character_name | No | Character name |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description bears full burden. It mentions the combination of image and optional voice and the return of a characterId, but does not disclose side effects, error handling, limitations on image URLs, or permission requirements.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is very concise, with three short, front-loaded sentences. No wasted words, but could be slightly more structured to separate purpose, behavior, and output.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has 4 parameters (2 required) and no output schema, the description is incomplete. It does not explain constraints on image URLs, how characterId is formatted, or how this tool fits into the video generation workflow.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, providing clear parameter descriptions. The tool description adds 'Combines image + optional voice' and 'Returns characterId', slightly extending schema info but not substantially.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb ('Create'), the resource ('reusable visual character for Gemini Omni video generation'), and distinguishes from siblings like 'create_omni_voice' by specifying that it combines image with optional voice.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for creating characters but does not explicitly state when to use this tool versus alternatives (e.g., generate_video) or provide prerequisites. Lack of exclusions limits guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_omni_voiceB
NEW — Create a reusable voice character for Gemini Omni video generation. Returns kieAudioId for use in generate_video audio_ids.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | Voice character name (max 210 chars) | |
| audio_id | Yes | Preset base voice (30 options). The created voice inherits this preset and is customized by voice_description. | |
| example_dialogue | No | Sample dialogue (max 120 chars), e.g. "Hello, I am Adam" | |
| voice_description | No | Detailed voice characteristics: timbre, style, rate, emotion (max 20000 chars) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It states that the tool creates a reusable voice character and returns a kieAudioId, but fails to explain mutation implications, authorization needs, rate limits, or any side effects. The description is too brief to be fully transparent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is extremely concise: two sentences covering purpose, newness, use case, and output. Every sentence adds value with no wasted words. Front-loaded with 'NEW' and the core action.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the moderate complexity of creating a voice character, the description covers the essential purpose and output linkage. However, it omits details on customization behavior and validation beyond what the schema provides. Still, it is largely complete for a well-schematized tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents each parameter. The description adds no new parameter-level meaning beyond the schema. Baseline 3 is appropriate as the description does not need to compensate for coverage gaps.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb (create), resource (reusable voice character), and context (Gemini Omni video generation). It also specifies the return value linking to generate_video. However, it does not explicitly differentiate from sibling tools like create_voice_clone, which creates a different type of voice resource.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description mentions a specific use case (video generation) and links to generate_video, but does not provide explicit guidance on when to use this tool versus alternatives such as create_voice_clone or prepare_voice_clone. No when-not or exclusion criteria are given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_voice_cloneA
EXPERIMENTAL (#20) — STEP 2 — after the voice owner records the verification phrase from prepare_voice_clone, submit that recording to finish the voice. On success returns a voiceId usable in generate_music. Unverified end-to-end.
| Name | Required | Description | Default |
|---|---|---|---|
| wait | No | ||
| style | No | Optional style hint, e.g. "Pop, Female Vocal" | |
| task_id | Yes | task_id from prepare_voice_clone (must be at wait_validating) | |
| verify_url | Yes | Public URL of the voice owner's recording of the verification phrase | |
| voice_name | Yes | Name for the custom voice | |
| description | No | Optional description | |
| max_wait_seconds | No | ||
| singer_skill_level | No | Optional |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations exist, so the description carries the full burden. It highlights experimental status and unverified end-to-end flow but does not disclose what happens on failure, required permissions, or state changes beyond returning a voiceId.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is extremely concise at one sentence, front-loading key info (EXPERIMENTAL, STEP 2). Every word adds value without fluff.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given 8 parameters and no output schema, the description is somewhat sparse. It omits details like verification URL format, timeout behavior for max_wait_seconds, and failure handling, leaving gaps for the agent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 75% schema coverage, most parameters already have descriptions. The tool description adds no extra meaning beyond what's in the schema, so it meets the baseline but does not improve understanding.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states it is the second step after prepare_voice_clone, requiring the verification recording. It specifies the outcome (voiceId) and usage context (generate_music), distinguishing it from siblings like prepare_voice_clone and regenerate_voice_clone.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly tells when to use (after prepare_voice_clone, after recording the verification phrase) but does not provide exclusions or alternatives for failure cases.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
download_resultC
Download a completed task result to kie/assets/raw/
| Name | Required | Description | Default |
|---|---|---|---|
| task_id | Yes | ||
| filename | No | ||
| download_dir | No | Absolute directory to save the file(s) into (created if missing). Defaults to the server's kie/assets/raw/. Must be absolute — the MCP server's working directory is not the caller's. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It does not mention error handling, authentication, or whether files are overwritten, leaving significant behavioral gaps.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, front-loaded sentence is concise with no wasted words. However, it may be too terse for important nuances.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With 3 parameters, no output schema, and zero annotations, the description is insufficient. It lacks details on return values, error cases, and workflow context among siblings.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 33% (only download_dir has a description). The description adds default directory info but does not explain task_id or filename, which are crucial but undocumented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action 'Download a completed task result' and the default target directory 'kie/assets/raw/', making the purpose apparent. However, it implies a fixed destination while the schema allows a custom directory, causing slight inconsistency.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this tool versus alternatives (e.g., upload_file) or prerequisites like task completion. The agent is left to infer usage context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
extend_musicB
Extend/continue an existing Suno track from a specific point. Requires audioId from a previous generate_music task.
| Name | Required | Description | Default |
|---|---|---|---|
| wait | No | Set false to submit and return immediately with the task_id (async mode) — then poll with check_task and fetch with download_result. Recommended for long generations to avoid client-side watchdog timeouts. | |
| model | No | V5 | |
| style | No | Style tags for the extension | |
| title | No | ||
| prompt | Yes | LYRICS for the extended section (Suno sings this text). Give real lyric lines; a short instruction like "add an outro" is rejected as malformed lyrics (error 531, refunded). | |
| audioId | Yes | Audio ID from a previous Suno generation (from sunoData) | |
| filename | No | ||
| continueAt | No | Timestamp in seconds to continue from | |
| download_dir | No | Absolute directory to save the file(s) into (created if missing). Defaults to the server's kie/assets/raw/. Must be absolute — the MCP server's working directory is not the caller's. | |
| defaultParamFlag | No | Use default params from original track | |
| max_wait_seconds | No | Override the blocking-mode polling budget in seconds (default: audio 300). Ignored when wait=false. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden, yet it only states the audioId prerequisite. It says nothing about cost/credit consumption, async vs blocking execution, refund behavior on malformed input, or mutation scope — behavior that matters for a paid, stateful generation call.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences with the core action front-loaded and the prerequisite second. Nothing is wasted, though the minimal content reflects under-specification rather than meaningful economy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an 11-parameter generator tool with no output schema and no annotations, the description is thin, omitting async/cost/refund context. The relatively rich per-parameter schema descriptions partially compensate, keeping this at the minimum-viable level rather than a gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 73%, so most parameters are already documented in the schema (wait, prompt, download_dir, max_wait_seconds). The description adds only the audioId requirement, which the schema already conveys, so it does not meaningfully extend parameter semantics. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (extend/continue) and resource (an existing Suno track) plus scoping detail (from a specific point). An agent can distinguish it from generate_music, but the description does not differentiate it from close siblings like replace_section or upload_extend_audio.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides one prerequisite (audioId from a previous generate_music task), which implies when the tool is usable. It gives no explicit rule for choosing between this and replace_section, upload_extend_audio, or add_instrumental, leaving alternatives to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
generate_cover_artA
NEW — Generate album cover art image for an existing Suno music track. One call per taskId only.
| Name | Required | Description | Default |
|---|---|---|---|
| taskId | Yes | Suno task ID from a previous music generation | |
| filename | No | Output filename. Auto-generated if omitted. | |
| download_dir | No | Absolute directory to save the file(s) into (created if missing). Defaults to the server's kie/assets/raw/. Must be absolute — the MCP server's working directory is not the caller's. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries full burden. It discloses only the 'one call per taskId' constraint, but fails to mention if the operation is destructive, what happens on repeated calls, authentication needs, or rate limits. The lack of behavioral context makes it risky for an AI agent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences, fully front-loaded with purpose in the first sentence and a key constraint in the second. No wasted words; every sentence adds value.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with 3 parameters and no output schema, the description provides core purpose and a constraint, but omits expected output type, file format, or error handling. Given the absence of output schema, more detail on return values would help completeness.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, but the description adds important context: 'Auto-generated if omitted' for filename and a crucial warning about absolute directory paths for download_dir. 'One call per taskId' also relates to parameter usage, enhancing understanding beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb 'generate' and the specific resource 'album cover art image for an existing Suno music track'. It distinguishes from sibling tools like 'generate_image' (general image generation) and 'create_music_video' (video). The constraint 'One call per taskId only' further clarifies the scope.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage context (existing music track) and a constraint (one call per taskId), but does not explicitly state when to use this tool over alternatives like 'generate_image' or what prerequisites are needed. There is no guidance on when not to use it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
generate_dialogueA
Generate multi-speaker dialogue using ElevenLabs Text-to-Dialogue V3 via kie.ai. Great for conversations between characters. Downloads to kie/assets/raw/.
| Name | Required | Description | Default |
|---|---|---|---|
| wait | No | Set false to submit and return immediately with the task_id (async mode) — then poll with check_task and fetch with download_result. Recommended for long generations to avoid client-side watchdog timeouts. | |
| dialogue | Yes | Array of dialogue lines with voice assignments | |
| filename | No | Output filename. Auto-generated if omitted. | |
| stability | No | Voice stability — kie accepts exactly 0 (creative), 0.5 (natural), or 1 (robust) | |
| download_dir | No | Absolute directory to save the file(s) into (created if missing). Defaults to the server's kie/assets/raw/. Must be absolute — the MCP server's working directory is not the caller's. | |
| language_code | No | Language code (e.g. "en") | |
| max_wait_seconds | No | Override the blocking-mode polling budget in seconds (defaults: image 600, video 900, audio 300, speech 300). Ignored when wait=false. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries full burden. It mentions the download destination and the service used, but lacks details on side effects, cost, rate limits, or error behavior. The schema provides async mode details, but the description itself is minimal.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two concise sentences, front-loaded with the action verb and resource. Every word adds value with no redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has 7 parameters and uses a complex async workflow, but the description is minimal. The schema compensates with thorough parameter documentation, yet the description does not cover return values or error handling. Adequate but not comprehensive.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all parameters well. The description adds no extra meaning beyond noting the download directory, so a baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool generates multi-speaker dialogue, mentions the underlying service (ElevenLabs Text-to-Dialogue V3 via kie.ai), and indicates the output location. This distinguishes it from sibling tools like generate_tts (single speaker) and generate_music (audio generation).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description says 'Great for conversations between characters,' implying a use case, but it does not provide explicit guidance on when to use this tool versus alternatives, nor does it mention prerequisites or constraints.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
generate_gemini_ttsA
NEW — Google Gemini native TTS via kie.ai: style-directed speech from natural-language direction, 30 named voices, up to 2 speakers, inline tone tags like [whispers]/[laughs] (flash model). ~4.2 credits per MINUTE of audio — cheaper than all ElevenLabs tiers. Simple mode: pass text (+ optional voice_name). Dialogue mode: pass speakers + dialogue_turns. model=flash is most expressive (keep expected audio <60s — quality degrades on long takes); model=pro is more stable for multi-minute narration. Downloads to kie/assets/raw/.
| Name | Required | Description | Default |
|---|---|---|---|
| text | No | Simple mode: the text to speak (single speaker). Inline tone tags like [whispers] work on flash. Ignored if dialogue_turns is set. | |
| wait | No | Set false to submit and return immediately with the task_id (async mode) — then poll with check_task and fetch with download_result. | |
| model | No | flash = Gemini 3.1 Flash TTS (most expressive, 200+ inline tags, best <60s; ~4.2 cr/min). flash-3.8 = Gemini 3.8 Flash TTS (NEW Sept 2026, same inputs, ~1.9 cr/min, ~40% cheaper per clip in live tests). flash-lite = Gemini 3.8 Flash Lite TTS (NEW, cheapest at ~1.3 cr/min; draft/bulk narration). pro = Gemini 2.5 Pro TTS (more stable long-form, ~4.2 cr/min). 3.8 prices are kie limited-time pricing until 2026-12-31. | flash |
| scene | No | Scene description, e.g. "A quiet, warm room with a fireplace crackling softly." | |
| filename | No | Output filename (default gemini-tts-<ts>.wav) | |
| speakers | No | Dialogue mode: 1-2 speakers as [{speaker_id: "Speaker 1", voice_name, audio_profile?, accent?, style?, pace?}]. accent: Neutral|American (Gen)|American (Valley)|American (South)|British (RP)|British (Brixton)|Transatlantic|Australian. style: Vocal Smile|Newscaster|Whisper|Empathetic|Promo/Hype|Deadpan. pace: Natural|Rapid Fire|The Drift|Staccato. | |
| voice_name | No | Simple mode voice (default Zephyr) | |
| temperature | No | Sampling temperature (default 1) | |
| download_dir | No | Absolute directory to save into (created if missing). Defaults to the server's kie/assets/raw/. | |
| dialogue_turns | No | Dialogue mode: [{speaker_id, text}] in order; text ≤10000 chars, may contain tone tags | |
| sample_context | No | Overall tone/direction, e.g. "Audiobook style narration. Tone is gentle and inviting." | |
| max_wait_seconds | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does well: it discloses cost (~4.2 credits/min, cheaper than ElevenLabs), the quality-degradation risk on long flash takes, the output destination (kie/assets/raw/), and limited-time pricing expiry. It omits nothing critical, though it says nothing about auth/permissions or concurrency limits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the identity and the two operating modes, and the model tradeoffs are tightly packed. Some promotional phrasing ('cheaper than all ElevenLabs tiers', repeated pricing) consumes space without helping invocation, keeping it from a 5.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 12-parameter tool with no annotations and no output schema, the description covers mode routing, model tradeoffs, cost, and save location well. It leaves speaker_id/pacing semantics and the async polling flow to the schema, which is acceptable given 92% coverage.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is already 92%, so the schema documents nearly every parameter (model enums, wait, speakers, temperature). The description largely restates schema content (2-speaker cap, tone tags, enum meaning) rather than adding new semantic detail, which is the baseline-3 case.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource (native Gemini TTS via kie.ai) and enumerates the distinguishing capabilities: style-directed speech, 30 named voices, up to 2 speakers, inline tone tags. It never names sibling tools like generate_tts or generate_dialogue, so the agent must infer the boundary rather than being routed explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear mode selection rules (simple mode = text+voice_name; dialogue mode = speakers+dialogue_turns) and model selection guidance keyed to clip length (flash for ~60s expressive takes, pro for stable multi-minute narration). However, it offers no when-not-to-use this tool versus the sibling TTS/dialogue tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
generate_imageA
Generate an image using kie.ai. TIP: for architecture/game-art/advertising/product-UI jobs, call profile_brief first — it returns the vertical's intake questions, routing, and prompt formulas. (60+ models). Downloads to kie/assets/raw/. MODEL GUIDE: Architecture/blueprints→gpt4o or nano-banana-2 (reasoning). Game art/3D→seedream/4.5 or 5-lite. Character sheets→ideogram/character. Text/logos→ideogram/v3 (best text). Photo editing→flux-kontext-pro. Newest OpenAI→gpt-image-2-5/flare-* (6cr @1K, fast default) or gpt-image-2-5/sunburst-* (premium polish); both 1K-4K + transparent background (NEW). Split any image into layers→seedream_layer_decompose tool (7cr/layer, NEW). Generate-then-refine by named region→grok-imagine-image-2-0/text-to-image (4cr, #2 Arena T2I+edit) then grok_segment_map (free) + grok_image_edit (4cr; also edits ANY uploaded image via image_urls mode). Anime→qwen (3cr cheapest); qwen2-1/* (4cr, NEW) adds transparent BG, mask inpainting, 10-ref compositing. Fast drafts→nano-banana-2-lite (4cr, ~4s, NEW). Upscale→recraft/crisp-upscale (0.5cr). BG removal→recraft/remove-background. Cheapest→z-image,qwen (3cr). Best quality→nano-banana-pro (24cr), flux-kontext-max (100cr). Use list_models filter="use-case" to explore.
| Name | Required | Description | Default |
|---|---|---|---|
| wait | No | Set false to submit and return immediately with the task_id (async mode) — then poll with check_task and fetch with download_result. Recommended for long generations to avoid client-side watchdog timeouts. | |
| model | No | Model ID. Use list_models to see all available models and their options. | gpt4o |
| prompt | Yes | Text prompt describing the image to generate | |
| filename | No | Output filename (saved to kie/assets/raw/). Auto-generated if omitted. | |
| image_urls | No | Reference/input image URLs for image-to-image models | |
| aspect_ratio | No | Aspect ratio (valid values depend on model — see list_models). Common: 1:1, 2:3, 3:2, 16:9, 9:16, 4:3, 3:4 | 2:3 |
| download_dir | No | Absolute directory to save the file(s) into (created if missing). Defaults to the server's kie/assets/raw/. Must be absolute — the MCP server's working directory is not the caller's. | |
| model_options | No | Model-specific options (quality, resolution, seed, negative_prompt, etc). Use list_models to see available options per model. | |
| max_wait_seconds | No | Override the blocking-mode polling budget in seconds (defaults: image 600, video 900, audio 300, speech 300). Ignored when wait=false. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does disclose real behavioral traits: output goes to kie/assets/raw/, per-model credit costs (3cr to 100cr), latency (~4s for nano-banana-2-lite), and quality tiers. It stops short of covering auth/permission needs, failure modes, or what happens when a model rejects an aspect_ratio, so it is strong but not exhaustive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The purpose and profile_brief tip are correctly front-loaded, but the body becomes a dense catalog of ~15 model IDs and prices in one paragraph, which is hard to scan and near the limit of what belongs in a tool description rather than a referenced resource. Information density is high, but structure suffers.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 9-parameter, no-output-schema tool this covers most of what an agent needs: output location, model selection guidance, and async handling is already documented in the schema's `wait` parameter. It lacks guidance on parameter interactions (e.g. aspect_ratio validity per model beyond a pointer) but is otherwise complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3, but the model guide meaningfully enriches the `model` parameter by mapping use-cases (blueprints, text/logos, anime, upscale, BG removal) to concrete model IDs and their tradeoffs. The remaining parameters get no semantic help in the description, so it does not reach 5.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening sentence gives a specific verb and resource ('Generate an image using kie.ai') and the surrounding model guide makes it immediately distinguishable from siblings like generate_video, grok_image_edit, and seedream_layer_decompose, each of which is explicitly redirected to. An agent can tell what this tool is and is not without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit routing rules are given: call profile_brief first for architecture/game-art/advertising/product-UI, use list_models filter="use-case" to explore, and delegating alternatives named for layer splitting, region refine, and editing. When-to-use and when-to-use-something-else are both stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
generate_lyricsB
Generate song lyrics from a prompt using Suno AI (max 200 characters). Returns text, no file download.
| Name | Required | Description | Default |
|---|---|---|---|
| prompt | Yes | Prompt for lyrics generation (max 200 chars) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations provided, so description must disclose all behavioral traits. Mentions output is text, no file download, and max 200 characters. Adequate but lacks details like rate limits, API key requirements, or whether generation is synchronous.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, no fluff. Every word is meaningful.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple single-parameter tool without output schema or annotations, the description sufficiently covers purpose, input constraints, and output type. Could mention that prompt should be lyrics-style or something, but not necessary.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% with description for prompt. Tool description reinforces the max length constraint but adds no new semantic information beyond what schema provides. Baseline score is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States it generates song lyrics from a prompt using Suno AI, which is specific. Does not explicitly differentiate from siblings like generate_music, but the resource 'lyrics' implies a distinct purpose.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use or when-not-to-use guidance. Does not mention alternatives among the many sibling tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
generate_mashupB
NEW — Mashup up to 2 Suno tracks into one new track. Provide audioIds from previous generations.
| Name | Required | Description | Default |
|---|---|---|---|
| wait | No | Set false to submit and return immediately with the task_id (async mode) — then poll with check_task and fetch with download_result. Recommended for long generations to avoid client-side watchdog timeouts. | |
| model | No | V5 | |
| prompt | No | Optional prompt for mashup direction | |
| taskId | No | Source task ID | |
| audioIds | Yes | Up to 2 audio IDs to mashup | |
| filename | No | ||
| download_dir | No | Absolute directory to save the file(s) into (created if missing). Defaults to the server's kie/assets/raw/. Must be absolute — the MCP server's working directory is not the caller's. | |
| max_wait_seconds | No | Override the blocking-mode polling budget in seconds (default: audio 300). Ignored when wait=false. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It discloses only that a new track is produced; it says nothing about credit consumption, whether the source tracks are preserved, or the blocking/async default behavior referenced only inside the schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, front-loaded with the core action and the required input. The 'NEW —' marker is minor noise, and otherwise every phrase earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an 8-parameter generative tool with no output schema and no annotations, this is thin. Missing are the cost/credit implications, model-version tradeoffs, and any hint of the async workflow, all of which an agent needs before invoking it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 75%, so most parameters are already documented in the schema (including the detailed wait/polling and download_dir notes). The description adds one genuinely useful semantic: audioIds must come from previous generations. Beyond that it adds nothing over the structured fields.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (mashup) and resource (up to 2 Suno tracks into one new track), which an agent can distinguish from sibling generation tools. However, it never contrasts itself with the closest siblings (extend_music, cover_audio, add_instrumental), so differentiation is left to inference.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The clause 'Provide audioIds from previous generations' gives one prerequisite (source audio must already exist), which is useful. But there is no guidance on when to mashup versus extending, covering, or layering a track, and no exclusions or alternative routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
generate_midiA
Export a Suno track to MIDI notation. Downloads .mid file to kie/assets/raw/.
| Name | Required | Description | Default |
|---|---|---|---|
| wait | No | Set false to submit and return immediately with the task_id (async mode) — then poll with check_task and fetch with download_result. | |
| taskId | Yes | Task ID of the Suno generation | |
| audioId | No | Audio ID from sunoData (optional) | |
| filename | No | ||
| download_dir | No | Absolute directory to save the file(s) into (created if missing). Defaults to the server's kie/assets/raw/. Must be absolute — the MCP server's working directory is not the caller's. | |
| max_wait_seconds | No | Override the blocking-mode polling budget in seconds (default: audio 300). Ignored when wait=false. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Discloses async behavior and download directory, but no annotations are provided. Missing details on file conflicts, naming, or authentication requirements, which would be needed for full transparency.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the main purpose. Efficient but could be slightly more structured with bullet points for parameters.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Covers core functionality and async mode but does not describe return values or error scenarios. Adequate for a simple export tool given no output schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 83%, and the description adds context for the wait parameter and directory, but most parameter meaning is already clear from the schema. Adds marginal value.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states 'Export a Suno track to MIDI notation' and specifies the output location, distinguishing it from sibling tools like generate_music or download_result.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No explicit guidance on when to use this tool versus alternatives, such as other export tools or polling mechanisms. The async mode hint is present but not contextualized.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
generate_musicA
Generate music using Suno via kie.ai. Supports V5.5 (custom style), V5 (best quality), V4.5+, V4.5, V4. Up to 8 minutes. Great for game music stems, ambient tracks, and jingles. Polls until done and downloads to kie/assets/raw/.
| Name | Required | Description | Default |
|---|---|---|---|
| wait | No | Set false to submit and return immediately with the task_id (async mode) — then poll with check_task and fetch with download_result. Recommended for long generations to avoid client-side watchdog timeouts. | |
| model | No | Suno model. V5_5=custom style, V5=best quality. Default: V5 | V5 |
| style | No | Style tags (e.g. "Celtic, orchestral, upbeat, fantasy, game music") | |
| title | No | Track title (optional) | |
| prompt | Yes | Music description (e.g. "upbeat Celtic fantasy adventure, flute and drums, heroic") | |
| filename | No | Output filename. Auto-generated if omitted. | |
| download_dir | No | Absolute directory to save the file(s) into (created if missing). Defaults to the server's kie/assets/raw/. Must be absolute — the MCP server's working directory is not the caller's. | |
| instrumental | No | No vocals when true (recommended for game music) | |
| max_wait_seconds | No | Override the blocking-mode polling budget in seconds (defaults: image 600, video 900, audio 300, speech 300). Ignored when wait=false. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, and it does disclose the key behavior: it 'polls until done and downloads to kie/assets/raw/'. However, it omits cost/credit implications, failure behavior, and the blocking-vs-async tradeoff (that detail lives only in the schema's wait param). Useful but incomplete for a paid, long-running generation call.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Short, front-loaded sentences with the core action stated first and useful details following; every clause adds something. Slightly listy but no filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, but the schema is fully documented and the description covers backend, model range, duration, purpose, and the polling/download destination, which is sufficient for an agent to invoke and interpret this tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all 9 parameters, setting the baseline at 3. The description adds only duration ('up to 8 minutes') and model tier hints that largely duplicate the model enum descriptions, so it does not meaningfully extend parameter semantics.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Generate music using Suno via kie.ai') and adds scope details (model range, up to 8 minutes) plus use cases (game music, ambient, jingles). It clearly separates itself from generative siblings like generate_sfx or generate_tts, though it never names an alternative explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Use cases ('great for game music stems, ambient tracks, and jingles') imply when this is appropriate, but there is no explicit routing against siblings such as extend_music, cover_audio, add_vocals, or generate_sfx, and no exclusions or prerequisites are given. Usage is inferable rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
generate_personaA
NEW — Create a Suno Persona (reusable music character) from an existing Suno track. Requires taskId from V3.6+ generation.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | Persona name | |
| style | No | Music style tag (e.g. "Electronic Pop") | |
| taskId | Yes | Task ID from a previous Suno generation (V3.6+) | |
| audioId | Yes | Audio ID from sunoData | |
| vocalEnd | No | End time in seconds (10-30s segment) | |
| vocalStart | No | Start time in seconds for vocal analysis | |
| description | Yes | Detailed Persona description (musical style, personality) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, description must disclose behavioral traits. It mentions 'reusable music character' and the prerequisite, but does not state side effects, auth needs, or what happens to the existing track. Minimal insight beyond basic action.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Extremely concise – two sentences. Front-loaded with 'NEW' to indicate freshness. No fluff; every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Description is functional but omits broader context: what a Suno Persona is, how to use it post-creation, or what the API returns. Given 7 parameters and no output schema, more context would aid completeness.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% (all parameters described). Description adds no extra meaning beyond the schema, e.g., the prerequisite is already in taskId description. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Explicitly states it creates a Suno Persona from an existing Suno track, with specific verb and resource. Clearly distinguishes from sibling generation tools like generate_music or cover_audio by focusing on persona creation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides a clear prerequisite (taskId from V3.6+ generation) but lacks guidance on when to use persona creation versus alternatives like cover_audio or generate_music. No explicit when-not or comparison to siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
generate_sfxA
Generate a sound effect from text via Suno V5 (kie.ai removed the ElevenLabs sound-effect model). Great for game sounds: UI clicks, magic spells, item pickups, explosions. For loop/BPM/key control use generate_sounds instead. Downloads to kie/assets/raw/.
| Name | Required | Description | Default |
|---|---|---|---|
| text | Yes | Sound description (e.g. "magical sparkle chime, fairy-like, short 0.5s") | |
| wait | No | Set false to submit and return immediately with the task_id (async mode) — then poll with check_task and fetch with download_result. Recommended for long generations to avoid client-side watchdog timeouts. | |
| filename | No | Output filename. Auto-generated if omitted. | |
| download_dir | No | Absolute directory to save the file(s) into (created if missing). Defaults to the server's kie/assets/raw/. Must be absolute — the MCP server's working directory is not the caller's. | |
| duration_seconds | No | Target duration hint, folded into the prompt (Suno has no hard duration control). | |
| max_wait_seconds | No | Override the blocking-mode polling budget in seconds (defaults: image 600, video 900, audio 300, speech 300). Ignored when wait=false. | |
| prompt_influence | No | Deprecated — ignored (no Suno equivalent). Kept for backward compatibility. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Discloses the underlying model (Suno V5) and removal of ElevenLabs model, mentions output directory (kie/assets/raw/), and via parameter descriptions explains async mode. However, with no annotations, the description does not fully cover safety, permissions, or guarantees. Provides some behavioral context but not comprehensive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Extremely concise: three sentences. Front-loaded with core action, followed by examples and a sibling distinction. No wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Describes model, use case, and alternative, but omits the return value/format (no output schema). While parameter descriptions explain async polling, the description lacks an overall picture of what the agent receives upon completion.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% with detailed parameter descriptions. The description adds value by suggesting use cases (game sounds) and directing to alternatives, but does not significantly augment the schema's per-parameter meaning.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Clearly states it generates a sound effect from text via Suno V5, with specific examples (game sounds: UI clicks, magic spells, etc.). Explicitly distinguishes from sibling generate_sounds for loop/BPM/key control.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides a clear when-to-use scenario (game sounds) and an explicit alternative (generate_sounds for loop/BPM/key control). However, lacks guidance on when not to use compared to other sound generation siblings (e.g., generate_tts, generate_dialogue).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
generate_soundsB
Generate loopable sound effects with BPM, key, and loop control via Suno. Downloads to kie/assets/raw/.
| Name | Required | Description | Default |
|---|---|---|---|
| wait | No | Set false to submit and return immediately with the task_id (async mode) — then poll with check_task and fetch with download_result. Recommended for long generations to avoid client-side watchdog timeouts. | |
| model | No | V5 | |
| prompt | Yes | Sound description (e.g. "ambient rain on a tin roof, soft thunder") | |
| filename | No | ||
| soundKey | No | Musical key (e.g. "C", "Am") | |
| soundLoop | No | Whether the sound should loop seamlessly | |
| grabLyrics | No | ||
| soundTempo | No | BPM for the sound | |
| download_dir | No | Absolute directory to save the file(s) into (created if missing). Defaults to the server's kie/assets/raw/. Must be absolute — the MCP server's working directory is not the caller's. | |
| max_wait_seconds | No | Override the blocking-mode polling budget in seconds (defaults: image 600, video 900, audio 300, speech 300). Ignored when wait=false. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It adds one genuine behavioral fact — results are downloaded to kie/assets/raw/ — but that destination is already stated in the download_dir schema description, so the marginal value is small. It does not disclose the blocking/async behavior, cost, or failure handling that a generation tool should surface.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, front-loaded with the action and resource, followed by the output location. No filler or redundancy; every clause carries information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 10-parameter tool with no annotations and no output schema, the description is only marginally sufficient. It covers purpose, three parameters, and the save destination, but omits sibling differentiation, model-selection guidance, and any behavioral context an agent would want before invoking a long-running generated-audio call.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 70%, above the midpoint, so the schema already documents prompt, soundKey, soundLoop, soundTempo, wait, download_dir, and max_wait_seconds. The description's mention of BPM, key, and loop control echoes those fields without adding syntax or constraints. Params like model, filename, and grabLyrics remain undocumented in both places, but the baseline 3 is appropriate given the schema does most of the work.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ("Generate loopable sound effects") plus the provider (Suno) and three controllable dimensions (BPM, key, loop). This is more than a restatement of the name. However, it does not distinguish itself from the very similar sibling generate_sfx (and generate_music), so an agent cannot tell which to pick from the description alone.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use or when-not-to-use guidance. With generate_sfx and generate_music in the sibling set, the agent gets no signal about when this loop-oriented sound tool is preferred over those alternatives. Usage must be inferred entirely from the name.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
generate_ttsA
Generate speech from text using ElevenLabs via kie.ai. Supports Turbo 2.5 (fast) and Multilingual V2 (high quality). Downloads to kie/assets/raw/.
| Name | Required | Description | Default |
|---|---|---|---|
| text | Yes | Text to synthesize into speech | |
| wait | No | Set false to submit and return immediately with the task_id (async mode) — then poll with check_task and fetch with download_result. Recommended for long generations to avoid client-side watchdog timeouts. | |
| model | No | turbo-2-5=fast, multilingual-v2=high quality with language support | turbo-2-5 |
| speed | No | Speech speed (0.7–1.2). Only for multilingual-v2. | |
| filename | No | Output filename. Auto-generated if omitted. | |
| voice_id | No | Voice name (e.g. "Bella", "Viking Bjorn", "Aria") or kie voice ID. kie.ai only accepts its curated ~67-voice set — arbitrary ElevenLabs voice IDs are rejected. An unknown value returns the full catalog. Optional — defaults to James. | |
| download_dir | No | Absolute directory to save the file(s) into (created if missing). Defaults to the server's kie/assets/raw/. Must be absolute — the MCP server's working directory is not the caller's. | |
| language_code | No | Language code for multilingual-v2 (e.g. "en", "es", "fr", "ja") | |
| max_wait_seconds | No | Override the blocking-mode polling budget in seconds (defaults: image 600, video 900, audio 300, speech 300). Ignored when wait=false. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It discloses model variants, async behavior, download location constraints, and voice_id limitations (curated set, rejection of arbitrary IDs). However, it does not explicitly state side effects (file creation) or auth requirements.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences that pack essential information with no redundancy. First sentence establishes core purpose and key differentiators; front-loaded for quick comprehension.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given 9 parameters and no output schema, the description covers models, async mode, and file location well. Missing details on return value (e.g., file path vs task ID) and output file format (e.g., .mp3). Adequate but not exhaustive.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 100% schema coverage, the description adds significant value beyond the schema: explains wait rationale, voice_id caveats, download_dir absolute path requirement, and model selection advice. Each parameter gets meaningful context.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb 'Generate speech from text', identifies the technology stack (ElevenLabs via kie.ai), and specifies output destination. This distinguishes it well from sibling tools like generate_music and generate_sfx.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides some usage context (model selection, async mode via wait parameter) but lacks explicit guidance on when to prefer this tool over alternatives like generate_gemini_tts. No comparative analysis with sibling tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
generate_videoA
Generate a video using kie.ai (85+ models). Downloads to kie/assets/raw/. MODEL GUIDE: Best cinematic→veo-3/text-to-video (50cr/s, audio). Fast+cheap→grok-imagine-video-1-5-preview (1.6-3cr/s, audio, NEW), wan/flash-image-to-video (6-8cr/s measured; alias of wan/2-6-flash). Budget cinematic→hailuo-standard (4cr/s). First→last-frame or anything-from-anything refs→gemini-omni/flash-1-1 (NEW, est. ~63cr per 4s clip). Budget multimodal refs→bytedance/seedance-2-mini (9.5cr/s @480p). 30s single takes→bytedance/seedance-2-5 (NEW). Budget all-rounder w/ audio+templates+extend→pixverse-v6 family (4-9.6cr/s, NEW; I2V is its strength; transition=first/last-frame morph). Multilingual lip-synced dialogue→happyhorse-1-1 T2V/I2V/R2V (NEW). 2K + stereo audio→minimax-h3 (8cr/s @768P, price halved Sept 2026). Per-shot scripted multi-shot→kling-3-omni (14cr/s @720p, NEW; transformation=restyle existing video). Next-gen Wan draft→wan/3-0-video (8cr/s @480P, NEW). Fast Kling→kling/v3-turbo (18cr/s, audio, NEW). Image-to-video→veo-3/image-to-video (include a sound cue like "SFX: room tone" — Veo I2V intermittently fails its audio pass without one; images must be served with their real Content-Type, upload via upload_file), kling/image-to-video. Avatar/talking head→omnihuman-1-5 (premium, NEW), kling/ai-avatar-pro, infinitalk/from-audio. Re-dub existing footage→volcengine/video-to-video-lip-sync (8cr/s, NEW). Motion control→kling/motion-control, wan/animate-move. Extend video→use veo_extend or runway_extend tools. NOTE: Sora 2 family removed (OpenAI API sunset Sept 2026). Use list_models filter="use-case" to explore.
| Name | Required | Description | Default |
|---|---|---|---|
| wait | No | Set false to submit and return immediately with the task_id (async mode) — then poll with check_task and fetch with download_result. Recommended for long generations to avoid client-side watchdog timeouts. | |
| model | No | Model ID (e.g. "veo-3/text-to-video", "kling/image-to-video", "wan/3-0-video") | veo-3/text-to-video |
| prompt | Yes | Video description prompt | |
| filename | No | Output filename (saved to kie/assets/raw/). Auto-generated if omitted. | |
| image_urls | No | Input image URLs for image-to-video models | |
| aspect_ratio | No | Aspect ratio: 16:9, 9:16, or 1:1 | 16:9 |
| download_dir | No | Absolute directory to save the file(s) into (created if missing). Defaults to the server's kie/assets/raw/. Must be absolute — the MCP server's working directory is not the caller's. | |
| model_options | No | Model-specific options (duration, resolution, mode, etc.) | |
| max_wait_seconds | No | Override the blocking-mode polling budget in seconds (defaults: image 600, video 900, audio 300, speech 300). Ignored when wait=false. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral burden and does well: it discloses per-second credit pricing, which models include audio, a documented failure mode ('Veo I2V intermittently fails its audio pass without one'), and a hard precondition (images must be uploaded via upload_file with real Content-Type). It doesn't cover error handling or what wait=true actually returns, so it stops short of a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The purpose is front-loaded, but the remainder is a single dense wall of arrow-delimited routing text rather than scannable structure. Most sentences carry model-selection value, though the volume partly overlaps with what list_models exists to provide.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 9-parameter tool with nested model_options and no output schema, the description covers the essentials: destination directory, model-selection guidance, upload prerequisite, and failure caveats. It leaves minor gaps around the blocking-mode return value, but nothing that would cause a wrong invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the guide adds real meaning to the `model` parameter that the schema's one-line example cannot (which model fits which use case and cost). It also enriches `image_urls` by mandating upload_file, going beyond the schema's bare 'Input image URLs'.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening sentence states a specific verb and resource ('Generate a video'), the backend (kie.ai, 85+ models), and the output location (kie/assets/raw/). It also explicitly routes adjacent work elsewhere ('Extend video→use veo_extend or runway_extend tools'), so an agent can separate it from the extend/upscale siblings without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The MODEL GUIDE is effectively a when-to-use decision tree: cinematic, budget, fast, lip-sync, avatar, motion control, multi-shot, and image-to-video each map to named models. It names alternatives and their selection conditions (e.g. 'Extend video→use veo_extend'), and points to list_models for exploration.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_timestamped_lyricsA
NEW — Get word-level timestamped lyrics from a Suno track. Useful for karaoke, captioning, or sync.
| Name | Required | Description | Default |
|---|---|---|---|
| taskId | Yes | Suno task ID | |
| audioId | Yes | Audio ID from sunoData |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of disclosing behavioral traits. It only states what the tool returns (timestamped lyrics) but does not mention side effects, idempotency, network calls, or any caveats. For a tool that likely performs an API fetch, the description lacks necessary behavioral transparency.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is extremely concise: one short sentence plus a 'NEW' prefix. It is front-loaded with the core functionality and uses, wasting no words. Every part earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no output schema, the description should compensate by explaining the return format (e.g., structure of timestamped words). It mentions use cases but omits what the agent will actually receive. For a simple 2-param tool, the missing output details reduce completeness.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% (both parameters have meaningful descriptions like 'Suno task ID' and 'Audio ID from sunoData'). The description adds no additional parameter information beyond the schema, so it meets the baseline of 3.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's function: 'Get word-level timestamped lyrics from a Suno track.' It uses a specific verb ('Get'), specifies the resource ('word-level timestamped lyrics'), and names the source ('Suno track'). This distinguishes it from siblings like 'generate_lyrics' (which creates lyrics) and 'speech_to_text' (which transcribes audio), though it does not explicitly contrast them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides context for when to use the tool: 'Useful for karaoke, captioning, or sync.' However, it does not offer guidance on when not to use it or mention alternative tools. Given the large set of sibling tools, explicit exclusions or comparisons would improve this dimension.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
grok_image_editA
Edit an image with Grok Imagine Image 2.0 (4 credits). TWO MODES: (1) Region mode — pass task_id (a prior Grok 2.0 generation) + mask_indexs from grok_segment_map (run it first, free, pick regions by NAME): only those regions change. (2) Whole-image mode (NEW Aug 2026) — pass image_urls (ANY uploaded/external image, e.g. from upload_file) + aspect_ratio + prompt: instruction-based edit of the full image, no segmentation. Returns a new full image; result task_ids chain back into segment/edit for iterative refinement. Downloads to kie/assets/raw/.
| Name | Required | Description | Default |
|---|---|---|---|
| prompt | Yes | Region mode: what the masked region(s) should become plus what to preserve. Whole-image mode: the edit instruction for the full image. | |
| task_id | No | Region mode: source task ID — a Grok Image 2.0 generation (or a previous grok_image_edit result). Mutually exclusive with image_urls. | |
| filename | No | Output filename. Auto-generated if omitted. | |
| image_urls | No | Whole-image mode: 1-5 public URL(s) of the image(s) to edit — any image, not just Grok generations (upload local files with upload_file first). Mutually exclusive with task_id. | |
| mask_indexs | No | Region mode only: region indices from grok_segment_map (e.g. [1] or [0, 2]). Field name matches kie's API spelling. | |
| aspect_ratio | No | Whole-image mode: required output aspect ratio (auto keeps the input's shape). | |
| download_dir | No | Absolute directory to save the file(s) into (created if missing). Defaults to the server's kie/assets/raw/. Must be absolute — the MCP server's working directory is not the caller's. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden and does much of it: it states the credit cost (4), the dependency chain on grok_segment_map, chaining semantics for returned task_ids, and the file download location. It does not cover auth requirements, rate limits, or failure behavior, leaving modest gaps for a zero-annotation mutation tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with action, cost, and the two modes as numbered branches; each sentence carries routing or behavioral payload. Density is justified by the two-mode complexity and nothing is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description compensates by describing the return (a new full image) and how returned task_ids chain into further segment/edit calls. It covers both modes and the prerequisite tool, but could say more about failure modes or what happens when neither mode's parameters are supplied.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents every parameter, including mutual exclusivity and the aspect_ratio default. The description adds mode-grouping context that helps interpret the parameters together, but no format or syntax detail beyond the schema, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Edit an image with Grok Imagine Image 2.0') and explicitly enumerates its two operating modes with the parameters that select each. An agent can distinguish it from siblings like grok_segment_map or generate_image without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly prescribes when to use each mode (region mode = task_id + mask_indexs; whole-image mode = image_urls + aspect_ratio + prompt) and gives an ordering dependency, instructing the agent to run grok_segment_map first and pick regions by NAME. This is a routing-quality guideline, not just context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
grok_segment_mapA
FREE (0 credits). Segment a Grok Imagine Image 2.0 generation into NAMED regions for targeted editing. Returns each region's index, semantic name (e.g. "red apple", "wooden table"), and mask PNG URL. Workflow: generate_image model="grok-imagine-image-2-0/text-to-image" → grok_segment_map (this, free) → grok_image_edit with the mask_indexs you want changed. Only works on task_ids from a Grok Image 2.0 generation.
| Name | Required | Description | Default |
|---|---|---|---|
| task_id | Yes | Task ID from a completed generate_image call with model grok-imagine-image-2-0/text-to-image |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are present, but the description discloses key behaviors: it returns region index, semantic name, and mask PNG URL, and notes it is free. This adds useful context beyond the bare schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is slightly verbose but well-structured, covering purpose, usage, output, and cost in a logical flow. It is not excessively redundant and remains readable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given only one parameter and a clear schema, the description sufficiently explains the tool's role in the workflow, its input constraints, and its output format, making it complete for an agent to decide when and how to use it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema already describes the task_id parameter in full (Task ID from a completed generate_image call with model grok-imagine-image-2-0/text-to-image). The description mainly repeats this requirement without adding new parameter-specific semantics.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Explicitly states the tool segments a Grok Imagine Image 2.0 generation into named regions for targeted editing, with a clear verb and resource. It is distinct from sibling tools like generate_image and grok_image_edit.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides clear usage conditions: only works on task_ids from a Grok Image 2.0 generation, and outlines a recommended workflow (generate_image -> grok_segment_map -> grok_image_edit).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_modelsA
List all available kie.ai models with their aspect ratios and model-specific options
| Name | Required | Description | Default |
|---|---|---|---|
| filter | No | Filter models by name (e.g. "flux", "gpt", "seedream") | |
| verbose | No | Show full option details for each model |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations provided, so description carries full burden. It clearly states the action (list) and output content (models with options). No mention of pagination or rate limits, but for a read-only list tool, this is sufficient and not misleading.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Single sentence that conveys purpose. It is concise but could be slightly more structured (e.g., mentioning that it returns a list). No wasted words, but could include more detail.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, so description should clarify return structure. It mentions 'aspect ratios and model-specific options', which gives a hint but not full details. For a simple list tool, it is adequate but not thorough.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% (both parameters are described). The description does not add new meaning beyond the schema; it echoes 'aspect ratios and options' which aligns with parameter 'verbose'. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool lists all available kie.ai models and includes aspect ratios and options. It distinguishes from sibling tools like list_tasks and check_credits by focusing on models.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No explicit guidance on when to use this tool versus alternatives. Usage is implied (e.g., before generation), but no conditions or exclusions are provided. Sibling tools are diverse, so guidance would be helpful.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_raw_assetsA
List all files in kie/assets/raw/ waiting to be processed
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries full burden. It implies a read-only, non-destructive operation. However, it does not disclose details like permissions required, behavior when the directory is empty, or if any filtering is applied.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Single sentence, no redundancy, front-loads key information. Every word serves a purpose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no output schema, the description should explain return format. It mentions 'list all files' but not the output structure (e.g., strings, objects). For a simple tool with no params, it is adequate but incomplete for a full understanding.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There are zero parameters, so the description's job is minimal. Baseline is 4. The description adds no extra meaning beyond the schema, but no value is needed.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('List all files'), the target resource ('in kie/assets/raw/'), and the context ('waiting to be processed'). It distinguishes from sibling 'list_tasks' by specifying a directory and state.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description lacks explicit guidance on when to use this tool versus alternatives. No mention of prerequisites, exclusion criteria, or comparison with sibling tools like 'list_tasks'.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_tasksB
List recent image generation tasks from this session
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description must disclose behavioral traits. It mentions listing recent tasks from the session, but does not clarify what 'recent' means, the nature of the session, or whether the operation has any side effects (though likely read-only). Minimal transparency.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, concise sentence with no extraneous information. It is front-loaded with the verb and resource, making it easy to scan.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple list tool with one parameter and no output schema, the description is minimally adequate but missing parameter explanation and clarification of 'recent' and 'session'. It does not reference return format or pagination. Given the simplicity, a score of 3 is appropriate.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description does not mention the 'limit' parameter at all. With 0% schema description coverage, the agent has no additional context about what the limit does or how it affects results beyond the schema default of 10.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb 'List' and the resource 'recent image generation tasks from this session'. It distinguishes from siblings like 'check_task' (which likely checks a single task) and various generation tools by specifying the scope of recent session tasks.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No explicit guidance on when to use this tool versus alternatives like 'check_task'. The description implies it is for listing, but does not state when not to use it or what to do if you need a different perspective.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
prepare_voice_cloneA
EXPERIMENTAL (#20) — STEP 1 of Suno custom-voice cloning (FREE). Submit a clean vocal sample; polls to wait_validating, after which kie sends a verification phrase to your KIE_CALLBACK_URL for the voice owner to read aloud (then use create_voice_clone). ⚠️ Only clone a voice you OWN or have explicit permission to use. The completion step is not verified end-to-end and needs a real callback URL.
| Name | Required | Description | Default |
|---|---|---|---|
| wait | No | Set false to submit and return the task_id immediately. | |
| language | No | Language of the vocal (e.g. "en") | en |
| voice_url | Yes | Public URL of a clean vocal sample (the voice to clone). Must be reachable by kie servers. | |
| vocal_end_s | No | End of the vocal segment, seconds | |
| vocal_start_s | No | Start of the vocal segment, seconds | |
| max_wait_seconds | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description fully carries behavioral disclosure: it notes the tool is experimental (#20), the two-phase process (submit + callback), and caveats about the completion step not being verified end-to-end and requiring a real callback URL. This is thorough for a tool with no annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three sentences with no wasted words. It front-loads the tool's identity (experimental, step 1, free) and efficiently covers the process, warnings, and limitations.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (6 parameters, no output schema), the description explains the workflow and key risks adequately. It could detail error scenarios or return values, but for a step-1 tool, the provided information is sufficient for an agent to use it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 83%, so baseline is 3. The description adds context beyond schema by explaining the polling behavior for 'wait' and 'max_wait_seconds', and the callback server expectation. It does not detail each parameter but provides workflow context that aids usage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states it is 'STEP 1 of Suno custom-voice cloning' and describes the action: submit a clean vocal sample, poll to wait_validating, and trigger a callback. It distinguishes from its sibling 'create_voice_clone' which is the next step.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly marks the tool as step 1 and directs to use 'create_voice_clone' afterward. It includes a warning about only cloning owned voices. However, it lacks explicit 'when not to use' or alternative tools for other cases.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
profile_briefA
Get a vertical playbook before generating: the intake questions a professional in that domain would ask, model routing per deliverable with live costs, per-model prompt formulas, and multi-tool workflows. Call with no args to list available profiles. Covers image, video, and audio verticals (architecture, game assets, advertising, product photography, film, brand, web product, editorial, short-form social video, audio branding). Call this FIRST when the user's request belongs to a known vertical — then ask the user only the unanswered intake questions, conversationally.
| Name | Required | Description | Default |
|---|---|---|---|
| profile | No | Profile id (omit to list profiles, or pass only `request` to auto-infer) | |
| request | No | The user's request verbatim — used to auto-infer the profile and echoed into the brief so intake can be filtered |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It discloses what is returned, that costs are live, that the profile auto-infers from `request`, and that zero-arg calls enumerate profiles — enough to imply a read-only, advisory tool. It stops short of explicitly stating there are no side effects or that the brief is a snapshot versus a live fetch, so a small gap remains.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core purpose and the when-to-call rule in the first two sentences; every sentence earns its place. The long parenthetical enumeration of verticals is the bulkiest part — useful scope information, but it bloats the text and could be trimmed.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only advisory tool with two fully documented optional params and no output schema, the description covers purpose, trigger condition, degenerate call, and return contents. Nothing critical is missing, though the absence of any note on whether the brief is cached or live keeps it short of complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds real meaning: it confirms that omitting everything lists profiles and that `request` drives auto-inference and is echoed back so intake can be filtered. That interaction semantics goes beyond the schema's field text.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Get') and a specific resource (a vertical playbook) and then enumerates exactly what that playbook contains: intake questions, model routing with live costs, per-model prompt formulas, multi-tool workflows. The vertical coverage list nails the scope, and it is clearly distinguishable from every sibling, which are all generation/listing tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives explicit ordering guidance ('Call this FIRST when the user's request belongs to a known vertical'), an explicit degenerate call ('Call with no args to list available profiles'), and a follow-up workflow ('ask the user only the unanswered intake questions, conversationally'). An agent knows precisely when to reach for this instead of going straight to generate_image/generate_video.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
regenerate_voice_cloneB
EXPERIMENTAL (#20) — retry a failed/incomplete custom-voice task by its task_id.
| Name | Required | Description | Default |
|---|---|---|---|
| task_id | Yes | Voice task_id to retry |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description bears full responsibility but only notes it's experimental (version #20) and does not disclose side effects, such as whether retrying overwrites previous results or the nature of the response.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence, front-loading the experimental nature and core action with no wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple tool with one parameter and no output schema, the description covers the purpose and parameter but omits details about return value, idempotency, or what happens on retry.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3. The description adds the qualifier 'failed/incomplete' to the parameter, providing marginal extra meaning beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb 'retry' and the resource 'custom-voice task', specifies it is for failed/incomplete tasks, and distinguishes it from sibling tools like create_voice_clone and prepare_voice_clone.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage when a custom-voice task failed or is incomplete, but does not explicitly state when not to use it or mention alternative tools for similar purposes.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
replace_sectionC
Replace a time range in a Suno track with new AI-generated content.
| Name | Required | Description | Default |
|---|---|---|---|
| tags | No | ||
| wait | No | Set false to submit and return immediately with the task_id (async mode) — then poll with check_task and fetch with download_result. Recommended for long generations to avoid client-side watchdog timeouts. | |
| title | No | ||
| prompt | Yes | Prompt for the replacement section | |
| taskId | Yes | Task ID of the original Suno generation | |
| audioId | Yes | Audio ID from sunoData | |
| filename | No | ||
| fullLyrics | No | Full lyrics for context | |
| infillEndS | Yes | End time in seconds | |
| download_dir | No | Absolute directory to save the file(s) into (created if missing). Defaults to the server's kie/assets/raw/. Must be absolute — the MCP server's working directory is not the caller's. | |
| infillStartS | Yes | Start time in seconds | |
| negativeTags | No | ||
| max_wait_seconds | No | Override the blocking-mode polling budget in seconds (default: audio 300). Ignored when wait=false. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries full burden for behavioral disclosure. It only states the basic operation ('replace a time range') without revealing side effects (e.g., whether the original track is modified, whether replacement is permanent, or any authentication or rate-limit considerations).
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, well-formed sentence that front-loads the core purpose. However, it could be slightly restructured to mention key parameters or usage notes without adding verbosity, so it earns a 4 rather than a 5.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the complexity of the tool (13 parameters, 5 required, no output schema, no annotations), the one-sentence description is insufficient. It does not explain return values, error states, or how the replacement process works (e.g., whether the rest of the track is preserved).
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description adds no additional meaning beyond the input schema. While schema description coverage is 69%, meaning some parameters are partly documented in the schema, the tool description itself does not explain the purpose or relationships of the parameters (e.g., how 'infillStartS' and 'infillEndS' define the target range).
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Replace'), names the resource ('a time range in a Suno track'), and specifies the action ('with new AI-generated content'). It clearly distinguishes this tool from siblings like 'extend_music' or 'cover_audio' which modify the full track or add sections rather than replace a segment.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives, nor any exclusions or prerequisites. For example, it does not mention that the original track must exist or that replacement may require specific credits.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
runway_extendB
Extend an existing Runway Aleph video with continuation content.
| Name | Required | Description | Default |
|---|---|---|---|
| prompt | Yes | Description of what happens in the extension | |
| quality | No | 720p | |
| task_id | Yes | Task ID from original Runway generation | |
| filename | No | ||
| download_dir | No | Absolute directory to save the file(s) into (created if missing). Defaults to the server's kie/assets/raw/. Must be absolute — the MCP server's working directory is not the caller's. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description must fully disclose behavioral traits. It only states the purpose without mentioning whether the operation is destructive, requires authorization, has rate limits, or any side effects. Minimal transparency for a mutation tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, well-formed sentence with no extraneous words. It is front-loaded with the action and resource. Could be slightly more informative but remains concise and clear.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has 5 parameters, no output schema, and sibling tools for other platforms, the description is incomplete. It lacks details on return values, error handling, or differentiation from similar extension tools like veo_extend. More information is needed for an agent to use it correctly without confusion.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema already describes most parameters (prompt, task_id, quality, download_dir) with coverage estimated at 60%. The description adds no additional meaning beyond what the schema provides. For filename, the schema lacks description, but the description does not compensate. Baseline of 3 is appropriate given moderate schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool extends an existing Runway Aleph video with continuation content, specifying the verb (extend) and resource (existing video). It distinguishes from sibling tools like generate_video (creation) and veo_extend (Veo platform).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No explicit when-to-use or when-not-to-use guidance is provided. The description implies usage only for extending existing Runway Aleph videos but does not mention alternatives or exclusions. The context of sibling tools helps, but the description itself lacks explicit guidelines.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
seedream_layer_decomposeA
Split ANY image into independent layers with Seedream 5.0 (billed per OUTPUT layer incl. base). tier="pro" (default): 7 cr @1K, 14 @2K, so a 3-layer split ≈ 21 cr. tier="flash" (NEW): 3.24 cr/layer at any size, so ≈ 10 cr. Works on any public image URL (upload local files with upload_file first). Describe which elements become layers in the prompt, optionally bounding them with x1 y1 x2 y2 tags. Downloads every layer image to kie/assets/raw/.
| Name | Required | Description | Default |
|---|---|---|---|
| size | No | auto, 1K, 1.5K or 2K. Pro: 7 cr/layer at 1K/1.5K, 14 at 2K. Flash: 3.24 at any size. | auto |
| tier | No | pro = Seedream 5.0 Pro (7 cr/layer @1K, 14 @2K; prompt required). flash = Seedream 5.0 Flash (NEW: 3.24 cr/layer at any size; prompt optional — but WITHOUT a prompt it splits out every element: a poster gave 12 layers = 38.9 cr. Name the elements you want to cap the layer count). | pro |
| prompt | No | Which elements to separate into layers, e.g. "Separate the title text <bbox>179 58 809 197</bbox> and the parrot <bbox>330 274 641 991</bbox> into independent layers" | |
| filename | No | Base output filename; layers get -2, -3... suffixes | |
| image_url | Yes | Public URL of the source image (singular — one image per call) | |
| download_dir | No | Absolute directory to save into (created if missing). Defaults to the server's kie/assets/raw/. | |
| output_format | No | png recommended for transparency | png |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does well: it discloses the per-output-layer credit model, the exact flash-vs-pro cost tradeoff, the 12-layer/38.9 cr failure mode of promptless flash, and the download destination (kie/assets/raw/). It omits call mechanics an agent needs, such as whether this is synchronous or returns a task id for check_task, and whether existing files are overwritten.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the action and cost, then tiers, then input/output mechanics. It is a dense single paragraph with heavy billing detail, but each clause carries operational information rather than filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Seven parameters are all covered and the cost model plus download location are explained, which matters since there is no output schema. The remaining gap is execution semantics (sync vs task-based, failure behavior), which an agent would need for a multi-credit generative call.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description restates the tier/size pricing and the <bbox> tag syntax already present in the schema's prompt description; its only additive detail is the '-2, -3...' filename suffix convention, which is marginal.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Split ANY image into independent layers with Seedream 5.0'. That is unambiguous and distinct from every sibling tool (audio, video, TTS, upscaling), so an agent can route to it without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives real routing guidance: local files must go through upload_file first, prompt is required for tier=pro but optional for tier=flash, and omitting the prompt on flash causes unbounded layer counts. It doesn't state hard exclusions (e.g. no video input), so it falls just short of a full when/when-not matrix.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
separate_vocalsB
Separate vocals from instrumentals, or split into individual stems. Downloads to kie/assets/raw/.
| Name | Required | Description | Default |
|---|---|---|---|
| type | No | split_stem_advanced = NEW finer multi-stem split (market API). separate_vocal=vocals+instrumental, split_stem=individual instruments | separate_vocal |
| wait | No | Set false to submit and return immediately with the task_id (async mode) — then poll with check_task and fetch with download_result. | |
| taskId | Yes | Task ID of the Suno generation | |
| audioId | Yes | Audio ID from sunoData | |
| filename | No | ||
| download_dir | No | Absolute directory to save the file(s) into (created if missing). Defaults to the server's kie/assets/raw/. Must be absolute — the MCP server's working directory is not the caller's. | |
| max_wait_seconds | No | Override the blocking-mode polling budget in seconds (default: audio 300). Ignored when wait=false. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It does disclose that output lands in kie/assets/raw/, which is genuine context, but it omits the long-running/async nature, polling expectations, and whether the source audio is altered. The schema covers wait/poll behavior, but the description adds only the output-location detail.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, front-loaded with the core operation. The trailing fragment 'Downloads to kie/assets/raw/.' is terse but useful; overall there is essentially no wasted text.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 7-parameter, no-annotation, no-output-schema tool, the description is thin: it never explains what the caller receives or references the check_task/download_result polling flow. The schema picks up most of the slack on parameters, but the description leaves behavioral and post-invocation context incomplete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 86%, well above the high-coverage threshold, so the schema already documents type, wait, download_dir, and max_wait_seconds thoroughly. The description adds no parameter-level detail beyond restating that files go to the default raw directory, so baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States specific verbs and resources: separating vocals from instrumentals or splitting into individual stems. It's clear what the tool does, but it never differentiates itself from the nearby sibling 'audio_isolation' or 'convert_to_wav', which an agent could easily confuse for the same job.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description tells you the two modes exist but gives no guidance on when to choose this over adjacent tools like audio_isolation. The mode distinctions it does mention (separate vocal vs split stems) are already spelled out in the type enum description, so no routing value is added.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
speech_to_textA
Transcribe audio to text using ElevenLabs Scribe v1. Supports diarization and audio event tagging. Returns transcription text.
| Name | Required | Description | Default |
|---|---|---|---|
| diarize | No | Identify different speakers | |
| audio_url | Yes | Audio URL to transcribe | |
| language_code | No | Language code (e.g. "en", "es"). Auto-detected if omitted. | |
| tag_audio_events | No | Tag non-speech audio events (laughter, music, etc.) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries full burden for behavioral disclosure. It mentions supported features (diarization, audio event tagging) but omits important details: required authentication, limits on audio length or file size, latency, and whether the operation is destructive or read-only. The description provides moderate transparency but lacks depth.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences long, entirely front-loaded with the core purpose. Every sentence adds value: first states the action and model, second lists supported features and output. No wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no output schema and no annotations, the description covers the main purpose and supported features. However, it lacks details on output format (e.g., structure of transcription text, timestamps), error handling, and constraints (e.g., audio length, supported languages). It is fairly complete for a straightforward transcription tool but could be richer.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3. The description mentions diarization and audio event tagging, which correspond to boolean parameters, adding slight context. However, it does not elaborate on parameter formats, defaults, or interactions beyond what the schema already provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb 'Transcribe' and the resource 'audio to text', specifying the model (ElevenLabs Scribe v1) and key features (diarization, audio event tagging). It effectively distinguishes this tool from sibling tools like audio_isolation or generate_tts.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus alternatives such as audio_isolation or separate_vocals. The description does not mention any prerequisites, constraints, or use cases, leaving the AI agent without decision-making context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
upload_extend_audioA
NEW — Extend uploaded audio (NOT a Suno track) with new AI-generated content. For Suno tracks, use extend_music instead.
| Name | Required | Description | Default |
|---|---|---|---|
| wait | No | Set false to submit and return immediately with the task_id (async mode) — then poll with check_task and fetch with download_result. Recommended for long generations to avoid client-side watchdog timeouts. | |
| model | No | V5 | |
| style | No | Style tags for the extension | |
| title | No | ||
| prompt | No | LYRICS for the extended section (sung as written); omit and set instrumental=true for an instrumental extension. | |
| filename | No | ||
| uploadUrl | Yes | URL of the audio file to extend | |
| continueAt | Yes | REQUIRED in practice: second where the extension starts (> 0 and < the upload's length). kie fails every request without it. | |
| download_dir | No | Absolute directory to save the file(s) into (created if missing). Defaults to the server's kie/assets/raw/. Must be absolute — the MCP server's working directory is not the caller's. | |
| instrumental | No | ||
| max_wait_seconds | No | Override the blocking-mode polling budget in seconds (default: audio 300). Ignored when wait=false. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden, yet it says nothing about this being an asynchronous task submission, expected generation latency, polling/retrieval flow, or failure modes. The async/polling behavior only appears incidentally inside parameter descriptions (wait, max_wait_seconds), not in the tool description.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, front-loaded with the differentiator and immediately followed by the alternative. No wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an 11-parameter, no-annotation, no-output-schema tool, the description covers only differentiation. Schema parameter text compensates for the async flow, but the description leaves the agent without any hint of the required-vs-optional reality of continueAt or the long-running nature of the call.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 11 parameters and only 64% schema description coverage, several inputs (model, title, filename, instrumental) have no description anywhere. The tool description contributes no parameter meaning at all — notably it never mentions that continueAt is effectively required and must fall inside the upload's duration.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (extend) and resource (uploaded audio, explicitly not a Suno track), and names the sibling it is not — extend_music — so an agent can route without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly gives both the when (uploaded audio) and the when-not (Suno tracks) plus the alternative tool to use in that case. Routing is fully determined.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
upload_fileA
Upload a file to kie.ai and get a public URL back. Image files are named by their REAL format (magic bytes), so a JPEG saved as .png is uploaded as .jpg — kie serves files with a Content-Type taken from the name, and Veo I2V rejects a mismatch. Use this to upload local images/audio/video before passing them to generation tools (image-to-image, image-to-video, reference/ingredient inputs). PREFER file_path for local files. Files expire after 3 days (kie temp storage).
| Name | Required | Description | Default |
|---|---|---|---|
| file_url | No | URL of file to upload — must be PUBLICLY reachable by kie.ai servers (no localhost/private IPs, no auth-gated or expired links). For local files use file_path | |
| file_name | No | Custom filename (optional) | |
| file_path | No | Absolute path to a local file on the machine running this MCP server (the normal case for stdio setups). The server reads and streams the bytes itself — reliable at any size, unlike base64_data. PREFERRED for local files. | |
| base64_data | No | Base64-encoded file data — raw base64 or a full data: URI. Whitespace and base64url are normalized and the data:<mime>;base64, prefix is stripped automatically (its MIME infers the extension if file_name is omitted). WARNING: payloads above ~10-12K chars (observed ceiling ~11.7K, #68) are silently truncated in transit as a tool argument — use file_path for local files; base64_data is a fallback for remote/HTTP-mode callers with small payloads. | |
| upload_path | No | Storage directory (e.g. "images", "audio", "video") | uploads |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and discloses important behavior: files expire after 3 days, image files are named by real magic-byte format rather than extension, Content-Type follows the name, and Veo I2V rejects mismatches. This is rich operational context beyond the schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with purpose and stays dense without obvious filler. It is slightly run-on and could be structured more clearly, but every sentence contributes relevant information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no output schema and no annotations, the description covers what matters for correct invocation: return value, expiry behavior, naming/Content-Type caveats, intended usage, and parameter preference. The fully documented schema covers the remaining parameter details.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents every parameter in detail. The description only restates the preference for file_path over other input methods, adding little meaning beyond what the schema provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: upload a file to kie.ai and get a public URL back. It distinguishes itself from generation siblings by framing itself as a preparatory upload utility used before passing files to generation tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It clearly says to use this before passing local images/audio/video to generation tools and explicitly prefers file_path for local files, which routes the agent among input parameters. It does not mention when not to use it or name a sibling upload alternative such as upload_extend_audio.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
veo_extendC
Extend an existing Veo 3.1 video with additional content. Requires taskId from a previous Veo generation.
| Name | Required | Description | Default |
|---|---|---|---|
| model | No | fast | |
| seeds | No | Random seed (10000-99999) for variation control | |
| prompt | Yes | Description of what happens in the extension | |
| task_id | Yes | Task ID from original Veo generation | |
| filename | No | ||
| download_dir | No | Absolute directory to save the file(s) into (created if missing). Defaults to the server's kie/assets/raw/. Must be absolute — the MCP server's working directory is not the caller's. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations provided, and the description does not disclose behavioral traits such as whether the original video is modified, permissions needed, output format, or error handling. The description carries full burden and fails to inform.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with clear front-loaded purpose. No extraneous content. However, key information is missing, reducing the value of conciseness.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no output schema, no annotations, and moderate complexity (6 params, many siblings), the description is incomplete. It lacks details on expected output, duration, effect on original, and error scenarios.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 67% (4 of 6 parameters described in schema). The description adds no additional parameter context beyond the prerequisite. It does not explain the meaning of 'prompt', 'model', 'seeds', 'filename', or 'download_dir' in relation to the extension process.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool extends an existing Veo 3.1 video with additional content, specifying the prerequisite (taskId). It distinguishes from generation tools but does not explicitly differentiate from other extension tools like runway_extend.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Only mentions the prerequisite of having a taskId from a previous Veo generation. No guidance on when to use this tool over alternatives (e.g., generate_video, runway_extend) or when not to use it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
veo_upscale_1080pA
Upscale a Veo 3.1 video to 1080p resolution. Requires taskId from a completed Veo generation.
| Name | Required | Description | Default |
|---|---|---|---|
| index | No | Video index if multiple outputs | |
| task_id | Yes | Task ID from completed Veo generation | |
| filename | No | ||
| download_dir | No | Absolute directory to save the file(s) into (created if missing). Defaults to the server's kie/assets/raw/. Must be absolute — the MCP server's working directory is not the caller's. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided. The description only states the action and prerequisite without disclosing behavioral traits like whether it is destructive, rate limits, or side effects. Carries minimal burden for transparency.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Single sentence that is short and to the point, containing no wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description explains the purpose and prerequisite, but lacks details on output/return values and behavioral impact. Suitable for a simple tool but still incomplete given no annotations or output schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 75%. The description adds the context that the task must be completed, which is not in the schema. For other parameters, it adds no additional meaning beyond the schema descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action (upscale), the target (Veo 3.1 video), the resolution (1080p), and a prerequisite (taskId from completed generation). It distinguishes from sibling tools like veo_upscale_4k.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly requires a task ID from a completed Veo generation, providing clear context for when to use. However, it does not explicitly contrast with veo_upscale_4k or mention when not to use.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
veo_upscale_4kA
Upscale a Veo 3.1 video to 4K resolution. Takes 5-10 minutes. Requires taskId from completed Veo generation.
| Name | Required | Description | Default |
|---|---|---|---|
| index | No | Video index if multiple outputs | |
| task_id | Yes | Task ID from completed Veo generation | |
| filename | No | ||
| download_dir | No | Absolute directory to save the file(s) into (created if missing). Defaults to the server's kie/assets/raw/. Must be absolute — the MCP server's working directory is not the caller's. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the burden of behavioral disclosure. It mentions the operation takes 5-10 minutes and requires a prerequisite taskId, providing some transparency about latency and dependency. However, it does not disclose potential side effects, authorization needs, or result format, leaving gaps.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is extremely concise, consisting of two sentences that deliver the core purpose, time expectation, and prerequisite. Every word serves a purpose with no redundancy or filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has four parameters and no output schema, the description is incomplete. It does not explain how to handle the output (e.g., where the upscaled video is saved) or the role of parameters like 'index' and 'download_dir'. While the purpose and prerequisite are clear, additional context about workflow integration is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema coverage is 75% (three of four parameters have descriptions), but the tool description adds no new information about how to use the parameters. For example, 'filename' lacks a schema description and is not explained in the description. The description does not compensate for missing parameter details, so value beyond schema is minimal.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states this tool upsamples a Veo 3.1 video to 4K resolution, distinguishing it from the sibling 'veo_upscale_1080p' which targets a different resolution. The verb 'Upscale' and resource 'Veo 3.1 video to 4K' are specific and actionable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides a prerequisite ('Requires taskId from completed Veo generation') and a time estimate, implying it should be used after a Veo generation. However, it does not explicitly contrast with alternatives like 'veo_upscale_1080p' or state when not to use this tool. The usage context is clear but lacks exclusion guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
10 tool updates
v5.5.0- Changed
add_instrumental1 field changed- changed
Input schema / properties / model / enumPrevious value: -[ - "V3_5", - "V4", - "V4_5", - "V4_5PLUS", - "V4_5ALL", - "V5", - "V5_5" -]New value: +[ + "V4", + "V4_5", + "V4_5PLUS", + "V4_5ALL", + "V5", + "V5_5", + "V6", + "V6_MINI", + "V6_WILD" +]
- Changed
add_vocals1 field changed- changed
Input schema / properties / model / enumPrevious value: -[ - "V3_5", - "V4", - "V4_5", - "V4_5PLUS", - "V4_5ALL", - "V5", - "V5_5" -]New value: +[ + "V4", + "V4_5", + "V4_5PLUS", + "V4_5ALL", + "V5", + "V5_5", + "V6", + "V6_MINI", + "V6_WILD" +]
- Changed
cover_audio1 field changed- changed
Input schema / properties / model / enumPrevious value: -[ - "V3_5", - "V4", - "V4_5", - "V4_5PLUS", - "V4_5ALL", - "V5", - "V5_5" -]New value: +[ + "V4", + "V4_5", + "V4_5PLUS", + "V4_5ALL", + "V5", + "V5_5", + "V6", + "V6_MINI", + "V6_WILD" +]
- Changed
extend_music2 fields changed- changed
Input schema / properties / model / enumPrevious value: -[ - "V3_5", - "V4", - "V4_5", - "V4_5PLUS", - "V4_5ALL", - "V5", - "V5_5" -]New value: +[ + "V4", + "V4_5", + "V4_5PLUS", + "V4_5ALL", + "V5", + "V5_5", + "V6", + "V6_MINI", + "V6_WILD" +] - changed
Input schema / properties / prompt / descriptionPrevious value: -"Prompt for the extension"New value: +"LYRICS for the extended section (Suno sings this text). Give real lyric lines; a short instruction like \"add an outro\" is rejected as malformed lyrics (error 531, refunded)."
- Changed
generate_mashup1 field changed- changed
Input schema / properties / model / enumPrevious value: -[ - "V3_5", - "V4", - "V4_5", - "V4_5PLUS", - "V4_5ALL", - "V5", - "V5_5" -]New value: +[ + "V4", + "V4_5", + "V4_5PLUS", + "V4_5ALL", + "V5", + "V5_5", + "V6", + "V6_MINI", + "V6_WILD" +]
- Changed
generate_music1 field changed- changed
Input schema / properties / model / enumPrevious value: -[ - "V3_5", - "V4", - "V4_5", - "V4_5PLUS", - "V4_5ALL", - "V5", - "V5_5" -]New value: +[ + "V4", + "V4_5", + "V4_5PLUS", + "V4_5ALL", + "V5", + "V5_5", + "V6", + "V6_MINI", + "V6_WILD" +]
- Changed
generate_sounds1 field changed- changed
Input schema / properties / model / enumPrevious value: -[ - "V3_5", - "V4", - "V4_5", - "V4_5PLUS", - "V4_5ALL", - "V5", - "V5_5" -]New value: +[ + "V4", + "V4_5", + "V4_5PLUS", + "V4_5ALL", + "V5", + "V5_5", + "V6", + "V6_MINI", + "V6_WILD" +]
- Changed
seedream_layer_decompose3 fields changed- changed
Input schema / properties / size / descriptionPrevious value: -"1K/1.5K = 7 cr per layer, 2K = 14"New value: +"auto, 1K, 1.5K or 2K. Pro: 7 cr/layer at 1K/1.5K, 14 at 2K. Flash: 3.24 at any size." - added
Input schema / properties / tierAdded value: +{ + "default": "pro", + "description": "pro = Seedream 5.0 Pro (7 cr/layer @1K, 14 @2K; prompt required). flash = Seedream 5.0 Flash (NEW: 3.24 cr/layer at any size; prompt optional — but WITHOUT a prompt it splits out every element: a poster gave 12 layers = 38.9 cr. Name the elements you want to cap the layer count).", + "enum": [ + "pro", + "flash" + ], + "type": "string" +} - changed
Input schema / requiredPrevious value: -[ - "image_url", - "prompt" -]New value: +[ + "image_url" +]
- Changed
separate_vocals2 fields changed- changed
Input schema / properties / type / descriptionPrevious value: -"separate_vocal=vocals+instrumental, split_stem=individual instruments"New value: +"split_stem_advanced = NEW finer multi-stem split (market API). separate_vocal=vocals+instrumental, split_stem=individual instruments" - changed
Input schema / properties / type / enumPrevious value: -[ - "separate_vocal", - "split_stem" -]New value: +[ + "separate_vocal", + "split_stem", + "split_stem_advanced" +]
- Changed
upload_extend_audio4 fields changed- changed
Input schema / properties / continueAt / descriptionPrevious value: -"Timestamp in seconds where to start the extension"New value: +"REQUIRED in practice: second where the extension starts (> 0 and < the upload's length). kie fails every request without it." - changed
Input schema / properties / model / enumPrevious value: -[ - "V3_5", - "V4", - "V4_5", - "V4_5PLUS", - "V4_5ALL", - "V5", - "V5_5" -]New value: +[ + "V4", + "V4_5", + "V4_5PLUS", + "V4_5ALL", + "V5", + "V5_5", + "V6", + "V6_MINI", + "V6_WILD" +] - changed
Input schema / properties / prompt / descriptionPrevious value: -"Description of the extension content"New value: +"LYRICS for the extended section (sung as written); omit and set instrumental=true for an instrumental extension." - changed
Input schema / requiredPrevious value: -[ - "uploadUrl" -]New value: +[ + "uploadUrl", + "continueAt" +]
5 tool updates
v5.3.0- Changed
generate_gemini_tts2 fields changed- changed
Input schema / properties / model / descriptionPrevious value: -"flash = Gemini 3.1 Flash TTS (most expressive, 200+ inline tags, best <60s); pro = Gemini 2.5 Pro TTS (more stable long-form). Same price."New value: +"flash = Gemini 3.1 Flash TTS (most expressive, 200+ inline tags, best <60s; ~4.2 cr/min). flash-3.8 = Gemini 3.8 Flash TTS (NEW Sept 2026, same inputs, ~1.9 cr/min, ~40% cheaper per clip in live tests). flash-lite = Gemini 3.8 Flash Lite TTS (NEW, cheapest at ~1.3 cr/min; draft/bulk narration). pro = Gemini 2.5 Pro TTS (more stable long-form, ~4.2 cr/min). 3.8 prices are kie limited-time pricing until 2026-12-31." - changed
Input schema / properties / model / enumPrevious value: -[ - "flash", - "pro" -]New value: +[ + "flash", + "flash-3.8", + "flash-lite", + "pro" +]
- Changed
generate_video1 field changed- changed
Input schema / properties / model / descriptionPrevious value: -"Model ID (e.g. \"veo-3/text-to-video\", \"sora/text-to-video\", \"kling/image-to-video\")"New value: +"Model ID (e.g. \"veo-3/text-to-video\", \"kling/image-to-video\", \"wan/3-0-video\")"
- Changed
grok_image_edit6 fields changed- added
Input schema / properties / aspect_ratioAdded value: +{ + "description": "Whole-image mode: required output aspect ratio (auto keeps the input's shape).", + "enum": [ + "auto", + "1:1", + "2:3", + "3:2", + "16:9", + "9:16" + ], + "type": "string" +} - added
Input schema / properties / image_urlsAdded value: +{ + "description": "Whole-image mode: 1-5 public URL(s) of the image(s) to edit — any image, not just Grok generations (upload local files with upload_file first). Mutually exclusive with task_id.", + "items": { + "type": "string" + }, + "maxItems": 5, + "type": "array" +} - changed
Input schema / properties / mask_indexs / descriptionPrevious value: -"Region indices to edit, from grok_segment_map (e.g. [1] or [0, 2]). Field name matches kie's API spelling."New value: +"Region mode only: region indices from grok_segment_map (e.g. [1] or [0, 2]). Field name matches kie's API spelling." - changed
Input schema / properties / prompt / descriptionPrevious value: -"What the masked region(s) should become, plus what to preserve (e.g. \"change the background to a sunset beach, keep the apple unchanged\")"New value: +"Region mode: what the masked region(s) should become plus what to preserve. Whole-image mode: the edit instruction for the full image." - changed
Input schema / properties / task_id / descriptionPrevious value: -"Source task ID — a Grok Image 2.0 generation (or a previous grok_image_edit result)"New value: +"Region mode: source task ID — a Grok Image 2.0 generation (or a previous grok_image_edit result). Mutually exclusive with image_urls." - changed
Input schema / requiredPrevious value: -[ - "task_id", - "prompt", - "mask_indexs" -]New value: +[ + "prompt" +]
- Added
profile_brief - Added
seedream_layer_decompose
3 tool updates
v4.8.0- Added
grok_image_edit - Added
grok_segment_map - Changed
upload_file3 fields changed- changed
Input schema / properties / base64_data / descriptionPrevious value: -"Base64-encoded file data — raw base64 or a full data: URI. Whitespace and base64url are normalized and the data:<mime>;base64, prefix is stripped automatically (its MIME infers the extension if file_name is omitted). NOTE: very large images can be truncated when passed as a tool argument — if you get a length/invalid error, prefer file_url with a public URL."New value: +"Base64-encoded file data — raw base64 or a full data: URI. Whitespace and base64url are normalized and the data:<mime>;base64, prefix is stripped automatically (its MIME infers the extension if file_name is omitted). WARNING: payloads above ~10-12K chars (observed ceiling ~11.7K, #68) are silently truncated in transit as a tool argument — use file_path for local files; base64_data is a fallback for remote/HTTP-mode callers with small payloads." - added
Input schema / properties / file_pathAdded value: +{ + "description": "Absolute path to a local file on the machine running this MCP server (the normal case for stdio setups). The server reads and streams the bytes itself — reliable at any size, unlike base64_data. PREFERRED for local files.", + "type": "string" +} - changed
Input schema / properties / file_url / descriptionPrevious value: -"URL of file to upload — must be PUBLICLY reachable by kie.ai servers (no localhost/private IPs, no auth-gated or expired links). For local files use base64_data"New value: +"URL of file to upload — must be PUBLICLY reachable by kie.ai servers (no localhost/private IPs, no auth-gated or expired links). For local files use file_path"
42 tool updates
v4.7.0- First observed
add_instrumental - First observed
add_vocals - First observed
audio_isolation - First observed
boost_style - First observed
check_credits - First observed
check_task - First observed
convert_to_wav - First observed
cover_audio - First observed
create_music_video - First observed
create_omni_character - First observed
create_omni_voice - First observed
create_voice_clone - First observed
download_result - First observed
extend_music - First observed
generate_cover_art - First observed
generate_dialogue - First observed
generate_gemini_tts - First observed
generate_image - First observed
generate_lyrics - First observed
generate_mashup - First observed
generate_midi - First observed
generate_music - First observed
generate_persona - First observed
generate_sfx - First observed
generate_sounds - First observed
generate_tts - First observed
generate_video - First observed
get_timestamped_lyrics - First observed
list_models - First observed
list_raw_assets - First observed
list_tasks - First observed
prepare_voice_clone - First observed
regenerate_voice_clone - First observed
replace_section - First observed
runway_extend - First observed
separate_vocals - First observed
speech_to_text - First observed
upload_extend_audio - First observed
upload_file - First observed
veo_extend - First observed
veo_upscale_1080p - First observed
veo_upscale_4k
TDQS
Scored across 46 tools
Many tools have overlapping purposes (e.g. generate_tts vs generate_gemini_tts vs generate_dialogue, generate_sfx vs generate_sounds, add_vocals vs add_instrumental vs cover_audio), but the descriptions explicitly differentiate models, use cases, and output types. An agent must read carefully, but selection is usually possible.
Mostly consistent snake_case with verb_noun patterns (generate_music, create_omni_voice, list_models), though some are noun-first (audio_isolation, profile_brief) or model-prefixed (veo_upscale_1080p, grok_segment_map). These are minor deviations from an otherwise predictable convention.
46 tools is far above the 3-15 well-scoped range; even for a broad multi-modal platform, the set is heavy and contains many model-specific variants that could be consolidated into fewer parameterized tools. It is not extreme enough for a 1, but it is clearly too many.
The surface covers image, video, music, TTS, voice cloning, stem separation, format conversion, and task/credit utilities, so core workflows are well supported. Minor gaps like asset deletion/cleanup and task cancellation exist but are unlikely to block most agent workflows.
Maintenance
Related MCP Connectors
A comprehensive Model Context Protocol (MCP) server that enables AI assistants to interact with yo…
Multi-model AI image and video generator. 14 models behind one OAuth-secured MCP endpoint.
Generate images, video, audio and short films with 140+ AI models from any MCP client.
MCP server for Clipkit — gives AI agents a video toolbox via the Clipkit schema.
Related MCP Servers
- AlicenseNot gradedqualityFmaintenanceA Model Context Protocol server that enables AI assistants to perform comprehensive video and audio editing operations including trimming, effects, overlays, audio processing, and YouTube downloads.16 PyPI27MIT
- AlicenseNot gradedqualityDmaintenanceUniversal MCP server for all Kie AI image generation and editing models, enabling text-to-image, image-to-image, and image editing via natural language.57 npm2MIT
- AlicenseCqualityDmaintenanceA FastMCP-based Model Context Protocol (MCP) server that provides unified access to multiple AI APIs including OpenAI GPT, Google Gemini, Anthropic Claude, and xAI Grok.58 npmMIT
- FlicenseNot gradedqualityBmaintenanceSelf-hosted MCP server that integrates Kie.ai image and video generation into GoHighLevel's AI tools, enabling AI agents to create media on demand.1-