kie-mcp
Server Configuration
Describes the environment variables required to run the server.
| Name | Required | Description | Default |
|---|---|---|---|
| KIE_API_KEY | Yes | Your kie.ai API key | |
| KIE_MCP_PORT | No | Port for HTTP mode (default: 3100) | |
| KIE_CALLBACK_URL | No | Callback URL sent with Suno generation requests (kie.ai requires the field; results are fetched by polling regardless). Defaults to an inert placeholder — set this only if you want to receive the callbacks yourself | |
| KIE_PROJECT_ROOT | No | Server-wide default for where generated files are saved (default: server cwd; files go to $KIE_PROJECT_ROOT/kie/assets/raw/). Per-call download_dir (absolute path) on any file-writing tool overrides this | |
| KIE_MAX_CONCURRENT | No | Max simultaneous task-creation calls (default 4). Excess parallel generations queue inside the server instead of hitting kie.ai's rate limits — parallel tool calls are safe | |
| KIE_POLL_BUDGET_AUDIO | No | Blocking-mode polling budget for audio tools, in seconds (default: 300) | |
| KIE_POLL_BUDGET_IMAGE | No | Blocking-mode polling budget for image tools, in seconds (default: 600) | |
| KIE_POLL_BUDGET_VIDEO | No | Blocking-mode polling budget for video tools, in seconds (default: 900) | |
| KIE_POLL_BUDGET_SPEECH | No | Blocking-mode polling budget for speech tools, in seconds (default: 300) |
Instructions
Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.
This server publishes no instructions, or was last inspected before Glama recorded them.
Capabilities
Features and capabilities supported by this server
Protocol revision2025-11-25
| Capability | Details |
|---|---|
| tools | {} |
| prompts | {} |
Tools
Functions exposed to the LLM to take actions
| Name | Description |
|---|---|
| generate_imageA | Generate an image using kie.ai. TIP: for architecture/game-art/advertising/product-UI jobs, call profile_brief first — it returns the vertical's intake questions, routing, and prompt formulas. (60+ models). Downloads to kie/assets/raw/. MODEL GUIDE: Architecture/blueprints→gpt4o or nano-banana-2 (reasoning). Game art/3D→seedream/4.5 or 5-lite. Character sheets→ideogram/character. Text/logos→ideogram/v3 (best text). Photo editing→flux-kontext-pro. Newest OpenAI→gpt-image-2-5/flare-* (6cr @1K, fast default) or gpt-image-2-5/sunburst-* (premium polish); both 1K-4K + transparent background (NEW). Split any image into layers→seedream_layer_decompose tool (7cr/layer, NEW). Generate-then-refine by named region→grok-imagine-image-2-0/text-to-image (4cr, #2 Arena T2I+edit) then grok_segment_map (free) + grok_image_edit (4cr; also edits ANY uploaded image via image_urls mode). Anime→qwen (3cr cheapest); qwen2-1/* (4cr, NEW) adds transparent BG, mask inpainting, 10-ref compositing. Fast drafts→nano-banana-2-lite (4cr, ~4s, NEW). Upscale→recraft/crisp-upscale (0.5cr). BG removal→recraft/remove-background. Cheapest→z-image,qwen (3cr). Best quality→nano-banana-pro (24cr), flux-kontext-max (100cr). Use list_models filter="use-case" to explore. |
| list_modelsA | List all available kie.ai models with their aspect ratios and model-specific options |
| check_taskA | Check the status of a kie.ai generation task by taskId |
| list_tasksB | List recent image generation tasks from this session |
| check_creditsB | Check remaining kie.ai account credits |
| download_resultC | Download a completed task result to kie/assets/raw/ |
| list_raw_assetsA | List all files in kie/assets/raw/ waiting to be processed |
| generate_videoA | Generate a video using kie.ai (85+ models). Downloads to kie/assets/raw/. MODEL GUIDE: Best cinematic→veo-3/text-to-video (50cr/s, audio). Fast+cheap→grok-imagine-video-1-5-preview (1.6-3cr/s, audio, NEW), wan/flash-image-to-video (6-8cr/s measured; alias of wan/2-6-flash). Budget cinematic→hailuo-standard (4cr/s). First→last-frame or anything-from-anything refs→gemini-omni/flash-1-1 (NEW, est. ~63cr per 4s clip). Budget multimodal refs→bytedance/seedance-2-mini (9.5cr/s @480p). 30s single takes→bytedance/seedance-2-5 (NEW). Budget all-rounder w/ audio+templates+extend→pixverse-v6 family (4-9.6cr/s, NEW; I2V is its strength; transition=first/last-frame morph). Multilingual lip-synced dialogue→happyhorse-1-1 T2V/I2V/R2V (NEW). 2K + stereo audio→minimax-h3 (8cr/s @768P, price halved Sept 2026). Per-shot scripted multi-shot→kling-3-omni (14cr/s @720p, NEW; transformation=restyle existing video). Next-gen Wan draft→wan/3-0-video (8cr/s @480P, NEW). Fast Kling→kling/v3-turbo (18cr/s, audio, NEW). Image-to-video→veo-3/image-to-video (include a sound cue like "SFX: room tone" — Veo I2V intermittently fails its audio pass without one; images must be served with their real Content-Type, upload via upload_file), kling/image-to-video. Avatar/talking head→omnihuman-1-5 (premium, NEW), kling/ai-avatar-pro, infinitalk/from-audio. Re-dub existing footage→volcengine/video-to-video-lip-sync (8cr/s, NEW). Motion control→kling/motion-control, wan/animate-move. Extend video→use veo_extend or runway_extend tools. NOTE: Sora 2 family removed (OpenAI API sunset Sept 2026). Use list_models filter="use-case" to explore. |
| generate_musicA | Generate music using Suno via kie.ai. Supports V5.5 (custom style), V5 (best quality), V4.5+, V4.5, V4. Up to 8 minutes. Great for game music stems, ambient tracks, and jingles. Polls until done and downloads to kie/assets/raw/. |
| generate_sfxA | Generate a sound effect from text via Suno V5 (kie.ai removed the ElevenLabs sound-effect model). Great for game sounds: UI clicks, magic spells, item pickups, explosions. For loop/BPM/key control use generate_sounds instead. Downloads to kie/assets/raw/. |
| generate_gemini_ttsA | NEW — Google Gemini native TTS via kie.ai: style-directed speech from natural-language direction, 30 named voices, up to 2 speakers, inline tone tags like [whispers]/[laughs] (flash model). ~4.2 credits per MINUTE of audio — cheaper than all ElevenLabs tiers. Simple mode: pass text (+ optional voice_name). Dialogue mode: pass speakers + dialogue_turns. model=flash is most expressive (keep expected audio <60s — quality degrades on long takes); model=pro is more stable for multi-minute narration. Downloads to kie/assets/raw/. |
| generate_ttsA | Generate speech from text using ElevenLabs via kie.ai. Supports Turbo 2.5 (fast) and Multilingual V2 (high quality). Downloads to kie/assets/raw/. |
| generate_dialogueA | Generate multi-speaker dialogue using ElevenLabs Text-to-Dialogue V3 via kie.ai. Great for conversations between characters. Downloads to kie/assets/raw/. |
| audio_isolationB | Isolate vocals or audio from background noise using ElevenLabs via kie.ai. Input an audio URL, get clean isolated audio back. |
| extend_musicB | Extend/continue an existing Suno track from a specific point. Requires audioId from a previous generate_music task. |
| cover_audioC | Create an AI cover from uploaded audio — custom vocals, style, and instrumentation via Suno. |
| add_instrumentalC | Add instrumental backing to uploaded vocal audio via Suno. |
| add_vocalsC | Add AI vocals to uploaded instrumental audio via Suno. |
| replace_sectionC | Replace a time range in a Suno track with new AI-generated content. |
| generate_lyricsB | Generate song lyrics from a prompt using Suno AI (max 200 characters). Returns text, no file download. |
| convert_to_wavC | Convert a Suno track to lossless WAV format. Downloads to kie/assets/raw/. |
| separate_vocalsB | Separate vocals from instrumentals, or split into individual stems. Downloads to kie/assets/raw/. |
| generate_midiA | Export a Suno track to MIDI notation. Downloads .mid file to kie/assets/raw/. |
| create_music_videoB | Generate an MP4 music video visualization from a Suno track. Downloads to kie/assets/raw/. |
| generate_soundsB | Generate loopable sound effects with BPM, key, and loop control via Suno. Downloads to kie/assets/raw/. |
| generate_personaA | NEW — Create a Suno Persona (reusable music character) from an existing Suno track. Requires taskId from V3.6+ generation. |
| generate_mashupB | NEW — Mashup up to 2 Suno tracks into one new track. Provide audioIds from previous generations. |
| boost_styleA | NEW — Convert concise style input (e.g. "Pop, Mysterious") into enhanced style description for music generation. |
| prepare_voice_cloneA | EXPERIMENTAL (#20) — STEP 1 of Suno custom-voice cloning (FREE). Submit a clean vocal sample; polls to |
| create_voice_cloneA | EXPERIMENTAL (#20) — STEP 2 — after the voice owner records the verification phrase from prepare_voice_clone, submit that recording to finish the voice. On success returns a voiceId usable in generate_music. Unverified end-to-end. |
| regenerate_voice_cloneB | EXPERIMENTAL (#20) — retry a failed/incomplete custom-voice task by its task_id. |
| get_timestamped_lyricsA | NEW — Get word-level timestamped lyrics from a Suno track. Useful for karaoke, captioning, or sync. |
| generate_cover_artA | NEW — Generate album cover art image for an existing Suno music track. One call per taskId only. |
| create_omni_voiceB | NEW — Create a reusable voice character for Gemini Omni video generation. Returns kieAudioId for use in generate_video audio_ids. |
| create_omni_characterB | NEW — Create a reusable visual character for Gemini Omni video generation. Combines image + optional voice. Returns characterId. |
| upload_extend_audioA | NEW — Extend uploaded audio (NOT a Suno track) with new AI-generated content. For Suno tracks, use extend_music instead. |
| speech_to_textA | Transcribe audio to text using ElevenLabs Scribe v1. Supports diarization and audio event tagging. Returns transcription text. |
| upload_fileA | Upload a file to kie.ai and get a public URL back. Image files are named by their REAL format (magic bytes), so a JPEG saved as .png is uploaded as .jpg — kie serves files with a Content-Type taken from the name, and Veo I2V rejects a mismatch. Use this to upload local images/audio/video before passing them to generation tools (image-to-image, image-to-video, reference/ingredient inputs). PREFER file_path for local files. Files expire after 3 days (kie temp storage). |
| profile_briefA | Get a vertical playbook before generating: the intake questions a professional in that domain would ask, model routing per deliverable with live costs, per-model prompt formulas, and multi-tool workflows. Call with no args to list available profiles. Covers image, video, and audio verticals (architecture, game assets, advertising, product photography, film, brand, web product, editorial, short-form social video, audio branding). Call this FIRST when the user's request belongs to a known vertical — then ask the user only the unanswered intake questions, conversationally. |
| seedream_layer_decomposeA | Split ANY image into independent layers with Seedream 5.0 (billed per OUTPUT layer incl. base). tier="pro" (default): 7 cr @1K, 14 @2K, so a 3-layer split ≈ 21 cr. tier="flash" (NEW): 3.24 cr/layer at any size, so ≈ 10 cr. Works on any public image URL (upload local files with upload_file first). Describe which elements become layers in the prompt, optionally bounding them with x1 y1 x2 y2 tags. Downloads every layer image to kie/assets/raw/. |
| grok_segment_mapA | FREE (0 credits). Segment a Grok Imagine Image 2.0 generation into NAMED regions for targeted editing. Returns each region's index, semantic name (e.g. "red apple", "wooden table"), and mask PNG URL. Workflow: generate_image model="grok-imagine-image-2-0/text-to-image" → grok_segment_map (this, free) → grok_image_edit with the mask_indexs you want changed. Only works on task_ids from a Grok Image 2.0 generation. |
| grok_image_editA | Edit an image with Grok Imagine Image 2.0 (4 credits). TWO MODES: (1) Region mode — pass task_id (a prior Grok 2.0 generation) + mask_indexs from grok_segment_map (run it first, free, pick regions by NAME): only those regions change. (2) Whole-image mode (NEW Aug 2026) — pass image_urls (ANY uploaded/external image, e.g. from upload_file) + aspect_ratio + prompt: instruction-based edit of the full image, no segmentation. Returns a new full image; result task_ids chain back into segment/edit for iterative refinement. Downloads to kie/assets/raw/. |
| veo_extendC | Extend an existing Veo 3.1 video with additional content. Requires taskId from a previous Veo generation. |
| veo_upscale_1080pA | Upscale a Veo 3.1 video to 1080p resolution. Requires taskId from a completed Veo generation. |
| veo_upscale_4kA | Upscale a Veo 3.1 video to 4K resolution. Takes 5-10 minutes. Requires taskId from completed Veo generation. |
| runway_extendB | Extend an existing Runway Aleph video with continuation content. |
Prompts
Interactive templates invoked by user choice
| Name | Description |
|---|---|
| architecture | Architecture & Interior Design — Exterior/interior renders, renovation previews, elevations, site plans, sketches, material boards — with the questions an architect would ask first. |
| game-assets | Video Game Assets — Character sheets, environment concepts, sprites, tileable textures, icons, and key art — engine-aware, style-bible-consistent. |
| advertising | Advertising & Marketing — Social posts, display ads, product heroes, campaign key visuals, and A/B variant sets — brand-safe, platform-sized, text-accurate, disclosure-aware. |
| web-product | Web & Software Product Imagery — In-product imagery for apps and sites: empty states, feature illustrations, icon sets, onboarding art, hero/og images, 404s — design-system-consistent. |
| film-storyboard | Film & Storyboarding — Storyboards, shot concepts, character/costume lookdev, set design, and mood frames — lens-language-aware, continuity-conscious. |
| product-photography | Product Photography & E-commerce — Packshots, lifestyle scenes, listing sets, and marketplace-compliant product imagery — built around the REAL product via reference photos. |
| brand-design | Brand & Graphic Design — Logo concepts, posters, typography-led pieces, patterns, and brand exploration — text-accuracy-first, print-aware, trademark-conscious. |
| editorial | Editorial & Publishing — Article art, book covers, spot illustrations, infographic bases, and series-consistent publication imagery — tone-calibrated, disclosure-aware. |
| social-video | Short-Form Social Video — TikTok/Reels/Shorts content: hook clips, talking heads, product showcases, b-roll loops, multi-scene stories — with voiceover, music, and SFX. |
| audio-branding | Audio Branding & Music — Brand music, jingles, podcast intros, voiceover/narration, multi-speaker dialogue, SFX, and sonic logos — with stems, extensions, and loudness discipline. |
Resources
Contextual data attached and managed by the client
| Name | Description |
|---|---|
No resources | |
TDQS
Scored across 46 tools
Many tools have overlapping purposes (e.g. generate_tts vs generate_gemini_tts vs generate_dialogue, generate_sfx vs generate_sounds, add_vocals vs add_instrumental vs cover_audio), but the descriptions explicitly differentiate models, use cases, and output types. An agent must read carefully, but selection is usually possible.
Mostly consistent snake_case with verb_noun patterns (generate_music, create_omni_voice, list_models), though some are noun-first (audio_isolation, profile_brief) or model-prefixed (veo_upscale_1080p, grok_segment_map). These are minor deviations from an otherwise predictable convention.
46 tools is far above the 3-15 well-scoped range; even for a broad multi-modal platform, the set is heavy and contains many model-specific variants that could be consolidated into fewer parameterized tools. It is not extreme enough for a 1, but it is clearly too many.
The surface covers image, video, music, TTS, voice cloning, stem separation, format conversion, and task/credit utilities, so core workflows are well supported. Minor gaps like asset deletion/cleanup and task cancellation exist but are unlikely to block most agent workflows.