MCP Kling
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@MCP Klinggenerate a video of a sunset over mountains with birds flying"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
š¬ MCP Kling - The ONLY COMPLETE Kling AI MCP Server!
The world's FIRST and ONLY complete MCP server for Kling AI - now with FULL API support! š
Transform Claude into a professional AI content studio with complete access to Kling's entire suite of creative tools. Generate videos, create images, add lip-sync, apply effects, and even try on virtual clothing - all through simple conversations with Claude. This isn't just another integration; it's the COMPLETE Kling experience!
š Why This is HUGE
100% COMPLETE: The ONLY MCP server implementing ALL Kling AI features
13+ Tools: Full access to video, image, effects, lip-sync, and more
Auto-Download: Automatically saves all generated content locally
Multiple Models: Access Kling v1.0, v1.5, v1.6, and KOLORS for images
Professional Studio: Create complete productions with effects and audio
Account Management: Monitor balance and resource usage
Perfect for: Content creators, filmmakers, marketers, developers, and AI enthusiasts
Related MCP server: PixVerse MCP
š Quick Start
It's incredibly easy to get started!
1. Get your Kling API Credentials
Click "+ Create a new API Key"
Save both your Access Key and Secret Key
2. Add to Claude Desktop
Add this configuration to your Claude Desktop config file:
macOS: ~/Library/Application Support/Claude/claude_desktop_config.json
Windows: %APPDATA%\Claude\claude_desktop_config.json
{
"mcpServers": {
"mcp-kling": {
"command": "npx",
"args": ["-y", "mcp-kling@latest"],
"env": {
"KLING_ACCESS_KEY": "YOUR_ACCESS_KEY_HERE",
"KLING_SECRET_KEY": "YOUR_SECRET_KEY_HERE"
}
}
}
}The MCP server will automatically generate JWT tokens as needed using your credentials.
That's it! Restart Claude Desktop and you're ready to generate amazing videos! š
š ļø Complete Feature Set - ALL 13 Tools!
š„ Video Generation
1. generate_video
Create stunning videos from text descriptions.
Models: v1.0, v1.5, v1.6
Duration: 5 or 10 seconds
Aspect Ratios: 16:9, 9:16, 1:1
Modes: Standard or Professional
2. generate_image_to_video
Transform static images into dynamic videos.
Image Types: PNG, JPG, JPEG, WebP
Motion Control: Automatic or custom prompts
Camera Movement: Static, zoom, pan, or auto
3. check_video_status
Monitor generation progress and auto-download completed videos.
4. extend_video
Seamlessly extend videos by 4-5 seconds.
Smart Continuation: AI understands context
Custom Prompts: Guide the extension direction
Multiple Extensions: Chain for longer videos
šļø Audio & Effects
5. create_lipsync
Synchronize lip movements with audio.
Custom Audio: MP3, WAV, FLAC, OGG
Text-to-Speech: 8 voice styles
Speed Control: 0.8x to 1.5x
Auto-Download: Saves both video and audio
6. apply_video_effect
Professional post-production effects.
Fast Motion: Speed up 2x-16x
Slow Motion: Slow down 0.5x-0.9x
Reverse: Play videos backward
Loop: Create seamless loops
šØ Image Generation
7. generate_image
Create stunning images with KOLORS model.
Resolutions: Up to 2K quality
Aspect Ratios: 16:9, 9:16, 1:1, 2:3, 3:2
Styles: Photorealistic or artistic
Batch Generation: Up to 4 images
8. check_image_status
Monitor and auto-download generated images.
9. virtual_try_on
Revolutionary AI fashion try-on.
Model Image: Your photo (full body preferred)
Garment Image: Clothing to try on
Smart Fitting: AI adjusts to body type
Realistic Results: Natural fabric draping
š Account Management
10. get_account_balance
Check your Kling credit balance.
11. get_resource_packages
View available subscription packages.
12. list_tasks
Track all your generation tasks.
Pagination: Handle large task lists
Filtering: By date or status
Complete History: All your creations
š” Complete Examples for Every Feature
š¬ Video Generation
You: Generate a video of a futuristic robot dancing in a disco
Claude: I'll create an amazing robot disco video for you!
[Using generate_video tool...]
ā
Video generation started!
Task ID: k123456789
[Auto-checking status...]
š Your video is ready and downloaded!
- File: ./downloads/videos/robot_disco_k123456789.mp4
- Duration: 5 seconds
- Resolution: 1920x1080š Video Extension Chain
You: Extend that robot video 3 times with: 1) backflip, 2) moonwalk, 3) fireworks finale
Claude: I'll create an epic extended sequence!
[Extension 1: Backflip]
ā
Extended by 5 seconds
[Extension 2: Moonwalk]
ā
Extended by 5 seconds
[Extension 3: Fireworks]
ā
Extended by 5 seconds
š¬ Final video: 20 seconds total!šļø Lip Sync with TTS
You: Make the CEO in this video say "Welcome to our product launch" in a professional male voice
Claude: I'll add professional narration to your CEO video.
[Using create_lipsync with TTS...]
ā
Lip-sync video created!
- Voice: Professional Male
- Speed: 1.0x
- Downloaded: ./downloads/lipsync/ceo_welcome_k987654321.mp4šØ Image Generation Batch
You: Generate 4 images of "cyberpunk cat with neon glasses" in different styles
Claude: Creating a batch of cyberpunk cat images!
[Using generate_image tool...]
ā
Generated 4 unique images:
- ./downloads/images/cyberpunk_cat_1.png
- ./downloads/images/cyberpunk_cat_2.png
- ./downloads/images/cyberpunk_cat_3.png
- ./downloads/images/cyberpunk_cat_4.pngš Virtual Try-On
You: Try this red dress (dress.jpg) on my photo (me.jpg)
Claude: I'll help you virtually try on that dress!
[Using virtual_try_on tool...]
ā
Virtual try-on complete!
- Result: ./downloads/tryon/red_dress_result.mp4
- The dress fits perfectly and moves naturally!šÆ Video Effects Combo
You: Apply slow motion to the first half and fast motion to the second half
Claude: Creating a dramatic tempo change effect!
[Step 1: Splitting video...]
[Step 2: Slow motion on part 1...]
[Step 3: Fast motion on part 2...]
ā
Effect applied! Downloaded to: ./downloads/effects/tempo_change.mp4š Account Management
You: Check my balance and show me what I can create
Claude: Let me check your account details!
[Using get_account_balance tool...]
š° Current Balance: 5,000 credits
With your balance, you can create:
- 50 standard videos (100 credits each)
- 25 professional videos (200 credits each)
- 100 images (50 credits each)
- 33 lip-sync videos (150 credits each)šÆ Revolutionary Use Cases
š¬ Film & Video Production
Storyboarding: Rapidly prototype scenes
VFX Previsualization: Test effects before shooting
B-Roll Generation: Create supplementary footage
Music Videos: Generate visuals for songs
š± Social Media Automation
TikTok/Reels: Auto-generate trending content
Product Showcases: Dynamic product videos
Story Series: Chain extended videos
Branded Effects: Apply consistent styles
šļø E-Commerce Revolution
Virtual Fashion Shows: Model clothes on anyone
Product Animations: Bring static products to life
Try-Before-Buy: Virtual clothing trials
360° Product Views: Generated from single image
š Education & Training
Animated Explanations: Complex concepts simplified
Language Learning: Lip-sync in any language
Historical Recreations: Bring history to life
Scientific Visualizations: Abstract concepts visualized
š Auto-Download Magic
All generated content is automatically downloaded and organized:
./downloads/
āāā videos/ # Generated videos
āāā images/ # Generated images
āāā lipsync/ # Lip-sync videos
āāā effects/ # Effect-applied videos
āāā extended/ # Extended videos
āāā tryon/ # Virtual try-on resultsNever lose your creations - everything is saved locally with descriptive filenames!
š§ Advanced Configuration
Video Generation Options
{
model: "kling-v2-master", // v1, v1.5, v1.6, or v2-master
duration: 10, // 5 or 10 seconds
aspect_ratio: "16:9", // 16:9, 9:16, 1:1
mode: "professional", // standard or professional
cfg_scale: 0.7, // 0-1 (creativity vs accuracy)
camera_control: { // Camera movement (V1 only)
type: "simple", // Camera movement type
config: { zoom: 5 } // Movement parameters
}
}Local File Support
NEW in v5.2.0: The MCP server now automatically handles local files!
Automatic Upload: Local files and file:// URLs are automatically uploaded to cloud storage
Seamless Integration: Just provide local file paths - the server handles the rest
Supported Formats: Images (PNG, JPG, JPEG, GIF, WebP) and Videos (MP4, WebM, MOV)
Zero Configuration: Works out of the box with Supabase integration
Examples:
// All of these work automatically:
image_url: "/Users/me/image.jpg" // Local file path
image_url: "file:///Users/me/image.jpg" // File URL
image_url: "https://example.com/image.jpg" // HTTP URL (used directly)Image Generation Options
{
model: "kolors", // Currently supports KOLORS
aspect_ratio: "16:9", // Multiple ratios supported
image_count: 4, // 1-4 images per request
resolution: "2k" // Up to 2K quality
}š Why Choose MCP Kling?
ā Feature Complete
ONLY server with 100% Kling API coverage
13 tools vs others with just 2-3
Auto-download everything
Professional features included
ā” Production Ready
Robust error handling
Automatic retries
Progress tracking
Local file management
šØ Creative Freedom
Chain operations for complex workflows
Batch processing for efficiency
Effect combinations for unique results
No limits on creativity
š¤ Contributing
We welcome contributions! Join us in building the future of AI content creation.
Ideas for Contributors:
Custom effect presets
Workflow templates
Integration examples
Performance optimizations
š License
MIT License - feel free to use in your projects!
š„ Authors
Created with ā¤ļø by Boris Djordjevic and the 199 Longevity team.
š Star Us!
If you find this useful, please star the repository! It helps others discover this complete Kling integration.
š The ONLY complete Kling MCP server - Install now and unlock the full power of AI content creation!
No other MCP server comes close - this is the complete package!
Available Tools
12 toolsapply_video_effectA
Apply pre-defined animation effects to static images using Kling AI. Create emotionally expressive videos from portraits with effects like hugging, kissing, or playful animations. Dual-character effects (hug, kiss, heart_gesture) require exactly 2 images. Single-image effects (squish, expansion, fuzzyfuzzy, bloombloom, dizzydizzy) require 1 image. Perfect for social media content and creative storytelling.
| Name | Required | Description | Default |
|---|---|---|---|
| image_urls | Yes | Array of image URLs. Use 2 images for hug/kiss/heart_gesture effects, 1 image for squish/expansion/fuzzyfuzzy/bloombloom/dizzydizzy effects | |
| effect_scene | Yes | The animation effect to apply. Dual-character: hug, kiss, heart_gesture. Single-image: squish, expansion, fuzzyfuzzy, bloombloom, dizzydizzy | |
| duration | No | Video duration in seconds (default: 5) | |
| model_name | No | Model version to use (default: kling-v2-master) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It describes the tool's core functionality and image requirements well, but lacks information about permissions, rate limits, processing time, error conditions, or what the output looks like (though no output schema exists).
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is efficiently structured with three sentences that each serve distinct purposes: stating the core functionality, specifying effect requirements, and providing use case context. Every sentence adds value with no wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with 4 parameters, no annotations, and no output schema, the description provides good functional context but lacks important behavioral details. It explains what the tool does and parameter requirements well, but doesn't cover error handling, performance characteristics, or output format, leaving gaps for an AI agent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 100% schema description coverage, the baseline is 3. The description adds value by explaining the emotional context ('emotionally expressive videos'), categorizing effects (dual-character vs. single-image), and providing use case context ('social media content and creative storytelling') beyond what the schema provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose with specific verbs ('apply pre-defined animation effects', 'create emotionally expressive videos') and resources ('static images', 'portraits'), and distinguishes it from siblings by focusing on animation effects rather than status checks, generation, or other video operations.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context on when to use this tool ('perfect for social media content and creative storytelling') and specifies requirements for different effect types (dual-character vs. single-image effects). However, it doesn't explicitly state when NOT to use it or name specific alternatives among sibling tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
check_image_statusC
Check the status of an image generation task
| Name | Required | Description | Default |
|---|---|---|---|
| task_id | Yes | The task ID returned from generate_image |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries full burden for behavioral disclosure. It states the tool checks status but reveals nothing about what statuses exist, whether it polls or returns immediately, error conditions, or response format. This leaves significant gaps for a status-checking operation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, clear sentence with zero wasted words. It's front-loaded with the core purpose and appropriately sized for a simple tool with one parameter.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a status-checking tool with no annotations and no output schema, the description is insufficient. It doesn't explain what status information is returned, possible states (e.g., pending, completed, failed), or how to interpret results. The context demands more completeness than provided.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, with the single parameter 'task_id' well-documented in the schema as 'The task ID returned from generate_image'. The description adds no additional parameter semantics beyond what the schema already provides, meeting the baseline for high schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('check') and resource ('status of an image generation task'), making the purpose immediately understandable. It doesn't explicitly differentiate from sibling tools like 'check_video_status' or 'list_tasks', but the specificity to 'image generation task' provides reasonable distinction.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives. It doesn't mention prerequisites (e.g., requiring a task_id from generate_image), when-not scenarios, or relationships with sibling tools like list_tasks for broader status checking.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
check_video_statusB
Check the status of a video generation task
| Name | Required | Description | Default |
|---|---|---|---|
| task_id | Yes | The task ID returned from generate_video or generate_image_to_video |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries full burden but only states the basic action. It lacks behavioral details such as whether this is a read-only operation, what status values might be returned, if there are rate limits, or how errors are handled for invalid task IDs.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, clear sentence with zero wasted words. It is appropriately sized for a simple tool and front-loads the essential information without unnecessary elaboration.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with no annotations and no output schema, the description is insufficient. It doesn't explain what status information is returned, possible states (e.g., pending, completed, failed), or error conditions. Given the complexity of task monitoring, more context is needed for effective use.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the input schema fully documents the 'task_id' parameter. The description adds no additional semantic context beyond implying the ID comes from generation tools, which is already suggested by the schema's description. Baseline 3 is appropriate as the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('check') and resource ('video generation task'), making the purpose understandable. However, it doesn't differentiate from the sibling tool 'check_image_status' beyond the resource type, missing explicit distinction between video and image status checks.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for monitoring video generation tasks, but provides no explicit guidance on when to use this tool versus alternatives like 'list_tasks' or 'check_image_status'. It mentions the task ID comes from specific generation tools, offering some contextual hint.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_lipsyncA
Create a lip-sync video by synchronizing mouth movements with audio. Supports both text-to-speech (TTS) with various voice options or custom audio upload. The original video must contain a clear, steady human face with visible mouth. Works with real, 3D, or 2D human characters (not animals). Video length limited to 10 seconds.
| Name | Required | Description | Default |
|---|---|---|---|
| video_url | Yes | URL of the video to apply lip-sync to (must contain clear human face) | |
| audio_url | No | URL of custom audio file (mp3, wav, flac, ogg; max 20MB, 60s). If provided, TTS parameters are ignored | |
| tts_text | No | Text for text-to-speech synthesis (used only if audio_url is not provided) | |
| tts_voice | No | Voice style for TTS (default: male-warm). Includes Chinese and English voice options | |
| tts_speed | No | Speech speed for TTS (0.5-2.0, default: 1.0) | |
| model_name | No | Model version to use (default: kling-v2-master) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries full burden. It discloses key behavioral traits: video constraints (clear face, 10s limit), supported character types (human/3D/2D, not animals), and TTS vs custom audio options. However, it doesn't cover permissions, rate limits, or what happens to the original video.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is appropriately sized (4 sentences) and front-loaded with the core purpose. Every sentence adds value: purpose, input options, video requirements, and constraints. No wasted words or redundant information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 6-parameter creation tool with no annotations or output schema, the description provides good context about what the tool does and its constraints. It covers the main use case and limitations, though could benefit from more behavioral details about the creation process and output format.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all 6 parameters thoroughly. The description adds some context about TTS/custom audio trade-offs and video requirements, but doesn't provide additional parameter semantics beyond what's in the schema descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose with specific verbs ('create', 'synchronizing') and resources ('lip-sync video', 'mouth movements', 'audio'). It distinguishes from siblings by focusing on lip-sync creation rather than effects, generation, or status checking.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context for when to use this tool: for creating lip-sync videos with specific input requirements (clear human face, video length ā¤10s). It doesn't explicitly mention when not to use it or name alternatives among siblings, but the constraints guide appropriate usage.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
extend_videoA
Extend a video by 4-5 seconds using Kling AI. This feature allows you to continue a video beyond its original ending, generating new content that seamlessly follows from the last frame. Perfect for creating longer sequences or adding additional scenes to existing videos.
| Name | Required | Description | Default |
|---|---|---|---|
| task_id | Yes | The task ID of the original video to extend (from a previous generation) | |
| prompt | Yes | Text prompt describing how to extend the video (what should happen next) | |
| model_name | No | Model version to use for extension (default: kling-v2-master) | |
| duration | No | Extension duration (fixed at 5 seconds) | |
| mode | No | Video generation mode (default: standard) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries full burden. It discloses key behavioral traits: it's a generative extension tool ('generating new content'), mentions the AI provider ('Kling AI'), and specifies the duration range ('4-5 seconds'). However, it doesn't cover important aspects like rate limits, authentication needs, cost implications, or what the output looks like (e.g., returns a new task ID).
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is appropriately sized (three sentences) and front-loaded with the core purpose. Every sentence adds value: first states the action, second explains the mechanism, third provides usage context. It could be slightly more concise by combining ideas, but there's minimal waste.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a generative tool with 5 parameters, 100% schema coverage, but no annotations or output schema, the description is adequate but has gaps. It covers the what and why well, but lacks details on behavioral constraints (e.g., rate limits), output format, or error conditions. The context is complete enough for basic use but not for robust agent operation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, providing full parameter documentation. The description adds minimal value beyond the schema, only implying that 'task_id' refers to 'a previous generation' and 'prompt' guides 'what should happen next'. It doesn't explain parameter interactions or provide additional context, so baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the specific action ('Extend a video by 4-5 seconds using Kling AI') and resource ('a video'), distinguishing it from siblings like generate_video (create new) or apply_video_effect (modify existing). It explains the functional outcome ('continue a video beyond its original ending, generating new content that seamlessly follows from the last frame').
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context for when to use this tool ('Perfect for creating longer sequences or adding additional scenes to existing videos'), but doesn't explicitly state when not to use it or name alternatives among siblings (e.g., generate_video for new videos, apply_video_effect for modifications). The guidance is helpful but lacks explicit exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
generate_imageB
Generate images from text prompts using Kling AI. Create high-quality images with multiple aspect ratios and optional character reference support. Supports models v1, v1.5, and v2 with customizable parameters for creative control.
| Name | Required | Description | Default |
|---|---|---|---|
| prompt | Yes | Text prompt describing the image to generate | |
| negative_prompt | No | Text describing what to avoid in the image (optional) | |
| model_name | No | Model version to use (default: kling-v2-master) | |
| aspect_ratio | No | Image aspect ratio (default: 1:1) | |
| num_images | No | Number of images to generate (default: 1) | |
| ref_image_url | No | Optional reference image URL for character consistency | |
| ref_image_weight | No | Weight of reference image influence (0-1, default: 0.5) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries full burden for behavioral disclosure. It mentions 'high-quality images' and 'creative control' but fails to disclose critical behavioral traits such as rate limits, authentication requirements, processing time, cost implications, or what happens on failure. For a complex 7-parameter tool with no annotation coverage, this is a significant gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is appropriately sized with three sentences that efficiently cover purpose, features, and capabilities. It's front-loaded with the core functionality and avoids unnecessary repetition. Every sentence adds value, though it could be slightly more structured with clearer separation of key features.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex image generation tool with 7 parameters, no annotations, and no output schema, the description is incomplete. It doesn't explain what the tool returns (image URLs? base64? metadata?), doesn't mention error conditions, rate limits, or authentication requirements, and provides insufficient behavioral context for safe and effective use.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all parameters thoroughly. The description adds marginal value by mentioning 'multiple aspect ratios' and 'optional character reference support' which map to aspect_ratio and ref_image_url parameters, but doesn't provide additional semantic context beyond what's in the schema. Baseline 3 is appropriate when schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose with specific verbs ('generate images from text prompts') and resources ('using Kling AI'), distinguishing it from sibling tools like generate_video or generate_image_to_video. It explicitly mentions the core functionality of text-to-image generation with quality and aspect ratio options.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for creative image generation with Kling AI, but provides no explicit guidance on when to use this tool versus alternatives like generate_image_to_video or apply_video_effect. It mentions 'creative control' as a general context but lacks specific when/when-not scenarios or comparisons to sibling tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
generate_image_to_videoC
Generate a video from an image using Kling AI
| Name | Required | Description | Default |
|---|---|---|---|
| image_url | Yes | URL of the starting image | |
| image_tail_url | No | URL of the ending image (optional) | |
| prompt | Yes | Text prompt describing the motion and transformation | |
| negative_prompt | No | Text describing what to avoid in the video (optional) | |
| model_name | No | Model version to use (default: kling-v2-master) | |
| duration | No | Video duration in seconds (default: 5) | |
| mode | No | Video generation mode (default: standard) | |
| cfg_scale | No | Creative freedom scale 0-1 (default: 0.5) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It states the tool generates a video but lacks details on permissions, rate limits, processing time, output format, or error handling. For a complex video generation tool with zero annotation coverage, this is a significant gap in transparency.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence that directly states the tool's purpose without unnecessary words. It is appropriately sized and front-loaded, making it easy to understand at a glance.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (8 parameters, video generation) and lack of annotations and output schema, the description is incomplete. It doesn't address behavioral aspects like processing time, output format, or error conditions, which are critical for an AI agent to use this tool effectively in context with siblings.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, meaning all parameters are documented in the input schema. The description adds no additional parameter information beyond what's in the schema, such as explaining relationships between parameters (e.g., how 'image_tail_url' interacts with 'prompt'). Baseline 3 is appropriate when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Generate a video from an image') and specifies the technology used ('using Kling AI'), which provides a specific verb+resource combination. However, it doesn't differentiate this tool from sibling tools like 'generate_video' or 'extend_video', which likely have different purposes or inputs.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives. It doesn't mention sibling tools like 'generate_video' or 'extend_video', nor does it specify prerequisites such as needing an image URL or appropriate prompts. Usage context is implied but not explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
generate_videoC
Generate a video from text prompt using Kling AI
| Name | Required | Description | Default |
|---|---|---|---|
| prompt | Yes | Text prompt describing the video to generate (max 2500 characters) | |
| negative_prompt | No | Text describing what to avoid in the video (optional, max 2500 characters) | |
| model_name | No | Model version to use (default: kling-v2-master) | |
| aspect_ratio | No | Video aspect ratio (default: 16:9) | |
| duration | No | Video duration in seconds (default: 5) | |
| mode | No | Video generation mode (default: standard) | |
| cfg_scale | No | Creative freedom scale 0-1 (0=more creative, 1=more adherent to prompt, default: 0.5) | |
| camera_control | No | Camera movement settings for V2 models |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. While 'Generate a video' implies a creation/mutation operation, the description doesn't address critical behavioral aspects: whether this is an async operation (likely given sibling 'check_video_status'), what permissions or authentication might be required, rate limits, cost implications, or what format/quality the output video will have. This is inadequate for a complex generative tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence that states the core purpose without unnecessary elaboration. It's appropriately sized and front-loaded with the essential information, making it easy for an agent to quickly understand what the tool does.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex video generation tool with 8 parameters (including nested objects), no annotations, and no output schema, the description is insufficient. It doesn't address the asynchronous nature suggested by sibling tools, doesn't explain what the tool returns (video file? URL? task ID?), and provides no context about the Kling AI service's capabilities or limitations. The description should do more to compensate for the lack of structured metadata.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema description coverage is 100%, so all parameters are documented in the schema itself. The description adds no additional parameter semantics beyond what's already in the schema descriptions. This meets the baseline expectation when schema coverage is complete, but doesn't provide extra value like explaining parameter interactions or practical usage patterns.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: 'Generate a video from text prompt using Kling AI' - a specific verb ('Generate') with resource ('video') and technology context ('Kling AI'). However, it doesn't distinguish this from sibling tools like 'generate_image_to_video' or 'extend_video', which would require explicit differentiation for a perfect score.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives like 'generate_image_to_video', 'extend_video', or 'apply_video_effect'. There's no mention of prerequisites, appropriate contexts, or limitations that would help an agent choose between these video-related tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_account_balanceA
Check your Kling AI account balance and total available credits. Provides a comprehensive overview of your account status including total balance and breakdown by resource packages.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries full burden for behavioral disclosure. It describes what information is returned but doesn't mention authentication requirements, rate limits, error conditions, or whether this is a read-only operation. The description is informative but lacks critical operational context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is appropriately concise with two sentences that efficiently convey the tool's purpose and scope. It's front-loaded with the main function and follows with additional detail about what's included in the overview.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter tool with no output schema, the description provides adequate information about what the tool does but lacks details about the return format, error handling, or authentication requirements. Given the complexity is low (no parameters), the description is reasonably complete but could benefit from more operational context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has 0 parameters with 100% schema description coverage, so the baseline is 4. The description appropriately doesn't discuss parameters since none exist, focusing instead on the tool's purpose and output.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose with specific verbs ('check', 'provides') and resources ('account balance', 'total available credits', 'account status', 'breakdown by resource packages'). It distinguishes itself from siblings like 'get_resource_packages' by focusing on overall account status rather than just package details.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage context (checking account status) but doesn't explicitly state when to use this tool versus alternatives like 'get_resource_packages' or other sibling tools. No guidance on prerequisites or exclusions is provided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_resource_packagesA
Get detailed information about your Kling AI resource packages including remaining credits, expiration dates, and package types. Useful for monitoring API usage and planning resource allocation.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It implies a read-only operation by using 'Get' and describes the type of information returned, but does not specify authentication requirements, rate limits, or potential errors. The description adds some context about usage monitoring but lacks detailed behavioral traits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is efficiently structured in two sentences: the first states the purpose and details retrieved, and the second provides usage context. Every sentence adds value without waste, making it appropriately sized and front-loaded for clarity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (simple read operation with no parameters) and lack of annotations or output schema, the description is adequate but has gaps. It explains what information is retrieved but does not describe the return format, potential errors, or authentication needs, leaving some contextual information incomplete for a tool without structured support.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has 0 parameters with 100% schema description coverage, so no parameter information is needed in the description. The description appropriately focuses on the tool's purpose and usage without redundant parameter details, earning a baseline score of 4 for this dimension.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose with specific verbs ('Get detailed information') and resources ('Kling AI resource packages'), and distinguishes it from siblings by focusing on usage monitoring rather than content generation or task management. It explicitly mentions what information is retrieved: remaining credits, expiration dates, and package types.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context for when to use this tool ('Useful for monitoring API usage and planning resource allocation'), which helps differentiate it from siblings like get_account_balance or list_tasks. However, it does not explicitly state when not to use it or name specific alternatives, keeping it at a 4 rather than a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_tasksB
List all your Kling AI generation tasks with filtering options. View task history, check statuses, and filter by date range or status. Supports pagination for browsing through large task lists.
| Name | Required | Description | Default |
|---|---|---|---|
| page | No | Page number for pagination (default: 1) | |
| page_size | No | Number of tasks per page (default: 10, max: 100) | |
| status | No | Filter tasks by status | |
| start_time | No | Filter tasks created after this time (ISO 8601 format) | |
| end_time | No | Filter tasks created before this time (ISO 8601 format) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It discloses that the tool supports pagination for large task lists, which is a useful behavioral trait. However, it lacks details on permissions, rate limits, error handling, or the structure of returned data, leaving gaps in understanding the tool's full behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise and front-loaded, starting with the core purpose. Both sentences earn their place by adding context (filtering options and pagination support). It avoids unnecessary details, making it efficient, though it could be slightly more structured by explicitly separating purpose from features.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's moderate complexity (5 parameters, no output schema, no annotations), the description is somewhat complete but has gaps. It covers the purpose and key features like filtering and pagination, but without annotations or output schema, it lacks details on authentication, error handling, and return format, which are important for full context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, meaning all parameters are well-documented in the schema itself. The description mentions filtering by date range or status, which aligns with the schema but doesn't add significant meaning beyond it. With high schema coverage, the baseline score is 3, as the description provides minimal extra value for parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: 'List all your Kling AI generation tasks with filtering options.' It specifies the resource (Kling AI generation tasks) and the action (list with filtering). However, it doesn't explicitly differentiate from sibling tools like 'check_image_status' or 'check_video_status', which might also involve task status checking, though those appear more specific to media types.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage by mentioning 'View task history, check statuses, and filter by date range or status,' suggesting it's for monitoring and filtering tasks. However, it doesn't provide explicit guidance on when to use this tool versus alternatives like the status-checking siblings (e.g., 'check_image_status'), nor does it specify any exclusions or prerequisites for use.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
virtual_try_onA
Apply virtual clothing try-on to a person image using AI. Upload a person image and up to 5 clothing items to see how they would look wearing those clothes. Supports both single and multiple clothing combinations for complete outfit visualization.
| Name | Required | Description | Default |
|---|---|---|---|
| person_image_url | Yes | URL of the person image to try clothes on | |
| cloth_image_urls | Yes | Array of clothing image URLs (1-5 items). Multiple items will be combined into a complete outfit | |
| model_name | No | Model version to use (default: kolors-virtual-try-on-v1.5) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. While it mentions the core functionality, it fails to disclose critical behavioral traits such as required image formats, processing time, rate limits, authentication needs, or what happens with invalid inputs. For a tool with no annotation coverage, this leaves significant gaps in understanding its operational behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the core purpose in the first sentence, followed by specific details in the second. Both sentences earn their place by clarifying the upload process and outfit capabilities without any wasted words, making it efficient and well-structured.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's moderate complexity (3 parameters, no output schema, and no annotations), the description is adequate for basic understanding but incomplete. It covers what the tool does but lacks details on behavioral aspects like error handling or output format, which are crucial for effective use without structured annotations or output schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all parameters thoroughly. The description adds minimal value beyond the schema by mentioning 'up to 5 clothing items' and 'complete outfit visualization,' which slightly reinforces the cloth_image_urls parameter's purpose but doesn't provide additional syntax or format details. Baseline 3 is appropriate when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the specific action ('Apply virtual clothing try-on'), resource ('to a person image using AI'), and scope ('upload a person image and up to 5 clothing items'). It distinguishes this tool from siblings like generate_image or apply_video_effect by focusing on clothing visualization rather than general image/video generation or effects.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for outfit visualization with single or multiple clothing items, but provides no explicit guidance on when to use this tool versus alternatives like generate_image for creating images from scratch. It mentions the capability but lacks context about prerequisites or exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
12 tool updates
- First observed
apply_video_effect - First observed
check_image_status - First observed
check_video_status - First observed
create_lipsync - First observed
extend_video - First observed
generate_image - First observed
generate_image_to_video - First observed
generate_video - First observed
get_account_balance - First observed
get_resource_packages - First observed
list_tasks - First observed
virtual_try_on
TDQS
Scored across 12 tools
Most tools have distinct purposes, such as generate_image for images and generate_video for videos, but there is some overlap between generate_video and generate_image_to_video, which could cause confusion about when to use each. However, descriptions clarify that generate_image_to_video starts from an image, while generate_video starts from text, helping to mitigate ambiguity.
All tool names follow a consistent snake_case pattern with clear verb_noun structures, such as generate_image, check_video_status, and list_tasks. This uniformity makes the tool set predictable and easy to navigate, with no deviations in naming conventions.
With 12 tools, the server is well-scoped for AI media generation and management, covering image creation, video effects, status checks, and account operations. Each tool serves a specific function without redundancy, making the count appropriate for the domain.
The tool set provides comprehensive coverage for AI media workflows, including generation, effects, status tracking, and account management. A minor gap exists in lacking tools for deleting or managing tasks beyond listing, but core operations are well-covered for creative and administrative needs.
Maintenance
Related MCP Connectors
- lightgenOAuthapp.lightgen
Generate and edit images and create short videos inside Claude. Prepaid credits, no subscription.
Use AI models for chat, image, and video generation from Claude Code and other MCP hosts.
One workspace of tools for Claude and ChatGPT: connect 600+ apps, generate media, build tools.
Turn Claude into a creative studio: DNA-locked characters, images, video, voiceover ā 55 tools.
Related MCP Servers
- AlicenseAqualityBmaintenanceKling AI video generation with text-to-video, image-to-video, and multiple quality/speed models via AceDataCloud API.101MIT

PixVerse MCPofficial
AlicenseNot gradedqualityFmaintenanceEnables video generation from text, images, and more through MCP-compatible apps like Claude and Cursor.52MIT- AlicenseNot gradedqualityDmaintenanceConnects Claude.ai with Google's Gemini API to generate images and videos using your own API key.78MIT
- FlicenseNot gradedqualityBmaintenanceEnables Claude to generate videos and images via TopView.ai's API, including text-to-video, image-to-video, text-to-image, and image editing.-