Media-infrastructure
This server provides AI-powered media generation and editing tools via the VAP API, along with an OpenAI-compatible API for coding agents (powered by Nemesis Deep Coder).
Image Capabilities
Generate images from text prompts (Flux2 Pro) with quality and aspect ratio options
Edit images using natural language instructions, including multi-image editing
Inpaint images to remove or replace specific objects/areas
Remove backgrounds from images
Upscale images by 2x or 4x
Video Capabilities
Generate videos from text prompts (Veo 3.1) with duration (4/6/8s), resolution (720p/1080p), aspect ratio, and optional AI audio
Trim videos to specific time ranges
Merge multiple video clips into a single continuous video
Music Capabilities
Generate music from text descriptions (Suno V5) with control over duration, format (MP3/WAV), instrumental mode, genre, mood, and loudness normalization
Task & Account Management
Check status and retrieve results for any generation or editing task
List recent tasks with optional status filtering
Check account balance (available, reserved, and usable)
Estimate costs before generating images, videos, or music
Utilizes FFmpeg for a media production pipeline that includes format conversion, audio normalization, and tools for trimming and merging video clips.
Provides tools for generating original music using Suno V5, enabling text-to-audio creation with support for custom prompts, cost estimation, and task status tracking.
VAP AI
Agent-native AI platform for AI Room, Media API, and Coding Plan API.
Website | Developer Hub | AI Room | Dashboard | Status
Products
AI Room
Conversational creative workspace where custom-trained AI agents use session context to create images, video, voice, and music.
Product URL: https://vapagent.com/new
Plans: Lite, Pro, and Max monthly Room plans
Media API
Unified generation API for image, video, and music workflows. Current public model surfaces include Pimo AI-Video, Aura Image Turbo, and Pira V5.5.
Developer Hub: https://vapagent.com/developer/
API base URL:
https://api.vapagent.com/api/v1Create generation:
POST /api/v1/generationsCreate operation:
POST /api/v1/operationsAuthentication: product-scoped VAP Media API key
MCP endpoint:
https://api.vapagent.com/mcp
Coding Plan API
OpenAI-compatible API for coding agents, IDEs, editors, and automation workflows. Powered by Nemesis Deep Coder with model ID vap-code.
Developer Hub: https://vapagent.com/developer/
Model page: https://vapagent.com/models/nemesis-deep-coder.html
Harness guide: https://vapagent.com/integrations/coding-harnesses.html
API base URL:
https://api.vapagent.com/v1Model ID:
vap-codeResponses endpoint:
POST /v1/responsesChat Completions endpoint:
POST /v1/chat/completionsAuthentication: product-scoped VAP Coding Plan API key
Related MCP server: Eversince MCP Server
Start From The Product Surface
Use these entry points for new integrations:
Need | Link |
Use AI Room | |
Generate a Media API key | |
Generate a Coding Plan API key | |
View Coding Plan API plans | |
Read Developer Hub | |
See Nemesis Deep Coder |
Media API Via MCP
This repository remains the public GitHub and MCP discovery surface for VAP Media API integrations. MCP is still supported for Claude Desktop, Claude Code, Cursor-compatible MCP clients, and other agent workflows that call VAP media tools.
Claude Desktop example:
{
"mcpServers": {
"vap": {
"url": "https://api.vapagent.com/mcp",
"headers": {
"Authorization": "Bearer YOUR_MEDIA_API_KEY"
}
}
}
}For clients that do not support headers directly, use the local proxy in mcp/vap_mcp_proxy.py and set VAP_API_KEY.
API Examples
Media API
curl -X POST https://api.vapagent.com/api/v1/generations \
-H "Authorization: Bearer YOUR_MEDIA_API_KEY" \
-H "Content-Type: application/json" \
-d '{"kind":"image","prompt":"a neon city at night"}'Coding Plan API
curl -X POST https://api.vapagent.com/v1/chat/completions \
-H "Authorization: Bearer YOUR_CODING_PLAN_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"vap-code","messages":[{"role":"user","content":"Review this patch."}]}'Current Product Contracts
Use these current public contracts from Developer Hub:
Media API:
https://api.vapagent.com/api/v1Media generation:
POST /api/v1/generationsMedia operations:
POST /api/v1/operationsMCP:
https://api.vapagent.com/mcpCoding Plan API:
https://api.vapagent.com/v1
Discovery Files
server.json: MCP Registry server metadatamcp.json: MCP package metadataglama.json: Glama MCP listing metadatamcp/tools.json: MCP tool schema metadatamcp/vap_mcp_proxy.py: local stdio/http proxy for MCP clients
Links
Website: https://vapagent.com/
Developer Hub: https://vapagent.com/developer/
AI Room: https://vapagent.com/new
MCP guide: https://vapagent.com/mcp.html
Webhooks: https://vapagent.com/webhooks.html
Status: https://vapagent.com/status.html
Support: support@vapagent.com
License
MIT License. See LICENSE.
Available Tools
16 toolsai_editA
Edit images using AI with natural language prompts. Supports single and multi-image editing. Cost: $0.15. Requires Tier 1+.
| Name | Required | Description | Default |
|---|---|---|---|
| media_url | Yes | URL of the primary image to edit | |
| prompt | Yes | Natural language instruction for editing | |
| additional_images | No | Up to 2 additional reference image URLs | |
| aspect_ratio | No | Output aspect ratio (default: match input image) | match_input_image |
| seed | No | Random seed for reproducible results |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden. It reveals cost ($0.15) and tier requirement (Tier 1+), but lacks details on output format, whether edits are destructive, or limits on image sizes.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is extremely concise: two sentences covering purpose, capability, cost, and tier. No unnecessary words, and the key information is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given 5 parameters and no output schema, the description covers purpose, capability, cost, and tier. Missing are output format, error handling, or rate limits, but overall sufficient for basic invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
All 5 parameters have clear descriptions in the input schema (100% coverage). The description adds minimal value beyond schema, only hinting at multi-image via 'Supports single and multi-image editing'.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states 'Edit images using AI with natural language prompts' and distinguishes from sibling tools like background_remove or upscale by emphasizing natural language editing and multi-image capability.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It mentions 'Supports single and multi-image editing' and cost/tier, but does not explicitly compare to alternatives like inpaint or background_remove, leaving the agent to infer when to use this tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
background_removeA
Remove the background from an image, leaving only the subject with transparency. Cost: $0.10. Requires Tier 1+.
| Name | Required | Description | Default |
|---|---|---|---|
| media_url | Yes | URL of the image to process |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description adds value by disclosing cost and tier requirements. However, it lacks information on rate limits, processing time, output format, or whether the original file is modified. This is adequate but not comprehensive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is extremely concise: one sentence for purpose, two short sentences for cost and tier. Every word contributes meaning with no redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's simplicity (1 parameter, no output schema), the description covers the core action, cost, and access requirement. It could mention supported image formats or output specifics, but is largely complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The only parameter (media_url) is fully described in the input schema. The tool description does not add additional meaning beyond the schema, so it scores baseline 3.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: removing background from an image and leaving transparency. It uses a specific verb-resource pair and distinguishes from siblings like 'inpaint' or 'ai_edit' which have different editing operations.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description mentions cost ($0.10) and a required tier (Tier 1+), providing context on when the tool can be used. However, it does not explicitly state when to use this tool over alternatives, so it gets a 4.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
check_balanceA
Check VAP account balance. Returns available, reserved, and usable balances.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It correctly indicates the tool returns balance data without side effects. While it does not explicitly state it is read-only, the nature of a balance check implies no destructive actions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences, front-loaded with the main action, and no extraneous information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple, parameterless tool with no output schema, the description adequately covers purpose and return values. No additional information is needed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There are no parameters (schema coverage 100%). The description adds value by explaining the return fields (available, reserved, usable), which goes beyond the empty schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states it checks a VAP account balance and lists the specific types of balances returned (available, reserved, usable). It distinguishes well from sibling tools which focus on generation and editing tasks.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description does not provide explicit guidance on when to use this tool versus siblings. However, the tool is simple and parameterless, and the context of siblings (generation/editing tools) makes the usage clear enough.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
estimate_costB
Estimate the cost of an image generation before executing. Cost: $0.18
| Name | Required | Description | Default |
|---|---|---|---|
| quality | No | Generation quality level | standard |
| num_outputs | No | Number of images to generate |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, and the description only mentions a fixed cost ($0.18), failing to disclose whether the tool makes an API call, is idempotent, or how parameters affect the cost.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence with no wasted words, though it could benefit from a bit more detail.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, and the description does not explain what the tool returns (e.g., a cost value), nor does it clarify how inputs affect the cost, leaving the agent uninformed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3. The description adds no additional meaning beyond the schema for 'quality' and 'num_outputs', missing a chance to explain their impact on cost.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool estimates the cost of image generation, but it does not explain how cost varies with parameters like quality or num_outputs, leaving some ambiguity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'before executing' implies use prior to generate_image, but there is no explicit guidance on when not to use it or how it compares to sibling estimation tools for music/video.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
estimate_music_costC
Estimate the cost of music generation. Cost: $0.68 (Suno V5)
| Name | Required | Description | Default |
|---|---|---|---|
| duration | No | Music duration in seconds | |
| audio_format | No | Output format. WAV adds +$0.10 | mp3 |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It states a cost but doesn't explain if this is per-use, includes taxes, requires authentication, has rate limits, or what the output format is. For a cost estimation tool with zero annotation coverage, this leaves critical behavioral traits unspecified.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is extremely concise—just one sentence with no wasted words. It front-loads the core purpose and includes essential cost information efficiently. Every element earns its place, making it easy for an agent to parse quickly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a cost estimation tool with no annotations and no output schema, the description is incomplete. It doesn't specify what the estimate returns (e.g., total cost, breakdown), how parameters affect the calculation beyond a vague format note, or any error conditions. Given the complexity of pricing and lack of structured data, more context is needed for effective use.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 100% description coverage, with clear documentation for both parameters (duration and audio_format). The description adds minimal value beyond the schema: it mentions a base cost and implies format affects price ('WAV adds +$0.10'), but doesn't fully explain how parameters influence the estimate. Given high schema coverage, a baseline score of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: 'Estimate the cost of music generation.' It specifies the verb ('estimate') and resource ('cost of music generation'), making the function unambiguous. However, it doesn't explicitly differentiate from sibling tools like 'estimate_cost' or 'estimate_video_cost', which prevents a perfect score.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives. It mentions a specific cost ('$0.68 (Suno V5)') but doesn't clarify if this is a fixed rate, how it relates to parameters, or when to choose this over 'estimate_cost' or 'estimate_video_cost'. Without usage context or exclusions, the agent lacks direction.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
estimate_video_costC
Estimate the cost of a video generation. Cost: $1.96 (Veo 3.1)
| Name | Required | Description | Default |
|---|---|---|---|
| duration | No | Video duration in seconds | |
| generate_audio | No | Whether audio will be generated | |
| resolution | No | Video resolution. 1080p adds +33% cost | 720p |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries full burden but provides minimal behavioral context. It states the base cost but doesn't explain how parameters affect the final estimate, whether this is a read-only operation, if it makes any external calls, or what the return format looks like. The description is insufficient for a tool with cost implications.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Extremely concise single sentence with zero waste. Every word earns its place by stating the core function and providing the base cost figure. No unnecessary elaboration or repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a cost estimation tool with no annotations and no output schema, the description is incomplete. It doesn't explain how the estimate is calculated from parameters, what format the response takes, or whether this is a real-time quote versus cached pricing. Users need more context to understand the tool's behavior.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so parameters are well-documented in the schema. The description adds minimal value beyond the schema - it mentions the base cost but doesn't explain how parameters like resolution affect pricing (though the schema notes 1080p adds +33% cost). Baseline 3 is appropriate when schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: estimating video generation cost with a specific price point. It distinguishes from siblings like 'estimate_cost' (generic) and 'estimate_music_cost' (music-specific) by specifying video generation. However, it doesn't explicitly mention the parameters that affect cost calculation beyond the base price.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this tool versus alternatives like 'estimate_cost' (which might be more generic) or 'generate_video' (which actually creates videos). The description doesn't mention prerequisites, dependencies, or typical usage scenarios beyond the basic function.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
generate_imageA
Generate an AI image from text prompt using VAP (Flux2 Pro). Returns a task ID for async tracking. Cost: $0.18
| Name | Required | Description | Default |
|---|---|---|---|
| prompt | Yes | Detailed description of the image to generate. Note: If aspect ratio is mentioned in the prompt (e.g., '16:9', 'widescreen', 'portrait'), also pass it in the aspect_ratio parameter for guaranteed correct dimensions. | |
| aspect_ratio | No | Output image aspect ratio. If the user mentions a specific ratio like '16:9' or 'widescreen' in their prompt, extract and pass it here explicitly for best results. | 1:1 |
| quality | No | Generation quality (high costs 1.5x) | standard |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations present, so description carries burden. It discloses async behavior and cost, but lacks detail on error handling, rate limits, or side effects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Extremely concise: two sentences and inline schema notes. Every piece adds value with no redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema, but description explains async tracking and cost. Adequate for a simple generation tool, though could mention how to retrieve results via get_task.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema covers all 3 parameters with descriptions. The description adds extra guidance on aspect ratio extraction and quality cost, enhancing understanding beyond schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Clearly states the action (generate), resource (AI image), model (VAP Flux2 Pro), and async return type (task ID). Distinct from siblings like ai_edit or inpaint.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No explicit guidance on when to use vs alternatives. Only cost is mentioned, but no comparison with other image or generation tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
generate_musicA
Generate AI music from text description using VAP (Suno V5). Returns a task ID for async tracking. Cost: $0.68.
IMPORTANT: Send ONLY the music description. Do NOT include any instructions or meta-text.
Describe: genre, mood, instruments, tempo, vocal style (or specify instrumental).
Example prompt: "Upbeat indie folk song with acoustic guitar, warm vocals, and light percussion. Feel-good summer vibes.
| Name | Required | Description | Default |
|---|---|---|---|
| prompt | Yes | Music description (200-500 chars recommended). Include genre, mood, instruments, tempo. | |
| instrumental | No | Generate without vocals (instrumental only) | |
| duration | No | Target duration in seconds (30-480, default 120 = 2 min) | |
| loudness_preset | No | Loudness normalization. streaming=-14 LUFS (YouTube/Spotify), apple=-16 LUFS, broadcast=-23 LUFS (TV/EBU R128) | streaming |
| audio_format | No | Output format. WAV for enterprise/lossless (+$0.10) | mp3 |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It effectively describes key behavioral traits: it returns a task ID for async tracking (implying asynchronous operation), states the cost ($0.68), and provides constraints like character recommendations (200-500 chars) and format instructions. It does not cover aspects like rate limits or error handling, but offers substantial context beyond basic functionality.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-structured and appropriately sized, with key information front-loaded (purpose, cost, async tracking). It uses bullet-like formatting for instructions and includes an example, which aids clarity. Some sentences could be more concise (e.g., the cost note is brief but clear), but overall, it avoids unnecessary verbosity and each sentence serves a purpose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (5 parameters, async operation, cost implications) and no output schema, the description is moderately complete. It covers the purpose, usage, cost, and async nature, but lacks details on output (e.g., what the task ID leads to, error cases, or links to sibling tools like get_task). With no annotations and no output schema, more context on behavioral outcomes would improve completeness.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema description coverage is 100%, so the baseline score is 3. The description adds some value by emphasizing the prompt parameter with an example and instructions, but it does not provide significant additional semantics for other parameters like instrumental, duration, loudness_preset, or audio_format beyond what the schema already documents. The description compensates minimally for the high schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: 'Generate AI music from text description using VAP (Suno V5).' It specifies the verb ('Generate'), resource ('AI music'), and technology ('VAP (Suno V5)'), distinguishing it from sibling tools like generate_image or generate_video. The description is specific and unambiguous about what the tool does.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context on when to use this tool by specifying it's for generating music from text descriptions, with an example prompt. It includes important usage instructions like 'Send ONLY the music description' and what to include in the description. However, it does not explicitly state when not to use it or compare it to alternatives like estimate_music_cost, which would be needed for a score of 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
generate_videoA
Generate an AI video from text prompt using VAP (Veo 3.1). Returns a task ID for async tracking. Cost: $1.96. IMPORTANT: Send ONLY the video description. Do NOT include any instructions, guidelines, or meta-text. Just the pure visual description.
| Name | Required | Description | Default |
|---|---|---|---|
| prompt | Yes | ONLY the visual description of the video. Do NOT include instructions or guidelines. Example: 'Cinematic aerial shot of a coastal cliff at golden hour, warm sunlight, gentle waves, camera slowly drifting forward' | |
| duration | No | Video duration in seconds (4, 6, or 8) | |
| aspect_ratio | No | Video aspect ratio. 16:9 for landscape/widescreen, 9:16 for portrait/vertical (TikTok, Reels). Extract from user's prompt if mentioned. | 16:9 |
| generate_audio | No | Generate audio with the video (costs more) | |
| resolution | No | Video resolution. 1080p recommended for enterprise (+33% cost) | 720p |
| negative_prompt | No | What to avoid in the video generation |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It effectively communicates key behavioral traits: the asynchronous nature ('Returns a task ID for async tracking'), cost implications ('Cost: $1.96'), and strict input requirements ('IMPORTANT: Send ONLY the video description...'). It doesn't mention rate limits, authentication needs, or error conditions, but provides substantial operational context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is perfectly front-loaded with the core purpose, followed by critical behavioral information (async tracking, cost), then essential usage instructions. Every sentence earns its place with zero waste. The structure moves from general to specific in a logical flow.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex video generation tool with 6 parameters and no annotations or output schema, the description provides substantial context about the operation's nature, cost, and input requirements. It effectively compensates for the lack of output schema by explaining the async task ID return. However, it doesn't address potential failure modes, quality expectations, or integration with sibling tools like get_task for tracking.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all 6 parameters thoroughly. The description reinforces the critical constraint for the 'prompt' parameter ('Send ONLY the video description...'), adding some semantic emphasis beyond the schema. However, it doesn't provide additional meaning for other parameters beyond what's already in their schema descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the specific action ('Generate an AI video from text prompt') and resource ('using VAP (Veo 3.1)'), distinguishing it from sibling tools like generate_image or generate_music. It provides a complete functional statement beyond just restating the tool name.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context about when to use this tool ('Generate an AI video from text prompt') and includes important usage instructions ('Send ONLY the video description...'). However, it doesn't explicitly differentiate when to choose this over alternatives like video_merge or video_trim, nor does it mention prerequisites like checking balance first.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_operationB
Get the status and result of an operation. Returns output URL when completed.
| Name | Required | Description | Default |
|---|---|---|---|
| operation_id | Yes | Operation UUID returned from an operation tool |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations provided; description only mentions output URL on completion, lacking details on pending states or error handling. Insufficient for a polling tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, no redundancy, front-loaded with purpose and key note about output URL.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema or annotations; description should cover status fields and error handling but only mentions completion scenario.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema covers parameter fully (100% coverage), description adds no extra meaning. Baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description clearly states it retrieves operation status and result, distinguishing it from sibling generation/editing tools. Specific verb+resource.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Implies use after obtaining an operation_id from an operation tool, but no explicit guidance on polling behavior or when not to use.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_taskA
Get the status and result of a generation task. Returns image URL when completed.
| Name | Required | Description | Default |
|---|---|---|---|
| task_id | Yes | Task UUID returned from generate_image |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations provided, so description must convey behavior. It explains that image URL is returned when completed, which is key. However, does not mention behavior for incomplete or failed tasks, or polling expectations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two brief sentences that are front-loaded with the main action. No unnecessary words; every sentence adds value.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple tool with one parameter and no output schema, the description adequately covers what it does (get status/result) and what it returns (image URL when completed). No gaps given the simplicity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Input schema covers the single parameter (task_id) with 100% coverage. The description adds no additional meaning beyond what the schema already provides, aligning with baseline of 3.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Clearly states it gets status and result of a generation task, and specifies return of image URL when completed. Distinguishes from sibling tools like generate_image (creation) and list_tasks (listing all tasks).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Implies use for checking completion of a generation task, but lacks explicit guidance on when to avoid or alternatives. Does not exclude other uses.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
inpaintA
Remove or replace objects in an image using AI inpainting. Cost: $0.15. Requires Tier 1+.
| Name | Required | Description | Default |
|---|---|---|---|
| media_url | Yes | URL of the image to edit | |
| prompt | Yes | What to remove, replace, or change in the image | |
| mask_url | No | Optional mask image URL (white = edit area, black = keep) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries full burden for behavioral disclosure. It reveals cost and tier requirement (positive) but omits other behaviors like processing time, image format limits, size constraints, or whether the operation is destructive. Partial transparency.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two concise sentences: first states core purpose, second adds cost and tier. No unnecessary words, front-loaded with key information. Efficient and scannable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers purpose and essential constraints (cost, tier), but for a tool with no output schema and no annotations, it lacks details on return format, image constraints, error handling, or typical behavior. Adequate but not fully complete for an AI agent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so baseline is 3. The description does not add additional meaning to parameters beyond what the schema provides (e.g., media_url, prompt, optional mask_url). No extra context or examples are given.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's action: 'Remove or replace objects in an image using AI inpainting.' It distinguishes itself from siblings like background_remove (removes backgrounds) and generate_image (creates new images), making its specific purpose immediate.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides cost ($0.15) and access tier (Tier 1+), which inform usage decisions, but does not explicitly state when to use this tool versus alternatives like ai_edit or background_remove. Usage is implied but no exclusions or comparative guidance given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_tasksC
List recent generation tasks with optional status filter.
| Name | Required | Description | Default |
|---|---|---|---|
| status | No | Filter by task status | |
| limit | No | Maximum number of tasks to return |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description bears full responsibility. It does not disclose ordering, pagination behavior, or what 'recent' means. No mention of user scope or whether it lists all tasks or only the caller's.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Single sentence, no fluff. Front-loaded with key action. Could be more informative but is appropriately concise.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema is provided, so description should explain return format or fields. It does not. Agent lacks information about what data is returned for each task.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3. Description adds 'optional status filter', clarifying the status parameter's role. No additional details on limit beyond schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool lists recent generation tasks, with the verb 'list' and object 'generation tasks'. It distinguishes from sibling 'get_task' which retrieves a single task. However, 'recent' is vague and could be more specific.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this tool vs siblings like 'get_task' or 'get_operation'. The agent must infer from context. No usage prerequisites or exclusions are mentioned.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
upscaleB
Upscale/enhance an image to higher resolution using AI. Cost: $0.15. Requires Tier 1+.
| Name | Required | Description | Default |
|---|---|---|---|
| media_url | Yes | URL of the image to upscale | |
| scale | No | Upscale factor (2x or 4x) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations provided; description only mentions cost and tier. Does not disclose behavior such as whether original is preserved, rate limits, or output format.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Very concise single sentence with essential info. No fluff, but lacks structure like bullet points.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Simple tool with 2 parameters; cost and tier add context. However, no description of output format (expected since no output schema) leaves minor gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%; description adds no extra meaning beyond schema. Baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Clearly states upscaling/enhancing an image using AI. Distinguishable from siblings like generate_image or background_remove, though could be more specific.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides cost and tier requirement ($0.15, Tier 1+), helpful context. No explicit when-to-use vs alternatives, but siblings imply different operations.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
video_mergeA
Merge multiple video clips into one continuous video. Cost: $0.05. Requires Tier 1+.
| Name | Required | Description | Default |
|---|---|---|---|
| media_urls | Yes | URLs of videos to merge (in playback order) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries full burden. It discloses the merge operation and cost/tier but lacks details on output format, failure modes, or processing behavior. Adequate but not thorough.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is one sentence with no wasted words. It efficiently conveys the core function and additional cost/tier info.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the simplicity of the tool (one parameter, no output schema), the description covers the essential behavior and access constraints. Minor gap: no mention of return value or error handling, but acceptable for a straightforward operation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, and the schema property description already includes 'in playback order'. The tool description adds no further meaning to the parameter, so baseline score of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states it merges multiple video clips into one continuous video. It is specific to the video_merge operation and distinguishes it from siblings like video_trim.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description mentions cost and tier requirement, providing some context, but it does not explicitly guide when to use this tool versus alternatives like video_trim or other merge-related tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
video_trimC
Trim a video to a specific time range. Cost: $0.05. Requires Tier 1+.
| Name | Required | Description | Default |
|---|---|---|---|
| media_url | Yes | URL of the video to trim | |
| start_time | Yes | Start time in seconds | |
| end_time | Yes | End time in seconds |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description must carry the burden. It discloses cost and tier requirement but does not mention side effects (e.g., original video preserved?), output format, or any other behavioral traits beyond basic operation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is very short (two sentences) and front-loaded with action and cost. However, it is minimal and could benefit from more structure or detail without becoming verbose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given three required simple parameters and no output schema or annotations, the description is incomplete. It lacks information on return values, error handling, or prerequisites beyond tier. The cost and tier notes add value but are insufficient for full agent decision-making.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% with descriptions for all parameters. The description adds 'time range' context but no additional meaning beyond what the schema already provides. Baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The verb 'trim' and resource 'video' are clear, and the description specifies 'to a specific time range'. It distinguishes from siblings like video_merge (merge) and ai_edit (edit) by focusing on trimming, though it does not explicitly differentiate.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives. It mentions cost and tier requirement but no context for selection among siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
16 tool updates
v1.12.6- First observed
ai_edit - First observed
background_remove - First observed
check_balance - First observed
estimate_cost - First observed
estimate_music_cost - First observed
estimate_video_cost - First observed
generate_image - First observed
generate_music - First observed
generate_video - First observed
get_operation - First observed
get_task - First observed
inpaint - First observed
list_tasks - First observed
upscale - First observed
video_merge - First observed
video_trim
TDQS
Scored across 16 tools
Most tools have distinct purposes, but there is notable overlap between get_operation and get_task, which both retrieve task/operation statuses, potentially causing confusion. Additionally, the three estimate_* tools are clearly differentiated by media type, but their naming and purpose similarity could lead to misselection if not carefully read.
The naming is mostly consistent with a verb_noun pattern (e.g., generate_image, check_balance, list_tasks), but there are minor deviations like ai_edit (noun_verb) and background_remove (noun_verb). These inconsistencies are few and do not severely impact readability, but they break the overall pattern.
With 16 tools, the count is reasonable for a media infrastructure server covering image, video, and music generation and editing, plus account management. It is slightly on the higher side but well within a manageable scope, as each tool serves a specific function in the domain.
The tool set provides good coverage for media generation and editing, including CRUD-like operations (generate, edit, upscale, list, get status) and cost estimation. However, there are minor gaps, such as no explicit tool for deleting tasks or managing media assets beyond generation and editing, which agents might need to work around.
Maintenance
Related MCP Connectors
FFmpeg as a service for AI agents: typed video editing tools, async jobs, downloadable outputs.
Video, audio, and image processing for AI agents: convert, transcribe, upscale - 150+ operations.
- RendobarOAuthcom.rendobar
Transform video, audio and images, and generate media from prompts. FFmpeg, captions, models.
A real timeline video editor for AI agents: journaled edits, FFmpeg/MLT rendering, exports
Related MCP Servers
- AlicenseBqualityCmaintenanceExecution control layer for AI agents - Reserve, execute, burn/refund pattern for media generation162MIT

Eversince MCP Serverofficial
AlicenseNot gradedqualityDmaintenanceA creative agent that plans and executes across image, video, and audio. Uses 30+ tools, orchestrates 20+ AI models, and does agentic timeline editing.MIT- AlicenseBqualityDmaintenanceEnables AI video automation pipeline: ComfyUI image-to-video, FFmpeg processing, After Effects template rendering, review, and multi-platform publishing preparation.9MIT
- AlicenseAqualityCmaintenanceProvides 30+ FFmpeg video and audio editing tools via MCP, enabling AI assistants to perform operations like trimming, transcoding, overlays, and composition directly.274MIT