media-mcp
Provides tools for generating images, videos, music, and speech using Google Gemini's AI models.
Click on "Install Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@media-mcpgenerate an image of a cute puppy in a garden"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
media-mcp
MCP server for AI-powered media generation using Google Gemini. Generate images, videos, music, and speech directly from your AI agent.
Features
Image Generation — Create and edit images using Gemini's Nano Banana models with support for multiple aspect ratios, resolutions up to 4K, and reference images
Video Generation — Generate videos with native audio, dialogue, and sound effects using Veo models (text-to-video, image-to-video, video extension)
Music Generation — Create instrumental music with weighted text prompts using Lyria RealTime (genre, instrument, mood control with BPM and scale)
Speech Generation — Convert text to speech with voice selection, multi-speaker support, and natural language style control using Gemini TTS
Related MCP server: Gemini Gen MCP
Installation
Using uvx (recommended)
uvx media-mcpUsing pip
pip install media-mcpPrerequisites
Python 3.10+
Configuration
Set your Gemini API key as an environment variable:
export GEMINI_API_KEY="your-gemini-api-key"Environment Variables
Variable | Required | Description |
| Yes | Google Gemini API key for authentication |
| No | Directory path for saving generated media files (see below) |
Output behavior
When MEDIA_OUTPUT_DIR is set, every generated file is saved to that directory and the tool returns only the file path — no binary data is included in the response. This is the recommended setup because MCP messages are stored in the conversation history, and large base64 payloads pollute context and waste tokens.
When MEDIA_OUTPUT_DIR is not set, the server has no filesystem target, so it returns the raw base64-encoded data directly in the response. This works for quick experiments but is not recommended for production use.
MCP Client Setup
Claude Desktop
Add to your claude_desktop_config.json:
macOS: ~/Library/Application Support/Claude/claude_desktop_config.json
Windows: %APPDATA%\Claude\claude_desktop_config.json
{
"mcpServers": {
"media-mcp": {
"command": "uvx",
"args": ["media-mcp"],
"env": {
"GEMINI_API_KEY": "your-gemini-api-key",
"MEDIA_OUTPUT_DIR": "/path/to/media/output"
}
}
}
}Claude Code
claude mcp add media-mcp --transport stdio -- uvx media-mcpOr add manually to .mcp.json:
{
"mcpServers": {
"media-mcp": {
"type": "stdio",
"command": "uvx",
"args": ["media-mcp"],
"env": {
"GEMINI_API_KEY": "${GEMINI_API_KEY}",
"MEDIA_OUTPUT_DIR": "/path/to/media/output"
}
}
}
}Tools
generate_image
Generate or edit images using Gemini's Nano Banana models.
Parameter | Type | Required | Default | Description |
| string | Yes | — | Text description of the image to generate |
| enum | No |
|
|
| enum | No |
|
|
| enum | No |
|
|
| list[str] | No | — | Base64-encoded reference images |
| enum | No |
|
|
| bool | No |
| Enable Google Search grounding |
Example prompt: "A watercolor painting of a cozy cabin in the mountains during autumn"
generate_video
Generate videos with native audio using Veo models.
Parameter | Type | Required | Default | Description |
| string | Yes | — | Text description including dialogue, sound effects, camera directions |
| enum | No |
|
|
| enum | No |
|
|
| enum | No | — |
|
| str | No | — | Base64-encoded image for first frame |
| str | No | — | Base64-encoded image for last frame |
| list[str] | No | — | Up to 3 base64-encoded reference images |
Example prompt: "A slow dolly shot through a neon-lit alley at night, rain falling, 'Where are you going?' whispered softly, footsteps echoing"
generate_music
Generate instrumental music using Lyria RealTime with weighted prompts.
Parameter | Type | Required | Default | Description |
| list[dict] | Yes | — | Weighted prompts, e.g. |
| int | No | — | Tempo in beats per minute |
| float | No |
| Randomness/creativity control |
| str | No | — | Musical scale constraint (e.g. |
| int | No |
| Duration of the output clip |
Example prompts: [{"text": "Piano", "weight": 2.0}, {"text": "Meditation", "weight": 0.5}]
generate_speech
Convert text to speech with voice and style control.
Parameter | Type | Required | Default | Description |
| string | Yes | — | Text to speak. For multi-speaker, format as dialogue with speaker names. |
| enum | No |
|
|
| str | No | — | Voice name: |
| bool | No |
| Enable multi-speaker mode |
| list[dict] | No | — | Speaker-to-voice mapping, e.g. |
| str | No | — | Style guidance, e.g. "Read in a calm, slow pace" |
Example: Text: "Welcome to the show!" with voice_name: "Kore" and style_instructions: "Say cheerfully"
Troubleshooting
"GEMINI_API_KEY environment variable is not set"
Set the environment variable before starting the server:
export GEMINI_API_KEY="your-key-here"When using Claude Desktop or Claude Code, pass the key via the env block in your MCP configuration (see MCP Client Setup).
"Authentication failed" or 401 errors
Your API key may be invalid or expired. Verify it at Google AI Studio.
"Rate limit or quota exceeded" or 429 errors
Wait a moment and retry. Check your API quota at Google AI Studio.
"Content blocked by safety filter"
Modify your prompt to avoid restricted content. The Gemini API applies safety filters to all generated media.
Python version errors
media-mcp requires Python 3.10 or later. Check your version:
python --versionLicense
MIT
Available Tools
4 toolsgenerate_imageB
Generate or edit images using Google's Gemini image generation models.
Supports conversational image creation/editing, multi-turn workflows, images with embedded text, infographics, and interleaved text+image output.
| Name | Required | Description | Default |
|---|---|---|---|
| model | No | nano-banana-2 | |
| prompt | Yes | ||
| image_size | No | 1K | |
| aspect_ratio | No | 1:1 | |
| thinking_level | No | minimal | |
| reference_images | No | ||
| use_image_search | No | ||
| use_google_search | No | ||
| response_modalities | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Without annotations, the description adds behavioral context by noting support for multi-turn and interleaved output, but lacks detail on safety, permissions, or side effects. The description is adequate but not comprehensive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is brief (two sentences) and front-loaded with the core purpose, followed by supporting capabilities. Every sentence earns its place with no redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the complexity (9 parameters, no output schema, no annotations), the description is insufficient. It does not explain return values or how parameters like model or reference_images affect behavior, leaving significant gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, meaning no parameter descriptions exist in the schema. The description does not explain any of the 9 parameters, leaving the agent to infer from names and enums alone.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool generates or edits images using Gemini models and lists specific capabilities like conversational workflows, embedded text, and infographics. It distinguishes from sibling tools (speech, video, music) by focusing on images.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description describes what the tool does but provides no explicit guidance on when to use it versus the sibling tools (generate_speech, generate_video, generate_music). No when-not or alternative scenarios are mentioned.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
generate_musicB
Generate instrumental music from weighted text prompts using Google's Lyria model.
Each prompt has a 'text' describing genre/instrument/mood and a 'weight' for emphasis. The server manages the streaming session internally.
| Name | Required | Description | Default |
|---|---|---|---|
| bpm | No | ||
| scale | No | ||
| prompts | Yes | ||
| temperature | No | ||
| duration_seconds | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
There are no annotations provided, so the description must fully disclose behavioral traits. It mentions the tool generates music and internally manages streaming, but it does not disclose whether the operation is destructive, requires authentication, has rate limits, or what the output format (e.g., audio file URL) is. This is insufficient for an AI agent to understand side effects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is only two sentences, making it very concise. It front-loads the primary purpose and then adds a clarifying technical detail. However, it could be slightly more structured by separating parameter explanation from behavioral notes.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the complexity of five parameters and no output schema, the description is incomplete. It fails to mention what the tool returns (e.g., an audio file URL or stream), and it does not cover the optional parameters. This leaves significant gaps for an AI agent to correctly invoke the tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must add meaning to parameters. It clarifies that the 'prompts' array contains objects with 'text' and 'weight' fields, which aids the agent. However, it does not explain other parameters like bpm, scale, temperature, or duration_seconds, leaving them underdocumented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states it generates instrumental music from weighted text prompts using Google's Lyria model, specifying the verb (generate), resource (music), and technology, distinguishing it from sibling tools like generate_image or generate_speech.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains the prompt structure and mentions server streaming, but it does not provide explicit guidance on when to use this tool versus alternatives, nor does it include prerequisites or exclusions. The context signals indicate sibling tools are different modalities, so the usage is partially implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
generate_speechA
Generate speech audio from text using Gemini TTS models.
Supports single-speaker and multi-speaker modes with voice selection and natural language style control.
| Name | Required | Description | Default |
|---|---|---|---|
| text | Yes | ||
| model | No | flash-tts | |
| speakers | No | ||
| voice_name | No | ||
| multi_speaker | No | ||
| style_instructions | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It discloses modes, voice selection, and style control, which are key behavioral traits. However, it omits details like supported audio formats, latency, or that it uses specific models (though model enum is in schema). Missing some transparency.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two concise sentences, front-loaded with purpose, no unnecessary words. Efficiently conveys key capabilities.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given 6 parameters, no output schema, and no annotations, the description is incomplete. It lacks details on the 'speakers' parameter structure, relationship between voice_name and speakers, and return format. Agent would need to infer or guess for effective use.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so description must add meaning. It hints at model, multi-speaker mode, voice selection, and style instructions, but does not explicitly map to parameters or explain complex ones like 'speakers' array. Adds some value but insufficient for full understanding.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states it generates speech audio from text using Gemini TTS models, and mentions modes and voice selection. It distinguishes from sibling tools (image, video, music) implicitly by modality, but does not explicitly differentiate.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context for when to use the tool: for text-to-speech with optional multi-speaker mode, voice selection, and style control. It does not explicitly state when not to use or list alternatives, which is acceptable given sibling tools cover different modalities.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
generate_videoB
Generate videos from text prompts or reference images using Google's Veo models.
Supports text-to-video, image-to-video, video extension, and frame-specified generation. Generation is asynchronous.
| Name | Required | Description | Default |
|---|---|---|---|
| model | No | veo-3.1 | |
| prompt | Yes | ||
| resolution | No | ||
| aspect_ratio | No | 16:9 | |
| extend_video_id | No | ||
| last_frame_image | No | ||
| reference_images | No | ||
| first_frame_image | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Discloses that generation is asynchronous, which is a key behavioral trait. However, with no annotations, the description does not cover other aspects like auth needs, rate limits, or return format, leaving gaps in transparency.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three concise sentences front-loading key information about purpose, capabilities, and async nature. No wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Missing crucial details for an async tool, such as how to retrieve generated videos, polling mechanism, or expected output format. With no output schema and 8 parameters, the description is insufficient for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description does not explain the purpose of parameters like extend_video_id, reference_images, etc. It only lists modes without mapping them to specific parameters, providing minimal additive value.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Clearly states that the tool generates videos from text prompts or reference images using Google's Veo models, and lists specific modes (text-to-video, image-to-video, etc.), distinguishing it from sibling tools like generate_image.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Does not provide explicit guidance on when to use this tool vs. alternatives (generate_image, generate_speech, generate_music). The description implies usage for video generation but lacks when-not-to-use or prerequisites.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
4 tool updates
v0.2.7- First observed
generate_image - First observed
generate_music - First observed
generate_speech - First observed
generate_video
TDQS
Scored across 4 tools
Each tool targets a distinct medium (image, speech, video, music) with no overlap. Descriptions clearly differentiate their purposes and capabilities.
All tools follow a consistent 'generate_<medium>' pattern (e.g., generate_image, generate_speech), making naming predictable and intuitive.
Four tools cover the core media generation types without unnecessary duplication. The count is appropriate for the server's focused purpose.
The tool surface covers generation for all major media types (image, speech, video, music) with support for editing and customization, leaving no obvious gaps.
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Connectors
MCP server for Google Veo AI video generation
MCP server for OpenAI Sora AI video generation
MCP server for Grok Imagine AI video generation
MCP server for Luma Dream Machine AI video generation
Related MCP Servers
- AlicenseBqualityBmaintenanceA MCP server that provides AI-powered image generation capabilities through Google's Gemini 2.5 Flash Image model.4399MIT
- AlicenseBqualityCmaintenanceMCP server for generating images and audio using Google's Gemini AI models.22MIT
- AlicenseBqualityBmaintenanceAn MCP server for AI-powered image generation, editing, analysis, and transformation using Google's Gemini and Imagen 4 models.192AGPL 3.0
- AlicenseAqualityDmaintenanceMCP server for Google Gemini image generation with configurable model support, enabling text-to-image generation, image editing, and iterative refinement.641MIT