Skip to main content
Glama

media-mcp

MCP server for AI-powered media generation using Google Gemini. Generate images, videos, music, and speech directly from your AI agent.

Features

  • Image Generation — Create and edit images using Gemini's Nano Banana models with support for multiple aspect ratios, resolutions up to 4K, and reference images

  • Video Generation — Generate videos with native audio, dialogue, and sound effects using Veo models (text-to-video, image-to-video, video extension)

  • Music Generation — Create instrumental music with weighted text prompts using Lyria RealTime (genre, instrument, mood control with BPM and scale)

  • Speech Generation — Convert text to speech with voice selection, multi-speaker support, and natural language style control using Gemini TTS

Related MCP server: Gemini Gen MCP

Installation

uvx media-mcp

Using pip

pip install media-mcp

Prerequisites

Configuration

Set your Gemini API key as an environment variable:

export GEMINI_API_KEY="your-gemini-api-key"

Environment Variables

Variable

Required

Description

GEMINI_API_KEY

Yes

Google Gemini API key for authentication

MEDIA_OUTPUT_DIR

No

Directory path for saving generated media files (see below)

Output behavior

When MEDIA_OUTPUT_DIR is set, every generated file is saved to that directory and the tool returns only the file path — no binary data is included in the response. This is the recommended setup because MCP messages are stored in the conversation history, and large base64 payloads pollute context and waste tokens.

When MEDIA_OUTPUT_DIR is not set, the server has no filesystem target, so it returns the raw base64-encoded data directly in the response. This works for quick experiments but is not recommended for production use.

MCP Client Setup

Claude Desktop

Add to your claude_desktop_config.json:

macOS: ~/Library/Application Support/Claude/claude_desktop_config.json Windows: %APPDATA%\Claude\claude_desktop_config.json

{
  "mcpServers": {
    "media-mcp": {
      "command": "uvx",
      "args": ["media-mcp"],
      "env": {
        "GEMINI_API_KEY": "your-gemini-api-key",
        "MEDIA_OUTPUT_DIR": "/path/to/media/output"
      }
    }
  }
}

Claude Code

claude mcp add media-mcp --transport stdio -- uvx media-mcp

Or add manually to .mcp.json:

{
  "mcpServers": {
    "media-mcp": {
      "type": "stdio",
      "command": "uvx",
      "args": ["media-mcp"],
      "env": {
        "GEMINI_API_KEY": "${GEMINI_API_KEY}",
        "MEDIA_OUTPUT_DIR": "/path/to/media/output"
      }
    }
  }
}

Tools

generate_image

Generate or edit images using Gemini's Nano Banana models.

Parameter

Type

Required

Default

Description

prompt

string

Yes

Text description of the image to generate

model

enum

No

nano-banana-2

nano-banana-2, nano-banana-pro, nano-banana

aspect_ratio

enum

No

1:1

1:1, 9:16, 16:9, 3:2, 4:3, and more

image_size

enum

No

1K

512px, 1K, 2K, 4K

reference_images

list[str]

No

Base64-encoded reference images

thinking_level

enum

No

minimal

minimal, high

use_google_search

bool

No

false

Enable Google Search grounding

Example prompt: "A watercolor painting of a cozy cabin in the mountains during autumn"

generate_video

Generate videos with native audio using Veo models.

Parameter

Type

Required

Default

Description

prompt

string

Yes

Text description including dialogue, sound effects, camera directions

model

enum

No

veo-3.1

veo-3.1, veo-3

aspect_ratio

enum

No

16:9

16:9 (landscape), 9:16 (portrait)

resolution

enum

No

720p, 1080p, 4K

first_frame_image

str

No

Base64-encoded image for first frame

last_frame_image

str

No

Base64-encoded image for last frame

reference_images

list[str]

No

Up to 3 base64-encoded reference images

Example prompt: "A slow dolly shot through a neon-lit alley at night, rain falling, 'Where are you going?' whispered softly, footsteps echoing"

generate_music

Generate instrumental music using Lyria RealTime with weighted prompts.

Parameter

Type

Required

Default

Description

prompts

list[dict]

Yes

Weighted prompts, e.g. [{"text": "minimal techno", "weight": 1.0}]

bpm

int

No

Tempo in beats per minute

temperature

float

No

1.0

Randomness/creativity control

scale

str

No

Musical scale constraint (e.g. C_MAJOR_A_MINOR)

duration_seconds

int

No

30

Duration of the output clip

Example prompts: [{"text": "Piano", "weight": 2.0}, {"text": "Meditation", "weight": 0.5}]

generate_speech

Convert text to speech with voice and style control.

Parameter

Type

Required

Default

Description

text

string

Yes

Text to speak. For multi-speaker, format as dialogue with speaker names.

model

enum

No

flash-tts

flash-tts, pro-tts

voice_name

str

No

Voice name: Kore, Puck, Charon, Fenrir, Aoede, Leda, Orus, Zephyr

multi_speaker

bool

No

false

Enable multi-speaker mode

speakers

list[dict]

No

Speaker-to-voice mapping, e.g. [{"name": "Alice", "voice_name": "Kore"}]

style_instructions

str

No

Style guidance, e.g. "Read in a calm, slow pace"

Example: Text: "Welcome to the show!" with voice_name: "Kore" and style_instructions: "Say cheerfully"

Troubleshooting

"GEMINI_API_KEY environment variable is not set"

Set the environment variable before starting the server:

export GEMINI_API_KEY="your-key-here"

When using Claude Desktop or Claude Code, pass the key via the env block in your MCP configuration (see MCP Client Setup).

"Authentication failed" or 401 errors

Your API key may be invalid or expired. Verify it at Google AI Studio.

"Rate limit or quota exceeded" or 429 errors

Wait a moment and retry. Check your API quota at Google AI Studio.

"Content blocked by safety filter"

Modify your prompt to avoid restricted content. The Gemini API applies safety filters to all generated media.

Python version errors

media-mcp requires Python 3.10 or later. Check your version:

python --version

License

MIT

Available Tools

4 tools
generate_imageB

Generate or edit images using Google's Gemini image generation models.

Supports conversational image creation/editing, multi-turn workflows, images with embedded text, infographics, and interleaved text+image output.

ParametersJSON Schema
NameRequiredDescriptionDefault
modelNonano-banana-2
promptYes
image_sizeNo1K
aspect_ratioNo1:1
thinking_levelNominimal
reference_imagesNo
use_image_searchNo
use_google_searchNo
response_modalitiesNo

TDQS

B3.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Without annotations, the description adds behavioral context by noting support for multi-turn and interleaved output, but lacks detail on safety, permissions, or side effects. The description is adequate but not comprehensive.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is brief (two sentences) and front-loaded with the core purpose, followed by supporting capabilities. Every sentence earns its place with no redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity (9 parameters, no output schema, no annotations), the description is insufficient. It does not explain return values or how parameters like model or reference_images affect behavior, leaving significant gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, meaning no parameter descriptions exist in the schema. The description does not explain any of the 9 parameters, leaving the agent to infer from names and enums alone.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool generates or edits images using Gemini models and lists specific capabilities like conversational workflows, embedded text, and infographics. It distinguishes from sibling tools (speech, video, music) by focusing on images.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description describes what the tool does but provides no explicit guidance on when to use it versus the sibling tools (generate_speech, generate_video, generate_music). No when-not or alternative scenarios are mentioned.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

generate_musicB

Generate instrumental music from weighted text prompts using Google's Lyria model.

Each prompt has a 'text' describing genre/instrument/mood and a 'weight' for emphasis. The server manages the streaming session internally.

ParametersJSON Schema
NameRequiredDescriptionDefault
bpmNo
scaleNo
promptsYes
temperatureNo
duration_secondsNo

TDQS

B3.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

There are no annotations provided, so the description must fully disclose behavioral traits. It mentions the tool generates music and internally manages streaming, but it does not disclose whether the operation is destructive, requires authentication, has rate limits, or what the output format (e.g., audio file URL) is. This is insufficient for an AI agent to understand side effects.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is only two sentences, making it very concise. It front-loads the primary purpose and then adds a clarifying technical detail. However, it could be slightly more structured by separating parameter explanation from behavioral notes.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity of five parameters and no output schema, the description is incomplete. It fails to mention what the tool returns (e.g., an audio file URL or stream), and it does not cover the optional parameters. This leaves significant gaps for an AI agent to correctly invoke the tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must add meaning to parameters. It clarifies that the 'prompts' array contains objects with 'text' and 'weight' fields, which aids the agent. However, it does not explain other parameters like bpm, scale, temperature, or duration_seconds, leaving them underdocumented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states it generates instrumental music from weighted text prompts using Google's Lyria model, specifying the verb (generate), resource (music), and technology, distinguishing it from sibling tools like generate_image or generate_speech.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explains the prompt structure and mentions server streaming, but it does not provide explicit guidance on when to use this tool versus alternatives, nor does it include prerequisites or exclusions. The context signals indicate sibling tools are different modalities, so the usage is partially implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

generate_speechA

Generate speech audio from text using Gemini TTS models.

Supports single-speaker and multi-speaker modes with voice selection and natural language style control.

ParametersJSON Schema
NameRequiredDescriptionDefault
textYes
modelNoflash-tts
speakersNo
voice_nameNo
multi_speakerNo
style_instructionsNo

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It discloses modes, voice selection, and style control, which are key behavioral traits. However, it omits details like supported audio formats, latency, or that it uses specific models (though model enum is in schema). Missing some transparency.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two concise sentences, front-loaded with purpose, no unnecessary words. Efficiently conveys key capabilities.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given 6 parameters, no output schema, and no annotations, the description is incomplete. It lacks details on the 'speakers' parameter structure, relationship between voice_name and speakers, and return format. Agent would need to infer or guess for effective use.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so description must add meaning. It hints at model, multi-speaker mode, voice selection, and style instructions, but does not explicitly map to parameters or explain complex ones like 'speakers' array. Adds some value but insufficient for full understanding.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states it generates speech audio from text using Gemini TTS models, and mentions modes and voice selection. It distinguishes from sibling tools (image, video, music) implicitly by modality, but does not explicitly differentiate.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear context for when to use the tool: for text-to-speech with optional multi-speaker mode, voice selection, and style control. It does not explicitly state when not to use or list alternatives, which is acceptable given sibling tools cover different modalities.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

generate_videoB

Generate videos from text prompts or reference images using Google's Veo models.

Supports text-to-video, image-to-video, video extension, and frame-specified generation. Generation is asynchronous.

ParametersJSON Schema
NameRequiredDescriptionDefault
modelNoveo-3.1
promptYes
resolutionNo
aspect_ratioNo16:9
extend_video_idNo
last_frame_imageNo
reference_imagesNo
first_frame_imageNo

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Discloses that generation is asynchronous, which is a key behavioral trait. However, with no annotations, the description does not cover other aspects like auth needs, rate limits, or return format, leaving gaps in transparency.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three concise sentences front-loading key information about purpose, capabilities, and async nature. No wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Missing crucial details for an async tool, such as how to retrieve generated videos, polling mechanism, or expected output format. With no output schema and 8 parameters, the description is insufficient for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, and the description does not explain the purpose of parameters like extend_video_id, reference_images, etc. It only lists modes without mapping them to specific parameters, providing minimal additive value.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Clearly states that the tool generates videos from text prompts or reference images using Google's Veo models, and lists specific modes (text-to-video, image-to-video, etc.), distinguishing it from sibling tools like generate_image.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Does not provide explicit guidance on when to use this tool vs. alternatives (generate_image, generate_speech, generate_music). The description implies usage for video generation but lacks when-not-to-use or prerequisites.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 4 tool updatesv0.2.7
    • First observedgenerate_image
    • First observedgenerate_music
    • First observedgenerate_speech
    • First observedgenerate_video

TDQS

A3.7/5.0

Scored across 4 tools

Disambiguation5/5

Each tool targets a distinct medium (image, speech, video, music) with no overlap. Descriptions clearly differentiate their purposes and capabilities.

Naming Consistency5/5

All tools follow a consistent 'generate_<medium>' pattern (e.g., generate_image, generate_speech), making naming predictable and intuitive.

Tool Count5/5

Four tools cover the core media generation types without unnecessary duplication. The count is appropriate for the server's focused purpose.

Completeness5/5

The tool surface covers generation for all major media types (image, speech, video, music) with support for editing and customization, leaving no obvious gaps.

Maintenance

ActivitySlowing
ResponsivenessNo issues

Resources

Unclaimed servers have limited discoverability.

Looking for Admin?

If you are the server author, to access and configure the admin panel.

Related MCP Connectors

Related MCP Servers