Music Media MCP Server
Click on "Install Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Music Media MCP Servercreate a music video from https://example.com/sunset.jpg with chill lo-fi beats"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
🎵 Music Media MCP Server
An MCP (Model Context Protocol) server that generates AI-powered music videos. Give it an image or video and it will analyze the visual content, compose a matching soundtrack using Google's Lyria 3 model, merge everything with FFmpeg, and return a playable video artifact.
Pipeline
Source Media (image/video URL)
→ Gemini Vision analyzes the visual content (if no prompt given)
→ Lyria 3 generates a 30-second AI music track
→ FFmpeg merges audio + media into a single .mp4
→ Uploads to Google Cloud Storage
→ Returns an HTML artifact with an inline video playerRelated MCP server: media-mcp
Features
Auto music prompting — If no music description is provided, Gemini Vision analyzes the image/video and generates a fitting music prompt automatically
Multiple media types — Supports images (.jpg, .png, .webp) and videos (.mp4, .mov)
Smart video handling — Images loop for 30s, short videos loop to fill, long videos trim to 30s
HTML artifact output — Returns a styled video player that MCP-compatible chatbots render inline
Cloud Run ready — Deploys to Google Cloud Run with a single command
Prerequisites
Python 3.10+
FFmpeg installed and on
PATH# macOS brew install ffmpeg # Ubuntu/Debian sudo apt install ffmpegGoogle Cloud project with:
Vertex AI API enabled (Lyria
lyria-002+ Geminigemini-2.0-flash-001)A GCS bucket for output storage (with public read access or signed URLs)
Application Default Credentials:
gcloud auth application-default login
Setup
Clone and install:
git clone https://github.com/joshndala/music-media-mcp.git cd music-media-mcp python -m venv .venv source .venv/bin/activate pip install -e .Configure environment:
cp .env.example .env # Edit .env with your GCP project ID and GCS bucket nameSet up GCS CORS (required for video playback in chatbot artifacts):
# Create cors.json echo '[{"origin":["*"],"method":["GET"],"responseHeader":["Content-Type","Content-Length","Range"],"maxAgeSeconds":3600}]' > cors.json gsutil cors set cors.json gs://YOUR_BUCKET_NAME
Running Locally
# stdio transport (for Claude Desktop and other MCP desktop clients)
python server.py
# SSE transport (for web-based MCP clients)
python server.py --transport sse --port 8000
# Test with MCP Inspector
npx @modelcontextprotocol/inspector
# Then connect to http://localhost:8000/sseDeploying to Cloud Run
# Build the container
gcloud builds submit \
--tag us-central1-docker.pkg.dev/YOUR_PROJECT/YOUR_REPO/music-media-server \
--project YOUR_PROJECT
# Deploy
gcloud run deploy music-media-server \
--image us-central1-docker.pkg.dev/YOUR_PROJECT/YOUR_REPO/music-media-server \
--region us-central1 \
--platform managed \
--allow-unauthenticated \
--set-env-vars "GCP_PROJECT_ID=YOUR_PROJECT,GCS_BUCKET_NAME=YOUR_BUCKET,GCP_LOCATION=us-central1" \
--memory 2Gi \
--timeout 300 \
--project YOUR_PROJECTYour SSE endpoint will be at: https://YOUR_SERVICE_URL/sse
MCP Client Configuration
Claude Desktop (claude_desktop_config.json)
{
"mcpServers": {
"music-media": {
"command": "/path/to/.venv/bin/python",
"args": ["/path/to/server.py", "--transport", "stdio"],
"env": {
"GCP_PROJECT_ID": "your-project-id",
"GCS_BUCKET_NAME": "your-bucket-name",
"GCP_LOCATION": "us-central1"
}
}
}
}Web/Chatbot (SSE)
Point your MCP client to your deployed Cloud Run URL:
https://your-service-url.run.app/sseTool Reference
generate_and_merge_media
Parameter | Type | Required | Description |
|
| ✅ | Direct URL to a source image or video |
|
| ❌ | Music style description (auto-generated if omitted) |
Returns: A complete HTML document with an inline video player.
Example prompts:
"Upbeat electronic dance music with synth arpeggios""Calm ambient piano piece evoking a misty morning""Cinematic orchestral score with soaring strings"(omit for automatic AI analysis)
Environment Variables
Variable | Required | Default | Description |
| ✅ | — | Google Cloud project ID |
| ✅ | — | GCS bucket for video uploads |
| ❌ |
| Vertex AI region |
License
MIT
Available Tools
1 toolgenerate_and_merge_mediaA
Generates a music video by combining AI-generated audio with the source image or video.
Args: source_media_url: A publicly accessible URL to a source image or video. music_prompt: Optional music description. If omitted, Gemini generates one.
IMPORTANT: The output of this tool is valid HTML. You MUST display this output to the user by generating a generic HTML Artifact containing the exact code returned by the tool.
| Name | Required | Description | Default |
|---|---|---|---|
| source_media_url | Yes | ||
| music_prompt | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description must cover behavioral traits. It reveals that the output is valid HTML and instructs the agent to display it in an HTML Artifact. This is helpful but could be more explicit about processing time or resource usage. Still, it adequately informs the agent of the output format and expected action.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise and well-structured. It starts with a clear purpose statement, then breaks down parameters, and ends with a critical usage instruction. Every sentence provides valuable information with no redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's two parameters and the presence of an output schema, the description is complete. It covers input requirements, optional parameters, output format (HTML), and the required user-facing action. No gaps are evident.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has 0% description coverage, so the description must compensate. It explains 'source_media_url' as a publicly accessible URL to source media, and 'music_prompt' as an optional music description that defaults to generation by Gemini. This adds meaningful context beyond the schema's type-only definitions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: 'Generates a music video by combining AI-generated audio with the source image or video.' This is a specific verb+resource pair, and since there are no sibling tools, differentiation is not needed.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides guidance on the optional music_prompt parameter, stating it can be omitted for Gemini to generate one. It also includes a crucial usage instruction: the output must be displayed via an HTML Artifact. However, it does not discuss when to use this tool versus alternatives (none exist) or any prerequisites beyond public URL accessibility.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
TDQS
Only one tool exists, so there is no possibility of confusion with other tools.
With a single tool, naming consistency is trivially maintained. The name uses snake_case and is descriptive.
A single tool for a 'Music Media MCP Server' is too few; the server likely requires multiple tools for different media operations.
The tool surface is severely incomplete, lacking separate functionalities for audio generation, video processing, or result retrieval.
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Connectors
Generate AI images and videos from any compatible MCP client.
AI image, video, voice and music generation over MCP, routed to Veo 3.1, Seedance 2.0 and more.
Generate AI images, video, music, and sound effects, and upscale them, from any MCP client.
Generate Suno AI music (v5.5) from any MCP client. Async; billed only on success.
Related MCP Servers
- AlicenseBqualityDmaintenanceMCP server that exposes Google's Veo2 video generation capabilities, allowing clients to generate videos from text prompts or images.732MIT
- AlicenseAqualityCmaintenanceAn MCP server for AI-powered media generation using Google Gemini, enabling creation of images, videos, music, and speech directly from AI agents.4MIT
- AlicenseBqualityDmaintenanceMCP server for generating and editing images using OpenAI, and creating videos using OpenAI Sora and Google Veo. Enables fetching media from URLs or disk with smart output placement.14289MIT
- AlicenseNot gradedqualityDmaintenanceEnables generating AI-powered soundtracks for videos via Muzaic AI from any MCP client.MIT
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/joshndala/music-media-mcp'
If you have feedback or need assistance with the MCP directory API, please join our Discord server