manim-mcp
Provides AI voiceover for generated math videos using ElevenLabs text-to-speech service.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@manim-mcpcreate a video on the Pythagorean theorem"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
manim-mcp
MCP server that generates 3Blue1Brown-style math videos with Manim directly inside Claude Desktop. Videos render with AI-narrated voiceover via ElevenLabs and play inline in the chat.
https://github.com/user-attachments/assets/placeholder
Quick Start
1. Get an ElevenLabs API key
Sign up at elevenlabs.io and copy your API key from the API Keys page.
2. Add to Claude Desktop
Open Settings > Developer > Edit Config and add:
{
"mcpServers": {
"manim": {
"command": "npx",
"args": ["-y", "manim-mcp", "--stdio"],
"env": {
"ELEVEN_API_KEY": "sk_your_key_here"
}
}
}
}3. Restart Claude Desktop
That's it. Ask Claude to "create a video on the Pythagorean theorem" and watch it render.
Related MCP server: Video MCP
Requirements
Node.js 18+ — for running the MCP server
Python 3.9+ — for Manim rendering (auto-installed into
~/.manim-mcp/.venvon first run)ffmpeg — for video concatenation (
brew install ffmpegon macOS)ElevenLabs API key — for AI voiceover
Python dependencies (manim, manim-voiceover, elevenlabs) are automatically installed on first run. This takes ~20 seconds and only happens once.
How it Works
You: "create a video on eigenvalues"
Claude plans 2-4 scenes with narration scripts
-> writes Manim Python code for each scene
-> calls render_video with all scenes
-> server auto-fixes common issues (wrong TTS service, LaTeX, colors)
-> scenes render in parallel with ElevenLabs voiceover
-> concatenated into one video
-> inline video player appears in chatThe server provides comprehensive Manim reference and 3Blue1Brown style guidelines in its instructions, so Claude generates good animation code. A fixGeneratedCode step silently corrects common mistakes before rendering:
Wrong TTS service (GTTSService, etc.) -> ElevenLabsService with correct config
LaTeX classes (MathTex, Tex) -> Text() with Unicode
Invalid colors (CYAN) -> TEAL
Invented APIs (set_speech_synthesizer) -> correct set_speech_service
Development
git clone https://github.com/zcsabbagh/manim-mcp.git
cd manim-mcp
npm install
npm run build
# Point Claude Desktop at local build:
# "command": "node",
# "args": ["/path/to/manim-mcp/dist/index.js", "--stdio"]License
MIT
Available Tools
2 toolsrender_videoRender Manim VideoA
Render one or more Manim scenes in parallel, concatenate them, and return one combined video inline. Each scene has complete Python code with voiceover baked in via manim-voiceover. Scenes render concurrently — voice is generated and synced during rendering automatically. The server auto-fixes common issues: wrong TTS service → ElevenLabs, CYAN → TEAL, MathTex → Text.
| Name | Required | Description | Default |
|---|---|---|---|
| scenes | Yes | Array of scenes to render in parallel. Order matters for final video. | |
| quality | No | Video quality: l=480p (default), m=720p, h=1080p | l |
Output Schema
| Name | Required | Description |
|---|---|---|
| scenes | Yes | |
| overview | Yes | |
| videoUri | Yes | |
| durationSeconds | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description takes full responsibility for disclosing behavior. It reveals that scenes render concurrently, voiceover is auto-synced, and the server auto-fixes common issues like wrong TTS, color, and MathTex usage. This goes beyond basic functionality and helps set expectations, though it doesn't cover all potential failure modes or side effects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is four sentences, each serving a purpose: main action, scene requirements, concurrency behavior, and auto-fix details. There is no redundancy or fluff. It is front-loaded with the primary action and remains skimmable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's moderate complexity (2 params, one enum, nested scene objects) and the presence of an output schema, the description covers all necessary behavioral context: parallel rendering, auto-fixing, and voiceover handling. It does not need to explain return values because an output schema exists, and the input schema already documents parameter constraints.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3. The description adds meaning beyond the schema by explaining that each scene's code must have voiceover baked in via manim-voiceover, and that scenes render concurrently. This helps the agent understand the intent behind the code parameter and the parallel nature of the scenes array.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb+resource+outcome: 'Render one or more Manim scenes in parallel, concatenate them, and return one combined video inline.' This clearly distinguishes it from the sibling tool show_demo_video, which presumably displays a pre-made demo rather than rendering new content.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage: you use this tool to render and combine Manim scenes into a single video. It doesn't explicitly mention alternatives or exclusions, but the clear purpose and the existence of a sibling tool provide enough context for an agent to decide. Since it lacks explicit 'when not to use' guidance, it misses the top score.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
show_demo_videoShow Demo VideoA
Show a pre-rendered demo video inline to test the MCP video player. No generation needed.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| scenes | Yes | |
| overview | Yes | |
| videoUri | Yes | |
| durationSeconds | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are present, so the description carries the burden. It discloses that the video is pre-rendered and that no generation is required, implying a non-destructive display action. However, it does not explicitly state read-only behavior, side effects, or failure modes.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence with no filler. Every phrase contributes meaningful information about the tool's purpose and behavior.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a no-parameter demo tool with an output schema, the description provides enough context: what it shows, where it shows it, and that it does not generate. It could add an explicit side-effect statement, but overall this is adequate.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters and schema coverage is 100%, so there is no parameter documentation needed. The description adds no parameter-specific meaning, but none is required.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Show') and resource ('pre-rendered demo video inline'), and it clearly distinguishes the tool from its sibling by adding 'No generation needed.' This makes the tool's purpose immediately clear.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description states the tool is for testing the MCP video player and contrasts itself with generation via 'No generation needed.' It does not explicitly name the alternative tool (render_video) or provide exclusion criteria, but the usage context is clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
2 tool updates
v0.1.0- First observed
render_video - First observed
show_demo_video
TDQS
Scored across 2 tools
render_video and show_demo_video are clearly distinct: one creates a video from user-provided scenes, the other displays a pre-rendered demo. There is no overlap in purpose or output.
Both tool names follow a consistent verb_noun pattern (render_video, show_demo_video), making the API predictable and easy to navigate.
With only two tools, the server feels minimal for a video rendering domain, but it targets a narrow use case (rendering with voiceover and demoing), so it is borderline acceptable.
The server lacks operations like listing available scenes, rendering individual scenes without concatenation, or retrieving prior renders, leaving notable gaps for agents needing more granular control over the rendering process.
Maintenance
Related MCP Connectors
MCP server for OpenAI Sora AI video generation
Hosted MCP server connecting claude.ai, ChatGPT and other AI apps to your own computer
MCP server for Grok Imagine AI video generation
Educational MCP server with 17 math/stats tools, visualizations, and persistent workspace
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceAn MCP server that executes Manim Python scripts to generate and return rendered animations. It enables users to create mathematical and programmatic videos dynamically through natural language interfaces.3MIT
- FlicenseNot gradedqualityDmaintenanceA local MCP server that gives Claude Desktop full video editing capabilities via FFmpeg, Whisper, and yt-dlp.-
- AlicenseNot gradedqualityDmaintenanceAn MCP server that uses Google's Gemini API to analyze videos and convert them to text descriptions that Claude Code can understand and act upon.33 npmMIT
- FlicenseNot gradedqualityBmaintenanceMCP server that lets Claude Code create explainer videos from plain English requests by writing and rendering manim animations locally, then publishing the result as a shareable artifact.-