manim-mcp
Server Configuration
Describes the environment variables required to run the server.
| Name | Required | Description | Default |
|---|---|---|---|
| ELEVEN_API_KEY | Yes | Your ElevenLabs API key for AI voiceover. |
Instructions
Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.
This server publishes no instructions, or was last inspected before Glama recorded them.
Capabilities
Features and capabilities supported by this server
Protocol revision2025-11-25
| Capability | Details |
|---|---|
| tools | {
"listChanged": true
} |
| resources | {
"listChanged": true
} |
Tools
Functions exposed to the LLM to take actions
| Name | Description |
|---|---|
| render_videoA | Render one or more Manim scenes in parallel, concatenate them, and return one combined video inline. Each scene has complete Python code with voiceover baked in via manim-voiceover. Scenes render concurrently — voice is generated and synced during rendering automatically. The server auto-fixes common issues: wrong TTS service → ElevenLabs, CYAN → TEAL, MathTex → Text. |
| show_demo_videoA | Show a pre-rendered demo video inline to test the MCP video player. No generation needed. |
Prompts
Interactive templates invoked by user choice
| Name | Description |
|---|---|
No prompts | |
Resources
Contextual data attached and managed by the client
| Name | Description |
|---|---|
| Manim Video Player |
TDQS
Scored across 2 tools
render_video and show_demo_video are clearly distinct: one creates a video from user-provided scenes, the other displays a pre-rendered demo. There is no overlap in purpose or output.
Both tool names follow a consistent verb_noun pattern (render_video, show_demo_video), making the API predictable and easy to navigate.
With only two tools, the server feels minimal for a video rendering domain, but it targets a narrow use case (rendering with voiceover and demoing), so it is borderline acceptable.
The server lacks operations like listing available scenes, rendering individual scenes without concatenation, or retrieving prior renders, leaving notable gaps for agents needing more granular control over the rendering process.