Skip to main content
Glama

manim-mcp

MCP server that generates 3Blue1Brown-style math videos with Manim directly inside Claude Desktop. Videos render with AI-narrated voiceover via ElevenLabs and play inline in the chat.

https://github.com/user-attachments/assets/placeholder

Quick Start

1. Get an ElevenLabs API key

Sign up at elevenlabs.io and copy your API key from the API Keys page.

2. Add to Claude Desktop

Open Settings > Developer > Edit Config and add:

{
  "mcpServers": {
    "manim": {
      "command": "npx",
      "args": ["-y", "manim-mcp", "--stdio"],
      "env": {
        "ELEVEN_API_KEY": "sk_your_key_here"
      }
    }
  }
}

3. Restart Claude Desktop

That's it. Ask Claude to "create a video on the Pythagorean theorem" and watch it render.

Related MCP server: Video MCP

Requirements

  • Node.js 18+ — for running the MCP server

  • Python 3.9+ — for Manim rendering (auto-installed into ~/.manim-mcp/.venv on first run)

  • ffmpeg — for video concatenation (brew install ffmpeg on macOS)

  • ElevenLabs API key — for AI voiceover

Python dependencies (manim, manim-voiceover, elevenlabs) are automatically installed on first run. This takes ~20 seconds and only happens once.

How it Works

You: "create a video on eigenvalues"

Claude plans 2-4 scenes with narration scripts
  -> writes Manim Python code for each scene
  -> calls render_video with all scenes
  -> server auto-fixes common issues (wrong TTS service, LaTeX, colors)
  -> scenes render in parallel with ElevenLabs voiceover
  -> concatenated into one video
  -> inline video player appears in chat

The server provides comprehensive Manim reference and 3Blue1Brown style guidelines in its instructions, so Claude generates good animation code. A fixGeneratedCode step silently corrects common mistakes before rendering:

  • Wrong TTS service (GTTSService, etc.) -> ElevenLabsService with correct config

  • LaTeX classes (MathTex, Tex) -> Text() with Unicode

  • Invalid colors (CYAN) -> TEAL

  • Invented APIs (set_speech_synthesizer) -> correct set_speech_service

Development

git clone https://github.com/zcsabbagh/manim-mcp.git
cd manim-mcp
npm install
npm run build

# Point Claude Desktop at local build:
# "command": "node",
# "args": ["/path/to/manim-mcp/dist/index.js", "--stdio"]

License

MIT

Available Tools

2 tools
render_videoRender Manim VideoA

Render one or more Manim scenes in parallel, concatenate them, and return one combined video inline. Each scene has complete Python code with voiceover baked in via manim-voiceover. Scenes render concurrently — voice is generated and synced during rendering automatically. The server auto-fixes common issues: wrong TTS service → ElevenLabs, CYAN → TEAL, MathTex → Text.

ParametersJSON Schema
NameRequiredDescriptionDefault
scenesYesArray of scenes to render in parallel. Order matters for final video.
qualityNoVideo quality: l=480p (default), m=720p, h=1080pl

Output Schema

ParametersJSON Schema
NameRequiredDescription
scenesYes
overviewYes
videoUriYes
durationSecondsYes

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description takes full responsibility for disclosing behavior. It reveals that scenes render concurrently, voiceover is auto-synced, and the server auto-fixes common issues like wrong TTS, color, and MathTex usage. This goes beyond basic functionality and helps set expectations, though it doesn't cover all potential failure modes or side effects.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is four sentences, each serving a purpose: main action, scene requirements, concurrency behavior, and auto-fix details. There is no redundancy or fluff. It is front-loaded with the primary action and remains skimmable.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's moderate complexity (2 params, one enum, nested scene objects) and the presence of an output schema, the description covers all necessary behavioral context: parallel rendering, auto-fixing, and voiceover handling. It does not need to explain return values because an output schema exists, and the input schema already documents parameter constraints.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description adds meaning beyond the schema by explaining that each scene's code must have voiceover baked in via manim-voiceover, and that scenes render concurrently. This helps the agent understand the intent behind the code parameter and the parallel nature of the scenes array.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb+resource+outcome: 'Render one or more Manim scenes in parallel, concatenate them, and return one combined video inline.' This clearly distinguishes it from the sibling tool show_demo_video, which presumably displays a pre-made demo rather than rendering new content.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage: you use this tool to render and combine Manim scenes into a single video. It doesn't explicitly mention alternatives or exclusions, but the clear purpose and the existence of a sibling tool provide enough context for an agent to decide. Since it lacks explicit 'when not to use' guidance, it misses the top score.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

show_demo_videoShow Demo VideoA

Show a pre-rendered demo video inline to test the MCP video player. No generation needed.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
scenesYes
overviewYes
videoUriYes
durationSecondsYes

TDQS

A4.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are present, so the description carries the burden. It discloses that the video is pre-rendered and that no generation is required, implying a non-destructive display action. However, it does not explicitly state read-only behavior, side effects, or failure modes.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence with no filler. Every phrase contributes meaningful information about the tool's purpose and behavior.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a no-parameter demo tool with an output schema, the description provides enough context: what it shows, where it shows it, and that it does not generate. It could add an explicit side-effect statement, but overall this is adequate.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters and schema coverage is 100%, so there is no parameter documentation needed. The description adds no parameter-specific meaning, but none is required.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Show') and resource ('pre-rendered demo video inline'), and it clearly distinguishes the tool from its sibling by adding 'No generation needed.' This makes the tool's purpose immediately clear.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description states the tool is for testing the MCP video player and contrasts itself with generation via 'No generation needed.' It does not explicitly name the alternative tool (render_video) or provide exclusion criteria, but the usage context is clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 2 tool updatesv0.1.0
    • First observedrender_video
    • First observedshow_demo_video

TDQS

A4.2/5.0

Scored across 2 tools

Disambiguation5/5

render_video and show_demo_video are clearly distinct: one creates a video from user-provided scenes, the other displays a pre-rendered demo. There is no overlap in purpose or output.

Naming Consistency5/5

Both tool names follow a consistent verb_noun pattern (render_video, show_demo_video), making the API predictable and easy to navigate.

Tool Count3/5

With only two tools, the server feels minimal for a video rendering domain, but it targets a narrow use case (rendering with voiceover and demoing), so it is borderline acceptable.

Completeness3/5

The server lacks operations like listing available scenes, rendering individual scenes without concatenation, or retrieving prior renders, leaving notable gaps for agents needing more granular control over the rendering process.

Maintenance

ActivityInactive
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers