grok-video-mcp
grok-media-mcp
MCP server that gives AI agents xAI Grok media generation — images and video. Submit a prompt, get a real file back. Works with OpenCode, Claude Desktop, Cursor, VS Code, and any MCP client.
Companion to vision-mcp: one agent can generate a clip or image and verify it — a full media loop.
Why
Text-based agents can't generate media. grok-media-mcp exposes xAI's
grok-imagine-video and grok-imagine-image models as plain MCP tools so any
agent can produce real images and video clips from a prompt — no shell
scripts, no manual API calls, no hand-rolled polling loops.
Tools
Tool | What it does |
| Submit a generation → returns |
| Poll: |
| Download the finished clip to disk |
| Submit + poll + download in one call (agent-friendly) |
| Generate an image — synchronous, returns the saved file path (~10-30s, ~$0.06) |
Requirements
Node.js ≥ 18
An xAI API key — https://console.x.ai (or docs.x.ai)
Install
npx from GitHub (recommended)
npx -y github:pongsakornp/grok-media-mcpnpx clones the repo, installs deps, auto-builds via the
preparescript, and runs the server over stdio.
From source
git clone https://github.com/pongsakornp/grok-media-mcp.git
cd grok-media-mcp
npm install
npm run buildUsage
OpenCode (opencode.jsonc)
{
"mcp": {
"grok-media-mcp": {
"type": "local",
"command": ["npx", "-y", "github:pongsakornp/grok-media-mcp"],
"environment": {
"XAI_API_KEY": "xai-..."
},
"enabled": true
}
}
}Claude Desktop (claude_desktop_config.json)
{
"mcpServers": {
"grok-media-mcp": {
"command": "npx",
"args": ["-y", "github:pongsakornp/grok-media-mcp"],
"env": {
"XAI_API_KEY": "xai-..."
}
}
}
}VS Code / Cursor (.vscode/mcp.json)
{
"servers": {
"grok-media-mcp": {
"type": "stdio",
"command": "npx",
"args": ["-y", "github:pongsakornp/grok-media-mcp"],
"environment": {
"XAI_API_KEY": "xai-..."
}
}
}
}Keys live in the MCP config — no shell profile edits needed.
Configuration
Env var | Default | Description |
| — | required — xAI API key |
|
| Model ( |
|
| Image model ( |
|
| Where media is saved |
|
| Max wait for a generation |
|
| Initial poll interval |
|
| Max poll interval (×1.5 backoff) |
How it works
Video — xAI's video API is async:
POST /v1/videos/generations → { request_id }
GET /v1/videos/{request_id} → poll until status: "done"
GET video.url → mp4 bytesgenerate_and_wait encapsulates submit → poll (5s→30s backoff, progress
reported) → download → save, returning the file path.
Images — synchronous, one call:
POST /v1/images/generations → { data: [{ url, mime_type }], usage }
GET image.url → jpeg/png bytesgenerate_image encapsulates generate → download → save in one call.
(Verified live: 1248×832 output, ~30s, ~$0.06/image.)
Pricing: video roughly $0.005 per second ($0.04 for an 8s clip),
images **$0.06 each** — an order of magnitude cheaper than Google Veo Lite
($0.05–0.08/s).
Development
npm run build # TypeScript → dist/
npm test # 28 tests (vitest)
npm run typecheck # tsc --noEmitTest coverage: config parsing, video + image request/response mapping (mocked fetch), polling backoff/timeout/FAILED handling, and the full MCP stdio protocol.
License
MIT — see LICENSE.