Google Veo 3.1 MCP Server
Server Configuration
Describes the environment variables required to run the server.
| Name | Required | Description | Default |
|---|---|---|---|
| DEBUG | No | Enable debug logging (default: false) | false |
| OUTPUT_DIR | No | Default output directory (default: ./output) | ./output |
| GOOGLE_API_KEY | Yes | Google API key (required) | |
| VIDEO_POLL_INTERVAL | No | Polling interval in ms (default: 15000) | 15000 |
| VIDEO_MAX_POLL_ATTEMPTS | No | Max polling attempts (default: 120) | 120 |
Instructions
Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.
This server publishes no instructions, or was last inspected before Glama recorded them.
Capabilities
Features and capabilities supported by this server
Protocol revision2025-11-25
| Capability | Details |
|---|---|
| tools | {} |
Tools
Functions exposed to the LLM to take actions
| Name | Description |
|---|---|
| generate_videoA | Generate a video using Google Veo 3.1 API. Supports:
Models:
Pricing (per second, audio always included):
|
| extend_videoA | Extend an existing video by 7 seconds using Veo 3.1 API. The extension continues from the last second of the input video. Requirements:
Estimated cost: ~$2.80 per extension (7 seconds x $0.40/sec) |
| interpolate_framesA | Generate a video that smoothly transitions between two keyframes using Veo 3.1 API. Creates a video that starts at the first frame and ends at the last frame with AI-generated motion in between. Pricing: Same as generate_video based on duration (audio always included) |
| get_video_statusA | Check the status of a video generation operation. Returns the current status (done/pending) and video URL if completed. Can also download the completed video (use after generate_video with wait: false). |
Prompts
Interactive templates invoked by user choice
| Name | Description |
|---|---|
No prompts | |
Resources
Contextual data attached and managed by the client
| Name | Description |
|---|---|
No resources | |
TDQS
Scored across 4 tools
Each tool targets a distinct video operation: generating from text/image, extending an existing video, interpolating between keyframes, and checking status. The descriptions clearly differentiate inputs and purposes, leaving no ambiguity for an agent.
All tool names follow a consistent verb_noun pattern: generate_video, extend_video, interpolate_frames, get_video_status. The pattern is predictable and mixes well with the domain.
The server has exactly four tools, which is well-scoped for a video generation service. It covers the core generation capabilities plus status checking without unnecessary clutter.
The tool surface covers generation, extension, interpolation, and status retrieval, which are the primary workflows. A minor gap is the lack of a cancel or list operations, but agents can work around this with the existing status tool.