Google AI Edge Gallery Video MCP Server
Enables on-device video generation and management on Android, including running natively via Termux, bridging from a host PC via ADB, monitoring battery/temperature/thermal throttling, and exporting videos to Android's DCIM/Gallery through MediaStore.
Integrates with the Google AI Edge Gallery ecosystem to generate on-device videos using supported models such as Gemma 2, Gemma 3n, MobileDiffusion, and MediaPipe tasks.
Automatically synchronizes and indexes generated videos into Google Photos via Android MediaStore scan intents, making them appear in the user's photo gallery.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Google AI Edge Gallery Video MCP ServerGenerate a 15s storyboard video of a robot dancing in neon city and save to Google Photos"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Google AI Edge Gallery Video MCP Server ๐ฌ๐ฑ
A high-performance Model Context Protocol (MCP) server enabling AI agents (Claude Desktop, Cursor, Antigravity, LLM frontends) to direct, generate, and compile on-device videos using the Google AI Edge Gallery ecosystem (google-ai-edge/gallery).
Built from the ground up to run natively on Android devices (via Termux), through an ADB Bridge (USB or Wi-Fi), or as a Standalone Desktop Engine, with automatic synchronization to Android's DCIM / Google Photos Gallery and hardware-aware battery & thermal throttling protection.
๐ Key Highlights
On-Device Edge Video Generation: Turn raw prompts and narrative ideas into complete, multi-scene video stories with cinematic keyframes, camera motions, color grades, and audio score.
Native Android & Termux Support: Zero native C++ compilation hassles. One-line installer on Termux with automatic storage permissions and Android MediaStore scanning.
ADB Host-to-Device Bridge: Run the MCP server on your PC/Mac while seamlessly streaming and indexing generated videos directly onto a connected Android smartphone.
Hardware-Aware Telemetry: Monitors Android battery charge, device temperature, and thermal throttling states to automatically adapt video resolution (720p vs 1080p) and frame rates.
Google AI Edge Gallery Integration: Built to interface with models supported by Google AI Edge Gallery (Gemma 2, Gemma 3n, MobileDiffusion, MediaPipe tasks).
Instant MediaStore Registration: Broadcasts
ACTION_MEDIA_SCANNER_SCAN_FILEintents so exported videos immediately appear in the phone's native Gallery and Google Photos.
Related MCP server: autoglm-mcp-server
๐๏ธ Architecture Overview
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โ MCP Client โ
โ (Claude Desktop / Cursor / Antigravity) โ
โโโโโโโโโโโโโโโโโโโโโโโโฌโโโโโโโโโโโโโโโโโโโโโโโโ
โ JSON-RPC (stdio / SSE)
โผ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โ Google AI Edge Gallery Video MCP Server โ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฌโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโค
โ Core Engine โ Android Subsystem โ
โ โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ โ โโโโโโโโโโโโโโโโโโโโ โโโโโโโโโโโโโโโโโโโโโโโโโ โ
โ โ Storyboard Director โ โ โ Native Termux โ โ ADB Bridge โ โ
โ โ (Gemma Scene Planning & Timing) โ โ โ (On-Device Host) โ โ (Remote Device Push) โ โ
โ โโโโโโโโโโโโโโโโโโฌโโโโโโโโโโโโโโโโโโ โ โโโโโโโโโโฌโโโโโโโโโโ โโโโโโโโโโโโโฌโโโโโโโโโโโโ โ
โ โผ โ โ โ โ
โ โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ โ โผ โผ โ
โ โ Keyframe Generator & Styler โ โ โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ โ
โ โ (Diffusion / Edge SVG Engine) โ โ โ Unified Device Manager โ โ
โ โโโโโโโโโโโโโโโโโโฌโโโโโโโโโโโโโโโโโโ โ โ - Battery Level & Temp Monitoring โ โ
โ โผ โ โ - Thermal Throttling Mitigation โ โ
โ โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ โ โโโโโโโโโโโโโโโโโโโโโโโโฌโโโโโโโโโโโโโโโโโโโโโโโ โ
โ โ MediaPipe FX & Filter Pipeline โ โ โ โ
โ โ (Ken Burns Pan/Zoom, Grade LUTs) โ โ โผ โ
โ โโโโโโโโโโโโโโโโโโฌโโโโโโโโโโโโโโโโโโ โ โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ โ
โ โผ โ โ Gallery Sync Manager โ โ
โ โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ โ โ - Destination: /sdcard/DCIM/GoogleEdgeAI โ โ
โ โ Hardware-Aware Video Compiler โโโผโโบ โ - MediaStore Intent Broadcast โ โ
โ โ (FFmpeg H.264 + Ambient Audio) โ โ โ - Google Photos / Gallery Auto-Indexing โ โ
โ โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ โ โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ โ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโดโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ๐ Quick Start Guide
Option 1: Native on Android (Termux)
Run this one-line command inside Termux:
pkg install -y curl && curl -sSL https://raw.githubusercontent.com/hayhihey/google-ai-edge-gallery-video-mcp/main/scripts/termux-install.sh | bashOnce installed, simply start the server:
edge-video-mcpAll created videos will automatically appear in your phone's Gallery under GoogleEdgeAI (/sdcard/DCIM/GoogleEdgeAI)!
Option 2: Host PC with Android Phone Connected (ADB Bridge)
Enable Developer Options and USB Debugging (or Wireless Debugging) on your Android phone.
Connect your phone to your computer via USB or Wi-Fi (
adb connect <phone_ip>:5555).Clone and build the project:
git clone https://github.com/google-ai-edge/google-ai-edge-gallery-video-mcp.git cd google-ai-edge-gallery-video-mcp npm install npm run buildVerify Android device connectivity:
bash scripts/adb-setup.sh
Option 3: Desktop Standalone (No phone required)
The server automatically detects when no Android device is attached and runs the edge simulation pipeline locally on Windows, macOS, or Linux.
Prerequisite: Ensure ffmpeg is installed:
macOS:
brew install ffmpegUbuntu/Debian:
sudo apt install -y ffmpegWindows:
winget install Gyan.FFmpegorchoco install ffmpeg
git clone https://github.com/google-ai-edge/google-ai-edge-gallery-video-mcp.git
cd google-ai-edge-gallery-video-mcp
npm install
npm run build
npm start๐ Connecting to MCP Clients
Claude Desktop Configuration
Add the following to your claude_desktop_config.json:
{
"mcpServers": {
"google-ai-edge-video": {
"command": "node",
"args": [
"/path/to/google-ai-edge-gallery-video-mcp/dist/index.js"
],
"env": {
"EDGE_MCP_OUTPUT_DIR": "/path/to/custom/output"
}
}
}
}On Windows, use escaped backslashes: C:\\Users\\<Username>\\...\\dist\\index.js
๐ ๏ธ MCP Tools Reference
Tool Name | Parameters | Description |
|
| Full end-to-end video synthesis pipeline. Directs scenes, synthesizes keyframes, applies camera motions & color grading, encodes MP4, and indexes in Android Gallery. |
|
| Directs multi-shot scene breakdowns, camera motions ( |
|
| Generates high-fidelity keyframe image assets for each scene. |
|
| Generates hardware-optimized filtergraphs for Ken Burns motions, vignette, and cinematic color palettes. |
|
| Low-level assembler for stitching shot clips, transitions, and audio beds into an H.264 MP4. |
| (None) | Inspects real-time battery charge, temperature, thermal throttling state, and acceleration delegates. |
|
| Moves any video into |
|
| Catalogs models compatible with Google AI Edge Gallery (Gemma 2, Gemma 3n, MobileDiffusion, MediaPipe). |
๐ฆ MCP Resources & Prompts
Resources
edge://device/telemetry: Real-time hardware telemetry (battery, temperature, storage, encoder support).edge://gallery/videos: Catalog of generated video files with sizes and timestamps.edge://models/inventory: List of edge models available on device.
Prompts
ai_video_director: Creative director persona to brainstorm screenplays tailored to edge constraints.social_shorts_creator: High-engagement 9:16 vertical short template designed for TikTok, Reels, and Shorts.
๐งช Testing & Verification
Run the comprehensive unit test suite:
npm testTest coverage includes:
Storyboard director planning across aspect ratios (9:16, 16:9, 1:1).
Mobile-constrained thermal throttling logic.
Procedural SVG keyframe rendering and XML escaping.
Ken Burns motion expressions and color grading filtergraphs.
Unified device detection (Termux / ADB / Local).
๐ณ Docker Deployment
A lightweight multi-stage Docker image with built-in FFmpeg and Android tools:
# Build and run with docker compose
docker compose up -d
# Or run directly with docker
docker build -t edge-video-mcp .
docker run --rm -v $(pwd)/output:/app/output edge-video-mcp๐ข Deploying to GitHub
To publish this repository to your GitHub account:
# 1. Initialize git repository
git init -b main
# 2. Add files and make initial commit
git add .
git commit -m "feat: initial release of Google AI Edge Gallery Video MCP Server"
# 3. Create a new repository on GitHub (e.g. google-ai-edge-gallery-video-mcp)
# 4. Link remote and push:
git remote add origin https://github.com/<your-username>/google-ai-edge-gallery-video-mcp.git
git push -u origin mainThe pre-configured GitHub Actions CI/CD workflows (.github/workflows/ci.yml and release.yml) will automatically:
Run automated tests across Ubuntu and Windows on Node.js 18, 20, and 22.
Package releases upon pushing tags (e.g.
git tag v1.0.0 && git push origin v1.0.0).
๐ License
Licensed under the Apache License, Version 2.0.
Available Tools
8 toolsandroid_device_statusA
Checks Android device battery percentage, thermal throttling status, available storage, and acceleration mode (Termux, ADB bridge, or Host).
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full disclosure burden. 'Checks' implies a non-destructive read on a zero-parameter tool, which is inherently low risk, and it enumerates the checked properties (including the three acceleration modes). However, it says nothing about permissions, cost, or whether repeated polling is safe.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single dense sentence that front-loads the verb and resource, then lists the returned signals in a compact parallel structure. No filler, no repetition of the tool name's meaning.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and no parameters, the description compensates by enumerating the four values the caller receives, which is the key information needed to use the result. It stops short of explaining formats or units, a minor gap for such a simple probe.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters and the schema is empty, so there are no parameter semantics to clarify. Baseline 4 applies for a parameterless tool.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb ('Checks') and a concrete resource ('Android device') and enumerates the exact signals returned: battery percentage, thermal throttling, storage, and acceleration mode. That makes it clearly distinct from the video/gallery siblings, though it never references an alternative tool by name.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to call this tool versus alternatives, no preconditions, and no exclusions. An agent must infer that this is a diagnostic probe for device health, but the description never states that intent.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
android_gallery_exportA
Exports any video file to Android DCIM folder and triggers Android MediaStore scanner so it instantly shows up in Google Photos / Gallery.
| Name | Required | Description | Default |
|---|---|---|---|
| customTitle | No | Display title | |
| videoFilePath | Yes | Path to video file |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does disclose a non-obvious side effect: it triggers the MediaStore scanner so the file appears in Google Photos/Gallery. It still omits permission/scoped-storage requirements, whether the source is copied or moved, and overwrite behavior on name collisions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence that states the action, the destination, and the observable effect with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple 2-param, 1-required tool with no output schema, the description covers purpose, destination, and post-condition adequately. Only operational details (permissions, copy vs move, failure handling) are missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% (both params documented), so the schema already carries parameter meaning. The description adds nothing about videoFilePath format/constraints or how customTitle affects the gallery entry, matching the baseline 3.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb (exports) plus resource (video file) plus destination (Android DCIM folder), and it distinguishes itself from the edge_* video-creation siblings by being the export/publish step. An agent can identify it without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The destination ('Android DCIM') implies the use case, but the description never states when to call this versus the sibling tools or what precondition must hold (e.g. a completed render). Usage is inferable rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
edge_apply_video_fxC
Generates hardware-optimized filtergraphs for Ken Burns camera motion, color balance, and cinematic vignette.
| Name | Required | Description | Default |
|---|---|---|---|
| fps | No | ||
| motion | No | zoom_in | |
| colorGrade | No | cinematic_teal_orange | |
| aspectRatio | No | 9:16 | |
| durationSeconds | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It clarifies that the tool generates filtergraphs rather than applying them, which is useful, but it omits whether it returns text or a file, any side effects, authentication needs, or how the output should be consumed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler. Every phrase contributes to identifying the tool's output domain.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 5-parameter tool with no annotations, no output schema, and no parameter descriptions, the definition is materially incomplete. It does not explain the generated filtergraph format, how to use the result, or what each input controls.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, but it only loosely maps to motion and colorGrade via 'Ken Burns camera motion' and 'color balance'. It says nothing about fps, aspectRatio, durationSeconds, or the enum values, and it references a cinematic vignette that is not an input parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Generates) and resource (filtergraphs) and names the effect areas it covers, so the agent can tell it apart from create/compile siblings. However, it does not explicitly differentiate itself from other video-effect or keyframe siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance, no preconditions, and no alternatives are mentioned. The description only says what it produces, leaving the agent to infer when this tool should be chosen over edge_compile_video or edge_generate_keyframes.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
edge_compile_videoC
Compiles an existing storyboard, keyframes, transitions, and audio into final MP4 video file.
| Name | Required | Description | Default |
|---|---|---|---|
| storyboard | Yes | Storyboard object | |
| applyMotionFx | No | ||
| outputFileName | No | Custom output file name (.mp4) | |
| addBackgroundScore | No | ||
| exportToAndroidGallery | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are supplied, so the description carries the full burden. It states the output is an MP4, but says nothing about whether compilation is long-running, what permissions or device access it needs, whether it overwrites prior output, or how the Android gallery export side effect is triggered.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single efficient sentence that front-loads the action and the result with no filler. It is arguably too terse given the tool's complexity, but there is no structural waste.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a five-parameter compile/render tool with no annotations, no output schema, and undocumented boolean flags that materially change behavior, the description leaves too much unstated. An agent would not know the cost, side effects, or flag semantics before invoking it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 40% across five parameters. The description mentions keyframes, transitions, and audio, but those are not parameters, while the four behavioral flags (applyMotionFx, outputFileName, addBackgroundScore, exportToAndroidGallery) receive no explanation of their effect on the compiled output.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb ("Compiles") and names the concrete inputs and output ("storyboard, keyframes, transitions, and audio into final MP4 video file"), which clearly separates it from generation-oriented siblings like edge_generate_keyframes and edge_apply_video_fx. It stops short of explicitly naming which sibling to run first or instead, so it is clear but not fully differentiated.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no statement of when to use this tool versus edge_video_create, edge_apply_video_fx, or android_gallery_export. The phrase "existing storyboard" hints that a storyboard must already exist, but no prerequisites or ordering guidance are given explicitly.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
edge_generate_keyframesC
Generates visual keyframe cards and assets for each shot in a storyboard.
| Name | Required | Description | Default |
|---|---|---|---|
| storyboard | Yes | Storyboard JSON object from edge_storyboard_plan |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden, yet it discloses nothing about side effects, whether the storyboard is mutated, output artifacts, cost, or latency. One sentence of pure purpose leaves core behavioral questions unanswered.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler. It is efficient, though perhaps overly terse for a generation tool whose behavior is otherwise undisclosed.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no annotations, no output schema, and a nested required object, the definition should say more about what the tool produces and any ordering/permission constraints. The one-liner leaves meaningful gaps for a generation operation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the single parameter's description ('Storyboard JSON object from edge_storyboard_plan') already documents both its shape and its source. The prose adds no parameter detail beyond that, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'Generates visual keyframe cards and assets for each shot in a storyboard.' An agent can distinguish this from edge_storyboard_plan (which produces the storyboard) and edge_compile_video. It does not explicitly name siblings, so it falls short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no when-to-use guidance, no prerequisites, and no ordering relative to siblings. The only hint that this runs after edge_storyboard_plan lives in the schema parameter text, not the description.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
edge_models_managerC
Inspects and manages Google AI Edge Gallery models (Gemma 2, Gemma 3n, MobileDiffusion, MediaPipe tasks).
| Name | Required | Description | Default |
|---|---|---|---|
| task | No | ||
| action | No | list |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description must carry the full behavioral burden. It states 'Inspects and manages' but doesn't clarify whether actions are read-only, destructive, or require authentication. It also doesn't describe side effects, return formats, or limitations, leaving critical behavioral traits undisclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence that front-loads the tool's purpose. It avoids unnecessary words and structure is clear, though it could benefit from additional sentences for guidelines or behavior.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no annotations, no output schema, and 0% parameter description coverage, the description is insufficient. It fails to explain how to use the parameters, what the tool returns, or how it differs from siblings. An agent would struggle to invoke it correctly without trial and error.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description provides no information about the parameters. The input schema defines 'task' and 'action' with enums, but the description doesn't explain what these parameters do or how they affect behavior. For example, it doesn't mention that 'action' controls list/get_recommended/verify operations or that 'task' filters by capability.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a verb ('Inspects and manages') and a resource ('Google AI Edge Gallery models'), which is clearer than a tautology. However, the description is too broad to distinguish it from siblings like edge_storyboard_plan or edge_generate_keyframes, which likely interact with the same models. The specific examples (Gemma, MobileDiffusion, MediaPipe) hint at scope but don't clarify the tool's unique role.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit guidance on when to use this tool versus alternatives. The description mentions model names but doesn't explain scenarios, prerequisites, or exclusions. With seven sibling tools in the edge_* space, this leaves the agent guessing about appropriate context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
edge_storyboard_planB
Plans a director-level screenplay and storyboard breakdown from a raw prompt, specifying scene shots, camera dynamics, transitions, and timing.
| Name | Required | Description | Default |
|---|---|---|---|
| title | No | Optional project title | |
| prompt | Yes | The creative concept or script | |
| shotsCount | No | Custom shot count (optional) | |
| aspectRatio | No | 9:16 | |
| targetDurationSeconds | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It discloses the output's content (shots, camera dynamics, transitions, timing), which is useful, but says nothing about whether this is a costly generation call, whether it is deterministic, or what permissions/limits apply. Adequate but incomplete.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler or repetition. 'director-level' is mild marketing color rather than a functional constraint, but the whole definition is free of bloat.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 5-parameter tool with no annotations and no output schema, the description usefully sketches the return content, which partially compensates. However, it omits pipeline positioning relative to edge_generate_keyframes and edge_compile_video and leaves two parameters unaddressed, so an agent lacks the full picture.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 60%: title/prompt/shotsCount are documented, while aspectRatio and targetDurationSeconds are not. The description's mention of 'scene shots' and 'timing' only loosely gestures at shotsCount and targetDurationSeconds and adds no format, range, or default-value guidance beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb ('Plans') and a concrete artifact ('director-level screenplay and storyboard breakdown'), then enumerates the constituents (shots, camera dynamics, transitions, timing). It is clearly distinguishable from siblings like edge_video_create or edge_compile_video, though it never explicitly names them as alternatives.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'from a raw prompt' implies this is an early planning stage, but there is no explicit when-to-use statement, no prerequisites, and no routing to or away from siblings such as edge_generate_keyframes. Usage is inferable but not stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
edge_video_createB
End-to-end prompt-to-video pipeline powered by Google AI Edge Gallery standards. Plans storyboard scenes, renders keyframes, applies Ken Burns motion and color grading, compiles into MP4, and syncs directly to Android DCIM Gallery / Google Photos.
| Name | Required | Description | Default |
|---|---|---|---|
| style | No | Aesthetic style and motion dynamics | social_reel |
| title | No | Optional custom title for the video | |
| prompt | Yes | The video concept, prompt, or script | |
| aspectRatio | No | Aspect ratio: 9:16 (vertical reel/short), 16:9 (landscape), 1:1 (square) | 9:16 |
| applyMotionFx | No | Apply dynamic Ken Burns camera motion (pan/zoom) | |
| addBackgroundScore | No | Synthesize ambient Edge AI background soundtrack | |
| targetDurationSeconds | No | Total video length in seconds (3 - 120s) | |
| exportToAndroidGallery | No | Export and index into Android MediaStore / Google Photos |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It does disclose real side effects beyond the schema - writing output to Android DCIM Gallery / Google Photos and indexing into MediaStore - which is genuinely useful. However, it omits runtime/latency expectations, permissions or environment requirements, whether existing files are overwritten, and any failure/re-run behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single sentence, front-loaded with what the tool is, followed by the pipeline stages in execution order - it is dense and every clause carries information. Minor waste in the branding phrase 'powered by Google AI Edge Gallery standards', which does not help an agent act.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an 8-parameter, multi-stage generation tool with no output schema, the description covers the workflow and destination well but never says what the call returns (file path, URI, job id, or whether it is synchronous). An agent cannot tell how to retrieve or verify the resulting video, and there is no output schema to compensate.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so every parameter (style, aspectRatio, targetDurationSeconds range, motion/score/gallery toggles) is already documented in the schema with defaults and enums. The description adds no parameter-level syntax, constraints, or interactions, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb+resource ('prompt-to-video pipeline') and enumerates the concrete stages it performs (storyboard, keyframes, motion/grading, MP4 compile, gallery sync). This implicitly distinguishes it from the step-level siblings (edge_storyboard_plan, edge_generate_keyframes, edge_apply_video_fx, edge_compile_video), but it never names them or explicitly frames itself as the orchestrating shortcut.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied: the phrase 'End-to-end ... pipeline' signals this is the all-in-one path versus the granular sibling tools, but there is no explicit when-to-use, when-not-to-use, or alternative-naming guidance. An agent must infer the routing decision rather than read it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
8 tool updates
v1.0.0- First observed
android_device_status - First observed
android_gallery_export - First observed
edge_apply_video_fx - First observed
edge_compile_video - First observed
edge_generate_keyframes - First observed
edge_models_manager - First observed
edge_storyboard_plan - First observed
edge_video_create
TDQS
Scored across 8 tools
edge_video_create is an end-to-end pipeline that subsumes the individual step tools (edge_storyboard_plan, edge_generate_keyframes, edge_apply_video_fx, edge_compile_video), and its gallery-sync step overlaps with android_gallery_export. The modular vs. monolithic paths are explicitly described, so an agent can reason about it, but boundaries remain blurry.
Names are readable snake_case with thematic prefixes (edge_/android_), but the verb/noun ordering is mixed: edge_video_create and edge_storyboard_plan are noun_verb while edge_generate_keyframes, edge_apply_video_fx, and edge_compile_video are verb_noun. The two-domain prefix scheme is a reasonable choice but not fully predictable.
Eight tools is a well-scoped set for a video-generation pipeline, covering planning, keyframes, FX, compile, orchestration, export, device status, and model management without obvious padding.
The surface covers the full prompt-to-MP4 lifecycle plus device/export/model management. Minor gaps exist: edge_compile_video references audio but there is no audio-generation tool, and there is no cleanup/delete for produced assets.
Maintenance
Related MCP Connectors
Orccut: a real timeline video editor for AI agents - journaled edits, FFmpeg/MLT rendering
Create and edit AI videos from chat: plan shots, generate scenes, and export stories and ads.
AI image + video generation for agents: --flag prompt DSL, async generate/poll, x402 pay-per-use.
AI video editor for agents and humans: timeline, captions, color, audio and generation as MCP tools.
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceEnables AI agents to directly control Android devices via Termux, providing 120+ tools for screen manipulation, file management, app control, and system operations with layered loading and security gating.1MIT
- AlicenseAqualityDmaintenanceEnables AI models to control Android devices via ADB through natural language commands, supporting screen analysis and automated actions.339 npmMIT
- AlicenseBqualityDmaintenanceEnables AI agents to control Android devices via ADB, supporting gestures, input, screenshots, UI analysis, and app management.1917 npmISC
- AlicenseNot gradedqualityBmaintenanceEnables AI assistants to remotely control an Android device from a PC, using Termux API for camera, location, SMS, notifications, and shell commands over SSH or SSE.MIT