Skip to main content
Glama
hayhihey

Google AI Edge Gallery Video MCP Server

by hayhihey

Google AI Edge Gallery Video MCP Server ๐ŸŽฌ๐Ÿ“ฑ

CI License Platform MCP

A high-performance Model Context Protocol (MCP) server enabling AI agents (Claude Desktop, Cursor, Antigravity, LLM frontends) to direct, generate, and compile on-device videos using the Google AI Edge Gallery ecosystem (google-ai-edge/gallery).

Built from the ground up to run natively on Android devices (via Termux), through an ADB Bridge (USB or Wi-Fi), or as a Standalone Desktop Engine, with automatic synchronization to Android's DCIM / Google Photos Gallery and hardware-aware battery & thermal throttling protection.


๐ŸŒŸ Key Highlights

  • On-Device Edge Video Generation: Turn raw prompts and narrative ideas into complete, multi-scene video stories with cinematic keyframes, camera motions, color grades, and audio score.

  • Native Android & Termux Support: Zero native C++ compilation hassles. One-line installer on Termux with automatic storage permissions and Android MediaStore scanning.

  • ADB Host-to-Device Bridge: Run the MCP server on your PC/Mac while seamlessly streaming and indexing generated videos directly onto a connected Android smartphone.

  • Hardware-Aware Telemetry: Monitors Android battery charge, device temperature, and thermal throttling states to automatically adapt video resolution (720p vs 1080p) and frame rates.

  • Google AI Edge Gallery Integration: Built to interface with models supported by Google AI Edge Gallery (Gemma 2, Gemma 3n, MobileDiffusion, MediaPipe tasks).

  • Instant MediaStore Registration: Broadcasts ACTION_MEDIA_SCANNER_SCAN_FILE intents so exported videos immediately appear in the phone's native Gallery and Google Photos.


Related MCP server: autoglm-mcp-server

๐Ÿ—๏ธ Architecture Overview

                        โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
                        โ”‚              MCP Client                      โ”‚
                        โ”‚   (Claude Desktop / Cursor / Antigravity)   โ”‚
                        โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
                                               โ”‚ JSON-RPC (stdio / SSE)
                                               โ–ผ
โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
โ”‚                        Google AI Edge Gallery Video MCP Server                         โ”‚
โ”œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ค
โ”‚           Core Engine                โ”‚              Android Subsystem                  โ”‚
โ”‚ โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ” โ”‚ โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”  โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ” โ”‚
โ”‚ โ”‚ Storyboard Director              โ”‚ โ”‚ โ”‚ Native Termux    โ”‚  โ”‚ ADB Bridge            โ”‚ โ”‚
โ”‚ โ”‚ (Gemma Scene Planning & Timing)  โ”‚ โ”‚ โ”‚ (On-Device Host) โ”‚  โ”‚ (Remote Device Push)  โ”‚ โ”‚
โ”‚ โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜ โ”‚ โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜  โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜ โ”‚
โ”‚                  โ–ผ                   โ”‚          โ”‚                        โ”‚             โ”‚
โ”‚ โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ” โ”‚          โ–ผ                        โ–ผ             โ”‚
โ”‚ โ”‚ Keyframe Generator & Styler      โ”‚ โ”‚ โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ” โ”‚
โ”‚ โ”‚ (Diffusion / Edge SVG Engine)    โ”‚ โ”‚ โ”‚ Unified Device Manager                      โ”‚ โ”‚
โ”‚ โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜ โ”‚ โ”‚ - Battery Level & Temp Monitoring           โ”‚ โ”‚
โ”‚                  โ–ผ                   โ”‚ โ”‚ - Thermal Throttling Mitigation             โ”‚ โ”‚
โ”‚ โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ” โ”‚ โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜ โ”‚
โ”‚ โ”‚ MediaPipe FX & Filter Pipeline   โ”‚ โ”‚                        โ”‚                        โ”‚
โ”‚ โ”‚ (Ken Burns Pan/Zoom, Grade LUTs) โ”‚ โ”‚                        โ–ผ                        โ”‚
โ”‚ โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜ โ”‚ โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ” โ”‚
โ”‚                  โ–ผ                   โ”‚ โ”‚ Gallery Sync Manager                        โ”‚ โ”‚
โ”‚ โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ” โ”‚ โ”‚ - Destination: /sdcard/DCIM/GoogleEdgeAI    โ”‚ โ”‚
โ”‚ โ”‚ Hardware-Aware Video Compiler    โ”‚โ”€โ”ผโ”€โ–บ โ”‚ - MediaStore Intent Broadcast             โ”‚ โ”‚
โ”‚ โ”‚ (FFmpeg H.264 + Ambient Audio)   โ”‚ โ”‚ โ”‚ - Google Photos / Gallery Auto-Indexing     โ”‚ โ”‚
โ”‚ โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜ โ”‚ โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜ โ”‚
โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜

๐Ÿš€ Quick Start Guide

Option 1: Native on Android (Termux)

Run this one-line command inside Termux:

pkg install -y curl && curl -sSL https://raw.githubusercontent.com/hayhihey/google-ai-edge-gallery-video-mcp/main/scripts/termux-install.sh | bash

Once installed, simply start the server:

edge-video-mcp

All created videos will automatically appear in your phone's Gallery under GoogleEdgeAI (/sdcard/DCIM/GoogleEdgeAI)!


Option 2: Host PC with Android Phone Connected (ADB Bridge)

  1. Enable Developer Options and USB Debugging (or Wireless Debugging) on your Android phone.

  2. Connect your phone to your computer via USB or Wi-Fi (adb connect <phone_ip>:5555).

  3. Clone and build the project:

    git clone https://github.com/google-ai-edge/google-ai-edge-gallery-video-mcp.git
    cd google-ai-edge-gallery-video-mcp
    npm install
    npm run build
  4. Verify Android device connectivity:

    bash scripts/adb-setup.sh

Option 3: Desktop Standalone (No phone required)

The server automatically detects when no Android device is attached and runs the edge simulation pipeline locally on Windows, macOS, or Linux.

Prerequisite: Ensure ffmpeg is installed:

  • macOS: brew install ffmpeg

  • Ubuntu/Debian: sudo apt install -y ffmpeg

  • Windows: winget install Gyan.FFmpeg or choco install ffmpeg

git clone https://github.com/google-ai-edge/google-ai-edge-gallery-video-mcp.git
cd google-ai-edge-gallery-video-mcp
npm install
npm run build
npm start

๐Ÿ”Œ Connecting to MCP Clients

Claude Desktop Configuration

Add the following to your claude_desktop_config.json:

{
  "mcpServers": {
    "google-ai-edge-video": {
      "command": "node",
      "args": [
        "/path/to/google-ai-edge-gallery-video-mcp/dist/index.js"
      ],
      "env": {
        "EDGE_MCP_OUTPUT_DIR": "/path/to/custom/output"
      }
    }
  }
}

On Windows, use escaped backslashes: C:\\Users\\<Username>\\...\\dist\\index.js


๐Ÿ› ๏ธ MCP Tools Reference

Tool Name

Parameters

Description

edge_video_create

prompt, title?, aspectRatio (9:16, 16:9, 1:1), targetDurationSeconds, style, addBackgroundScore, applyMotionFx, exportToAndroidGallery

Full end-to-end video synthesis pipeline. Directs scenes, synthesizes keyframes, applies camera motions & color grading, encodes MP4, and indexes in Android Gallery.

edge_storyboard_plan

prompt, title?, aspectRatio, targetDurationSeconds, shotsCount?

Directs multi-shot scene breakdowns, camera motions (zoom_in, pan_right, tilt_up), transitions, and captions.

edge_generate_keyframes

storyboard

Generates high-fidelity keyframe image assets for each scene.

edge_apply_video_fx

motion, colorGrade, aspectRatio, durationSeconds, fps

Generates hardware-optimized filtergraphs for Ken Burns motions, vignette, and cinematic color palettes.

edge_compile_video

storyboard, outputFileName?, addBackgroundScore, applyMotionFx, exportToAndroidGallery

Low-level assembler for stitching shot clips, transitions, and audio beds into an H.264 MP4.

android_device_status

(None)

Inspects real-time battery charge, temperature, thermal throttling state, and acceleration delegates.

android_gallery_export

videoFilePath, customTitle?

Moves any video into /sdcard/DCIM/GoogleEdgeAI and broadcasts Android's MEDIA_SCANNER_SCAN_FILE intent.

edge_models_manager

action (list, get_recommended), task?

Catalogs models compatible with Google AI Edge Gallery (Gemma 2, Gemma 3n, MobileDiffusion, MediaPipe).


๐Ÿ“ฆ MCP Resources & Prompts

Resources

  • edge://device/telemetry: Real-time hardware telemetry (battery, temperature, storage, encoder support).

  • edge://gallery/videos: Catalog of generated video files with sizes and timestamps.

  • edge://models/inventory: List of edge models available on device.

Prompts

  • ai_video_director: Creative director persona to brainstorm screenplays tailored to edge constraints.

  • social_shorts_creator: High-engagement 9:16 vertical short template designed for TikTok, Reels, and Shorts.


๐Ÿงช Testing & Verification

Run the comprehensive unit test suite:

npm test

Test coverage includes:

  • Storyboard director planning across aspect ratios (9:16, 16:9, 1:1).

  • Mobile-constrained thermal throttling logic.

  • Procedural SVG keyframe rendering and XML escaping.

  • Ken Burns motion expressions and color grading filtergraphs.

  • Unified device detection (Termux / ADB / Local).


๐Ÿณ Docker Deployment

A lightweight multi-stage Docker image with built-in FFmpeg and Android tools:

# Build and run with docker compose
docker compose up -d

# Or run directly with docker
docker build -t edge-video-mcp .
docker run --rm -v $(pwd)/output:/app/output edge-video-mcp

๐Ÿšข Deploying to GitHub

To publish this repository to your GitHub account:

# 1. Initialize git repository
git init -b main

# 2. Add files and make initial commit
git add .
git commit -m "feat: initial release of Google AI Edge Gallery Video MCP Server"

# 3. Create a new repository on GitHub (e.g. google-ai-edge-gallery-video-mcp)
# 4. Link remote and push:
git remote add origin https://github.com/<your-username>/google-ai-edge-gallery-video-mcp.git
git push -u origin main

The pre-configured GitHub Actions CI/CD workflows (.github/workflows/ci.yml and release.yml) will automatically:

  • Run automated tests across Ubuntu and Windows on Node.js 18, 20, and 22.

  • Package releases upon pushing tags (e.g. git tag v1.0.0 && git push origin v1.0.0).


๐Ÿ“„ License

Licensed under the Apache License, Version 2.0.

Available Tools

8 tools
android_device_statusA

Checks Android device battery percentage, thermal throttling status, available storage, and acceleration mode (Termux, ADB bridge, or Host).

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full disclosure burden. 'Checks' implies a non-destructive read on a zero-parameter tool, which is inherently low risk, and it enumerates the checked properties (including the three acceleration modes). However, it says nothing about permissions, cost, or whether repeated polling is safe.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single dense sentence that front-loads the verb and resource, then lists the returned signals in a compact parallel structure. No filler, no repetition of the tool name's meaning.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema and no parameters, the description compensates by enumerating the four values the caller receives, which is the key information needed to use the result. It stops short of explaining formats or units, a minor gap for such a simple probe.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters and the schema is empty, so there are no parameter semantics to clarify. Baseline 4 applies for a parameterless tool.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a specific verb ('Checks') and a concrete resource ('Android device') and enumerates the exact signals returned: battery percentage, thermal throttling, storage, and acceleration mode. That makes it clearly distinct from the video/gallery siblings, though it never references an alternative tool by name.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to call this tool versus alternatives, no preconditions, and no exclusions. An agent must infer that this is a diagnostic probe for device health, but the description never states that intent.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

edge_apply_video_fxC

Generates hardware-optimized filtergraphs for Ken Burns camera motion, color balance, and cinematic vignette.

ParametersJSON Schema
NameRequiredDescriptionDefault
fpsNo
motionNozoom_in
colorGradeNocinematic_teal_orange
aspectRatioNo9:16
durationSecondsNo

TDQS

C2.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden. It clarifies that the tool generates filtergraphs rather than applying them, which is useful, but it omits whether it returns text or a file, any side effects, authentication needs, or how the output should be consumed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler. Every phrase contributes to identifying the tool's output domain.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 5-parameter tool with no annotations, no output schema, and no parameter descriptions, the definition is materially incomplete. It does not explain the generated filtergraph format, how to use the result, or what each input controls.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, but it only loosely maps to motion and colorGrade via 'Ken Burns camera motion' and 'color balance'. It says nothing about fps, aspectRatio, durationSeconds, or the enum values, and it references a cinematic vignette that is not an input parameter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Generates) and resource (filtergraphs) and names the effect areas it covers, so the agent can tell it apart from create/compile siblings. However, it does not explicitly differentiate itself from other video-effect or keyframe siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance, no preconditions, and no alternatives are mentioned. The description only says what it produces, leaving the agent to infer when this tool should be chosen over edge_compile_video or edge_generate_keyframes.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

edge_compile_videoC

Compiles an existing storyboard, keyframes, transitions, and audio into final MP4 video file.

ParametersJSON Schema
NameRequiredDescriptionDefault
storyboardYesStoryboard object
applyMotionFxNo
outputFileNameNoCustom output file name (.mp4)
addBackgroundScoreNo
exportToAndroidGalleryNo

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are supplied, so the description carries the full burden. It states the output is an MP4, but says nothing about whether compilation is long-running, what permissions or device access it needs, whether it overwrites prior output, or how the Android gallery export side effect is triggered.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single efficient sentence that front-loads the action and the result with no filler. It is arguably too terse given the tool's complexity, but there is no structural waste.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a five-parameter compile/render tool with no annotations, no output schema, and undocumented boolean flags that materially change behavior, the description leaves too much unstated. An agent would not know the cost, side effects, or flag semantics before invoking it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 40% across five parameters. The description mentions keyframes, transitions, and audio, but those are not parameters, while the four behavioral flags (applyMotionFx, outputFileName, addBackgroundScore, exportToAndroidGallery) receive no explanation of their effect on the compiled output.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a specific verb ("Compiles") and names the concrete inputs and output ("storyboard, keyframes, transitions, and audio into final MP4 video file"), which clearly separates it from generation-oriented siblings like edge_generate_keyframes and edge_apply_video_fx. It stops short of explicitly naming which sibling to run first or instead, so it is clear but not fully differentiated.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no statement of when to use this tool versus edge_video_create, edge_apply_video_fx, or android_gallery_export. The phrase "existing storyboard" hints that a storyboard must already exist, but no prerequisites or ordering guidance are given explicitly.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

edge_generate_keyframesC

Generates visual keyframe cards and assets for each shot in a storyboard.

ParametersJSON Schema
NameRequiredDescriptionDefault
storyboardYesStoryboard JSON object from edge_storyboard_plan

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden, yet it discloses nothing about side effects, whether the storyboard is mutated, output artifacts, cost, or latency. One sentence of pure purpose leaves core behavioral questions unanswered.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler. It is efficient, though perhaps overly terse for a generation tool whose behavior is otherwise undisclosed.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no annotations, no output schema, and a nested required object, the definition should say more about what the tool produces and any ordering/permission constraints. The one-liner leaves meaningful gaps for a generation operation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and the single parameter's description ('Storyboard JSON object from edge_storyboard_plan') already documents both its shape and its source. The prose adds no parameter detail beyond that, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: 'Generates visual keyframe cards and assets for each shot in a storyboard.' An agent can distinguish this from edge_storyboard_plan (which produces the storyboard) and edge_compile_video. It does not explicitly name siblings, so it falls short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no when-to-use guidance, no prerequisites, and no ordering relative to siblings. The only hint that this runs after edge_storyboard_plan lives in the schema parameter text, not the description.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

edge_models_managerC

Inspects and manages Google AI Edge Gallery models (Gemma 2, Gemma 3n, MobileDiffusion, MediaPipe tasks).

ParametersJSON Schema
NameRequiredDescriptionDefault
taskNo
actionNolist

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description must carry the full behavioral burden. It states 'Inspects and manages' but doesn't clarify whether actions are read-only, destructive, or require authentication. It also doesn't describe side effects, return formats, or limitations, leaving critical behavioral traits undisclosed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, efficient sentence that front-loads the tool's purpose. It avoids unnecessary words and structure is clear, though it could benefit from additional sentences for guidelines or behavior.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no annotations, no output schema, and 0% parameter description coverage, the description is insufficient. It fails to explain how to use the parameters, what the tool returns, or how it differs from siblings. An agent would struggle to invoke it correctly without trial and error.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, and the description provides no information about the parameters. The input schema defines 'task' and 'action' with enums, but the description doesn't explain what these parameters do or how they affect behavior. For example, it doesn't mention that 'action' controls list/get_recommended/verify operations or that 'task' filters by capability.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a verb ('Inspects and manages') and a resource ('Google AI Edge Gallery models'), which is clearer than a tautology. However, the description is too broad to distinguish it from siblings like edge_storyboard_plan or edge_generate_keyframes, which likely interact with the same models. The specific examples (Gemma, MobileDiffusion, MediaPipe) hint at scope but don't clarify the tool's unique role.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit guidance on when to use this tool versus alternatives. The description mentions model names but doesn't explain scenarios, prerequisites, or exclusions. With seven sibling tools in the edge_* space, this leaves the agent guessing about appropriate context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

edge_storyboard_planB

Plans a director-level screenplay and storyboard breakdown from a raw prompt, specifying scene shots, camera dynamics, transitions, and timing.

ParametersJSON Schema
NameRequiredDescriptionDefault
titleNoOptional project title
promptYesThe creative concept or script
shotsCountNoCustom shot count (optional)
aspectRatioNo9:16
targetDurationSecondsNo

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It discloses the output's content (shots, camera dynamics, transitions, timing), which is useful, but says nothing about whether this is a costly generation call, whether it is deterministic, or what permissions/limits apply. Adequate but incomplete.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler or repetition. 'director-level' is mild marketing color rather than a functional constraint, but the whole definition is free of bloat.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 5-parameter tool with no annotations and no output schema, the description usefully sketches the return content, which partially compensates. However, it omits pipeline positioning relative to edge_generate_keyframes and edge_compile_video and leaves two parameters unaddressed, so an agent lacks the full picture.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 60%: title/prompt/shotsCount are documented, while aspectRatio and targetDurationSeconds are not. The description's mention of 'scene shots' and 'timing' only loosely gestures at shotsCount and targetDurationSeconds and adds no format, range, or default-value guidance beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a specific verb ('Plans') and a concrete artifact ('director-level screenplay and storyboard breakdown'), then enumerates the constituents (shots, camera dynamics, transitions, timing). It is clearly distinguishable from siblings like edge_video_create or edge_compile_video, though it never explicitly names them as alternatives.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'from a raw prompt' implies this is an early planning stage, but there is no explicit when-to-use statement, no prerequisites, and no routing to or away from siblings such as edge_generate_keyframes. Usage is inferable but not stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

edge_video_createB

End-to-end prompt-to-video pipeline powered by Google AI Edge Gallery standards. Plans storyboard scenes, renders keyframes, applies Ken Burns motion and color grading, compiles into MP4, and syncs directly to Android DCIM Gallery / Google Photos.

ParametersJSON Schema
NameRequiredDescriptionDefault
styleNoAesthetic style and motion dynamicssocial_reel
titleNoOptional custom title for the video
promptYesThe video concept, prompt, or script
aspectRatioNoAspect ratio: 9:16 (vertical reel/short), 16:9 (landscape), 1:1 (square)9:16
applyMotionFxNoApply dynamic Ken Burns camera motion (pan/zoom)
addBackgroundScoreNoSynthesize ambient Edge AI background soundtrack
targetDurationSecondsNoTotal video length in seconds (3 - 120s)
exportToAndroidGalleryNoExport and index into Android MediaStore / Google Photos

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden. It does disclose real side effects beyond the schema - writing output to Android DCIM Gallery / Google Photos and indexing into MediaStore - which is genuinely useful. However, it omits runtime/latency expectations, permissions or environment requirements, whether existing files are overwritten, and any failure/re-run behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single sentence, front-loaded with what the tool is, followed by the pipeline stages in execution order - it is dense and every clause carries information. Minor waste in the branding phrase 'powered by Google AI Edge Gallery standards', which does not help an agent act.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For an 8-parameter, multi-stage generation tool with no output schema, the description covers the workflow and destination well but never says what the call returns (file path, URI, job id, or whether it is synchronous). An agent cannot tell how to retrieve or verify the resulting video, and there is no output schema to compensate.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so every parameter (style, aspectRatio, targetDurationSeconds range, motion/score/gallery toggles) is already documented in the schema with defaults and enums. The description adds no parameter-level syntax, constraints, or interactions, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb+resource ('prompt-to-video pipeline') and enumerates the concrete stages it performs (storyboard, keyframes, motion/grading, MP4 compile, gallery sync). This implicitly distinguishes it from the step-level siblings (edge_storyboard_plan, edge_generate_keyframes, edge_apply_video_fx, edge_compile_video), but it never names them or explicitly frames itself as the orchestrating shortcut.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is only implied: the phrase 'End-to-end ... pipeline' signals this is the all-in-one path versus the granular sibling tools, but there is no explicit when-to-use, when-not-to-use, or alternative-naming guidance. An agent must infer the routing decision rather than read it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 8 tool updatesv1.0.0
    • First observedandroid_device_status
    • First observedandroid_gallery_export
    • First observededge_apply_video_fx
    • First observededge_compile_video
    • First observededge_generate_keyframes
    • First observededge_models_manager
    • First observededge_storyboard_plan
    • First observededge_video_create

TDQS

B3.1/5.0

Scored across 8 tools

Disambiguation3/5

edge_video_create is an end-to-end pipeline that subsumes the individual step tools (edge_storyboard_plan, edge_generate_keyframes, edge_apply_video_fx, edge_compile_video), and its gallery-sync step overlaps with android_gallery_export. The modular vs. monolithic paths are explicitly described, so an agent can reason about it, but boundaries remain blurry.

Naming Consistency3/5

Names are readable snake_case with thematic prefixes (edge_/android_), but the verb/noun ordering is mixed: edge_video_create and edge_storyboard_plan are noun_verb while edge_generate_keyframes, edge_apply_video_fx, and edge_compile_video are verb_noun. The two-domain prefix scheme is a reasonable choice but not fully predictable.

Tool Count5/5

Eight tools is a well-scoped set for a video-generation pipeline, covering planning, keyframes, FX, compile, orchestration, export, device status, and model management without obvious padding.

Completeness4/5

The surface covers the full prompt-to-MP4 lifecycle plus device/export/model management. Minor gaps exist: edge_compile_video references audio but there is no audio-generation tool, and there is no cleanup/delete for produced assets.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers