flow-studio
Automates Google Flow (Veo, Nano Banana, Omni) in the user's signed-in browser session to create and manage video/image generation projects. Provides tools for creating and opening projects, uploading and listing grid assets, inspecting live Flow options (modes, models, aspect ratios, resolution, duration, variants and credit price), generating images/videos/frames from prompts with named references, building reusable characters and voices, running shot lists, extending/editing scenes and exporting them as mp4, using Flow's built-in and community tools, and upscaling downloads to 1080p/4K — all gated behind an explicit price quote that must be confirmed before credits are spent.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@flow-studiogenerate a 5-second Veo video of a cat, show me the price first"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
mcp-google-flow
An MCP server that lets any MCP client produce videos and images in Google Flow (Veo, Nano Banana, Omni) through your own signed-in Chrome, with spending guards and a sandboxed file vault.
Independent project,not affiliated with, endorsed or maintained by Google. "Google Flow" and "Veo" are trademarks of Google. It automates Flow's web interface on your account. Automation may go against Google's Terms of Service and could lead to limits or suspension of the account. Use it at your own risk, ideally with an account separate from your main one. Flow changes its interface without notice; when it does, a tool may stop working until its labels are adjusted (see When Flow's interface changes).
Showcase

15 visual styles generated in Google Flow: food, cinematic, anime, fight scene, brand story, music video, social hook, 3D, motion design, cartoon, comic, product ad, fashion, real estate and product 360. ▶ Watch the 62-second reel on YouTube.
Related MCP server: google-flow-mcp
What it does
Projects and media: creates and opens projects, uploads images and videos, lists what is in the grid, waits for renders and downloads with predictable names.
Generation with Flow's real options: reads the modes (Image, Video, Frames), models (Veo 3.1, Omni, Nano Banana…), aspect ratios, resolution, duration, variants and the credit price Flow itself shows, live.
References by name, like typing
@name: characters, voices, avatars, images and videos from the library.Frames: video that starts and/or ends on a chosen image.
Reusable characters, created from Flow's presets or from a description.
Shot lists: several shots with the same character, with the whole list checked against the budget before the first credit is spent.
Scenes: builds the timeline, extends the video, edits a clip with a prompt and exports the whole scene as a single mp4.
Flow tools (Grid Architect, Stringout Creator, Storyboard Studio, Video Resizer and community tools): opens any of them, reads their controls, fills and clicks them.
Upscale to 1080p and 4K on download.
Shot planner: turns a script into ready shot prompts, with the camera chosen by each shot's job, natural gestures, a voice and accent kept identical across clips, and single-take prompts.
Guided slash commands for 9 base-image jobs, script writing and voice design.
Pro version with ready-made video formats (see Pro version).
Nothing spends credits without an approved quote. Every paid action takes two calls: the first prepares everything, shows the price and returns a quote_id; only the second, with confirm: true and that id, spends.
Requirements
Node.js 20.11+
Google Chrome
A Google account with access to Flow
Installation
git clone https://github.com/felipedamacenoteodoro/mcp-google-flow.git
cd mcp-google-flow
npm ci
npm run buildRegister it in your MCP client as a stdio server. Most clients accept this configuration:
{
"mcpServers": {
"flow-studio": {
"command": "node",
"args": ["/absolute/path/to/mcp-google-flow/dist/main.js"]
}
}
}First run
Ask the agent to "call flow_sign_in". It opens a normal Chrome window, with no automation attached, on the project's own profile. Google refuses to sign in inside automated browsers, so the sign-in happens here.
Sign in to Google yourself in that window. The server never sees or types your password.
Quit that Chrome completely (Cmd+Q on Mac; closing the window is not enough there).
Ask for anything else, e.g. "flow_new_project". The server reopens the same profile, already signed in, and takes it from there.
The login is kept in the dedicated profile (~/.flow-studio-mcp/chrome-profile). Your everyday Chrome is never touched or copied.
Flow can be in Portuguese or English out of the box. If your Google account uses another language, set FLOW_MCP_LANGUAGE=en and Flow opens in English for the server. Don't use browser translation on Flow (see Troubleshooting).
Tools
Area | Tool | Spends? | What it does |
Session |
| no | Connection, sign-in state, current URL and credit budget. Never launches Chrome. |
| no | Opens Flow in the dedicated profile so you can sign in. | |
| no | Lists projects, newest first. | |
| no | Creates a project and turns off the "Agent" that rewrites prompts. | |
| no | Opens a project (only | |
| no | Screenshot for debugging. | |
Grid |
| no | Project items: index, kind (video/image/scene), name, ready or not. |
| no | Uploads an image or video from an allowed input folder. | |
| no | Waits until the project holds N finished items. | |
| 1080p/4K only | Saves into the output folder. | |
| no | Creates a scene from a video. | |
Generation |
| no | Modes, models, ratios, resolution, duration, variants and current price. |
| yes | Image, video or frames, with references and attachments. | |
| yes | A sequence of shots with shared settings. | |
Library |
| no | Finds characters, voices, avatars, images and videos by name. |
| no | Attaches a resource to the prompt box. | |
| no | Character presets and the image model in use. | |
| yes | Creates a reusable character. | |
Scenes |
| no | Opens a scene and reports whether it can be extended. |
| yes | Extends the scene (Veo-generated clips only). | |
| yes | Edits the clip with a prompt. | |
| no | Exports the whole scene as one mp4. | |
| no | Back to the grid. | |
Planning |
| no | Layouts, camera moves by purpose, gestures, prompt rules and the review checklist. |
| no | Script → shots with line, framing, camera, prompt and attachments. | |
| no | Base-image prompt: avatar, identity sheet, outfit, angle, product, app… | |
Tools |
| no | Your tools, Google templates or community tools. |
| no | Opens a tool by name; reuses your copy when one exists. | |
| no | Reads the open tool's controls and their values. | |
| no | Fills a field. | |
| when it looks like generating | Clicks a control. Buttons such as "Generate" require a quote. | |
| no | Back to the grid. |
How spending is approved
The agent calls, for example,
flow_generatewithconfirm: false. Flow gets configured, the prompt is typed, and the reply carries the price Flow shows plus aquotewith aquoteId.You review it. If you approve, the agent calls again with the same request,
confirm: trueandquote_id.The server refuses when there is no quote, the quote belongs to another request, more than 10 minutes have passed, it was already used, or the price went up. When Flow shows no price (upscale, extend, tools) a conservative estimate is charged, never zero.
Example conversation
Create a project, upload
~/FlowStudio/input/barista.png, make a character from it and generate three 9:16 shots of her presenting the new seasonal coffee. Show me the price first.
The agent calls flow_new_project → flow_upload → flow_create_character (quote → confirm) → flow_run_shot_list with attach_resources (quote → you approve → confirm) → flow_wait → flow_download.
Guided commands
The server ships step-by-step flows that clients supporting MCP prompts show as slash commands, e.g. /mcp__flow-studio__image_avatar.
Image (the base image is what keeps the face identical in every clip)
Command | What it does |
| Creates a realistic person to be the base of the videos |
| Nine views of the same face, to keep the character consistent |
| Changes outfit or background, keeping the person |
| Changes one detail of the appearance (hair, glasses, age) |
| The same scene from another camera angle |
| The person holds the product, label untouched |
| A real app screenshot on the phone, in two steps |
| Before and after of the same person |
| A different person in the same pose, place and light |
Support
Command | What it does |
| Writes a short script, one sentence per shot |
| Builds a voice description to repeat in every clip |
The agent leads the conversation one question at a time, shows the quote, generates only after your "yes" and tells you the library name of the result so it can be reused as a reference. flow_image_prompt can also be called directly; it does not touch Flow and spends nothing.
Shot planner
flow_plan_shots turns a script into one shot per spoken sentence and writes each prompt in a fixed order: scene → camera → gesture → line → voice → accent → "one continuous take, no cuts". You pick a layout:
Layout | On screen |
| One person talking to the lens |
| Both people in every shot; one talks, the other listens with their mouth closed |
| Each shot shows only whoever is speaking, looking toward the other person |
| Nobody talks on camera; shots illustrate a narration |
What it takes care of:
One spoken sentence per shot. Several sentences in one clip lead to cuts mid-speech.
The camera follows the shot's job: hook, argument, key line or call to action, e.g. locked on the key line, slow zoom on the face at the close. Framings and camera moves can be overridden per role (see
flow_directing_guide).Voice and visual description repeated word for word in every clip of the same person; the plan warns when a voice is missing.
Two people in frame: the speaker is named by side (left/right) and only they get a voice. All clips can be animated from one base image with both people (frames mode).
Long scripts: the plan warns that lip-sync degrades and suggests narration.
Script markers:
Leo: I never know which coffee to order.
Ana: [smiles] * Start with the seasonal blend. [nods]Name: sets the speaker, [gesture] at the start or end of a line adds a gesture (at most one on each side, written in English), and * marks the key line. Spoken lines can be in any language. Each person's voice goes in the cast, as text or as {gender, age, pitch, texture, delivery}.
Pro version
The reel above shows what Google Flow can produce. The Pro version gets you there faster: it adds ready-made video formats on top of this server, each with its own guided slash command: selfie testimonial (UGC), podcast, dualcast, voiceover, product demo, skincare, app demo, fashion, animated product, trend and story. Every format brings its casting, framings, camera plan, rules and the questions to ask, so a complete video comes out of a single conversation.

All 11 Pro formats, generated in Google Flow with the Pro guided commands. ▶ Watch the 44-second reel on YouTube.
Interested? Get in touch:
Email: felipe.devops@gmail.com
WhatsApp: +55 21 97274-5771
Configuration
Everything is set through environment variables, validated at startup:
Variable | Default | Purpose |
|
| Folders uploads are accepted from (separate with |
|
| Where downloads are saved. |
|
| Credit budget per server session. |
|
| Estimate per variant when Flow shows no price. |
|
| Cap on paid actions per hour. |
|
| Minimum pause between two paid actions (shot lists wait automatically). |
|
|
|
| — | Forces Flow's interface language, e.g. |
|
| Maximum upload size. |
|
| Dedicated Chrome profile. |
| auto-detected | Path to Chrome. |
| — | Connect to an already running Chrome (loopback only). |
| — | JSON with interface labels (see below). |
Security
What the server guarantees, in short (details in SECURITY.md):
Dedicated profile: never uses or copies your personal Chrome profile.
DevTools on loopback only, random port: Chrome picks the port. Endpoints outside
127.0.0.1/localhostare refused.Navigation restricted to
https://flow.google.com.File vault: uploads only from allowed folders, with the real path resolved (symlinks cannot escape), content checked by its leading bytes rather than its extension, and a size limit. Downloads only into the output folder, always as a new file (never overwrites, never follows a planted symlink).
Spending: single-use quote bound to the exact request and price, a credit budget per session, an hourly limit, and attachments re-checked right before the click.
No code built from user text: prompts are typed as keystrokes; values reach the browser as arguments, never as script source.
Logs on stderr only, with emails, tokens and cookies masked. Internal errors never leak stack traces or paths to the agent.
Tools run one at a time (there is a single tab), so concurrent calls cannot interleave.
When Flow's interface changes
Every piece of interface text the server depends on lives in src/infrastructure/flow/ui-labels.ts. If Flow changes a label, or your interface is in another language, create a JSON file with only what changed and point FLOW_MCP_UI_LABELS to it:
{ "submit": "Generate", "tileMenu": "More options" }Unknown keys are rejected, so a typo never passes silently. See ui-labels.example.json.
Every label accepts alternatives separated by |, and the defaults carry both Portuguese and English (e.g. "Iniciar geração|Start generation"), so the same install works in either language. An override can add a third language the same way.
The defaults were checked against both languages on 2026-10-08: home, project, navigation, settings panel (modes, ratios, resolution, duration, models, price), start/end frames, resource picker, tile menu, download qualities, characters and tools. In English the scene builder (add clip, extend, download scene) and the video-rights dialog are not yet validated.
Troubleshooting
"We noticed unusual activity… You have not been charged" ("Notamos uma atividade incomum"). This comes from Google's abuse protection, not from the server, and Flow does not charge for it. The server does not try to get around it. What helps:
Don't translate the Flow page. Browser translation and other extensions that change pages are the most common trigger. Turn translation off for flow.google.com and use
FLOW_MCP_LANGUAGE=eninstead if your account is in another language.Slow down. Generate fewer videos back to back; the defaults already pace paid actions (
FLOW_MCP_MIN_SECONDS_BETWEEN_PAID,FLOW_MCP_PAID_PER_HOUR).Use Flow normally for a while in the same Chrome profile, and avoid VPNs.
Click generate yourself: with
FLOW_MCP_SUBMIT=manualthe server still prepares everything (mode, model, references, prompt, quote) and brings the Flow window to the front; you press the generate arrow, then the agent continues withflow_wait. Shot lists need automatic clicks, so in this mode generate one shot at a time.
"Couldn't sign you in — this browser or app may not be secure." You signed in inside the automated window. Call flow_sign_in again: it now opens a plain Chrome window for signing in.
"Chrome did not open a DevTools port" right after signing in. The sign-in Chrome is still running. Quit it completely (Cmd+Q on Mac) and try again.
A generation fails with "Failed to generate audio". Flow could not voice the line; it does not charge for it. Try a shorter, simpler line, or the same shot without dialogue.
Architecture
src/
├── domain/ pure rules: prompt, settings, credits, shot planner, directing craft, image jobs, scenes, tools
├── application/ use cases, ports (one interface per area of Flow) and the spend guard
├── infrastructure/ adapters: Chrome/CDP, one class per area of Flow, file vault, logger, config
├── interface/mcp/ tool schemas and registration, guided commands
└── main.ts composition root: the only place that instantiates concrete classesDependencies always point inward: the domain imports nothing, the application imports only the domain, and the infrastructure implements the ports. That is why the use cases are tested against a fake Flow, with no browser.
npm test # domain, use cases and security guards
npm run typecheckLicense
MIT © Felipe D. Teodoro
This server cannot be deployed
Maintenance
Related MCP Connectors
AI image, video, voice and music generation over MCP, routed to Veo 3.1, Seedance 2.5 and more.
Remote streamable-HTTP MCP server running on a single Cloudflare Worker. Your assistant gets live Airbnb, Amazon, Booking.com, Google Flights, Maps and Reddit data, social search on X, Instagram and TikTok, the Meta Ad Library, and image/video generation without any keys. Connect your own accounts to let it send WhatsApp or Telegram messages, work an IMAP inbox, manage Meta Ads campaigns and publish to X and LinkedIn. OAuth 2.1 with PKCE; stored credentials are AES-256-GCM encrypted.
Browser MCP for logged-in tasks. Uses your Chrome — credentials stay local. Zero-token replay.
Plan, compare, price, generate, and recover AI video from compatible MCP clients.
Related MCP Servers
- AlicenseAqualityDmaintenanceEnables AI agents to drive Google Flow through a real Chrome profile to generate images, videos, characters, and scenes without sharing credentials.1911 npm1MIT
- AlicenseAqualityAmaintenanceEnables AI agents to programmatically generate images and videos through the authenticated Google Flow web interface via a direct Chrome DevTools Protocol connection, exposing tools for media generation, project management, status checks, and asset downloads without requiring official API keys.273MIT
- AlicenseAqualityAmaintenanceEnables AI agents to generate Google Flow videos and images and automate Scene Builder clip extensions through a user's own Chrome session.1174 npm4MIT
- AlicenseCqualityCmaintenanceEnables Claude, Cursor, or any MCP client to create consistent AI characters, generate Veo 3.1 and Omni videos, edit clips, build scenes, export MP4s, and chat with Google Flow's creative agent. Automates AI video production workflows such as reels, storyboards, and multi-shot stories through natural language, with credit guardrails and dry-run pricing.232MIT