Flow Workspace MCP
Integrates with Google Flow's web generation workspace to read available models and parameters, submit image and video generation tasks, track job progress, download generated assets, and optionally upscale videos using the user's Google Flow account and quota.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Flow Workspace MCP读取我 Flow 账号可用的视频模型,用横版生成一个纸艺动画镜头,完成后下载到 D:\clips"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Flow Workspace MCP
把 Google Flow 的网页生成流程,接入可追踪、可重复的本地视频制作流程。
适合需要持续制作教学视频、科普视频和多个分镜素材的创作者。通过 MCP,AI 助手可以读取账号实际可用的模型与参数、提交生成任务、查询进度,并把对应素材下载到本地。无需另购视频生成 API,但仍使用你的 Google Flow 权限和额度。
开发状态:正在适配新版
flow.google.com。登录会话复用和工作区读取已实测;单张图片和一个 4 秒视频已生成、下载并通过解码检查。测试期间修复了问题并恢复原任务,尚未验证修复后新任务全程无干预运行;多输出和放大也未验收。当前不应视作稳定生产版。最新证据见 验证记录。
解决什么问题
做一条视频往往需要多个片段。重复复制提示词、选择模型和比例、等待生成、确认素材、下载改名,会把制作过程拆成很多人工操作。让 AI 助手逐次通过截图点击,也需要反复读取界面,耗时且容易受页面变化影响。
本项目将这些动作收束为有明确输入、任务编号和输出文件的工具调用。提示词来自你自己的脚本或分镜;Flow 负责生成;本项目负责连接工作区、执行参数、跟踪任务和交付素材。
Related MCP server: google-flow-mcp
有什么优势
能力 | 对制作流程的意义 |
复用本地登录会话 | 会话有效时继续工作,减少重复登录 |
记住上次工作区 | 重启后回到已有项目,减少反复创建空项目 |
读取实际模型和参数 | 以账号当前界面为准,不把写死的模型列表当成可用能力 |
持久化任务编号 | 等待、查询和下载围绕同一任务进行,避免排队时重复提交 |
匹配具体生成素材 | 不简单下载图库里“最后一个”或“最新的”文件 |
本地文件与参数记录 | 素材能交给 FFmpeg 或其他剪辑流程,便于按镜头管理 |
截图与控件诊断 | 页面改版或识别失败时留下证据,方便定位问题 |
相比 AI 助手逐步操作浏览器,明确参数和简短任务状态通常能减少界面读取及交互次数;具体 Token 和时间收益需要按同一任务实测。项目不提升视频模型本身的画质,也不减少 Flow 对生成任务收取的额度。
如何工作
flowchart LR
A[脚本与分镜提示词] --> B[AI 助手调用 MCP]
B --> C[本地 Flow 工作区适配器]
C --> D[Google Flow 生成]
D --> E[任务与素材身份跟踪]
E --> F[本地视频及参数清单]
F --> G[配音、字幕与剪辑流程]连接账号:浏览器扩展将你已有的 Google 登录会话发送到本机回环地址。账号使用隔离的本地浏览器配置;不会要求把密码写入项目。
识别工作区:校验当前 Flow 地址、项目与编辑器控件,读取实际显示的模型、比例和输出数量。适配器同时兼容已有控件和新版 Angular Material 控件。
执行任务:每次生成都有独立任务编号与状态记录。请求必须明确允许消耗额度;项目不购买订阅或额外额度。
确定素材:生成前记录素材基线,之后按素材标识与来源匹配新增输出。匹配不明确时返回错误,避免交付错误视频。
保存结果:下载到指定绝对路径,记录文件大小与 SHA-256。有可用 FFprobe 时,还记录时长、分辨率和编码信息。
它通过真实网页工作区执行操作,并非 Google 官方生成 API。因此网页改版仍可能需要更新适配器;任务记录和诊断能帮助维护,但不能保证永远免维护。
快速开始
需要 Node.js 20+、已安装的 Chrome/Chromium,以及可以访问 Flow 的 Google 账号。Windows 是目前实际调试的平台。
git clone https://github.com/QIANLING-0831/flow-workspace-mcp.git
cd flow-workspace-mcp
npm ci
npm run check使用绝对路径注册到 Codex,例如:
codex mcp add flow-workspace -- node "C:\tools\flow-workspace-mcp\dist\index.js"其他支持 stdio 的 MCP 客户端也可以配置 node 加绝对路径 dist/index.js。首次连接步骤见 安装说明。后续先读取账号状态,已有有效连接就直接使用;仅在会话失效、退出登录或更换账号时重新连接。
主要工具
工具 | 用途 |
| 查询已连接账号与默认账号 |
| 验证工作区、读取可用能力和诊断 |
| 首次连接或重新连接账号 |
| 提交生成请求,返回任务状态 |
| 按原任务编号查询进度 |
| 下载该任务对应的素材 |
| 使用账号实际提供的升级分辨率选项,可能额外消耗额度 |
工具保留 flow_* 名称,以便复用已有客户端和制作流程。
例如对 AI 助手说:
读取我的 Flow 账号可用的视频模型。用我确认的模型生成一个横版纸艺动画镜头,只生成一个输出,不升级分辨率,保存到我指定的素材目录。如果仍在生成,查询同一任务,不要再次提交。
边界
这是素材生成自动化工具,不是完整的脚本、配音、字幕或剪辑软件。
模型、时长、参考素材和升级分辨率是否可用,以账号当前能力为准。
自动化不绕过账号访问限制或验证码,不承诺免费无限生成。
默认保留本地账号配置、任务和诊断;不要将这些运行数据上传到仓库。
生成等待超时不等于任务失败。查询已有任务,避免重复生成和重复消耗额度。
开发与来源
npm run check
npm pack --dry-run本项目由 QIANLING-0831 维护,基于 Google Flow MCP 的 MIT 代码发展而来,保留原始贡献历史。我们重点维护新版工作区适配、会话复用、任务可追踪性和本地视频制作交接。来源与改动说明见 SOURCE.md,许可证见 LICENSE。
本项目与 Google 无官方关联。
Available Tools
11 toolsflow_begin_account_connectionBegin Google Flow Account ConnectionAIdempotent
FIRST CONNECTION STEP. Returns the exact user instructions and extension path immediately; it never blocks waiting for a click. After this response, STOP and tell the user to open Flow Login Bridge in their normal signed-in Chromium, click Connect Flow, and reply when the popup says Session sent. Do not call the completion tool until the user replies. Use this tool for Google Flow; never substitute generic browser/computer-use automation on the Flow website.
| Name | Required | Description | Default |
|---|---|---|---|
| label | No | Optional intended local label, used only in the returned connection plan. | |
| accountId | No | Optional intended local account ID, used only to make the returned instructions specific. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the annotations (readOnlyHint=false, idempotentHint=true), the description discloses the key non-blocking behavior: it returns instructions immediately and never waits for a click. It does not describe state mutation details or auth requirements, but adds meaningful operational context that the annotations alone do not convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The critical instruction ('FIRST CONNECTION STEP') is front-loaded and every sentence serves the workflow. However, the closing stop-and-wait directives slightly overlap ('STOP and tell the user...' vs 'Do not call the completion tool until the user replies'), adding minor redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Even without an output schema, the description states what the response contains (instructions and extension path) and lays out the full begin → user action → completion lifecycle. Nothing an agent needs to invoke and sequence this tool correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the schema fully documents both optional parameters. The description does not add syntax or format detail for label/accountId, so the baseline of 3 is appropriate when the schema carries the parameter semantics.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb+resource marker ('FIRST CONNECTION STEP') and states exactly what the call returns (user instructions and extension path). It is clearly distinguished from the sibling flow_complete_account_connection by naming it as the later step that must not be called yet.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives explicit when-to-use ('first connection step'), when-not-to-call ('Do not call the completion tool until the user replies'), and an alternative-avoidance rule ('never substitute generic browser/computer-use automation on the Flow website'). The sequencing constraint is unambiguous.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
flow_complete_account_connectionComplete Google Flow Account ConnectionAIdempotent
SECOND CONNECTION STEP. Call only after the user confirms they clicked Connect Flow and the extension popup says Session sent. Before calling, tell the user that a temporary Google chooser will open and they should click the desired existing account. Requires the begin-step connectionId and explicit user confirmation; if the session was not sent it fails immediately instead of blocking. Use this tool for Google Flow; never substitute generic browser/computer-use automation on the Flow website.
| Name | Required | Description | Default |
|---|---|---|---|
| label | No | Optional local label. Omit to use the transferred Chromium profile name. | |
| accountId | No | Optional local ID. Omit for the simplest setup; a safe ID is created automatically. | |
| connectionId | Yes | Required token returned by flow_begin_account_connection. Never invent or reuse it. | |
| chooseGoogleAccount | No | Keep true to show Google's chooser populated with the transferred existing accounts. | |
| userConfirmedSessionSent | Yes | Must be true only after the user explicitly says the extension popup displayed Session sent. | |
| waitForAccountSelectionSeconds | No | Time for the user to click an existing account in the temporary Google chooser and for Flow to load. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare openWorldHint=true, idempotentHint=true and destructiveHint=false, so the safety profile is mostly covered. The description still adds real behavioral context beyond the structured fields: a temporary Google chooser will appear, the user must select an account, and the call fails immediately (non-blocking) if the session wasn't sent. It could go further on rate limits or what happens on partial success, but it is well above the baseline.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The critical routing information ('SECOND CONNECTION STEP', call only after confirmation) is front-loaded, and every sentence carries requirements or user-facing instructions. It is slightly denser than necessary, with some redundancy between the confirmation precondition and the userConfirmedSessionSent semantics, but nothing is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With annotations covering the mutation/open-world profile and no output schema to explain, the description still supplies the full orchestration story: ordering relative to the begin step, the user interaction required mid-call, the timeout behavior, and the ban on falling back to generic browser automation. Nothing an agent needs to call this correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents every parameter including connectionId provenance, the const on userConfirmedSessionSent, and the wait timeout. The description reinforces that connectionId comes from the begin step and that user confirmation is required, but adds no new syntax or format detail, making 3 the correct baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with 'SECOND CONNECTION STEP' and names the specific verb+resource (complete the Google Flow account connection), explicitly positioning it against its sibling flow_begin_account_connection by requiring the begin-step connectionId. An agent can distinguish it from flow_begin_account_connection and the flow_* automation tools without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It states the exact precondition ('Call only after the user confirms they clicked Connect Flow and the extension popup says Session sent') and an explicit exclusion ('never substitute generic browser/computer-use automation on the Flow website'). The when-to-use condition is fully specified rather than implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
flow_download_jobDownload an Existing Google Flow JobA
The exclusive path to download an existing Flow job. Uses the exact asset identities recorded for that job and fails instead of guessing from gallery order. Never download through generic browser/computer-use tools. Saves to the configured absolute directory, validates files, writes manifests, and does not start a generation.
| Name | Required | Description | Default |
|---|---|---|---|
| jobId | Yes | Exact UUID returned by the video or image generation whose asset must be downloaded. Never substitute a different or guessed job ID. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare non-readOnly, openWorld, non-idempotent, non-destructive. The description adds substantive behavior: saves to a configured absolute directory, validates files, writes manifests, fails rather than guessing a mismatched asset, and does not trigger generation. It does not state overwrite/idempotency behavior on repeat calls, which is the main remaining gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the tool's role and single-sentence constraint, then behavior. Slightly redundant in restating exactness ('exact asset identities' / 'fails instead of guessing'), but every sentence carries usable information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter tool with no output schema, the description covers scope, side effects (directory writes, manifests), failure mode, and the non-generation guarantee. An agent can call it correctly; only idempotency/overwrite semantics on repeated downloads are unaddressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the single jobId parameter is fully documented in the schema itself (UUID format, pattern, 'exact UUID returned by the generation'). The description reinforces exactness but adds no syntax or format detail beyond the schema, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('download an existing Flow job') and immediately differentiates the mechanism from siblings/categories: it uses recorded asset identities rather than gallery order, and explicitly rules out generic browser/computer-use tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a clear condition for use ('the exclusive path to download an existing Flow job') and an exclusion ('never download through generic browser/computer-use tools', 'does not start a generation'). It does not name the specific flow_generate_image/flow_generate_video siblings that produce the jobId, but the surrounding guidance is strong.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
flow_generate_imageGenerate and Download a Google Flow ImageADestructive
THE REQUIRED AND EXCLUSIVE PATH for every Google Flow image request. Submits one persistent job, waits briefly, and returns. If status=processing, poll flow_job_status with the SAME job ID; never generate again. It generates or edits through the verified workspace and downloads locally. Never open or control Flow with browser/computer-use tools. Requires explicit credit authorization.
| Name | Required | Description | Default |
|---|---|---|---|
| model | No | Normalized image model ID returned by flow_inspect_account, e.g. nano-banana-pro, nano-banana-2, or nano-banana-2-lite. Exact visible labels are also accepted. Use ui-default to keep the selected model. | ui-default |
| prompt | Yes | Detailed image prompt or edit instruction in any language. | |
| outputs | No | Number of image outputs requested. | |
| download | No | Download ready outputs when true; otherwise return the Flow job without downloading. | |
| fileName | No | Optional safe file stem. The job ID and downloaded extension are added automatically. | |
| accountId | No | Omit accountId to use the most recently verified connected account. Supply it only when the user explicitly chose a different connected account returned by flow_list_accounts. | |
| aspectRatio | No | Exact image aspect ratio returned by flow_inspect_account, or ui-default. | ui-default |
| flowProject | No | Existing Flow project name to open. If omitted, the current project is reused or a new project is created. | |
| referenceFiles | No | Absolute paths to optional local images or videos to attach as Flow references/ingredients/frames. | |
| timeoutSeconds | No | Maximum seconds to wait for this generation or upscale. A timeout leaves a persistent job that can be polled. | |
| outputDirectory | Yes | Absolute directory where downloads and .flow.json manifests are saved, e.g. C:\project\public\generated\flow. | |
| confirmCreditSpend | Yes | Must be true. Confirms the user explicitly authorized this operation to consume Google Flow/AI credits. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare destructive/openWorld/non-idempotent, and the description materially extends them: one persistent job that returns briefly and must be polled rather than re-invoked, local downloading, credit consumption, and the prohibition on browser automation. It discloses the destructive risk (credit spend, non-idempotent job creation) in operational terms the annotations cannot express.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Dense, front-loaded imperative prose; the exclusivity claim and the polling rule lead, and every sentence carries operational information. No filler or restatement of the title.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 12-parameter, credit-consuming, non-idempotent mutation with no output schema, the description covers the full job lifecycle (submit, wait briefly, poll, download), the auth/credit gate, and the forbidden alternative approach. Nothing an agent needs to invoke it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the schema itself documents all 12 parameters and the baseline would be 3. The description still adds workflow-level meaning by tying the job-ID reuse and the 'explicit credit authorization' requirement to invocation, so it earns a step above baseline without duplicating field-level detail.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource ('generates or edits' images), declares itself 'THE REQUIRED AND EXCLUSIVE PATH for every Google Flow image request,' and implicitly separates itself from the video sibling (flow_generate_video). An agent can identify the tool's remit without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit routing rules are given: on status=processing, poll flow_job_status with the SAME job ID and 'never generate again,' and never use browser/computer-use tools for Flow. It also states the credit-authorization prerequisite. This is a textbook when/when-not/alternative block.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
flow_generate_videoGenerate and Download a Google Flow VideoADestructive
THE REQUIRED AND EXCLUSIVE PATH for every Google Flow video request. Submits one persistent job, waits briefly, and returns. If status=processing, poll flow_job_status with the SAME job ID; never generate again. It optionally upscales and downloads locally. Never open or control Flow with browser/computer-use tools. Requires explicit credit authorization.
| Name | Required | Description | Default |
|---|---|---|---|
| model | No | Normalized video model ID returned by flow_inspect_account, e.g. omni-flash, veo-3-1-lite, veo-3-1-fast, or veo-3-1-quality. Exact visible labels are also accepted. Use ui-default to keep the selected model. | ui-default |
| prompt | Yes | Video prompt in any language, describing subject, action, setting, camera, lighting, style, and audio as desired. | |
| outputs | No | Number of generated video outputs requested from Flow. Credits are typically charged per generation. | |
| upscale | No | Exact normalized upscale ID returned by flow_inspect_account (for example 1080p, 2x, or 4k), none, or highest_available. Unsupported and unavailable choices fail explicitly. | none |
| download | No | When true, download ready outputs immediately. When false, leave them in Flow and return the job. | |
| fileName | No | Optional safe file stem. The job ID and downloaded extension are added automatically. | |
| accountId | No | Omit accountId to use the most recently verified connected account. Supply it only when the user explicitly chose a different connected account returned by flow_list_accounts. | |
| aspectRatio | No | Exact video aspect ratio returned by flow_inspect_account, or ui-default. | ui-default |
| flowProject | No | Existing Flow project name to open. If omitted, the current project is reused or a new project is created. | |
| referenceFiles | No | Absolute paths to optional local images or videos to attach as Flow references/ingredients/frames. | |
| timeoutSeconds | No | Maximum seconds to wait for this generation or upscale. A timeout leaves a persistent job that can be polled. | |
| durationSeconds | No | Requested clip length only when flow_inspect_account reports that exact value in visibleDurations. Omit when the current Flow Agent UI exposes no duration control. | |
| outputDirectory | Yes | Absolute directory where downloads and .flow.json manifests are saved, e.g. C:\project\public\generated\flow. | |
| confirmCreditSpend | Yes | Must be true. Confirms the user explicitly authorized this operation to consume Google Flow/AI credits. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true and openWorldHint=true, so the safety profile is covered. The description adds genuinely new behavior: one persistent job, a brief wait then return, timeout leaving a pollable job, and credit spend requiring authorization. It stops short of describing the returned job/manifest shape, which the missing output schema would otherwise require.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four tight sentences, front-loaded with the exclusivity claim and lifecycle, then routing, then the prohibition. Every sentence carries a distinct instruction with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 14-parameter, credit-consuming, non-idempotent tool with no output schema, the description covers the job lifecycle, the polling escape hatch, the timeout failure mode, and the authorization gate. It could say a bit more about what is returned on the initial call (job ID/manifest), but the essentials are present.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so each of the 14 parameters is already documented in the schema (model IDs, upscale IDs, durationSeconds visibility rules, etc.). The description adds no parameter-level detail beyond that, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('submits one persistent job' for a Google Flow video) and explicitly positions itself as 'THE REQUIRED AND EXCLUSIVE PATH for every Google Flow video request,' which cleanly separates it from flow_generate_image and flow_upscale_video. The submit-wait-return lifecycle is named up front.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit when/when-not guidance: poll flow_job_status with the SAME job ID on status=processing, 'never generate again,' and 'Never open or control Flow with browser/computer-use tools.' It also names the precondition (explicit credit authorization), leaving almost nothing to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
flow_helpWhat Can Google Flow MCP Do?ARead-onlyIdempotent
Call this when the user asks what Google Flow MCP can do, how to use it, or for example requests. Returns a concise multilingual-ready capability guide without opening a browser or spending credits.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false and openWorldHint=false, so safety is covered. The description adds genuinely new context beyond them: no browser is launched and no credits are consumed, which lets an agent treat this as a free, side-effect-free fallback.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single sentence with the trigger front-loaded and the return-value behavior trailing it; every clause earns its place and nothing is padded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-param, annotation-covered helper with no output schema, the definition tells the agent when to call it and roughly what comes back. Only the shape/content of the guide itself is left unspecified, which is a minor gap given the tool's simplicity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so the baseline of 4 applies; there is no parameter surface for the description to clarify or for the schema to omit.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a concrete return ('a concise multilingual-ready capability guide') and implicitly separates itself from all siblings by noting it works 'without opening a browser or spending credits' — i.e., every other flow_* tool does one of those. An agent can identify this as the zero-cost help entry point without inspecting any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit trigger conditions ('when the user asks what Google Flow MCP can do, how to use it, or for example requests'), which is exactly the routing information needed. No exclusion or alternative is named, but there is no sibling help/status tool to route against, so a 4 is appropriate.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
flow_inspect_accountInspect Google Flow Account UIARead-onlyIdempotent
MANDATORY before generation. Verifies the real Flow workspace and returns a language-independent live capability map. A landing page returns workspaceAvailable=false and a stop instruction; never browse or scroll it. Use this tool for Google Flow; never substitute generic browser/computer-use automation on the Flow website.
| Name | Required | Description | Default |
|---|---|---|---|
| accountId | No | Omit accountId to use the most recently verified connected account. Supply it only when the user explicitly chose a different connected account returned by flow_list_accounts. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false and openWorldHint, so the safety profile is covered. The description adds genuinely new behavioral content: the landing-page case returns workspaceAvailable=false with a stop instruction, and browsing/scrolling it is prohibited. It stops short of describing the capability map's contents or latency, keeping it at a strong 4.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, no filler, and the highest-value constraint (MANDATORY before generation) is front-loaded. Each sentence carries a distinct instruction: mandate, return value, and exclusion.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description compensates by naming the return shape (live capability map) and one key field (workspaceAvailable). For a zero-required-parameter read-only inspection tool this is close to complete, though it omits failure modes other than the landing page.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the single optional accountId is fully documented in the schema, including the default-to-most-recent behavior and when to supply it. The description adds nothing about parameters, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action (verify the real Flow workspace) and a specific output (a language-independent live capability map), and identifies the tool as a mandatory pre-generation step. It also distinguishes itself from siblings and from generic automation, so an agent can tell what this tool is for without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit 'MANDATORY before generation' establishes when to call it, and 'never substitute generic browser/computer-use automation on the Flow website' names the excluded alternative. The landing-page branch (stop, never browse or scroll) is a concrete when-not instruction.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
flow_job_statusCheck a Google Flow JobAIdempotent
The exclusive status path for jobs returned by a Flow generation tool. Poll the SAME job ID until completed or failed; never resubmit generation. It safely finalizes any download already authorized by the original request. Never inspect Flow with generic browser/computer-use tools.
| Name | Required | Description | Default |
|---|---|---|---|
| jobId | Yes | UUID returned by flow_generate_video or flow_generate_image. | |
| waitSeconds | No | Seconds to poll before returning; use 0 for an immediate snapshot. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=false, idempotentHint=true, openWorldHint=true and destructiveHint=false; the description adds the key missing context for why this poll is not read-only — it 'safely finalizes any download already authorized by the original request' — plus the idempotent polling expectation. It does not cover failure modes, rate limits, or how waitSeconds affects behavior, so it stops short of a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four short sentences, front-loaded with the routing rule and zero filler. The 'safely finalizes any download' sentence is slightly ambiguous about whether it retrieves content, which slightly muddies otherwise tight phrasing.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With only two parameters and no output schema, the description carries more return-value burden, and it does name the terminal states ('completed or failed'). It stops short of describing the status response shape or what intermediate states exist, but for a polling tool this is close to complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%: jobId and waitSeconds are both fully documented in the schema, including format, defaults, and bounds. The description reinforces the job ID's origin ('jobs returned by a Flow generation tool') but adds no syntax or format detail beyond what the schema already provides, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource (check/poll a job's status) and scopes it as 'the exclusive status path for jobs returned by a Flow generation tool.' It clearly separates this tool from the generation siblings (flow_generate_video/flow_generate_image) and from generic inspection tools, so an agent can distinguish it without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit when-to-use ('Poll the SAME job ID until completed or failed'), explicit when-not ('never resubmit generation'), and an explicit negative alternative ('Never inspect Flow with generic browser/computer-use tools'). This covers the main failure modes an agent would otherwise fall into.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
flow_list_accountsList Google Flow AccountsARead-onlyIdempotent
MANDATORY FIRST STEP for Google Flow work. Returns readyForGeneration and the verified default account. When readyForGeneration=true, use defaultAccountId and DO NOT call either account-connection tool. Use this tool for Google Flow; never substitute generic browser/computer-use automation on the Flow website.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint and destructiveHint=false, so the safety profile is covered. The description adds genuine value beyond that: it discloses the two returned fields and the routing decision that follows from readyForGeneration. It stops short of describing what the caller should do when readyForGeneration=false, which is the one behavioral branch left implicit.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, zero filler, and the imperative 'MANDATORY FIRST STEP' is front-loaded so the agent sees the gating role first. Every sentence carries either a precondition, a routing rule, or a prohibition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter read with no output schema, the description covers the essential contract: when to call it, what it returns, and how to branch on the result. The only gap is the negative branch (readyForGeneration=false), which must be inferred from the mere mention of the account-connection tools.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so the baseline is 4. The description correctly references defaultAccountId as an output value rather than an input, so there is no risk of an agent trying to pass it as an argument.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('List Google Flow Accounts') and immediately names what it returns: readyForGeneration and the verified default account. This distinguishes it from the sibling account-connection and inspection tools without requiring the schema to be opened.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly declares itself the mandatory first step for Google Flow work, specifies the branch condition (when readyForGeneration=true, use defaultAccountId and DO NOT call either account-connection tool), and names an anti-pattern to avoid (generic browser/computer-use automation on the Flow website). This is about as prescriptive as usage guidance gets.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
flow_login_bridge_statusCheck Flow Login BridgeARead-onlyIdempotent
Login diagnostic only. Reports whether the localhost Flow Login Bridge is waiting for the Chromium extension. Do not use this for generation and do not open Flow with browser tools. Normally call flow_list_accounts first.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, non-destructive, closed-world, so the safety profile is covered. The description adds genuinely non-obvious operational context: it is diagnostic-only, it probes a localhost bridge, and the agent must not open Flow with browser tools. It stops short of describing what the reported state looks like, but that is a minor gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four short sentences, front-loaded with the key framing ('Login diagnostic only.'). Every sentence carries a distinct constraint or instruction, with only mild repetition between 'diagnostic only' and 'do not use this for generation'.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter, read-only diagnostic with no output schema, the description covers purpose, constraints, and ordering. It could note what the reported bridge state values mean, but nothing critical for correct invocation is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so the schema is trivially complete and there is nothing for the description to clarify. Baseline 4 applies since no parameter semantics are needed.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific resource and scope: it reports whether the localhost Flow Login Bridge is waiting for the Chromium extension. It also explicitly separates itself from generation siblings, so an agent can distinguish it from flow_generate_image/video.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives when-to-use ('login diagnostic'), when-not ('do not use this for generation', 'do not open Flow with browser tools'), and a sequencing prerequisite ('Normally call flow_list_accounts first'). Alternatives are implied and the preconditions are explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
flow_upscale_videoUpscale an Existing Google Flow Video JobADestructive
The exclusive path for real Flow video upscaling. Uses live available options and rejects unavailable/upgrade-only choices. Never right-click or control Flow with generic browser/computer-use tools. May consume credits and requires explicit authorization.
| Name | Required | Description | Default |
|---|---|---|---|
| jobId | Yes | UUID of an existing video generation job. | |
| factor | Yes | Exact available upscale ID returned by flow_inspect_account, such as 1080p or 2x, or highest_available. Missing/unavailable options fail explicitly. | |
| timeoutSeconds | No | Maximum seconds to wait for this generation or upscale. A timeout leaves a persistent job that can be polled. | |
| confirmCreditSpend | Yes | Must be true. Confirms the user explicitly authorized this operation to consume Google Flow/AI credits. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, idempotentHint=false and openWorldHint=true, so the safety profile is partly covered. The description adds genuinely new context beyond the annotations: it may consume credits, requires explicit authorization, and rejects unavailable/upgrade-only options rather than silently downgrading.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short, front-loaded sentences with no filler; the exclusivity claim, the guardrail, and the cost warning each carry distinct information. Minor slack in adjectives like 'real' that don't add selection value.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with a 100%-covered schema and no output schema, it covers the key agent-facing risks: credit consumption, authorization, exclusivity, and rejection of bad options. It stops short of describing the return/polling flow (though timeout behavior is noted in the schema), which would fully close the loop.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents jobId, factor, timeoutSeconds and confirmCreditSpend in detail. The description adds no parameter-level syntax or format guidance (e.g. where factor values come from is only stated in the schema, not the description), so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Flow video upscaling') and frames itself as 'the exclusive path,' which separates it from flow_generate_video and similar siblings. It does not name any sibling explicitly, so differentiation is achieved via the exclusivity claim rather than routing language.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a clear negative guidance rule ('Never right-click or control Flow with generic browser/computer-use tools') and a precondition (explicit authorization for credit spend). It implies the alternative path (this tool vs. generic automation) but never names a specific sibling tool to use instead in edge cases.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
11 tool updates
v0.1.0- First observed
flow_begin_account_connection - First observed
flow_complete_account_connection - First observed
flow_download_job - First observed
flow_generate_image - First observed
flow_generate_video - First observed
flow_help - First observed
flow_inspect_account - First observed
flow_job_status - First observed
flow_list_accounts - First observed
flow_login_bridge_status - First observed
flow_upscale_video
TDQS
Scored across 11 tools
Generation, job-status, download, and upscale tools target clearly distinct actions on distinct resources. However, three account/workspace-state tools (flow_list_accounts, flow_login_bridge_status, flow_inspect_account) overlap in inspecting readiness, and the two-step connection pair could be confused with them, though the descriptions give strong ordering cues to separate them.
Every tool uses a consistent flow_ prefix with snake_case verb/noun structure (flow_generate_image, flow_list_accounts, flow_download_job, etc.). The pattern is predictable and uniform across all 11 tools.
11 tools is a well-scoped count for a media-generation server. Each tool covers a distinct part of the connect-generate-poll-download lifecycle, and nothing appears redundant or padded.
The core lifecycle is covered: account connection (begin/complete), readiness inspection, image/video generation, upscaling, polling, and downloading. Minor gaps remain, such as no job-history/list operation or job cancellation, which agents may have to work around by retaining job IDs themselves.
Maintenance
Related MCP Connectors
LLM chat, text tools, image generation, editing, batch image jobs, and asynchronous video generation
- FlowNodeOAuthio.flownode
Generate images, video, audio and 3D with FlowNode; results land in your asset library.
Design, save, and run outcome-aligned AI workflows and verifiers, with reliable image output.
- RendobarOAuthcom.rendobar
Transform video, audio and images, and generate media from prompts. FFmpeg, captions, models.
Related MCP Servers
- AlicenseAqualityDmaintenanceEnables AI agents to drive Google Flow through a real Chrome profile to generate images, videos, characters, and scenes without sharing credentials.1912 npmMIT
- AlicenseAqualityAmaintenanceEnables AI agents to programmatically generate images and videos through the authenticated Google Flow web interface via a direct Chrome DevTools Protocol connection, exposing tools for media generation, project management, status checks, and asset downloads without requiring official API keys.271MIT
- AlicenseAqualityAmaintenanceEnables AI agents to generate Google Flow videos and images and automate Scene Builder clip extensions through a user's own Chrome session.1165 npm1MIT
- AlicenseNot gradedqualityCmaintenanceEnables generating images and videos in Google Flow through an automated browser session using your Google AI Pro subscription, without API keys or third-party credit fees.116 npm2MIT