Google Flow MCP
Generates images and video on Google Flow (labs.google/flow) via browser automation using an existing signed-in Google account and AI Pro subscription. Provides tools to connect to and check the Flow session/account, report status and credit balance, generate images (Nano Banana / Imagen) and video (Veo 3.1 / Omni Flash) with model, aspect ratio, duration and reference/first/last-frame/ingredient options, manage characters and scenes, download generated media, and submit/track generation jobs through a serialized job queue.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Google Flow MCPgenerate an image of a sleepy cat on a windowsill at golden hour"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
google-flow-mcp
MCP server that generates images and video on Google Flow through browser automation, so you can use your own Google AI Pro subscription instead of paying per-credit services. Ships with a Claude Code skill.
Validated end-to-end: images (Nano Banana / Imagen) and video (Veo 3.1 / Omni Flash) are generated in a real Flow project and downloaded to disk.
Adapted and hardened for the current agent-first Flow UI (and Windows) from TMSSS05/google-flow-browser-mcp.
What it does
Playwright connects over the Chrome DevTools Protocol to a dedicated Chrome that is logged into your Google account. It drives Flow's agent to generate media and downloads the result through the authenticated session. No API keys, no password handling — it uses your existing browser session.
Tools (17): flow_connect, flow_status, flow_account_check, flow_discover_ui,
flow_generate_image, flow_generate_video, flow_download_latest, character/scene
tools, flow_use_grid_architect, flow_screenshot, flow_queue_status, …
Related MCP server: google-flow-mcp
⚠️ Terms of Service
This is unofficial browser automation. There is no official Google API for Flow.
The launcher starts Chrome directly so navigator.webdriver is false, which is an
explicit anti-bot measure. Automating Google properties can violate Google's Terms of
Service and may put your account at risk. Use at your own risk, on your own account.
Requirements
Node.js ≥ 22 (with npm)
Google Chrome or Microsoft Edge (Chrome 149+ needs Playwright ≥ 1.61.1, already pinned)
A Google account with access to Flow (Google AI Pro recommended)
Install as a Claude Code plugin
This repository is both a Claude Code plugin and a plugin marketplace. The plugin bundles the
MCP server and the google-flow-generate skill.
claude plugin marketplace add minhezluv/Google-flow-mcp-generate
claude plugin install google-flow@google-flow-mcpOr inside a session: /plugin marketplace add minhezluv/Google-flow-mcp-generate, then
/plugin install google-flow@google-flow-mcp.
When the plugin is enabled, Claude Code asks for two optional settings:
Setting | Default | Purpose |
Google account | empty | Gmail address that must be signed in to Flow; empty skips the check |
Browser debugging port | 9222 | CDP port of the dedicated browser |
On the first tool call the server copies itself to the plugin data directory
(~/.claude/plugins/data/<plugin-id>/), runs npm ci there (about a minute) and creates
home/config/flow.config.json, where the remaining settings live. Outputs, logs and the job
store are kept under that home/ directory, so they survive plugin updates.
Then ask Claude for an image or a video. The skill starts the dedicated browser with
scripts/ensure-flow-chrome.mjs (Windows, macOS, Linux). On first use, sign in to Google in
that window and click "Sign in to Flow"; the profile in ~/.google-flow-mcp/browser-profile
keeps the session.
To try a local checkout without the marketplace: claude --plugin-dir .
Use with any MCP client (npm)
The server is published on npm as @minhezluv/google-flow-mcp and listed in the
MCP Registry as io.github.minhezluv/google-flow.
Add it to your client's MCP configuration (Claude Desktop, Cursor, VS Code, …):
{
"mcpServers": {
"google-flow": {
"command": "npx",
"args": ["-y", "@minhezluv/google-flow-mcp"],
"env": { "FLOW_EXPECTED_ACCOUNT": "you@gmail.com" }
}
}
}Variable | Default | Purpose |
| unset | Gmail address that must be signed in to Flow |
| 9222 | CDP port of the dedicated browser |
|
| Config, outputs, logs and job store |
Before generating, start the dedicated browser and sign in to Google and Flow in it once:
npx -y -p @minhezluv/google-flow-mcp google-flow-chromeManual setup (from a checkout)
npm install
cp config/flow.config.example.json config/flow.config.json
# edit config/flow.config.json → set expectedAccount and chromeUserDataDirStart the dedicated Chrome (idempotent — launches only if needed):
node scripts/ensure-flow-chrome.mjs # any OS; --port, --profile, --browser optional
powershell -File scripts/ensure-flow-chrome.ps1 # Windows-only originalFirst run: in that Chrome window, sign in to your Google account and click "Sign in to Flow" on labs.google (Flow uses a separate sign-in). The session is saved in the dedicated profile and reused.
Register the server with your MCP client (Claude Code, etc.):
{
"mcpServers": {
"google-flow": { "type": "stdio", "command": "node", "args": ["<path>/src/index.js"] }
}
}Restart the client afterwards (the server loads into memory at startup).
Daemon
Only one process may drive the Flow Chrome. src/daemon/main.js owns it, keeps a serial job
queue in data/jobs.json and listens on 127.0.0.1:47821 (daemonPort). The MCP server starts
the daemon on first use and forwards every tool call to it; other programs (for example the
Hypit provider) submit generation jobs over HTTP.
npm run daemon
curl http://127.0.0.1:47821/healthRequests other than /health need authorization: Bearer <config/daemon-token>; the token is
created on first start.
Route | Purpose |
| Chrome, sign-in and queue state |
| Current credit balance; authenticated and serialized under the browser lock |
| Reference image bytes (PNG/JPEG/WebP) → |
|
|
|
|
| Generated file |
| Run one MCP tool under the browser lock |
Every generation job needs confirmCredits: true. A repeated idempotencyKey returns the
existing queued, running or succeeded job instead of spending credits again.
Models: nano-banana-2, nano-banana-pro, nano-banana-2-lite, veo-3.1-lite, veo-3.1-fast,
veo-3.1-quality, omni-flash (Omni 1.1 Flash). Images accept ratios 16:9, 4:3, 1:1, 3:4, 9:16;
videos 16:9 and 9:16. Reference images (image jobs), first/last frames and ingredients (video jobs)
are uploaded through Flow's ingredient picker.
The driver targets https://flow.google.com/, finds controls by their Material icon names, sets
model/ratio/count in Flow's settings panel before each job and switches Flow's "confirm before
generating" option to "never" — the daemon's confirmCredits flag is the spending gate. It refuses
to run when the signed-in account differs from expectedAccount. Flow adds a visible AI watermark
in some regions.
The prompt names the exact model and forbids substitution. After rendering, the driver reloads
the finished project's persisted history and checks its model_display_name against the media
id before downloading. It reads the UI's observed GN0Bre batchexecute response; unknown or
ambiguous metadata fails with UI_CHANGED, and a substituted model fails with UNSUPPORTED_INPUT.
Neither failure submits another generation. Verified signed media URLs stay in a bounded in-memory
cache so download retries do not rely on opaque thumbnails after reload.
The check runs after Flow has already charged for the render. If Flow changes that undocumented
response and every job starts failing with UI_CHANGED, set "verifyModel": false in
config/flow.config.json and restart the daemon: results are then downloaded without the model
check, so a silent model substitution would go unnoticed until the check is repaired.
Live measurements on 2026-10-01: Nano Banana 2/Pro/2 Lite images cost 0; Veo Lite/Fast/Quality 8-second text videos cost 10/20/100; Omni Flash text at 6 seconds costs 10, and its 8-second first+last-frame or single-ingredient modes cost 12. Flow silently changed requested Lite frame-pair/ingredient jobs to Omni in earlier tests; request Omni explicitly for these modes. See capabilities for the measured scope.
After updating, restart your MCP client so it loads the proxy version of the server.
Notes that matter
Images are effectively free against the monthly Flow credit pool; video consumes credits (Veo 3.1 Lite ~10, Fast ~20, Quality ~100; Omni Flash 10 for 6s text / 12 for 8s references, out of ~1000/month). Video shows a credit-confirmation dialog which the server approves.
Model/duration must be a valid combo or Flow's agent asks for clarification and nothing generates (e.g. Veo 3.1 Lite is 8s-only on the Pro plan; measured Omni modes are 6s text and 8s references).
Flow is agent-first: prompts are wrapped imperatively so the agent generates directly instead of asking questions.
The UI language follows your Google account; navigation selectors cover IT/FR/EN.
Claude Code skill
The plugin installs the skill automatically. Without the plugin, copy
skills/google-flow-generate/ to ~/.claude/skills/ and replace ${CLAUDE_PLUGIN_ROOT} and
${user_config.cdp_port} in its SKILL.md with your checkout path and port.
Releasing
Keep these versions equal: package.json, server.json (version and packages[0].version)
and .claude-plugin/plugin.json. Then:
npm test
claude plugin validate --strict .
npm publish --access public # npm package used by npx and the MCP Registry
mcp-publisher login github # once; the account must be minhezluv
mcp-publisher publish # reads server.json
git push # plugin marketplace users get the updateLicense
MIT — see LICENSE.
videoCooldownMs (default 30000) spaces video jobs: each video waits that long after the previous
one finished. Three Veo jobs submitted back to back were each rejected within ~27 s with Flow's
generic "something went wrong" tile, and the failures still consumed credits. Images do not wait.
Available Tools
17 toolsflow_account_checkA
Verify the logged-in Google account matches the configured expected email (Profile 3).
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description must fully disclose behavioral traits. It states the verification intention but does not describe what the tool actually returns (e.g., boolean, message), whether it is read-only, or what happens when the accounts do not match. This lack of behavioral detail is a significant gap given the absence of annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence, perfectly concise and front-loaded with the main action ('Verify'). It contains no filler or redundant information. Every word contributes to the tool's purpose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with no parameters, no output schema, and no annotations, the description provides a basic but adequate purpose. However, it lacks critical operational context such as what the tool returns, whether it is a safe read-only operation, and any failure indications. Given the simplicity of the tool, this makes it minimally viable but not fully complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has zero parameters and 100% schema description coverage, so there are no parameter semantics to explain. The description adds context about the expected email (Profile 3) which clarifies the tool's purpose. The baseline for zero parameters is 4, and the description does not require further elaboration.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool verifies that the logged-in Google account matches a configured expected email. The verb 'verify' is specific, and the resource is precisely defined (logged-in Google account vs expected email). This distinguishes it from sibling tools like flow_connect or flow_status, which pertain to connections or general status rather than account verification.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit guidance on when to use this tool versus alternatives. The description does not mention any prerequisite conditions, exclusions, or cases where another tool would be more appropriate. An agent is left to infer that this is a pre-flight check, but that is not stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
flow_connectA
Launch Chrome with the configured Google profile, connect CDP, navigate to Google Flow, and verify account.
| Name | Required | Description | Default |
|---|---|---|---|
| headless | No | Launch in headless mode (not recommended, Google Flow needs visible browser). | |
| open_flow | No | Auto-navigate to Google Flow after connection. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden. It discloses the main side effect (launching Chrome) and the steps performed, but lacks detail on whether the connection is persistent, what happens on verification failure, or whether the browser closes after the operation. This is moderate transparency but not complete.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence that packs four specific actions without fluff. It is front-loaded with the primary action ('Launch Chrome') and efficiently communicates the full flow.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description does not mention return values (no output schema exists) or the fact that this tool likely establishes a required connection for other flow_* tools. Without annotations, it could benefit from noting that it is a prerequisite and what 'verify account' entails. This is adequate but not fully complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%: both 'headless' and 'open_flow' have clear descriptions in the schema, including a warning about headless mode. The tool description does not add additional parameter semantics, but since the schema already covers them, a baseline score of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's specific actions: launching Chrome, connecting CDP, navigating to Google Flow, and verifying the account. It distinguishes itself from sibling tools like flow_disconnect or flow_status by focusing on the connection setup.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description does not provide explicit guidance on when to use this tool relative to alternatives. It implies it is a connection step but does not mention prerequisites, exclusions, or alternative tools. For example, it does not clarify whether this should be run before flow_status or flow_discover_ui.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
flow_create_characterC
Create a new character in Google Flow Characters with name and description.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | Character name. | |
| campaign | No | Campaign identifier for project matching (e.g., "ete-2026", "nouvelle-collection"). | |
| description | Yes | Character description/prompt. | |
| project_name | No | Name for the project (will reuse existing project with same campaign, or create new). | |
| reference_images | No | Paths to reference images for character design. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries full burden for behavioral disclosure, but it only repeats the basic action. It does not mention project matching, handling of optional parameters, side effects, authentication requirements, or return behavior, leaving the agent uninformed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, clear sentence that is front-loaded and free of waste. It conveys the core purpose without padding, despite being minimal.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has 5 parameters, no output schema, and no annotations, but the description covers only the two required parameters. It omits behavior around project matching, reference images, return values, and edge cases, making it severely incomplete for reliable invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all parameters are documented in the schema. The description mentions only 'name and description,' which is redundant and adds no meaning beyond the schema, meriting the baseline score of 3.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Create') and resource ('character in Google Flow Characters'), clearly indicating its function. It distinguishes from flow_import_character by emphasizing 'new,' though it does not explicitly contrast with alternatives.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus alternatives like flow_import_character or flow_create_scene. The description only states what it does, without context for selection.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
flow_create_sceneB
Create a new scene in Google Flow Scenes with characters and prompt.
| Name | Required | Description | Default |
|---|---|---|---|
| prompt | Yes | Scene description/prompt. | |
| campaign | No | Campaign identifier for project matching (e.g., "ete-2026", "nouvelle-collection"). | |
| characters | No | Character names to include in the scene. | |
| project_name | No | Name for the project (will reuse existing project with same campaign, or create new). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It only states the action without covering side effects, dependencies, project reuse logic, or what happens after creation. This is a minimal disclosure for a tool that likely creates persistent state.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single sentence that is front-loaded with the verb and resource. Every word earns its place, and there is no redundant fluff or repetition of schema details. This is an effective level of conciseness.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with 4 parameters and no output schema, the description provides no information about return values, error conditions, or how the creation process behaves (e.g., whether it creates a new project or reuses an existing one via campaign). The description is too sparse to be considered complete for an agent invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description mentions 'characters and prompt,' which is partially redundant with the schema, and adds no extra meaning beyond what the property descriptions already provide. The campaign and project_name parameters are not discussed in the description, but the schema covers them adequately.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Create') and clearly identifies the resource ('a new scene in Google Flow Scenes'), with mention of characters and prompt. This distinguishes it from sibling tools like flow_create_character, which targets a different resource.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No context is provided about when to use this tool versus alternatives. The description only states what it does, without any mention of prerequisites, typical use cases, or exclusions relative to sibling tools like flow_create_character or flow_use_grid_architect.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
flow_disconnectA
Close the browser and clean up the MCP connection to Google Flow.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description declares the primary actions (closing browser, cleaning up connection) but does not disclose side effects such as terminating active sessions or preventing further operations. This is adequate for a simple cleanup tool but lacks deeper behavioral context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, well-structured sentence that directly states the tool's function without unnecessary words. It is perfectly concise and front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a parameterless tool with no output schema or annotations, the description provides the essential information about its role as a disconnect/cleanup operation. It lacks explicit usage context but is otherwise complete for its simplicity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters and an empty schema, so the description adds nothing about parameters. Since there are no parameters to explain, the baseline of 4 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses specific verbs 'Close' and 'clean up' to describe the resource (browser and MCP connection), clearly indicating this is the teardown counterpart to flow_connect. It distinguishes itself from sibling tools by its focus on disconnection and cleanup.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description does not explicitly state when to use the tool or contrast it with alternatives like flow_connect. Usage must be inferred from the name and context, so the agent receives no direct guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
flow_discover_uiA
Navigate to a Google Flow page and discover all interactive elements (buttons, inputs, links, headings). Updates the internal selectors map for robust automation.
| Name | Required | Description | Default |
|---|---|---|---|
| page | Yes | Page to discover. Options: main, image-generation, video-generation, characters, scenes, tools-gallery, grid-architect. | main |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the burden. It discloses navigation, discovery, and the internal side effect of updating the selectors map. However, it does not mention authentication, error behavior, potential page modifications, or what happens on failure, leaving some behavioral gaps.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two concise sentences with no redundant content. The action and purpose are front-loaded, and the secondary effect is stated efficiently.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter tool with no output schema, the description explains the core action and side effect, but does not clarify what the agent receives upon success or failure, nor any error handling. It is adequate but leaves room for more detail.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema already fully documents the sole 'page' parameter with a description and default value. The tool description adds limited extra meaning beyond implying navigation to that page, so baseline 3 applies due to high schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Clearly defines the action (navigate and discover) and the resource (Google Flow page), and distinctly differentiates itself from siblings by mentioning the internal selectors map update. Sibling tools like flow_generate_image or flow_download_latest have clearly different purposes.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Implies usage for preparing robust automation, but does not explicitly state when to use this tool over alternatives or mention any exclusions. The phrase 'for robust automation' gives context but lacks direct guidance on when to invoke it versus other flow tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
flow_download_latestB
Download the most recently generated file from Google Flow.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden of behavioral disclosure. It only states select 'most recently generated' but does not explain whether the file is returned as content, a URL, or saved locally, nor any side effects or permissions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, clear sentence with no filler. It is front-loaded with the main action and resource, making it highly concise and easy to process.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with no output schema and no annotations, the description is incomplete. It does not specify what the agent should expect as a result (file path, binary content, URL) or how this tool fits into the flow with sibling generation tools (e.g., used after flow_generate_image).
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has zero parameters, so the description need not elaborate on parameter meaning. The description implicitly communicates that no user input is required, aligning with the baseline for a parameterless tool.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb 'Download' and a clear resource 'the most recently generated file from Google Flow.' It distinguishes itself from sibling tools (e.g., flow_generate_image, flow_connect) by being the only download-focused tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus alternatives, prerequisites, or typical usage contexts. The description simply states the action without explaining when it is appropriate or what conditions must exist (e.g., a file must have been generated first).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
flow_generate_imageA
⚠️ CES IMAGES CONSOMMENT DES CRÉDITS. Par défaut (auto_confirm=false): remplit le prompt, sélectionne le modèle/ratio, prend un screenshot et retourne "ready_for_confirmation". NE clique PAS sur Generate. Quand auto_confirm=true: vérifie d'abord que l'interface est bien en mode IMAGE (pas Vidéo), que le modèle est un modèle image, prend un screenshot de vérification, PUIS clique Generate, attend les images et les télécharge. NAN/BANANA modèles image seulement.
| Name | Required | Description | Default |
|---|---|---|---|
| brand | No | Brand context for automatic model selection: premium, standard. | |
| model | No | Model: Nano Banana 2 (default), Nano Banana Pro, Nano Banana 2 Lite. | Nano Banana 2 |
| ratio | No | Aspect ratio: 1:1, 16:9, 9:16, 4:3, 3:4. | 1:1 |
| prompt | Yes | The text prompt for image generation. | |
| campaign | No | Campaign identifier for project matching (e.g., "ete-2026", "nouvelle-collection"). | |
| auto_confirm | No | ⚠️ CRÉDITS. Si false (défaut): prépare seulement, ne consomme rien. Si true: vérifie que le mode Image est actif, PUIS clique Generate (consomme des crédits). | |
| project_name | No | Name for the project (will reuse existing project with same campaign, or create new). | |
| reference_images | No | Paths to reference images (optional). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and delivers: it discloses credit consumption, the exact default sequence (fill prompt, pick model/ratio, screenshot, return 'ready_for_confirmation', no Generate click), the confirmation sequence, and that images are then downloaded. This is rich, non-obvious behavioral detail an agent cannot infer from the schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The credit warning is front-loaded, followed by the two operating modes in sequence, which matches the decision an agent must make. It is dense but nearly every clause earns its place; minor redundancy exists between the description and the auto_confirm schema text.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex 8-parameter mutation tool with no annotations and no output schema, the description covers credit cost, mode gating, model restriction, and the return signal ('ready_for_confirmation'). It could say more about what the download yields or how reference_images influence output, but the essentials for correct invocation are present.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds value by clarifying the model constraint ('NAN/BANANA modèles image seulement') and the credit cost tied specifically to auto_confirm. It reinforces rather than merely repeats schema semantics, though most parameter detail still lives in the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb+resource (generate image) and explicitly scopes it to image models, warning that the UI must be in IMAGE mode 'pas Vidéo'. This distinguishes it from the sibling flow_generate_video without the agent needing to open either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It clearly specifies when each mode applies: auto_confirm=false prepares and returns 'ready_for_confirmation' without consuming credits, while auto_confirm=true verifies mode/model then clicks Generate and consumes credits. The when-to-use guidance is strong, though it never explicitly names a sibling tool to use for the alternative (video) case.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
flow_generate_videoA
⚠️ CONSUMA CREDITI FLOW (Veo/Omni). Con auto_confirm=false (default): prepara il prompt video e si ferma senza generare (nessun credito). Con auto_confirm=true: invia, attende il render (minuti) e scarica l'mp4. Modello economico per test: "lite" (Veo 3.1 Lite).
| Name | Required | Description | Default |
|---|---|---|---|
| model | No | Model: lite, fast (default), quality, flash (Omni 1.1 Flash), or the exact Flow name. | fast |
| ratio | No | Aspect ratio: 16:9 or 9:16. | 16:9 |
| prompt | Yes | The text prompt for video generation. | |
| campaign | No | Campaign identifier for project matching. | |
| duration | No | Duration like "4s", "6s", "8s". | 4s |
| auto_confirm | No | ⚠️ CREDITI. false (default): prepara soltanto. true: genera davvero (consuma crediti Flow) e scarica. | |
| project_name | No | Name for the project (will reuse existing project with same campaign, or create new). | |
| reference_images | No | Paths to images the video must include as ingredients (optional). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full load and does well: it warns that the tool consumes Flow credits, discloses that auto_confirm=true waits minutes for render, and that the result is a downloaded mp4. It omits failure behavior, rate limits, and any account/permission requirements, which is why it is not a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the credit warning with a ⚠️ marker, then covers the two modes and the model hint in three tight sentences. There is minor redundancy with the auto_confirm parameter description already in the schema, but no wasted prose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a costly, non-read-only tool with no annotations and no output schema, the description adequately covers cost, mode behavior, timing, and the returned artifact (mp4). It stops short of covering failure modes and account/prerequisite setup, leaving a small gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% with detailed per-parameter descriptions, so the baseline is 3. The description adds genuine value beyond the schema by explaining the auto_confirm workflow end-to-end and flagging 'lite' (Veo 3.1 Lite) as the economical model for testing, which the schema alone does not convey.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource (generate video) and immediately distinguishes the two operating modes (auto_confirm false = prepare only, true = actually render and download). The video focus cleanly separates it from the sibling flow_generate_image.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly tells the agent when to use each mode: false for a no-cost dry run, true to actually spend credits and download, plus a recommendation to use the cheap 'lite' model for testing. It does not, however, explain prerequisites (e.g., whether flow_connect is required first) or when to prefer flow_generate_image instead.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
flow_import_characterB
Import a character from a saved JSON file into Google Flow.
| Name | Required | Description | Default |
|---|---|---|---|
| file_path | Yes | Path to character JSON file. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description must disclose behavioral traits, but it only states the operation. It does not mention side effects (e.g., overwriting an existing character), validation behavior, error outcomes, or any permissions required, leaving significant behavioral gaps.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single concise sentence that clearly states the action and its target. There is no redundancy or wasted words, making it appropriately front-loaded and efficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is simple with one parameter and no output schema, but the description does not explain what happens after import (e.g., whether a character is created or replaced, or how to confirm success). There are clear gaps in the expected outcome, so the description is merely adequate.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema covers 100% of the parameter (file_path) with a description 'Path to character JSON file.' The tool description reiterates 'saved JSON file' but adds no additional meaning beyond the schema, so the baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb 'Import' with a clear resource ('character'), source ('saved JSON file'), and destination ('Google Flow'). This clearly distinguishes it from sibling tools like flow_create_character, which creates a new character from scratch.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives such as flow_create_character. It does not mention prerequisites, intended scenarios, or exclusions, leaving the agent without explicit usage context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
flow_open_charactersA
Open the Google Flow Characters page and list existing characters.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It states the tool opens a page and lists characters, implying a read-only action, but it does not explicitly confirm it is read-only, mention any limitations, or describe what exactly 'list' returns (e.g., a formatted list, a visual UI, or a status message).
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence that is precise and immediately states the primary actions and target resource. Every word earns its place, with no redundant or filler content, making it both concise and front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple tool with no parameters and no output schema, the description is moderately complete. It conveys the main purpose ('open page' and 'list characters') but leaves ambiguity about the exact output format or how the agent receives the character list. Given the absence of an output schema, a slightly more explicit description would improve completeness.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, and the schema confirms this. The description adds no parameter-specific information, but with no parameters to describe, the baseline of 4 applies. The description's mention of 'listing existing characters' gives context to what the tool does without needing parameter details.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool opens the Google Flow Characters page and lists existing characters. It uses specific verbs ('open', 'list') and a resource ('Google Flow Characters page'), effectively distinguishing it from sibling tools like flow_create_character and flow_import_character.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the tool is used for viewing existing characters, but it does not explicitly state when to use it over alternatives or mention any prerequisites. There is no direct comparison with sibling tools such as flow_create_character or flow_import_character, so usage guidance is only implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
flow_open_tools_galleryA
Open the Google Flow Tools Gallery and list available tools.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It only states what the tool does (opens a gallery and lists tools) but does not disclose any side effects, permissions required, or whether it returns data in the response. For a simple tool this might be acceptable, but the lack of explicit behavioral context is a gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, compact sentence with no wasted words. It front-loads the action and resource clearly, making it immediately understandable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's simplicity (no parameters, no output schema, minimal annotations), the description covers the essential purpose. However, it leaves slight ambiguity about whether 'list available tools' means returning a list in the response or opening a UI gallery. This prevents a perfect score.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool accepts zero parameters, so the schema coverage is trivially 100%. The baseline for 0 parameters is 4, and the description adds no confusion about parameters. Since there are no parameters to clarify, the description does not need to provide additional semantics.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool opens the Google Flow Tools Gallery and lists available tools. This is a specific verb+resource combination that distinguishes it from sibling tools like flow_download_latest or flow_create_scene.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for discovering available tools, but does not explicitly state when to use this tool versus alternatives or mention any prerequisites. There are no exclusions or comparisons with siblings, so it relies on the name and context to convey intent.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
flow_queue_statusA
Check the job queue: active job, pending queue, completed and failed job history.
| Name | Required | Description | Default |
|---|---|---|---|
| history_limit | No | Number of recent history entries to return. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Since no annotations are provided, the description carries the full burden. It clearly indicates a read-only operation ('Check') and lists the information returned, but it does not disclose any potential side effects, authentication requirements, or details about how history_limit affects results. It is minimally transparent but not deeply descriptive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single concise sentence that front-loads the primary action and enumerates the specific aspects of the queue. There is no wasted information; every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple tool with one optional parameter and no output schema, the description adequately captures the main functionality. It names all categories of queue status without needing to explain return values extensively. The only minor gap is not explicitly stating that history_limit controls the number of history entries, but that is already in the schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema provides 100% coverage for the only parameter (history_limit) with a clear description. The tool description does not add additional meaning beyond the schema, but since schema coverage is high, the baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Check') and resource ('job queue'), and enumerates the exact information covered (active job, pending queue, completed and failed job history). This clearly distinguishes it from sibling tools like flow_status, which likely covers broader status, and other flow_* tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit guidance on when to use this tool versus alternatives such as flow_status. The description implies that it is for checking queue information, but it does not state exclusions or mention alternative tools, leaving the agent without clear decision criteria.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
flow_screenshotB
Take a screenshot of the current Google Flow page.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | Custom name for the screenshot file. | manual |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of disclosing behavioral details. It only states the action without mentioning where the screenshot is saved, whether it returns a file path, authentication requirements, or any side effects, leaving significant ambiguity for an agent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence with no filler, immediately conveying the tool's action. It is appropriately concise and front-loaded, making it easy for an agent to parse.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's simplicity (one optional param, no output schema, no annotations), the description is minimally adequate but incomplete. It fails to explain what happens after the screenshot is taken (e.g., file location, return value), leaving the agent without enough context to predict the tool's full behavior.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The single optional parameter 'name' is fully described in the input schema (100% coverage), so the baseline is 3. The description adds no additional semantic detail about the parameter, but the schema already provides sufficient information.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states a specific action ('Take a screenshot') on a specific resource ('the current Google Flow page'), making its purpose unambiguous. This distinguishes it from sibling tools like flow_download_latest or flow_create_scene, which serve different actions.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description does not provide explicit guidance on when to use this tool versus alternatives. However, the intended use is implied by the action itself — taking a screenshot when a visual capture is needed — but no exclusions or alternative comparisons are mentioned.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
flow_statusB
Check current connection status: browser connected, Flow page loaded, account verified, job queue state.
| Name | Required | Description | Default |
|---|---|---|---|
| full | No | Return full status with screenshot. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It transparently lists the exact status dimensions checked (browser, page, account, queue), which is useful. However, it doesn't disclose whether this is a read-only operation, how the response is structured, or whether it performs live checks or returns cached state.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence of 12 words, front-loaded with the primary action and resource. Every word adds value, with no redundancy or filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simplistic status-check tool with one fully described parameter, the description covers core purpose and scope. However, without an output schema or annotations, it omits details about return format, potential latency, and any safety characteristics, leaving the agent uncertain about how to interpret the result.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema description fully documents the 'full' parameter ('Return full status with screenshot'), achieving 100% coverage. The tool description adds no parameter-specific context, so the baseline score of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses the specific verb 'Check' and clearly enumerates the status components (browser connected, Flow page loaded, account verified, job queue state). It implies a broader scope than individual siblings like flow_account_check or flow_queue_status but does not explicitly name them for differentiation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus alternatives. It doesn't mention that flow_queue_status should be used for queue-only details or flow_connect for establishing a connection. The agent is left without explicit selection criteria.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
flow_use_grid_architectA
Open Grid Architect in Google Flow, fill theme prompt, shot prompts, engine, ratio, and visual logic settings. Supports batch shot generation for brand campaigns.
| Name | Required | Description | Default |
|---|---|---|---|
| ratio | No | Aspect ratio for all shots. | 16:9 |
| engine | No | Engine/model for the grid. | Nano Banana 2 |
| campaign | No | Campaign identifier for project matching (e.g., "ete-2026", "nouvelle-collection"). | |
| references | No | Paths to reference images. | |
| project_name | No | Name for the project (will reuse existing project with same campaign, or create new). | |
| shot_prompts | No | Array of individual shot prompts for the grid. | |
| theme_prompt | Yes | Overall theme prompt for the grid. | |
| visual_logic | No | Visual logic type: None, Colour Pop, Side by Side, etc. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It discloses that the tool opens a UI screen and fills settings, and mentions batch generation behavior, but it does not disclose prerequisites (e.g., needing an active Google Flow session), side effects like project creation/reuse, or what the output will be. This is more transparent than a minimal description but still lacks depth.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences, front-loaded with the main action and settings, with no fluff. Every sentence conveys essential information, and it is appropriately sized for the tool's complexity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has 8 parameters, no output schema, and no annotations, the description provides an adequate overview but omits details on expected return values, prerequisites, and project handling behavior. The schema covers parameter details, but the description could be more complete about invocation context and outcomes.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3. The description lists several parameter names and adds context for batch generation and campaigns, but it does not add substantial meaning beyond the schema's parameter descriptions, which are already detailed.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool opens Grid Architect in Google Flow and fills specific settings (theme prompt, shot prompts, engine, ratio, visual logic), with support for batch shot generation. It distinguishes itself from sibling tools like flow_generate_image by focusing on batch grid generation for brand campaigns.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'Supports batch shot generation for brand campaigns' provides clear context for when to use this tool (batch/brand campaigns). It does not explicitly state alternatives or when-not to use, but the batch context differentiates it from single-generation siblings like flow_generate_image.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
flow_use_toolA
Open any tool by name in Google Flow and optionally fill its configuration parameters.
| Name | Required | Description | Default |
|---|---|---|---|
| params | No | Optional configuration parameters for the tool. | |
| campaign | No | Campaign identifier for project matching (e.g., "ete-2026", "nouvelle-collection"). | |
| tool_name | Yes | Name of the tool to open (e.g. Grid Architect, Image Generation). | |
| project_name | No | Name for the project (will reuse existing project with same campaign, or create new). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It discloses that the tool opens a tool and optionally fills parameters, but it does not mention potential side effects like creating projects (suggested by project_name/campaign params) or whether it navigates the UI. This is minimal behavioral disclosure for a tool with state-changing potential.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, clear sentence with no wasted words. It earns its place by succinctly conveying the tool's core purpose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has 4 parameters and no output schema. While the description is concise, it doesn't explain the behavior of opening a tool in terms of project management (campaign/project_name) or what happens after opening. Considering the presence of sibling tools and the generic nature of this tool, more context would improve completeness.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, and the description mentions 'optionally fill its configuration parameters' which aligns with the params field. However, it adds no additional meaning beyond what the schema already provides, so the baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb ('Open') and resource ('any tool by name in Google Flow'), which distinguishes it from sibling tools that target specific tools like flow_use_grid_architect. It effectively communicates the generic, tool-agnostic scope.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies a generic use case ('any tool by name') but does not explicitly state when to choose this tool over specific sibling tools. No exclusions or alternative recommendations are provided, leaving usage guidance implied rather than explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
17 tool updates
v1.0.0- First observed
flow_account_check - First observed
flow_connect - First observed
flow_create_character - First observed
flow_create_scene - First observed
flow_disconnect - First observed
flow_discover_ui - First observed
flow_download_latest - First observed
flow_generate_image - First observed
flow_generate_video - First observed
flow_import_character - First observed
flow_open_characters - First observed
flow_open_tools_gallery - First observed
flow_queue_status - First observed
flow_screenshot - First observed
flow_status - First observed
flow_use_grid_architect - First observed
flow_use_tool
TDQS
Scored across 17 tools
Most tools target distinct operations, but there is some overlap between flow_status and flow_queue_status, and flow_use_tool potentially overlaps with specific tools like flow_use_grid_architect or flow_open_tools_gallery. However, the descriptions generally clarify the intended boundaries.
All tool names follow a consistent flow_ prefix with snake_case verb/noun structure (e.g., flow_connect, flow_generate_image, flow_create_character). There are no mixed conventions or confusing naming patterns.
At 17 tools, the set is slightly heavy but still reasonable for a browser automation server covering connection, UI discovery, generation, characters, scenes, and queue monitoring. Most tools appear to earn their place, though a few could potentially be consolidated.
Core workflows like connection, generation, and character/scene creation are covered, but there are notable gaps such as no update/delete operations for characters or scenes, no explicit scene listing, and no job cancellation or cleanup tool beyond disconnection. These omissions could limit full lifecycle management.
Maintenance
Related MCP Connectors
- FlowNodeOAuthio.flownode
Generate images, video, audio and 3D with FlowNode; results land in your asset library.
Image and video AI tools and your own pipelines, run from any AI assistant.
Generate AI images, video, voiceovers and music from Claude, ChatGPT or Cursor through 50+ models (Veo 3.1, Kling 3, Seedance, Nano Banana, GPT Image, ElevenLabs). Also image editing, upscaling, background removal, face swap, transcription, voice cloning and UGC-style video ads. Sign in with OAuth — no API key to paste. Tools are annotated (read-only vs. credit-spending); failed generations are refunded.
Build and run visual creative-production workflows from your AI agent.
Related MCP Servers
- AlicenseAqualityAmaintenanceEnables AI agents to programmatically generate images and videos through the authenticated Google Flow web interface via a direct Chrome DevTools Protocol connection, exposing tools for media generation, project management, status checks, and asset downloads without requiring official API keys.71MIT
- AlicenseNot gradedqualityCmaintenanceEnables generating images and videos on Google Flow through browser automation using your Google AI Pro subscription, with image generation free and video consuming Flow credits.116 npm1MIT
- AlicenseNot gradedqualityCmaintenanceEnables generating images and videos in Google Flow through an automated browser session using your Google AI Pro subscription, without API keys or third-party credit fees.116 npm2MIT
- AlicenseAqualityAmaintenanceDrives Google Flow from an agent or the terminal: Veo text-to-video, image-to-video and clip extension, plus Imagen image generation. Runs locally against your own Google account with Flow access; video generation bills that account's credits.161,013 PyPI251MIT