Google Flow Browser MCP
Drives Google Flow (flow.google.com) through a Chrome profile already signed in to Google, enabling an agent to generate images and videos, create reusable characters and scenes, reference them by @name, list live models/ratios/durations, download generated assets, and manage Flow projects — all without the agent ever handling Google credentials.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Google Flow Browser MCPgenerate an image of a sunset over a mountain lake"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Google Flow Browser MCP
Originally created by Mitanshp5. This repository is a maintained fork — created and maintained by DarthAdry. All original credit for the project's design and implementation belongs to Mitanshp5, who authored the original 45 commits. Released under the MIT License (see LICENSE).
An MCP (Model Context Protocol) server that lets an AI agent drive
Google Flow (flow.google.com)
through your own logged-in Chrome profile — generating images, videos,
characters, and scenes — without ever sharing your Google credentials with
the agent.
This server supports Windows, macOS, and Linux out of the box using unified npm commands that automatically dispatch platform-specific execution.
How it works
Chrome is launched with
--remote-debugging-port(Chrome DevTools Protocol / CDP) using a real Chrome profile that's already signed in to your Google account.The MCP server (
src/index.js) connects to that Chrome instance via Playwright'sconnectOverCDPand drives the Google Flow web app — filling prompts, clicking buttons, reading results.Your AI agent (Claude / OpenCode / any MCP client) talks to the MCP server over stdio and calls tools like
flow_generate_image,flow_generate_video,flow_create_character, etc.
Because the agent never touches your Google password — it only sends commands to a browser that's already logged in — there's no credential sharing involved.
Related MCP server: Google Flow Browser MCP
Prerequisites
Node.js >= 20.11 (see
enginesinpackage.json; LTS recommended)Google Chrome installed normally (Playwright connects to your real Chrome via CDP — it does not need its own bundled browser for this)
A Google account already signed in to a Chrome profile (e.g. "Default" or "Profile 1" — any profile works, you just need to tell the config which one)
(Optional) OpenCode if you want to register this MCP server with it
Setup
1. Clone and install dependencies
git clone https://github.com/DarthAdry/Google-Flow_MCP.git
cd Google-Flow_MCP
npm installnpm install will also download Playwright's browser binaries. This project
doesn't use Playwright's bundled Chromium though — it connects to your real
installed Chrome.
2. Find your Chrome profile
You need a Chrome profile that's already signed in to the Google account you want Flow to use.
Open Chrome and go to
chrome://versionLook at Profile Path — note the User Data directory and Profile folder name (e.g.
Default,Profile 1,Profile 2).
Default locations by platform:
Windows:
Executable:
C:\Program Files\Google\Chrome\Application\chrome.exe(Auto-detected)User Data:
C:\Users\<username>\AppData\Local\Google\Chrome\User Data
macOS:
Executable:
/Applications/Google Chrome.app/Contents/MacOS/Google Chrome(Auto-detected)User Data:
/Users/<username>/Library/Application Support/Google/Chrome
Linux:
Executable:
/opt/google/chrome/chromeor/usr/bin/google-chrome(Auto-detected)User Data:
/home/<username>/.config/google-chrome
If you don't have a profile signed in yet, sign in to your Google account in
any Chrome profile first (Settings → You and Google → Sign in).
⚠️ Close all running Chrome windows for that profile before using this tool — Chrome locks its profile directory while running, so the MCP server's "direct + CDP" launch mode copies the profile into a temporary directory to avoid conflicts. If Chrome is running with that profile while you start the scripts, you may see a "profile already in use" issue.
3. Create your config file
Copy the example config and edit it:
Windows (PowerShell):
copy config\flow.config.example.json config\flow.config.jsonmacOS / Linux:
cp config/flow.config.example.json config/flow.config.jsonOpen config/flow.config.json and set at minimum:
{
"expectedAccount": "your-email@gmail.com",
"chromeProfile": "Default"
}Note: You only need to set chromeUserDataDir and chromeExecutable if Chrome is installed in a custom location, as the server will otherwise auto-detect them for your current platform.
Unified Commands
This project uses a Node.js dispatcher script under the hood, allowing you to run the same npm commands on Windows, macOS, or Linux. The runner will automatically execute the correct scripts (.ps1 for Windows, .sh for Unix).
1. Start Chrome with CDP debugging
npm run start-browserThis launches Chrome with your configured profile and --remote-debugging-port=9222. A Chrome window will open — leave it running.
2. Start the MCP server (optional manual test)
In a separate terminal:
npm run start-mcpThis checks Node is installed, checks Chrome's CDP port is responding, and runs the MCP server. Press Ctrl+C to stop it.
3. Connect to AI Agents
You can automatically configure OpenCode and Gemini CLI in one step:
npm run registerThis will automatically find their configuration files and add the Google-Flow MCP server.
(You can also target specific clients: npm run register -- --opencode or --gemini)
Claude Desktop
Currently, Claude Desktop requires manual configuration:
Add the following to your claude_desktop_config.json (run npm run build first — dist/ is what ships):
{
"mcpServers": {
"Google-Flow": {
"command": "node",
"args": ["/absolute/path/to/Google-Flow_MCP/dist/index.js"]
}
}
}Cursor (Codex) Currently, Cursor requires manual UI configuration:
Open Cursor Settings > Features > MCP.
Click + Add New MCP Server.
Name:
Google-FlowType:
commandCommand:
node /absolute/path/to/Google-Flow_MCP/dist/index.js(runnpm run buildfirst;src/index.jsworks for dev only)
4. Verify everything end-to-end
With Chrome running (step 1), run unit tests plus a live smoke test:
npm test # vitest unit tests (validation, queue, sanitizers)
npm run test:smoke # flow_connect over stdio (needs Chrome running)
npm run build # bundles src/ -> dist/ (required before register)Typical workflow
npm run start-browser— once per session, leave the Chrome window openYour AI agent (with this MCP registered) calls:
flow_connect— attach to the running Chrome (also refreshes the live model catalog)flow_account_check— confirm the right Google account is signed inflow_list_models— list live video/image models, ratios, durations (call once per session before generating)flow_generate_image— prepare (or withauto_confirm:true, generate) images in a Flow projectflow_create_character— create reusable charactersflow_generate_video— prepare (or withconfirm_generate:true, paid-generate) videos, optionally referencing previously generated images/characters viaingredients(@namereferences)flow_list_mention_options— see what images/characters are available to reference by nameflow_create_scene— create a new scene, optionally referencing characters
The MCP server automatically creates a Google Flow project on first use and reuses it for the rest of the session (tracked by project ID), so everything ends up in one place instead of a new project per request.
Available tools
Tool | Purpose |
| Connect to / launch Chrome via CDP, optionally open Flow |
| Close the browser connection |
| Navigate the connected browser to a Flow URL without reconnecting |
| Report current connection/page status |
| Verify the signed-in Google account matches |
| Navigate to a Flow page and dump interactive elements (debugging) |
| Generate image(s) from a prompt in the current project (prepare-only by default; |
| Generate video(s), optionally with |
| List video/image models, ratios, durations from Flow UI (live) or cache |
| Download the most recently generated asset |
| Create a new character (name + description + reference images) |
| Import a character from a saved JSON file |
| Open the Characters page for a project |
| List images/characters available for |
| Create a new scene, optionally referencing characters |
| Open Flow's Tools gallery |
| Open the Grid Architect tool |
| Use an arbitrary tool from the Tools gallery |
| Take a debug screenshot of the current page |
| Check the status of queued generation jobs |
| Forcefully reset the job queue if a job is stuck in "running" state |
Configuration reference
All settings live in config/flow.config.json (copy from flow.config.example.json). Key fields:
Field | Description | Default |
| Google account email | (required) |
| Full path to | auto-detected |
| Path to Chrome's "User Data" folder | auto-detected per-OS |
| Profile folder name (e.g. |
|
| Chrome DevTools Protocol port |
|
| Base Google Flow URL |
|
| Run Chrome headless (not recommended — Flow needs visible browser) |
|
| Watchdog: max time a job may stay |
|
| Bounded queue history kept in |
|
| Delay between automated UI steps |
|
| Agent dialog window / generation wait |
|
| Wait for download to complete |
|
| Display name → internal model ID maps (merged with live discovery) | see example config |
| Allowed generation parameters | see example config |
Env overrides win over the file: FLOW_URL, FLOW_EXPECTED_ACCOUNT, FLOW_CHROME_PROFILE, FLOW_CDP_PORT, FLOW_HEADLESS, FLOW_JOB_TIMEOUT_MS, FLOW_CHROME_EXECUTABLE, FLOW_CHROME_USER_DATA_DIR, FLOW_HOME.
Troubleshooting
"Chrome not found"
Make sure Chrome is installed normally. If it is in a custom path, set chromeExecutable in config/flow.config.json.
"Chrome profile not found"
Double-check chromeUserDataDir + chromeProfile against chrome://version in your browser.
"CDP port 9222 not responding"
Run npm run start-browser first and leave that Chrome window open.
Google sign-in gets blocked ("This browser may not be secure")
This is why the server connects to your existing signed-in Chrome profile. Ensure chromeUserDataDir / chromeProfile points to the profile you logged in with.
The agent creates a new Flow project every time
The server tracks one project per session by ID and reuses it. If this happens, check the logs for Session project no longer reachable and start a fresh session.
Project layout
config/
flow.config.example.json # template — copy to flow.config.json (gitignored)
flow.models.json # live model catalog cache (gitignored, regenerated)
selectors.map.json # self-healing selector cache from flow_discover_ui (gitignored)
scripts/
run.js # OS-independent dispatcher (prefers pwsh on Windows)
start-browser.ps1 / .sh # launch Chrome with CDP debugging
start-mcp.ps1 / .sh / .bat # run the MCP server (manual check)
register.ps1 / .sh # register this server (prefers dist/, falls back to src/)
test-flow-image.ps1 / .sh # quick flow_connect smoke test (npm run test:smoke)
src/
index.js # MCP server entry point (validated with zod, fail-closed errors)
browser/ # Chrome/CDP connection (non-destructive reuse, liveness)
navigation/ # project registry, @ mention references, live model discovery
queue/ # single-job queue with watchdog + bounded history
tools/ # one file per MCP tool
utils/ # config, logger (stderr-only), screenshots, sanitize, validate,
# prompt QC, dynamic model universe, selector registry
tests/ # vitest unit tests: sanitize, validate, job-queue,
# file-manager, video-quality, dynamic-catalog, prompt-barSafety notes
This server only automates a browser you already control and are signed into — it does not store, transmit, or need your Google credentials.
Separate from credentials: the server deliberately minimizes automation fingerprints — it launches Chrome with
--disable-blink-features=AutomationControlled(sonavigator.webdriverreads false) and runs your signed-in profile from a temp copy. That is bot-detection evasion against Google's own sign-in/abuse checks (the "this browser may not be secure" block), distinct from the credential story above. By using this tool you accept that tradeoff relative to Google's Terms of Service.Image and video generation consume Google Flow credits. All mutating tools default to safe "prepare only" behavior (
auto_confirm: false; video live-generate additionally requiresconfirm_generate: true).flow_create_characteralso defaults tofalse.flow_disconnectnever kills your personal Chrome tabs — it detaches unless the server launched Chrome itself.flow_account_checkis fail-closed: without positive evidence it returnsverified:false, needsManualCheck:true(neverassumed:true).config/flow.config.jsonis gitignored — do not commit it. Runtime caches (flow.models.json,selectors.map.json,flow.projects.json) are gitignored too.
Recent changes (v1.0.1)
Probe reliability —
flow_discover_uinow waits for panel content, not just the container element, so it no longer returns an empty element list while the panel is still rendering.Import hygiene — added a test that fails if
src/imports fromdist/, keeping the bundle boundary clean.Lint gate — ESLint config (
no-unused-vars,no-undef) added, clearing 19 findings; run withnpm run lint.Debug-trail retention —
flow_connectkeeps recent debug context on failure so it's available inflow_statusafter an error, instead of being swept immediately.CI — GitHub Actions workflow runs
npm ci,npm test, andnpm run buildon push and pull requests.
Test suite: 127 tests across 21 files (npm test).
Contributing
Fork the repo and create a branch:
git checkout -b my-changeMake your change, then verify:
npm run lint && npm test && npm run buildCommit and open a pull request
CI runs the same three commands, so a green local run means a green PR.
Please don't commit config/flow.config.json or anything containing real Google account details — both are gitignored by design.
License
This project is licensed under the MIT License.
Available Tools
21 toolsflow_account_checkA
Verify the logged-in Google account matches the configured expected email (Default).
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It only says 'Verify' but does not disclose what happens on success/failure, whether it returns a boolean, throws an error, or has side effects. For a check tool, this lack of behavioral detail is a significant gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, well-structured sentence with leading verb and no fluff. It conveys the core purpose efficiently without wasting words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool is simple with no parameters and no output schema, the description is partially complete. It explains what it does but not what the agent should expect as a result (e.g., does it return a Boolean, throw an exception, or print a message?). The mention of 'Default' also lacks context about how the expected email is configured, leaving some ambiguity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so the description does not need to explain any. The baseline for 0 params is 4, and the description correctly avoids adding unnecessary parameter details since the schema is empty.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states a specific action (verify) and the exact subject (logged-in Google account) against an expected configured email. This distinguishes it from sibling tools like flow_status or flow_connect, which have different purposes.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the tool is used to ensure the correct Google account is active, but it does not explicitly state when to use it versus alternatives. For example, it doesn't mention using it after flow_connect or before other operations. Usage context is implied rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
flow_connectA
Launch Chrome with the configured Google profile, connect CDP, navigate to Google Flow, and verify account.
| Name | Required | Description | Default |
|---|---|---|---|
| headless | No | Launch in headless mode (not recommended, Google Flow needs visible browser). | |
| open_flow | No | Auto-navigate to Google Flow after connection. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden. It discloses the main side effect (launching Chrome) and the steps performed, but lacks detail on whether the connection is persistent, what happens on verification failure, or whether the browser closes after the operation. This is moderate transparency but not complete.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence that packs four specific actions without fluff. It is front-loaded with the primary action ('Launch Chrome') and efficiently communicates the full flow.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description does not mention return values (no output schema exists) or the fact that this tool likely establishes a required connection for other flow_* tools. Without annotations, it could benefit from noting that it is a prerequisite and what 'verify account' entails. This is adequate but not fully complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%: both 'headless' and 'open_flow' have clear descriptions in the schema, including a warning about headless mode. The tool description does not add additional parameter semantics, but since the schema already covers them, a baseline score of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's specific actions: launching Chrome, connecting CDP, navigating to Google Flow, and verifying the account. It distinguishes itself from sibling tools like flow_disconnect or flow_status by focusing on the connection setup.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description does not provide explicit guidance on when to use this tool relative to alternatives. It implies it is a connection step but does not mention prerequisites, exclusions, or alternative tools. For example, it does not clarify whether this should be run before flow_status or flow_discover_ui.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
flow_create_characterA
Create a new character in Google Flow Characters with name and description. By default (auto_confirm=false): fills the description and returns "ready_for_confirmation" without submitting (no credits). When auto_confirm=true: fills and clicks submit, then renames "Untitled Character" to the given name.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | Character name. | |
| campaign | No | Campaign identifier for project matching (e.g., "summer-2026", "new-collection"). | |
| description | Yes | Character description/prompt. | |
| auto_confirm | No | If false (default): only fills the description and returns "ready_for_confirmation" without submitting. If true: submits (may use credits). | |
| project_name | No | Name for the project (will reuse existing project with same campaign, or create new). | |
| reference_images | No | Paths to reference images for character design. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does well: it discloses the non-submitting default, the credit consumption risk on submit, and the post-submit rename of 'Untitled Character'. It omits auth/permission requirements and error behavior, but the mutation semantics and side effects are clearly stated.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences, front-loaded with the core action and then the branching behavior. Dense but every clause carries information; only the repeated auto_confirm restatement is mildly redundant with the schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 6-parameter mutation tool with no annotations and no output schema, the description covers the critical unknown (what happens on submit vs. not, including credits and the rename) and reveals the returned status token. The remaining params are fully documented in the schema, so the agent has enough to invoke correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so name, description, campaign, project_name, reference_images and auto_confirm are already documented. The description's explanation of auto_confirm largely restates the schema, adding no format or syntax detail beyond it; baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Create a new character in Google Flow Characters') and the required inputs, which cleanly separates it from siblings like flow_import_character and flow_create_scene.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explains the two operating modes and their consequences (default fills only and returns ready_for_confirmation with no credits; auto_confirm=true submits), which is exactly the when-to-use decision an agent must make. It does not name sibling alternatives such as flow_import_character for existing characters.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
flow_create_sceneB
Create a new scene in Google Flow Scenes with characters and prompt. By default (auto_confirm=false) prepares only; set auto_confirm=true to actually create.
| Name | Required | Description | Default |
|---|---|---|---|
| prompt | Yes | Scene description/prompt. | |
| campaign | No | Campaign identifier for project matching (e.g., "summer-2026", "new-collection"). | |
| characters | No | Character names to include in the scene. | |
| auto_confirm | No | If false (default): prepare only. If true: click New Scene and create. | |
| project_name | No | Name for the project (will reuse existing project with same campaign, or create new). | |
| reference_image | No | Path to a reference image (optional). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It does disclose one genuinely valuable trait — that the default is a non-mutating 'prepare only' state and auto_confirm=true performs the real creation — but says nothing about permissions, side effects, idempotency, or what 'prepares only' actually produces or how it is later confirmed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two compact sentences with the core action, then the mode caveat, both front-loaded. Slightly awkward parenthetical phrasing keeps it from being a model of efficiency, but there is no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a six-parameter creation tool with no output schema and no annotations, the description leaves important questions open: what a 'prepared only' invocation returns, how project/campaign reuse interacts with scene creation, and what happens on failure. The rich schema compensates for parameter gaps, but the workflow behavior is under-explained overall.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all six parameters (prompt, campaign, characters, auto_confirm, project_name, reference_image) are already documented in the schema. The description restates characters/prompt and the auto_confirm behavior, adding no syntax, format, or constraint detail beyond what the schema already supplies; baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'Create a new scene in Google Flow Scenes', and names the key inputs (characters, prompt). It is clearly distinguishable from non-creative siblings like flow_status or flow_list_models, though it does not explicitly contrast itself with the nearby flow_create_character.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains the auto_confirm mode switch (prepare-only vs. actually create), which implies a two-step workflow, but gives no guidance on when to prefer this tool over flow_generate_video, flow_create_character, or the other creation siblings. Usage context is only partially inferred.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
flow_disconnectA
Close the browser and clean up the MCP connection to Google Flow.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description declares the primary actions (closing browser, cleaning up connection) but does not disclose side effects such as terminating active sessions or preventing further operations. This is adequate for a simple cleanup tool but lacks deeper behavioral context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, well-structured sentence that directly states the tool's function without unnecessary words. It is perfectly concise and front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a parameterless tool with no output schema or annotations, the description provides the essential information about its role as a disconnect/cleanup operation. It lacks explicit usage context but is otherwise complete for its simplicity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters and an empty schema, so the description adds nothing about parameters. Since there are no parameters to explain, the baseline of 4 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses specific verbs 'Close' and 'clean up' to describe the resource (browser and MCP connection), clearly indicating this is the teardown counterpart to flow_connect. It distinguishes itself from sibling tools by its focus on disconnection and cleanup.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description does not explicitly state when to use the tool or contrast it with alternatives like flow_connect. Usage must be inferred from the name and context, so the agent receives no direct guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
flow_discover_uiB
Navigate to a Google Flow page and discover all interactive elements (buttons, inputs, links, headings). Updates the internal selectors map for robust automation.
| Name | Required | Description | Default |
|---|---|---|---|
| url | No | Exact https://flow.google.com/... URL to navigate to, overriding page-based resolution. Other hosts are rejected. | |
| page | No | Page to discover. Options: main, image-generation, video-generation, characters, scenes, images, all-media, tools-gallery, grid-architect. "characters" is project-scoped (/project/{id}/characters). "scenes"/"images"/"all-media" are sidebar tabs within the project (same URL). project_name/campaign target a specific project, or the current/active project is used. | main |
| campaign | No | Campaign identifier for project matching, used when page is "characters", "scenes", "images", or "all-media". | |
| project_name | No | Project name, used when page is "characters", "scenes", "images", or "all-media" (will reuse existing project with same campaign, or create new). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does disclose an important non-obvious side effect: it 'Updates the internal selectors map.' That mutation info is valuable. But it omits auth requirements, browser-state effects, error behavior, and how the result is surfaced, leaving substantial gaps.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences with the action front-loaded and the side effect trailing. Every clause earns its place with no redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
All four parameters are fully documented in the schema, but with no output schema the description does not explain what 'discover' returns (a list of elements? just the internal map?) or how the agent accesses the discovered elements. Adequate but leaves the return/usage outcome underspecified.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents url, page, campaign, and project_name in detail, including enum-like options and host restrictions. The description adds no parameter meaning beyond that, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb+resource: 'Navigate to a Google Flow page and discover all interactive elements (buttons, inputs, links, headings),' which clearly describes the action and its scope. However it does not distinguish itself from sibling flow_open, which also appears to navigate to Flow pages, leaving the agent to guess which to select.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description hints at purpose ('for robust automation') but gives no explicit when-to-use, when-not-to-use, or alternatives. It never explains how it relates to flow_open or when discovering UI should precede another sibling like flow_generate_image.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
flow_download_latestB
Download the most recently generated file from Google Flow.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden of behavioral disclosure. It only states select 'most recently generated' but does not explain whether the file is returned as content, a URL, or saved locally, nor any side effects or permissions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, clear sentence with no filler. It is front-loaded with the main action and resource, making it highly concise and easy to process.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with no output schema and no annotations, the description is incomplete. It does not specify what the agent should expect as a result (file path, binary content, URL) or how this tool fits into the flow with sibling generation tools (e.g., used after flow_generate_image).
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has zero parameters, so the description need not elaborate on parameter meaning. The description implicitly communicates that no user input is required, aligning with the baseline for a parameterless tool.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb 'Download' and a clear resource 'the most recently generated file from Google Flow.' It distinguishes itself from sibling tools (e.g., flow_generate_image, flow_connect) by being the only download-focused tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus alternatives, prerequisites, or typical usage contexts. The description simply states the action without explaining when it is appropriate or what conditions must exist (e.g., a file must have been generated first).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
flow_generate_imageA
⚠️ THESE IMAGES CONSUME CREDITS. By default (auto_confirm=false): fills the prompt, selects model/ratio, takes a screenshot and returns "ready_for_confirmation". Does NOT click Generate. When auto_confirm=true: first verifies the UI is in IMAGE mode (not Video), that the model is an image model, takes a verification screenshot, THEN clicks Generate and waits for the images. NANO/BANANA image models only.
| Name | Required | Description | Default |
|---|---|---|---|
| model | No | Model to use: Nano Banana Pro, Nano Banana 2, or Imagen 4. | Nano Banana 2 |
| ratio | No | Aspect ratio: 1:1, 16:9, 9:16, 4:3, 3:4. | 1:1 |
| prompt | Yes | The text prompt for image generation. | |
| campaign | No | Campaign identifier for project matching (e.g., "summer-2026", "new-collection"). | |
| use_scene | No | Name of a single project scene to reference via "@name" (added in addition to ingredients). | |
| ingredients | No | Names of existing project images/characters to reference via "@name" (e.g., ["Bob the Astronaut", "Image 3"]). Use flow_list_mention_options to discover available names. | |
| auto_confirm | No | ⚠️ CREDITS. If false (default): only prepares, consumes nothing. If true: verifies that Image mode is active, THEN clicks Generate (consumes credits). | |
| project_name | No | Name for the project (will reuse existing project with same campaign, or create new). | |
| use_character | No | Name of a single project character to reference via "@name" (added in addition to ingredients). | |
| reference_images | No | Paths to local reference images to upload (optional). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden and does so well: it discloses credit consumption, the exact steps taken in each mode (fill prompt, select model/ratio, screenshot, return 'ready_for_confirmation'), that Generate is NOT clicked by default, the pre-flight verification of mode and model, and the wait-for-images behavior. This is far richer than a typical schema would convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The credit warning is front-loaded, and the two modes are laid out as parallel clauses rather than repetition. Density is high with virtually no filler despite covering a complex flow.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 10-parameter, side-effectful generation tool with no annotations and no output schema, the description covers the critical unknowns: cost implications, what is and isn't executed per mode, and the partial return value ('ready_for_confirmation'). Nothing essential to safe invocation is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so every parameter is already documented with defaults, formats, and examples. The description reinforces auto_confirm's semantics but adds no parameter meaning beyond the schema. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource (generate image) and explicitly scopes the model family ('NANO/BANANA image models only'), which cleanly separates it from sibling flow_generate_video. The two-mode behavior is front-loaded so an agent knows immediately what the tool does.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly contrasts the two operating modes: auto_confirm=false prepares only, auto_confirm=true verifies and clicks Generate. It also specifies prerequisites (UI must be in IMAGE mode, model must be an image model). It does not name sibling alternatives for video generation, but the scoping note covers the main branch point.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
flow_generate_videoB
Prepare a Veo video (prepare-only, no spend). Fast tests, Quality finals; ingredients need 8s.
| Name | Required | Description | Default |
|---|---|---|---|
| model | No | Video model name — call flow_list_models for the live list — or auto. | auto |
| ratio | No | 16:9 or 9:16. | 16:9 |
| shots | No | Multi-shot plan, max 6: [{prompt, duration}]. | |
| style | No | Look/mood/lighting. No extra subjects in refs. | |
| camera | No | Camera move: push-in, pull-back, pan, tilt, orbit, tracking shot. | |
| prompt | Yes | Scene + action. Optics stay in UI selectors. | |
| quality | No | Resolution 360p/720p. | |
| campaign | No | Campaign id (doubles as project name when name absent). | |
| duration | No | 4s, 6s, 8s, 10s. Ingredients force 8s. | 4s |
| quantity | No | Variants 1-4 for compare-pick-refine. | |
| end_frame | No | Last-frame image path. Shares light/framing with start. | |
| use_scene | No | Name of a single project scene to reference via "@name" (added in addition to ingredients). | |
| hero_frame | No | Alias for start_frame. | |
| ingredients | No | Project images/characters to reference via "@name" (e.g., ["Bob the Astronaut", "Image 3"]). Discover names via flow_list_mention_options. | |
| project_url | No | Direct Flow /project/ URL — wins over name matching. | |
| start_frame | No | First-frame image path (hero frame anchor). | |
| anchor_frame | No | Prior strongest frame carried for continuity. | |
| auto_confirm | No | Always prepare-only today. Reserved. | |
| project_name | No | Project name (remembered; reopens instead of duplicating). | |
| use_character | No | Name of a single project character to reference via "@name" (added in addition to ingredients). | |
| confirm_generate | No | LIVE PAID CLICK: verify, click Generate (credits), poll video. Default false. | |
| reference_images | No | Paths to local reference images to upload (optional). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It usefully discloses the default no-spend prepare behavior, which counteracts the misleading 'generate' name, but it says nothing about the paid confirm_generate path, required auth, or polling/credit consumption beyond what the schema param already states.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three front-loaded fragments lead with the single most important behavior (prepare-only, no spend) and waste no words. The terseness makes 'Fast tests, Quality finals' slightly cryptic, but nothing is padding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 22-parameter, annotation-free tool with no output schema, the description covers the core safety profile but is thin on the surrounding workflow (how it pairs with flow_list_models, flow_list_mention_options, or the confirm/poll path). The rich schema compensates for parameters, but behavioral coverage is minimal.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all 22 parameters are already documented, giving a baseline of 3. The description adds only a marginal cross-parameter rule ('ingredients need 8s', quality tiers) beyond what the individual param descriptions already say.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Prepare) and resource (Veo video) and immediately disambiguates against its own name by noting 'prepare-only, no spend'. An agent can tell it produces a video plan rather than a video, though it never names the sibling flow_generate_image or explains how it relates to the image tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Fast tests, Quality finals' implies a quality-tier decision and 'ingredients need 8s' gives a conditional constraint, but there is no explicit when-to-use vs alternatives, no exclusion of the live-generate path, and no prerequisites (auth, project) stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
flow_import_characterA
Import a character from a saved JSON file into Google Flow.
| Name | Required | Description | Default |
|---|---|---|---|
| campaign | No | Campaign identifier for project matching (e.g., "summer-2026", "new-collection"). | |
| file_path | Yes | Path to character JSON file. | |
| project_name | No | Name for the project (will reuse existing project with same campaign, or create new). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description must carry the full burden of behavioral disclosure. It only states the high-level action, without detailing side effects (e.g., overwriting existing characters), validation of the JSON, error handling, or any prerequisites. This is a significant gap for an import tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence with no filler or redundant information. Every word earns its place, making it highly concise and easy to parse.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The core purpose is clear and the schema covers parameters, but the description lacks usage guidance relative to sibling tools and does not disclose behavioral traits (e.g., whether it overwrites or merges). For a tool with no annotations and no output schema, this is adequate but leaves gaps for an agent deciding between import and create.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema provides descriptions for all three parameters (campaign, file_path, project_name), giving 100% schema coverage. The description itself adds no parameter-specific detail, so the baseline of 3 is appropriate since the schema handles parameter semantics.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Import') with a clear resource ('character') and source/destination ('saved JSON file' into 'Google Flow'). This clearly distinguishes it from siblings like flow_create_character (which creates from scratch) and flow_open_characters (which opens existing characters).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage when a saved JSON character file exists, but it offers no explicit when-to-use or when-not-to-use guidance, and does not mention alternative tools such as flow_create_character. The context is inferred from the phrase 'saved JSON file' rather than stated directly.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
flow_list_mention_optionsA
List the images and characters in the current project that can be referenced via "@name" in an image/video prompt (opens Flow's "@" reference popup and reads its options).
| Name | Required | Description | Default |
|---|---|---|---|
| campaign | No | Campaign identifier for project matching (e.g., "summer-2026", "new-collection"). | |
| project_name | No | Name for the project (will reuse existing project with same campaign, or create new). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden of disclosure. It transparently states the side effect of opening Flow's '@' reference popup and reading its options, which goes beyond a generic 'list' description. It does not detail safety permissions, but the read-only nature is inferred and no harmful behavior is hidden.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence that is both concise and information-dense. It includes the action, target, mechanism, and context without redundancy or filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description sufficiently explains the tool's purpose and behavior for a listing tool. It lacks an explicit return format, but that is partially mitigated by the absence of an output schema and the simplicity of the feature. It is complete enough for an agent to invoke correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema already provides full descriptions for both 'campaign' and 'project_name' (100% coverage). The tool description adds no additional parameter meaning, so the baseline score of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('List'), the resource ('images and characters'), and the context ('current project' and 'via @name in an image/video prompt'). It distinguishes this from sibling tools like flow_open_characters by specifying the popup-reading mechanism.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context on when to use the tool (when needing @name references for image/video prompts) and implies the popup interaction. It does not explicitly name alternatives or exclusions, so it stops short of a 5 but is clear enough for an agent.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
flow_list_modelsA
List video/image models, ratios, durations from Flow UI (live) or cache. Call once per session before generating.
| Name | Required | Description | Default |
|---|---|---|---|
| refresh | No | Force re-discovery from the open UI instead of cache. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It discloses the live-vs-cache duality and the once-per-session constraint, but omits whether a connected/open Flow UI is required, what the call costs, and whether results are exhaustive. For a read-only discovery tool this is minimally adequate rather than rich.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, no filler. The resource list is front-loaded and the calling constraint follows immediately, so an agent can act after one read.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter listing tool with full schema coverage and no output schema, the description covers what is returned (models, ratios, durations) and when to call it. The remaining gap is the undeclared dependency on a live Flow UI session, which matters for a browser-driven tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the single 'refresh' parameter is fully documented in the schema. The description's 'live or cache' phrasing mirrors that schema text rather than adding syntax or format detail, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('List') and concrete resources (video/image models, ratios, durations), and scopes the source (Flow UI live or cache). It does not name a sibling explicitly, but 'before generating' clearly separates it from the flow_generate_* tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit timing guidance: 'Call once per session before generating.' That is a clear when-to-use directive tied to the generate workflow, though it names no exclusions or alternatives (e.g., what to do if the cache is stale besides refresh).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
flow_openA
Navigate the already-connected browser to a Google Flow URL without reconnecting. Only https://flow.google.com/... URLs are allowed (other hosts rejected).
| Name | Required | Description | Default |
|---|---|---|---|
| url | No | Exact Flow URL to navigate to. Defaults to the configured flowUrl when omitted. | |
| waitFor | No | Extra ms to wait after load (clamped 0-15000). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations the description carries the full burden, and it does disclose real behavior: navigation reuses the existing connection and a host allowlist rejects anything that is not https://flow.google.com/... It omits what happens if no session exists or if navigation fails, so it is not exhaustive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, zero filler, with the core action front-loaded and the host restriction immediately after. Every clause carries information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a low-complexity navigation tool with only two well-documented parameters, the description covers the action, the precondition, and the input constraint. It does not say what the call returns or how failures surface, which is the only remaining gap given there is no output schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% so baseline is 3, but the description adds a genuine constraint on the url parameter (only https://flow.google.com/... accepted, other hosts rejected) that the schema does not state. The waitFor parameter is left entirely to the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb (navigate) plus resource (already-connected browser to a Flow URL), and the phrase 'without reconnecting' implicitly separates it from flow_connect and flow_open_tools_gallery. An agent can select it correctly without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The precondition 'already-connected browser' and 'without reconnecting' tell the agent this is for an existing session rather than establishing one, which routes it away from flow_connect. It never names the alternative explicitly, so the routing is inferred rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
flow_open_charactersB
Open the Characters page for the current/target project and list existing characters.
| Name | Required | Description | Default |
|---|---|---|---|
| campaign | No | Campaign identifier for project matching (e.g., "summer-2026", "new-collection"). | |
| project_name | No | Name for the project (will reuse existing project with same campaign, or create new). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries full burden. It states the tool opens a page and lists characters but does not disclose potential side effects, such as creating a new project if none matches (as implied by the schema's project_name description). Lacks details on permissions or reversibility.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Single sentence, front-loaded with action and purpose. No wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description is adequate for a simple navigation/listing tool but lacks context on project selection mechanism and output nature. With no output schema or annotations, it could benefit from explaining the relationship between campaign and project_name and whether the list is returned as data.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema has 100% coverage, so baseline is 3. The description adds no parameter-specific information; it relies entirely on schema descriptions for campaign and project_name.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description uses specific verb 'open' and resource 'Characters page', and further clarifies it lists existing characters. This clearly distinguishes it from sibling tools like flow_create_character and flow_import_character, which perform different operations.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this tool versus alternatives such as flow_create_character or flow_import_character. The description implies usage for viewing existing characters but doesn't state exclusions or prerequisites.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
flow_open_tools_galleryA
Open the Google Flow Tools Gallery and list available tools.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It only states what the tool does (opens a gallery and lists tools) but does not disclose any side effects, permissions required, or whether it returns data in the response. For a simple tool this might be acceptable, but the lack of explicit behavioral context is a gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, compact sentence with no wasted words. It front-loads the action and resource clearly, making it immediately understandable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's simplicity (no parameters, no output schema, minimal annotations), the description covers the essential purpose. However, it leaves slight ambiguity about whether 'list available tools' means returning a list in the response or opening a UI gallery. This prevents a perfect score.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool accepts zero parameters, so the schema coverage is trivially 100%. The baseline for 0 parameters is 4, and the description adds no confusion about parameters. Since there are no parameters to clarify, the description does not need to provide additional semantics.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool opens the Google Flow Tools Gallery and lists available tools. This is a specific verb+resource combination that distinguishes it from sibling tools like flow_download_latest or flow_create_scene.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for discovering available tools, but does not explicitly state when to use this tool versus alternatives or mention any prerequisites. There are no exclusions or comparisons with siblings, so it relies on the name and context to convey intent.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
flow_queue_resetA
Forcefully reset the job queue. Use this if a job is permanently stuck in "running" state and blocking other generations.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description must convey behavioral traits. It states the action is 'forceful' and targets stuck jobs, implying a destructive/resetting behavior. However, it does not disclose potential side effects (e.g., loss of queued jobs, irreversibility, or required permissions), which are important for a reset operation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences long, with the primary action front-loaded in the first sentence. The second sentence adds valuable usage context without any fluff or redundancy. Every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple, zero-parameter tool, this description covers the purpose and the specific situation. There is no output schema, and the tool's return behavior is not described, but for a queue reset, the main point is the action itself. A small note about consequences (e.g., pending jobs may be cleared) would make it more complete, but it's adequate as is.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, and the input schema is an empty object. Per the rubric, 0 params is a baseline of 4. The description correctly implies no configuration is needed, so no additional parameter explanation is necessary.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Forcefully reset the job queue') and the specific resource (the job queue). It also provides a concrete scenario (job stuck in 'running' state blocking others), which distinguishes it from sibling tools like flow_queue_status that likely just query status.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly states when to use the tool: 'if a job is permanently stuck in "running" state and blocking other generations.' This gives clear context, though it does not mention when not to use it or explicitly name alternatives. Still, the use case is well-defined.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
flow_queue_statusA
Check the job queue: active job, pending queue, completed and failed job history.
| Name | Required | Description | Default |
|---|---|---|---|
| history_limit | No | Number of recent history entries to return. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Since no annotations are provided, the description carries the full burden. It clearly indicates a read-only operation ('Check') and lists the information returned, but it does not disclose any potential side effects, authentication requirements, or details about how history_limit affects results. It is minimally transparent but not deeply descriptive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single concise sentence that front-loads the primary action and enumerates the specific aspects of the queue. There is no wasted information; every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple tool with one optional parameter and no output schema, the description adequately captures the main functionality. It names all categories of queue status without needing to explain return values extensively. The only minor gap is not explicitly stating that history_limit controls the number of history entries, but that is already in the schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema provides 100% coverage for the only parameter (history_limit) with a clear description. The tool description does not add additional meaning beyond the schema, but since schema coverage is high, the baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Check') and resource ('job queue'), and enumerates the exact information covered (active job, pending queue, completed and failed job history). This clearly distinguishes it from sibling tools like flow_status, which likely covers broader status, and other flow_* tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit guidance on when to use this tool versus alternatives such as flow_status. The description implies that it is for checking queue information, but it does not state exclusions or mention alternative tools, leaving the agent without clear decision criteria.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
flow_screenshotB
Take a screenshot of the current Google Flow page.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | Custom name for the screenshot file. | manual |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of disclosing behavioral details. It only states the action without mentioning where the screenshot is saved, whether it returns a file path, authentication requirements, or any side effects, leaving significant ambiguity for an agent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence with no filler, immediately conveying the tool's action. It is appropriately concise and front-loaded, making it easy for an agent to parse.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's simplicity (one optional param, no output schema, no annotations), the description is minimally adequate but incomplete. It fails to explain what happens after the screenshot is taken (e.g., file location, return value), leaving the agent without enough context to predict the tool's full behavior.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The single optional parameter 'name' is fully described in the input schema (100% coverage), so the baseline is 3. The description adds no additional semantic detail about the parameter, but the schema already provides sufficient information.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states a specific action ('Take a screenshot') on a specific resource ('the current Google Flow page'), making its purpose unambiguous. This distinguishes it from sibling tools like flow_download_latest or flow_create_scene, which serve different actions.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description does not provide explicit guidance on when to use this tool versus alternatives. However, the intended use is implied by the action itself — taking a screenshot when a visual capture is needed — but no exclusions or alternative comparisons are mentioned.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
flow_statusB
Check current connection status: browser connected, Flow page loaded, account verified, job queue state.
| Name | Required | Description | Default |
|---|---|---|---|
| full | No | Return full status with screenshot. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It transparently lists the exact status dimensions checked (browser, page, account, queue), which is useful. However, it doesn't disclose whether this is a read-only operation, how the response is structured, or whether it performs live checks or returns cached state.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence of 12 words, front-loaded with the primary action and resource. Every word adds value, with no redundancy or filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simplistic status-check tool with one fully described parameter, the description covers core purpose and scope. However, without an output schema or annotations, it omits details about return format, potential latency, and any safety characteristics, leaving the agent uncertain about how to interpret the result.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema description fully documents the 'full' parameter ('Return full status with screenshot'), achieving 100% coverage. The tool description adds no parameter-specific context, so the baseline score of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses the specific verb 'Check' and clearly enumerates the status components (browser connected, Flow page loaded, account verified, job queue state). It implies a broader scope than individual siblings like flow_account_check or flow_queue_status but does not explicitly name them for differentiation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus alternatives. It doesn't mention that flow_queue_status should be used for queue-only details or flow_connect for establishing a connection. The agent is left without explicit selection criteria.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
flow_use_grid_architectB
Open Grid Architect in Google Flow, fill theme prompt, shot prompts, engine, ratio, and visual logic settings. Supports batch shot generation for brand campaigns. Prepare-only by default (auto_confirm=false).
| Name | Required | Description | Default |
|---|---|---|---|
| ratio | No | Aspect ratio for all shots. | 16:9 |
| engine | No | Engine/model for the grid. | Nano Banana 2 |
| campaign | No | Campaign identifier for project matching (e.g., "summer-2026", "new-collection"). | |
| references | No | Paths to reference images. | |
| auto_confirm | No | Reserved — currently prepare-only. | |
| project_name | No | Name for the project (will reuse existing project with same campaign, or create new). | |
| shot_prompts | No | Array of individual shot prompts for the grid. | |
| theme_prompt | Yes | Overall theme prompt for the grid. | |
| visual_logic | No | Visual logic type: None, Colour Pop, Side by Side, etc. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, and it does disclose one important trait: 'Prepare-only by default (auto_confirm=false)'. It also hints at project reuse via the schema, but it never says what happens after preparation, whether a separate confirm step exists, or what permissions/account state are required — significant gaps for an unattended UI-driving tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences with the core action front-loaded and the prepare-only caveat placed last where it matters most. No filler, though the middle sentence's field enumeration is largely redundant with the schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 9-parameter, no-annotation, no-output-schema tool, the description covers the primary flow and the safety default but omits the post-prepare lifecycle (what the caller does next) and any notion of the return value. It is minimally sufficient but leaves the agent guessing about the step after invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description restates theme prompt, shot prompts, engine, ratio and visual logic — all already documented in the schema — and adds no format or constraint detail (e.g., valid engine names, visual_logic enum values, reference image path expectations).
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Open Grid Architect in Google Flow') and enumerates the fields it fills, so the agent knows exactly what operation this performs. It implies batch shot generation, which differentiates it somewhat from flow_generate_image, but it never names a sibling or explicitly contrasts the two.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Supports batch shot generation for brand campaigns' gives an implied use case, but there is no explicit when-to-use/when-not guidance and no mention of alternatives such as flow_generate_image or flow_create_scene for single-shot work. The agent must infer routing from the batch/campaign phrasing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
flow_use_toolB
Open any tool by name in Google Flow and optionally fill its configuration parameters. Prepare-only by default (auto_confirm=false).
| Name | Required | Description | Default |
|---|---|---|---|
| params | No | Optional configuration parameters for the tool. | |
| campaign | No | Campaign identifier for project matching (e.g., "summer-2026", "new-collection"). | |
| tool_name | Yes | Name of the tool to open (e.g. Grid Architect, Image Generation). | |
| auto_confirm | No | Reserved — currently prepare-only. | |
| project_name | No | Name for the project (will reuse existing project with same campaign, or create new). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It does disclose one genuinely important trait: the tool is prepare-only unless auto_confirm is set, so the agent knows opening a tool is non-destructive by default. It omits other relevant behavior — whether a connected session/project is required, what state the opened tool leaves behind, and what happens to params that don't match the target tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences that are front-loaded with the core action, followed immediately by the most consequential behavioral caveat. The clause about optionally filling configuration parameters is mildly redundant with the schema, but there is no padding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 5-parameter tool (with a nested object) and no output schema and no annotations, the description is only minimally sufficient. It does not explain what an agent gets back after opening a tool, nor the relationship between project_name/campaign reuse and an already-open session — details the agent must have since there is no return schema to fall back on.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all five parameters, including the campaign/project_name reuse semantics and the reserved auto_confirm flag. The description's only parameter-related statement ('optionally fill its configuration parameters') restates the schema rather than adding syntax, format, or validation meaning. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb ('Open') and resource ('any tool by name in Google Flow'), plus the optional parameter-filling behavior. It implicitly distinguishes itself from the sibling flow_use_grid_architect (one specific tool) by saying 'any tool', but never names an alternative explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives one useful usage signal — 'Prepare-only by default (auto_confirm=false)' — which tells the agent the call will not execute anything by default. However, there is no guidance on when to prefer this generic opener over flow_use_grid_architect, flow_open_tools_gallery, or whether flow_connect must run first. Usage is only implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
21 tool updates
v1.0.1- First observed
flow_account_check - First observed
flow_connect - First observed
flow_create_character - First observed
flow_create_scene - First observed
flow_disconnect - First observed
flow_discover_ui - First observed
flow_download_latest - First observed
flow_generate_image - First observed
flow_generate_video - First observed
flow_import_character - First observed
flow_list_mention_options - First observed
flow_list_models - First observed
flow_open - First observed
flow_open_characters - First observed
flow_open_tools_gallery - First observed
flow_queue_reset - First observed
flow_queue_status - First observed
flow_screenshot - First observed
flow_status - First observed
flow_use_grid_architect - First observed
flow_use_tool
TDQS
Scored across 21 tools
Most tools have a clearly distinct resource or action (status, connect, generate_image, create_character, etc.), so an agent can usually tell them apart. However, flow_use_tool overlaps with specific wrappers like flow_use_grid_architect, and flow_status partially duplicates flow_queue_status and flow_account_check.
All tool names use a consistent flow_ prefix plus snake_case verb/noun structure. The convention is predictable throughout the set.
21 tools is heavy for a single browser-automation server and lands in the borderline range per the rubric. While many operations are justified, some could be consolidated or deferred, especially generic flow_use_tool versus specific tool wrappers.
Core lifecycle operations are present: connect/disconnect, navigate, discover UI, generate image, create characters/scenes, and queue status/reset. But there are notable gaps: flow_generate_video is prepare-only with no execute path, and there are no update/delete operations for characters/scenes or listing generated assets beyond the latest download.
Maintenance
Related MCP Connectors
Build and run visual creative-production workflows from your AI agent.
Stealth web automation for AI agents. Login, signup, navigate, screenshot.
Stealth web automation for AI agents. Login, signup, navigate, screenshot.
- FlowNodeOAuthio.flownode
Generate images, video, audio and 3D with FlowNode; results land in your asset library.
Related MCP Servers
- AlicenseAqualityDmaintenanceControls Google Flow for image and video generation from an AI agent. Enables generating images with models like Imagen 4, creating videos, managing characters and scenes via browser automation.17125 npm81MIT
- AlicenseAqualityDmaintenanceEnables AI agents to drive Google Flow through a real Chrome profile to generate images, videos, characters, and scenes without sharing credentials.1914 npmMIT
- AlicenseAqualityAmaintenanceEnables AI agents to programmatically generate images and videos through the authenticated Google Flow web interface via a direct Chrome DevTools Protocol connection, exposing tools for media generation, project management, status checks, and asset downloads without requiring official API keys.273MIT
- AlicenseAqualityAmaintenanceEnables AI agents to generate Google Flow videos and images and automate Scene Builder clip extensions through a user's own Chrome session.1174 npm4MIT