Skip to main content
Glama

Google Flow Browser MCP

Originally created by Mitanshp5. This repository is a maintained fork — created and maintained by DarthAdry. All original credit for the project's design and implementation belongs to Mitanshp5, who authored the original 45 commits. Released under the MIT License (see LICENSE).

An MCP (Model Context Protocol) server that lets an AI agent drive Google Flow (flow.google.com) through your own logged-in Chrome profile — generating images, videos, characters, and scenes — without ever sharing your Google credentials with the agent.

This server supports Windows, macOS, and Linux out of the box using unified npm commands that automatically dispatch platform-specific execution.


How it works

  1. Chrome is launched with --remote-debugging-port (Chrome DevTools Protocol / CDP) using a real Chrome profile that's already signed in to your Google account.

  2. The MCP server (src/index.js) connects to that Chrome instance via Playwright's connectOverCDP and drives the Google Flow web app — filling prompts, clicking buttons, reading results.

  3. Your AI agent (Claude / OpenCode / any MCP client) talks to the MCP server over stdio and calls tools like flow_generate_image, flow_generate_video, flow_create_character, etc.

Because the agent never touches your Google password — it only sends commands to a browser that's already logged in — there's no credential sharing involved.


Related MCP server: Google Flow Browser MCP

Prerequisites

  • Node.js >= 20.11 (see engines in package.json; LTS recommended)

  • Google Chrome installed normally (Playwright connects to your real Chrome via CDP — it does not need its own bundled browser for this)

  • A Google account already signed in to a Chrome profile (e.g. "Default" or "Profile 1" — any profile works, you just need to tell the config which one)

  • (Optional) OpenCode if you want to register this MCP server with it


Setup

1. Clone and install dependencies

git clone https://github.com/DarthAdry/Google-Flow_MCP.git
cd Google-Flow_MCP
npm install

npm install will also download Playwright's browser binaries. This project doesn't use Playwright's bundled Chromium though — it connects to your real installed Chrome.

2. Find your Chrome profile

You need a Chrome profile that's already signed in to the Google account you want Flow to use.

  1. Open Chrome and go to chrome://version

  2. Look at Profile Path — note the User Data directory and Profile folder name (e.g. Default, Profile 1, Profile 2).

Default locations by platform:

  • Windows:

    • Executable: C:\Program Files\Google\Chrome\Application\chrome.exe (Auto-detected)

    • User Data: C:\Users\<username>\AppData\Local\Google\Chrome\User Data

  • macOS:

    • Executable: /Applications/Google Chrome.app/Contents/MacOS/Google Chrome (Auto-detected)

    • User Data: /Users/<username>/Library/Application Support/Google/Chrome

  • Linux:

    • Executable: /opt/google/chrome/chrome or /usr/bin/google-chrome (Auto-detected)

    • User Data: /home/<username>/.config/google-chrome

If you don't have a profile signed in yet, sign in to your Google account in any Chrome profile first (Settings → You and Google → Sign in).

⚠️ Close all running Chrome windows for that profile before using this tool — Chrome locks its profile directory while running, so the MCP server's "direct + CDP" launch mode copies the profile into a temporary directory to avoid conflicts. If Chrome is running with that profile while you start the scripts, you may see a "profile already in use" issue.

3. Create your config file

Copy the example config and edit it:

Windows (PowerShell):

copy config\flow.config.example.json config\flow.config.json

macOS / Linux:

cp config/flow.config.example.json config/flow.config.json

Open config/flow.config.json and set at minimum:

{
  "expectedAccount": "your-email@gmail.com",
  "chromeProfile": "Default"
}

Note: You only need to set chromeUserDataDir and chromeExecutable if Chrome is installed in a custom location, as the server will otherwise auto-detect them for your current platform.


Unified Commands

This project uses a Node.js dispatcher script under the hood, allowing you to run the same npm commands on Windows, macOS, or Linux. The runner will automatically execute the correct scripts (.ps1 for Windows, .sh for Unix).

1. Start Chrome with CDP debugging

npm run start-browser

This launches Chrome with your configured profile and --remote-debugging-port=9222. A Chrome window will open — leave it running.

2. Start the MCP server (optional manual test)

In a separate terminal:

npm run start-mcp

This checks Node is installed, checks Chrome's CDP port is responding, and runs the MCP server. Press Ctrl+C to stop it.

3. Connect to AI Agents

You can automatically configure OpenCode and Gemini CLI in one step:

npm run register

This will automatically find their configuration files and add the Google-Flow MCP server.

(You can also target specific clients: npm run register -- --opencode or --gemini)

Claude Desktop Currently, Claude Desktop requires manual configuration: Add the following to your claude_desktop_config.json (run npm run build first — dist/ is what ships):

{
  "mcpServers": {
    "Google-Flow": {
      "command": "node",
      "args": ["/absolute/path/to/Google-Flow_MCP/dist/index.js"]
    }
  }
}

Cursor (Codex) Currently, Cursor requires manual UI configuration:

  1. Open Cursor Settings > Features > MCP.

  2. Click + Add New MCP Server.

  3. Name: Google-Flow

  4. Type: command

  5. Command: node /absolute/path/to/Google-Flow_MCP/dist/index.js (run npm run build first; src/index.js works for dev only)

4. Verify everything end-to-end

With Chrome running (step 1), run unit tests plus a live smoke test:

npm test            # vitest unit tests (validation, queue, sanitizers)
npm run test:smoke  # flow_connect over stdio (needs Chrome running)
npm run build       # bundles src/ -> dist/ (required before register)

Typical workflow

  1. npm run start-browser — once per session, leave the Chrome window open

  2. Your AI agent (with this MCP registered) calls:

    • flow_connect — attach to the running Chrome (also refreshes the live model catalog)

    • flow_account_check — confirm the right Google account is signed in

    • flow_list_models — list live video/image models, ratios, durations (call once per session before generating)

    • flow_generate_image — prepare (or with auto_confirm:true, generate) images in a Flow project

    • flow_create_character — create reusable characters

    • flow_generate_video — prepare (or with confirm_generate:true, paid-generate) videos, optionally referencing previously generated images/characters via ingredients (@name references)

    • flow_list_mention_options — see what images/characters are available to reference by name

    • flow_create_scene — create a new scene, optionally referencing characters

The MCP server automatically creates a Google Flow project on first use and reuses it for the rest of the session (tracked by project ID), so everything ends up in one place instead of a new project per request.


Available tools

Tool

Purpose

flow_connect

Connect to / launch Chrome via CDP, optionally open Flow

flow_disconnect

Close the browser connection

flow_open

Navigate the connected browser to a Flow URL without reconnecting

flow_status

Report current connection/page status

flow_account_check

Verify the signed-in Google account matches expectedAccount

flow_discover_ui

Navigate to a Flow page and dump interactive elements (debugging)

flow_generate_image

Generate image(s) from a prompt in the current project (prepare-only by default; auto_confirm:true spends credits)

flow_generate_video

Generate video(s), optionally with ingredients/use_character/use_scene references (prepare-only by default; confirm_generate:true spends credits)

flow_list_models

List video/image models, ratios, durations from Flow UI (live) or cache

flow_download_latest

Download the most recently generated asset

flow_create_character

Create a new character (name + description + reference images)

flow_import_character

Import a character from a saved JSON file

flow_open_characters

Open the Characters page for a project

flow_list_mention_options

List images/characters available for @name references

flow_create_scene

Create a new scene, optionally referencing characters

flow_open_tools_gallery

Open Flow's Tools gallery

flow_use_grid_architect

Open the Grid Architect tool

flow_use_tool

Use an arbitrary tool from the Tools gallery

flow_screenshot

Take a debug screenshot of the current page

flow_queue_status

Check the status of queued generation jobs

flow_queue_reset

Forcefully reset the job queue if a job is stuck in "running" state


Configuration reference

All settings live in config/flow.config.json (copy from flow.config.example.json). Key fields:

Field

Description

Default

expectedAccount

Google account email flow_account_check expects

(required)

chromeExecutable

Full path to chrome executable (optional, auto-detected)

auto-detected

chromeUserDataDir

Path to Chrome's "User Data" folder

auto-detected per-OS

chromeProfile

Profile folder name (e.g. Default, Profile 1)

Default

cdpPort

Chrome DevTools Protocol port

9222

flowUrl

Base Google Flow URL

https://flow.google.com/

headless

Run Chrome headless (not recommended — Flow needs visible browser)

false

jobTimeoutMs

Watchdog: max time a job may stay running before auto-fail

300000

jobHistoryLimit

Bounded queue history kept in flow_queue_status

50

actionDelayMs

Delay between automated UI steps

800

agentResponseTimeoutMs / generationTimeoutMs

Agent dialog window / generation wait

5000 / 120000

downloadWaitMs

Wait for download to complete

30000

imageModels / videoModels

Display name → internal model ID maps (merged with live discovery)

see example config

ratios / videoRatios / durations (4s,6s,8s,10s) / quantities

Allowed generation parameters

see example config

Env overrides win over the file: FLOW_URL, FLOW_EXPECTED_ACCOUNT, FLOW_CHROME_PROFILE, FLOW_CDP_PORT, FLOW_HEADLESS, FLOW_JOB_TIMEOUT_MS, FLOW_CHROME_EXECUTABLE, FLOW_CHROME_USER_DATA_DIR, FLOW_HOME.


Troubleshooting

"Chrome not found" Make sure Chrome is installed normally. If it is in a custom path, set chromeExecutable in config/flow.config.json.

"Chrome profile not found" Double-check chromeUserDataDir + chromeProfile against chrome://version in your browser.

"CDP port 9222 not responding" Run npm run start-browser first and leave that Chrome window open.

Google sign-in gets blocked ("This browser may not be secure") This is why the server connects to your existing signed-in Chrome profile. Ensure chromeUserDataDir / chromeProfile points to the profile you logged in with.

The agent creates a new Flow project every time The server tracks one project per session by ID and reuses it. If this happens, check the logs for Session project no longer reachable and start a fresh session.


Project layout

config/
  flow.config.example.json   # template — copy to flow.config.json (gitignored)
  flow.models.json           # live model catalog cache (gitignored, regenerated)
  selectors.map.json         # self-healing selector cache from flow_discover_ui (gitignored)
scripts/
  run.js                     # OS-independent dispatcher (prefers pwsh on Windows)
  start-browser.ps1 / .sh    # launch Chrome with CDP debugging
  start-mcp.ps1 / .sh / .bat # run the MCP server (manual check)
  register.ps1 / .sh         # register this server (prefers dist/, falls back to src/)
  test-flow-image.ps1 / .sh  # quick flow_connect smoke test (npm run test:smoke)
src/
  index.js                   # MCP server entry point (validated with zod, fail-closed errors)
  browser/                   # Chrome/CDP connection (non-destructive reuse, liveness)
  navigation/                # project registry, @ mention references, live model discovery
  queue/                     # single-job queue with watchdog + bounded history
  tools/                     # one file per MCP tool
  utils/                     # config, logger (stderr-only), screenshots, sanitize, validate,
                             # prompt QC, dynamic model universe, selector registry
tests/                       # vitest unit tests: sanitize, validate, job-queue,
                             # file-manager, video-quality, dynamic-catalog, prompt-bar

Safety notes

  • This server only automates a browser you already control and are signed into — it does not store, transmit, or need your Google credentials.

  • Separate from credentials: the server deliberately minimizes automation fingerprints — it launches Chrome with --disable-blink-features=AutomationControlled (so navigator.webdriver reads false) and runs your signed-in profile from a temp copy. That is bot-detection evasion against Google's own sign-in/abuse checks (the "this browser may not be secure" block), distinct from the credential story above. By using this tool you accept that tradeoff relative to Google's Terms of Service.

  • Image and video generation consume Google Flow credits. All mutating tools default to safe "prepare only" behavior (auto_confirm: false; video live-generate additionally requires confirm_generate: true). flow_create_character also defaults to false.

  • flow_disconnect never kills your personal Chrome tabs — it detaches unless the server launched Chrome itself.

  • flow_account_check is fail-closed: without positive evidence it returns verified:false, needsManualCheck:true (never assumed:true).

  • config/flow.config.json is gitignored — do not commit it. Runtime caches (flow.models.json, selectors.map.json, flow.projects.json) are gitignored too.


Recent changes (v1.0.1)

  • Probe reliability — flow_discover_ui now waits for panel content, not just the container element, so it no longer returns an empty element list while the panel is still rendering.

  • Import hygiene — added a test that fails if src/ imports from dist/, keeping the bundle boundary clean.

  • Lint gate — ESLint config (no-unused-vars, no-undef) added, clearing 19 findings; run with npm run lint.

  • Debug-trail retention — flow_connect keeps recent debug context on failure so it's available in flow_status after an error, instead of being swept immediately.

  • CI — GitHub Actions workflow runs npm ci, npm test, and npm run build on push and pull requests.

Test suite: 127 tests across 21 files (npm test).


Contributing

  1. Fork the repo and create a branch: git checkout -b my-change

  2. Make your change, then verify: npm run lint && npm test && npm run build

  3. Commit and open a pull request

CI runs the same three commands, so a green local run means a green PR.

Please don't commit config/flow.config.json or anything containing real Google account details — both are gitignored by design.


License

This project is licensed under the MIT License.

Available Tools

21 tools
flow_account_checkA

Verify the logged-in Google account matches the configured expected email (Default).

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A3.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It only says 'Verify' but does not disclose what happens on success/failure, whether it returns a boolean, throws an error, or has side effects. For a check tool, this lack of behavioral detail is a significant gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, well-structured sentence with leading verb and no fluff. It conveys the core purpose efficiently without wasting words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool is simple with no parameters and no output schema, the description is partially complete. It explains what it does but not what the agent should expect as a result (e.g., does it return a Boolean, throw an exception, or print a message?). The mention of 'Default' also lacks context about how the expected email is configured, leaving some ambiguity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, so the description does not need to explain any. The baseline for 0 params is 4, and the description correctly avoids adding unnecessary parameter details since the schema is empty.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states a specific action (verify) and the exact subject (logged-in Google account) against an expected configured email. This distinguishes it from sibling tools like flow_status or flow_connect, which have different purposes.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies the tool is used to ensure the correct Google account is active, but it does not explicitly state when to use it versus alternatives. For example, it doesn't mention using it after flow_connect or before other operations. Usage context is implied rather than stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

flow_connectA

Launch Chrome with the configured Google profile, connect CDP, navigate to Google Flow, and verify account.

ParametersJSON Schema
NameRequiredDescriptionDefault
headlessNoLaunch in headless mode (not recommended, Google Flow needs visible browser).
open_flowNoAuto-navigate to Google Flow after connection.

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the burden. It discloses the main side effect (launching Chrome) and the steps performed, but lacks detail on whether the connection is persistent, what happens on verification failure, or whether the browser closes after the operation. This is moderate transparency but not complete.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence that packs four specific actions without fluff. It is front-loaded with the primary action ('Launch Chrome') and efficiently communicates the full flow.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description does not mention return values (no output schema exists) or the fact that this tool likely establishes a required connection for other flow_* tools. Without annotations, it could benefit from noting that it is a prerequisite and what 'verify account' entails. This is adequate but not fully complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%: both 'headless' and 'open_flow' have clear descriptions in the schema, including a warning about headless mode. The tool description does not add additional parameter semantics, but since the schema already covers them, a baseline score of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's specific actions: launching Chrome, connecting CDP, navigating to Google Flow, and verifying the account. It distinguishes itself from sibling tools like flow_disconnect or flow_status by focusing on the connection setup.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description does not provide explicit guidance on when to use this tool relative to alternatives. It implies it is a connection step but does not mention prerequisites, exclusions, or alternative tools. For example, it does not clarify whether this should be run before flow_status or flow_discover_ui.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

flow_create_characterA

Create a new character in Google Flow Characters with name and description. By default (auto_confirm=false): fills the description and returns "ready_for_confirmation" without submitting (no credits). When auto_confirm=true: fills and clicks submit, then renames "Untitled Character" to the given name.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYesCharacter name.
campaignNoCampaign identifier for project matching (e.g., "summer-2026", "new-collection").
descriptionYesCharacter description/prompt.
auto_confirmNoIf false (default): only fills the description and returns "ready_for_confirmation" without submitting. If true: submits (may use credits).
project_nameNoName for the project (will reuse existing project with same campaign, or create new).
reference_imagesNoPaths to reference images for character design.

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does well: it discloses the non-submitting default, the credit consumption risk on submit, and the post-submit rename of 'Untitled Character'. It omits auth/permission requirements and error behavior, but the mutation semantics and side effects are clearly stated.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences, front-loaded with the core action and then the branching behavior. Dense but every clause carries information; only the repeated auto_confirm restatement is mildly redundant with the schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 6-parameter mutation tool with no annotations and no output schema, the description covers the critical unknown (what happens on submit vs. not, including credits and the rename) and reveals the returned status token. The remaining params are fully documented in the schema, so the agent has enough to invoke correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so name, description, campaign, project_name, reference_images and auto_confirm are already documented. The description's explanation of auto_confirm largely restates the schema, adding no format or syntax detail beyond it; baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Create a new character in Google Flow Characters') and the required inputs, which cleanly separates it from siblings like flow_import_character and flow_create_scene.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explains the two operating modes and their consequences (default fills only and returns ready_for_confirmation with no credits; auto_confirm=true submits), which is exactly the when-to-use decision an agent must make. It does not name sibling alternatives such as flow_import_character for existing characters.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

flow_create_sceneB

Create a new scene in Google Flow Scenes with characters and prompt. By default (auto_confirm=false) prepares only; set auto_confirm=true to actually create.

ParametersJSON Schema
NameRequiredDescriptionDefault
promptYesScene description/prompt.
campaignNoCampaign identifier for project matching (e.g., "summer-2026", "new-collection").
charactersNoCharacter names to include in the scene.
auto_confirmNoIf false (default): prepare only. If true: click New Scene and create.
project_nameNoName for the project (will reuse existing project with same campaign, or create new).
reference_imageNoPath to a reference image (optional).

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It does disclose one genuinely valuable trait — that the default is a non-mutating 'prepare only' state and auto_confirm=true performs the real creation — but says nothing about permissions, side effects, idempotency, or what 'prepares only' actually produces or how it is later confirmed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two compact sentences with the core action, then the mode caveat, both front-loaded. Slightly awkward parenthetical phrasing keeps it from being a model of efficiency, but there is no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a six-parameter creation tool with no output schema and no annotations, the description leaves important questions open: what a 'prepared only' invocation returns, how project/campaign reuse interacts with scene creation, and what happens on failure. The rich schema compensates for parameter gaps, but the workflow behavior is under-explained overall.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all six parameters (prompt, campaign, characters, auto_confirm, project_name, reference_image) are already documented in the schema. The description restates characters/prompt and the auto_confirm behavior, adding no syntax, format, or constraint detail beyond what the schema already supplies; baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: 'Create a new scene in Google Flow Scenes', and names the key inputs (characters, prompt). It is clearly distinguishable from non-creative siblings like flow_status or flow_list_models, though it does not explicitly contrast itself with the nearby flow_create_character.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explains the auto_confirm mode switch (prepare-only vs. actually create), which implies a two-step workflow, but gives no guidance on when to prefer this tool over flow_generate_video, flow_create_character, or the other creation siblings. Usage context is only partially inferred.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

flow_disconnectA

Close the browser and clean up the MCP connection to Google Flow.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description declares the primary actions (closing browser, cleaning up connection) but does not disclose side effects such as terminating active sessions or preventing further operations. This is adequate for a simple cleanup tool but lacks deeper behavioral context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, well-structured sentence that directly states the tool's function without unnecessary words. It is perfectly concise and front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a parameterless tool with no output schema or annotations, the description provides the essential information about its role as a disconnect/cleanup operation. It lacks explicit usage context but is otherwise complete for its simplicity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters and an empty schema, so the description adds nothing about parameters. Since there are no parameters to explain, the baseline of 4 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses specific verbs 'Close' and 'clean up' to describe the resource (browser and MCP connection), clearly indicating this is the teardown counterpart to flow_connect. It distinguishes itself from sibling tools by its focus on disconnection and cleanup.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description does not explicitly state when to use the tool or contrast it with alternatives like flow_connect. Usage must be inferred from the name and context, so the agent receives no direct guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

flow_discover_uiB

Navigate to a Google Flow page and discover all interactive elements (buttons, inputs, links, headings). Updates the internal selectors map for robust automation.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlNoExact https://flow.google.com/... URL to navigate to, overriding page-based resolution. Other hosts are rejected.
pageNoPage to discover. Options: main, image-generation, video-generation, characters, scenes, images, all-media, tools-gallery, grid-architect. "characters" is project-scoped (/project/{id}/characters). "scenes"/"images"/"all-media" are sidebar tabs within the project (same URL). project_name/campaign target a specific project, or the current/active project is used.main
campaignNoCampaign identifier for project matching, used when page is "characters", "scenes", "images", or "all-media".
project_nameNoProject name, used when page is "characters", "scenes", "images", or "all-media" (will reuse existing project with same campaign, or create new).

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does disclose an important non-obvious side effect: it 'Updates the internal selectors map.' That mutation info is valuable. But it omits auth requirements, browser-state effects, error behavior, and how the result is surfaced, leaving substantial gaps.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences with the action front-loaded and the side effect trailing. Every clause earns its place with no redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

All four parameters are fully documented in the schema, but with no output schema the description does not explain what 'discover' returns (a list of elements? just the internal map?) or how the agent accesses the discovered elements. Adequate but leaves the return/usage outcome underspecified.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents url, page, campaign, and project_name in detail, including enum-like options and host restrictions. The description adds no parameter meaning beyond that, so baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Specific verb+resource: 'Navigate to a Google Flow page and discover all interactive elements (buttons, inputs, links, headings),' which clearly describes the action and its scope. However it does not distinguish itself from sibling flow_open, which also appears to navigate to Flow pages, leaving the agent to guess which to select.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description hints at purpose ('for robust automation') but gives no explicit when-to-use, when-not-to-use, or alternatives. It never explains how it relates to flow_open or when discovering UI should precede another sibling like flow_generate_image.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

flow_download_latestB

Download the most recently generated file from Google Flow.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

B3.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden of behavioral disclosure. It only states select 'most recently generated' but does not explain whether the file is returned as content, a URL, or saved locally, nor any side effects or permissions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, clear sentence with no filler. It is front-loaded with the main action and resource, making it highly concise and easy to process.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with no output schema and no annotations, the description is incomplete. It does not specify what the agent should expect as a result (file path, binary content, URL) or how this tool fits into the flow with sibling generation tools (e.g., used after flow_generate_image).

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has zero parameters, so the description need not elaborate on parameter meaning. The description implicitly communicates that no user input is required, aligning with the baseline for a parameterless tool.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb 'Download' and a clear resource 'the most recently generated file from Google Flow.' It distinguishes itself from sibling tools (e.g., flow_generate_image, flow_connect) by being the only download-focused tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives, prerequisites, or typical usage contexts. The description simply states the action without explaining when it is appropriate or what conditions must exist (e.g., a file must have been generated first).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

flow_generate_imageA

⚠️ THESE IMAGES CONSUME CREDITS. By default (auto_confirm=false): fills the prompt, selects model/ratio, takes a screenshot and returns "ready_for_confirmation". Does NOT click Generate. When auto_confirm=true: first verifies the UI is in IMAGE mode (not Video), that the model is an image model, takes a verification screenshot, THEN clicks Generate and waits for the images. NANO/BANANA image models only.

ParametersJSON Schema
NameRequiredDescriptionDefault
modelNoModel to use: Nano Banana Pro, Nano Banana 2, or Imagen 4.Nano Banana 2
ratioNoAspect ratio: 1:1, 16:9, 9:16, 4:3, 3:4.1:1
promptYesThe text prompt for image generation.
campaignNoCampaign identifier for project matching (e.g., "summer-2026", "new-collection").
use_sceneNoName of a single project scene to reference via "@name" (added in addition to ingredients).
ingredientsNoNames of existing project images/characters to reference via "@name" (e.g., ["Bob the Astronaut", "Image 3"]). Use flow_list_mention_options to discover available names.
auto_confirmNo⚠️ CREDITS. If false (default): only prepares, consumes nothing. If true: verifies that Image mode is active, THEN clicks Generate (consumes credits).
project_nameNoName for the project (will reuse existing project with same campaign, or create new).
use_characterNoName of a single project character to reference via "@name" (added in addition to ingredients).
reference_imagesNoPaths to local reference images to upload (optional).

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden and does so well: it discloses credit consumption, the exact steps taken in each mode (fill prompt, select model/ratio, screenshot, return 'ready_for_confirmation'), that Generate is NOT clicked by default, the pre-flight verification of mode and model, and the wait-for-images behavior. This is far richer than a typical schema would convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The credit warning is front-loaded, and the two modes are laid out as parallel clauses rather than repetition. Density is high with virtually no filler despite covering a complex flow.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 10-parameter, side-effectful generation tool with no annotations and no output schema, the description covers the critical unknowns: cost implications, what is and isn't executed per mode, and the partial return value ('ready_for_confirmation'). Nothing essential to safe invocation is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so every parameter is already documented with defaults, formats, and examples. The description reinforces auto_confirm's semantics but adds no parameter meaning beyond the schema. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource (generate image) and explicitly scopes the model family ('NANO/BANANA image models only'), which cleanly separates it from sibling flow_generate_video. The two-mode behavior is front-loaded so an agent knows immediately what the tool does.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly contrasts the two operating modes: auto_confirm=false prepares only, auto_confirm=true verifies and clicks Generate. It also specifies prerequisites (UI must be in IMAGE mode, model must be an image model). It does not name sibling alternatives for video generation, but the scoping note covers the main branch point.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

flow_generate_videoB

Prepare a Veo video (prepare-only, no spend). Fast tests, Quality finals; ingredients need 8s.

ParametersJSON Schema
NameRequiredDescriptionDefault
modelNoVideo model name — call flow_list_models for the live list — or auto.auto
ratioNo16:9 or 9:16.16:9
shotsNoMulti-shot plan, max 6: [{prompt, duration}].
styleNoLook/mood/lighting. No extra subjects in refs.
cameraNoCamera move: push-in, pull-back, pan, tilt, orbit, tracking shot.
promptYesScene + action. Optics stay in UI selectors.
qualityNoResolution 360p/720p.
campaignNoCampaign id (doubles as project name when name absent).
durationNo4s, 6s, 8s, 10s. Ingredients force 8s.4s
quantityNoVariants 1-4 for compare-pick-refine.
end_frameNoLast-frame image path. Shares light/framing with start.
use_sceneNoName of a single project scene to reference via "@name" (added in addition to ingredients).
hero_frameNoAlias for start_frame.
ingredientsNoProject images/characters to reference via "@name" (e.g., ["Bob the Astronaut", "Image 3"]). Discover names via flow_list_mention_options.
project_urlNoDirect Flow /project/ URL — wins over name matching.
start_frameNoFirst-frame image path (hero frame anchor).
anchor_frameNoPrior strongest frame carried for continuity.
auto_confirmNoAlways prepare-only today. Reserved.
project_nameNoProject name (remembered; reopens instead of duplicating).
use_characterNoName of a single project character to reference via "@name" (added in addition to ingredients).
confirm_generateNoLIVE PAID CLICK: verify, click Generate (credits), poll video. Default false.
reference_imagesNoPaths to local reference images to upload (optional).

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It usefully discloses the default no-spend prepare behavior, which counteracts the misleading 'generate' name, but it says nothing about the paid confirm_generate path, required auth, or polling/credit consumption beyond what the schema param already states.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three front-loaded fragments lead with the single most important behavior (prepare-only, no spend) and waste no words. The terseness makes 'Fast tests, Quality finals' slightly cryptic, but nothing is padding.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 22-parameter, annotation-free tool with no output schema, the description covers the core safety profile but is thin on the surrounding workflow (how it pairs with flow_list_models, flow_list_mention_options, or the confirm/poll path). The rich schema compensates for parameters, but behavioral coverage is minimal.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all 22 parameters are already documented, giving a baseline of 3. The description adds only a marginal cross-parameter rule ('ingredients need 8s', quality tiers) beyond what the individual param descriptions already say.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Prepare) and resource (Veo video) and immediately disambiguates against its own name by noting 'prepare-only, no spend'. An agent can tell it produces a video plan rather than a video, though it never names the sibling flow_generate_image or explains how it relates to the image tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Fast tests, Quality finals' implies a quality-tier decision and 'ingredients need 8s' gives a conditional constraint, but there is no explicit when-to-use vs alternatives, no exclusion of the live-generate path, and no prerequisites (auth, project) stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

flow_import_characterA

Import a character from a saved JSON file into Google Flow.

ParametersJSON Schema
NameRequiredDescriptionDefault
campaignNoCampaign identifier for project matching (e.g., "summer-2026", "new-collection").
file_pathYesPath to character JSON file.
project_nameNoName for the project (will reuse existing project with same campaign, or create new).

TDQS

A3.5/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description must carry the full burden of behavioral disclosure. It only states the high-level action, without detailing side effects (e.g., overwriting existing characters), validation of the JSON, error handling, or any prerequisites. This is a significant gap for an import tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence with no filler or redundant information. Every word earns its place, making it highly concise and easy to parse.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The core purpose is clear and the schema covers parameters, but the description lacks usage guidance relative to sibling tools and does not disclose behavioral traits (e.g., whether it overwrites or merges). For a tool with no annotations and no output schema, this is adequate but leaves gaps for an agent deciding between import and create.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema provides descriptions for all three parameters (campaign, file_path, project_name), giving 100% schema coverage. The description itself adds no parameter-specific detail, so the baseline of 3 is appropriate since the schema handles parameter semantics.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Import') with a clear resource ('character') and source/destination ('saved JSON file' into 'Google Flow'). This clearly distinguishes it from siblings like flow_create_character (which creates from scratch) and flow_open_characters (which opens existing characters).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage when a saved JSON character file exists, but it offers no explicit when-to-use or when-not-to-use guidance, and does not mention alternative tools such as flow_create_character. The context is inferred from the phrase 'saved JSON file' rather than stated directly.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

flow_list_mention_optionsA

List the images and characters in the current project that can be referenced via "@name" in an image/video prompt (opens Flow's "@" reference popup and reads its options).

ParametersJSON Schema
NameRequiredDescriptionDefault
campaignNoCampaign identifier for project matching (e.g., "summer-2026", "new-collection").
project_nameNoName for the project (will reuse existing project with same campaign, or create new).

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the burden of disclosure. It transparently states the side effect of opening Flow's '@' reference popup and reading its options, which goes beyond a generic 'list' description. It does not detail safety permissions, but the read-only nature is inferred and no harmful behavior is hidden.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence that is both concise and information-dense. It includes the action, target, mechanism, and context without redundancy or filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description sufficiently explains the tool's purpose and behavior for a listing tool. It lacks an explicit return format, but that is partially mitigated by the absence of an output schema and the simplicity of the feature. It is complete enough for an agent to invoke correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema already provides full descriptions for both 'campaign' and 'project_name' (100% coverage). The tool description adds no additional parameter meaning, so the baseline score of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('List'), the resource ('images and characters'), and the context ('current project' and 'via @name in an image/video prompt'). It distinguishes this from sibling tools like flow_open_characters by specifying the popup-reading mechanism.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear context on when to use the tool (when needing @name references for image/video prompts) and implies the popup interaction. It does not explicitly name alternatives or exclusions, so it stops short of a 5 but is clear enough for an agent.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

flow_list_modelsA

List video/image models, ratios, durations from Flow UI (live) or cache. Call once per session before generating.

ParametersJSON Schema
NameRequiredDescriptionDefault
refreshNoForce re-discovery from the open UI instead of cache.

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It discloses the live-vs-cache duality and the once-per-session constraint, but omits whether a connected/open Flow UI is required, what the call costs, and whether results are exhaustive. For a read-only discovery tool this is minimally adequate rather than rich.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, no filler. The resource list is front-loaded and the calling constraint follows immediately, so an agent can act after one read.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-parameter listing tool with full schema coverage and no output schema, the description covers what is returned (models, ratios, durations) and when to call it. The remaining gap is the undeclared dependency on a live Flow UI session, which matters for a browser-driven tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the single 'refresh' parameter is fully documented in the schema. The description's 'live or cache' phrasing mirrors that schema text rather than adding syntax or format detail, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('List') and concrete resources (video/image models, ratios, durations), and scopes the source (Flow UI live or cache). It does not name a sibling explicitly, but 'before generating' clearly separates it from the flow_generate_* tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit timing guidance: 'Call once per session before generating.' That is a clear when-to-use directive tied to the generate workflow, though it names no exclusions or alternatives (e.g., what to do if the cache is stale besides refresh).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

flow_openA

Navigate the already-connected browser to a Google Flow URL without reconnecting. Only https://flow.google.com/... URLs are allowed (other hosts rejected).

ParametersJSON Schema
NameRequiredDescriptionDefault
urlNoExact Flow URL to navigate to. Defaults to the configured flowUrl when omitted.
waitForNoExtra ms to wait after load (clamped 0-15000).

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations the description carries the full burden, and it does disclose real behavior: navigation reuses the existing connection and a host allowlist rejects anything that is not https://flow.google.com/... It omits what happens if no session exists or if navigation fails, so it is not exhaustive.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, zero filler, with the core action front-loaded and the host restriction immediately after. Every clause carries information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a low-complexity navigation tool with only two well-documented parameters, the description covers the action, the precondition, and the input constraint. It does not say what the call returns or how failures surface, which is the only remaining gap given there is no output schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% so baseline is 3, but the description adds a genuine constraint on the url parameter (only https://flow.google.com/... accepted, other hosts rejected) that the schema does not state. The waitFor parameter is left entirely to the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Specific verb (navigate) plus resource (already-connected browser to a Flow URL), and the phrase 'without reconnecting' implicitly separates it from flow_connect and flow_open_tools_gallery. An agent can select it correctly without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The precondition 'already-connected browser' and 'without reconnecting' tell the agent this is for an existing session rather than establishing one, which routes it away from flow_connect. It never names the alternative explicitly, so the routing is inferred rather than stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

flow_open_charactersB

Open the Characters page for the current/target project and list existing characters.

ParametersJSON Schema
NameRequiredDescriptionDefault
campaignNoCampaign identifier for project matching (e.g., "summer-2026", "new-collection").
project_nameNoName for the project (will reuse existing project with same campaign, or create new).

TDQS

B3.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries full burden. It states the tool opens a page and lists characters but does not disclose potential side effects, such as creating a new project if none matches (as implied by the schema's project_name description). Lacks details on permissions or reversibility.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Single sentence, front-loaded with action and purpose. No wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description is adequate for a simple navigation/listing tool but lacks context on project selection mechanism and output nature. With no output schema or annotations, it could benefit from explaining the relationship between campaign and project_name and whether the list is returned as data.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema has 100% coverage, so baseline is 3. The description adds no parameter-specific information; it relies entirely on schema descriptions for campaign and project_name.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description uses specific verb 'open' and resource 'Characters page', and further clarifies it lists existing characters. This clearly distinguishes it from sibling tools like flow_create_character and flow_import_character, which perform different operations.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool versus alternatives such as flow_create_character or flow_import_character. The description implies usage for viewing existing characters but doesn't state exclusions or prerequisites.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

flow_queue_resetA

Forcefully reset the job queue. Use this if a job is permanently stuck in "running" state and blocking other generations.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description must convey behavioral traits. It states the action is 'forceful' and targets stuck jobs, implying a destructive/resetting behavior. However, it does not disclose potential side effects (e.g., loss of queued jobs, irreversibility, or required permissions), which are important for a reset operation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences long, with the primary action front-loaded in the first sentence. The second sentence adds valuable usage context without any fluff or redundancy. Every word earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple, zero-parameter tool, this description covers the purpose and the specific situation. There is no output schema, and the tool's return behavior is not described, but for a queue reset, the main point is the action itself. A small note about consequences (e.g., pending jobs may be cleared) would make it more complete, but it's adequate as is.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, and the input schema is an empty object. Per the rubric, 0 params is a baseline of 4. The description correctly implies no configuration is needed, so no additional parameter explanation is necessary.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('Forcefully reset the job queue') and the specific resource (the job queue). It also provides a concrete scenario (job stuck in 'running' state blocking others), which distinguishes it from sibling tools like flow_queue_status that likely just query status.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly states when to use the tool: 'if a job is permanently stuck in "running" state and blocking other generations.' This gives clear context, though it does not mention when not to use it or explicitly name alternatives. Still, the use case is well-defined.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

flow_queue_statusA

Check the job queue: active job, pending queue, completed and failed job history.

ParametersJSON Schema
NameRequiredDescriptionDefault
history_limitNoNumber of recent history entries to return.

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Since no annotations are provided, the description carries the full burden. It clearly indicates a read-only operation ('Check') and lists the information returned, but it does not disclose any potential side effects, authentication requirements, or details about how history_limit affects results. It is minimally transparent but not deeply descriptive.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single concise sentence that front-loads the primary action and enumerates the specific aspects of the queue. There is no wasted information; every word earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple tool with one optional parameter and no output schema, the description adequately captures the main functionality. It names all categories of queue status without needing to explain return values extensively. The only minor gap is not explicitly stating that history_limit controls the number of history entries, but that is already in the schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema provides 100% coverage for the only parameter (history_limit) with a clear description. The tool description does not add additional meaning beyond the schema, but since schema coverage is high, the baseline of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Check') and resource ('job queue'), and enumerates the exact information covered (active job, pending queue, completed and failed job history). This clearly distinguishes it from sibling tools like flow_status, which likely covers broader status, and other flow_* tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit guidance on when to use this tool versus alternatives such as flow_status. The description implies that it is for checking queue information, but it does not state exclusions or mention alternative tools, leaving the agent without clear decision criteria.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

flow_screenshotB

Take a screenshot of the current Google Flow page.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameNoCustom name for the screenshot file.manual

TDQS

B3.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of disclosing behavioral details. It only states the action without mentioning where the screenshot is saved, whether it returns a file path, authentication requirements, or any side effects, leaving significant ambiguity for an agent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence with no filler, immediately conveying the tool's action. It is appropriately concise and front-loaded, making it easy for an agent to parse.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's simplicity (one optional param, no output schema, no annotations), the description is minimally adequate but incomplete. It fails to explain what happens after the screenshot is taken (e.g., file location, return value), leaving the agent without enough context to predict the tool's full behavior.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The single optional parameter 'name' is fully described in the input schema (100% coverage), so the baseline is 3. The description adds no additional semantic detail about the parameter, but the schema already provides sufficient information.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states a specific action ('Take a screenshot') on a specific resource ('the current Google Flow page'), making its purpose unambiguous. This distinguishes it from sibling tools like flow_download_latest or flow_create_scene, which serve different actions.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description does not provide explicit guidance on when to use this tool versus alternatives. However, the intended use is implied by the action itself — taking a screenshot when a visual capture is needed — but no exclusions or alternative comparisons are mentioned.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

flow_statusB

Check current connection status: browser connected, Flow page loaded, account verified, job queue state.

ParametersJSON Schema
NameRequiredDescriptionDefault
fullNoReturn full status with screenshot.

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It transparently lists the exact status dimensions checked (browser, page, account, queue), which is useful. However, it doesn't disclose whether this is a read-only operation, how the response is structured, or whether it performs live checks or returns cached state.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence of 12 words, front-loaded with the primary action and resource. Every word adds value, with no redundancy or filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simplistic status-check tool with one fully described parameter, the description covers core purpose and scope. However, without an output schema or annotations, it omits details about return format, potential latency, and any safety characteristics, leaving the agent uncertain about how to interpret the result.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema description fully documents the 'full' parameter ('Return full status with screenshot'), achieving 100% coverage. The tool description adds no parameter-specific context, so the baseline score of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses the specific verb 'Check' and clearly enumerates the status components (browser connected, Flow page loaded, account verified, job queue state). It implies a broader scope than individual siblings like flow_account_check or flow_queue_status but does not explicitly name them for differentiation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives. It doesn't mention that flow_queue_status should be used for queue-only details or flow_connect for establishing a connection. The agent is left without explicit selection criteria.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

flow_use_grid_architectB

Open Grid Architect in Google Flow, fill theme prompt, shot prompts, engine, ratio, and visual logic settings. Supports batch shot generation for brand campaigns. Prepare-only by default (auto_confirm=false).

ParametersJSON Schema
NameRequiredDescriptionDefault
ratioNoAspect ratio for all shots.16:9
engineNoEngine/model for the grid.Nano Banana 2
campaignNoCampaign identifier for project matching (e.g., "summer-2026", "new-collection").
referencesNoPaths to reference images.
auto_confirmNoReserved — currently prepare-only.
project_nameNoName for the project (will reuse existing project with same campaign, or create new).
shot_promptsNoArray of individual shot prompts for the grid.
theme_promptYesOverall theme prompt for the grid.
visual_logicNoVisual logic type: None, Colour Pop, Side by Side, etc.

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden, and it does disclose one important trait: 'Prepare-only by default (auto_confirm=false)'. It also hints at project reuse via the schema, but it never says what happens after preparation, whether a separate confirm step exists, or what permissions/account state are required — significant gaps for an unattended UI-driving tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences with the core action front-loaded and the prepare-only caveat placed last where it matters most. No filler, though the middle sentence's field enumeration is largely redundant with the schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 9-parameter, no-annotation, no-output-schema tool, the description covers the primary flow and the safety default but omits the post-prepare lifecycle (what the caller does next) and any notion of the return value. It is minimally sufficient but leaves the agent guessing about the step after invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description restates theme prompt, shot prompts, engine, ratio and visual logic — all already documented in the schema — and adds no format or constraint detail (e.g., valid engine names, visual_logic enum values, reference image path expectations).

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Open Grid Architect in Google Flow') and enumerates the fields it fills, so the agent knows exactly what operation this performs. It implies batch shot generation, which differentiates it somewhat from flow_generate_image, but it never names a sibling or explicitly contrasts the two.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Supports batch shot generation for brand campaigns' gives an implied use case, but there is no explicit when-to-use/when-not guidance and no mention of alternatives such as flow_generate_image or flow_create_scene for single-shot work. The agent must infer routing from the batch/campaign phrasing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

flow_use_toolB

Open any tool by name in Google Flow and optionally fill its configuration parameters. Prepare-only by default (auto_confirm=false).

ParametersJSON Schema
NameRequiredDescriptionDefault
paramsNoOptional configuration parameters for the tool.
campaignNoCampaign identifier for project matching (e.g., "summer-2026", "new-collection").
tool_nameYesName of the tool to open (e.g. Grid Architect, Image Generation).
auto_confirmNoReserved — currently prepare-only.
project_nameNoName for the project (will reuse existing project with same campaign, or create new).

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden. It does disclose one genuinely important trait: the tool is prepare-only unless auto_confirm is set, so the agent knows opening a tool is non-destructive by default. It omits other relevant behavior — whether a connected session/project is required, what state the opened tool leaves behind, and what happens to params that don't match the target tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences that are front-loaded with the core action, followed immediately by the most consequential behavioral caveat. The clause about optionally filling configuration parameters is mildly redundant with the schema, but there is no padding.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 5-parameter tool (with a nested object) and no output schema and no annotations, the description is only minimally sufficient. It does not explain what an agent gets back after opening a tool, nor the relationship between project_name/campaign reuse and an already-open session — details the agent must have since there is no return schema to fall back on.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all five parameters, including the campaign/project_name reuse semantics and the reserved auto_confirm flag. The description's only parameter-related statement ('optionally fill its configuration parameters') restates the schema rather than adding syntax, format, or validation meaning. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb ('Open') and resource ('any tool by name in Google Flow'), plus the optional parameter-filling behavior. It implicitly distinguishes itself from the sibling flow_use_grid_architect (one specific tool) by saying 'any tool', but never names an alternative explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives one useful usage signal — 'Prepare-only by default (auto_confirm=false)' — which tells the agent the call will not execute anything by default. However, there is no guidance on when to prefer this generic opener over flow_use_grid_architect, flow_open_tools_gallery, or whether flow_connect must run first. Usage is only implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 21 tool updatesv1.0.1
    • First observedflow_account_check
    • First observedflow_connect
    • First observedflow_create_character
    • First observedflow_create_scene
    • First observedflow_disconnect
    • First observedflow_discover_ui
    • First observedflow_download_latest
    • First observedflow_generate_image
    • First observedflow_generate_video
    • First observedflow_import_character
    • First observedflow_list_mention_options
    • First observedflow_list_models
    • First observedflow_open
    • First observedflow_open_characters
    • First observedflow_open_tools_gallery
    • First observedflow_queue_reset
    • First observedflow_queue_status
    • First observedflow_screenshot
    • First observedflow_status
    • First observedflow_use_grid_architect
    • First observedflow_use_tool

TDQS

A3.6/5.0

Scored across 21 tools

Disambiguation4/5

Most tools have a clearly distinct resource or action (status, connect, generate_image, create_character, etc.), so an agent can usually tell them apart. However, flow_use_tool overlaps with specific wrappers like flow_use_grid_architect, and flow_status partially duplicates flow_queue_status and flow_account_check.

Naming Consistency5/5

All tool names use a consistent flow_ prefix plus snake_case verb/noun structure. The convention is predictable throughout the set.

Tool Count3/5

21 tools is heavy for a single browser-automation server and lands in the borderline range per the rubric. While many operations are justified, some could be consolidated or deferred, especially generic flow_use_tool versus specific tool wrappers.

Completeness3/5

Core lifecycle operations are present: connect/disconnect, navigate, discover UI, generate image, create characters/scenes, and queue status/reset. But there are notable gaps: flow_generate_video is prepare-only with no execute path, and there are no update/delete operations for characters/scenes or listing generated assets beyond the latest download.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers