Skip to main content
Glama

Gemini MCP - Google Gemini image generation, video and search grounding inside Claude

npm version MCP Registry Known Vulnerabilities License: Apache 2.0

I've had this Gemini MCP server running in my Claude Desktop setup for the best part of a year now. It's one of the few I leave switched on permanently. Not because Gemini replaces Claude (it doesn't), but because grounded search, image generation, SVG diagrams and video are things Gemini happens to do well, and having them as tools inside Claude beats flipping between browser tabs.

Fourteen tools, built around the models people actually come looking for: Nano Banana Pro (gemini-3-pro-image) and Nano Banana 2 for images, Veo 3.1 for video with synchronised audio, and Gemini 3.1 Pro for chat and deep research with Google Search grounding. Everything previews inline in Claude Desktop through MCP Apps rather than landing as a file path you have to go and open.

One npx command. That's it.


Quick Navigation

What it makes | Get started | What it does | Image output | Configuration | Tools | Models | Requirements


What it makes

Everything below came out of the tools in this repo, unretouched, on the afternoon I wrote this. Prompts are in the captions so you can judge for yourself.

Image generation with search grounding. One generate_image call with use_search=true. Gemini looked up the actual Met Office forecast for the week, then drew it. The dates and temperatures are real (well, as real as a forecast gets).

London five-day weather infographic, generated by Nano Banana Pro from live search data

Prompt: "A polished editorial infographic poster: London weather, this week. Use real forecast data from search for the next 5 days..." - gemini-3-pro-image, 16:9, 2K.

Text that's actually spelled correctly. This is the thing Nano Banana Pro does that the previous generation of image models couldn't. Every label here is straight out of the model.

Four-step infographic explaining an MCP tool call, with correctly rendered text

Prompt: "A clean, magazine-quality infographic titled How an MCP tool call works, with four numbered steps..." - gemini-3-pro-image.

Generate, then edit. Left is generate_image. Right is edit_image on that file with one sentence of instructions: change the track to Monza in daylight, swap the rim lighting for window light, make the pedals red. Same cockpit, same camera angle.

generate_image

edit_image

Sim racing cockpit at night, rainy Spa on the monitor

The same cockpit edited to daytime Monza with red pedals

Image to video with Veo 3.1. The night-time cockpit above, passed to generate_video as firstFrameImage. Eight seconds at 1080p with generated audio (engine note, tyre hiss) - the GIF below is silent and squashed for GitHub, the real file is a proper MP4.

Veo 3.1 video: the cockpit comes to life, hands turning the wheel through Eau Rouge

SVG that you can actually use. Not a picture of a diagram - real vector markup you can drop into a page, edit by hand or commit to a repo. Both of these are the raw .svg files the tool wrote to disk.

generate_svg style=technical

generate_svg style=data-viz

Architecture diagram of this MCP server as an SVG

Grouped bar chart of API latency by region as an SVG

A landing page from a paragraph. generate_landing_page with a brief, a company name and a brand colour. Self-contained HTML, inline CSS, the little chart in the hero is animated SVG. Screenshot of the file opened in Chrome, nothing else touched.

Generated SaaS landing page rendered in a browser

And the fast one. Nano Banana 2 (gemini-3.1-flash-image) for when you want volume rather than 4K. This took about ten seconds.

Flat-lay of a mechanical keyboard on a walnut desk, Nano Banana 2


Related MCP server: Gemini MCP Server

Get started in two minutes

Step 1: Get a Gemini API key

Go to Google AI Studio and create one.

A word on the free tier, because the defaults lean on a paid model. Google's pricing page (as of 24 September 2026) gives Gemini 3 Flash Preview a free tier, but Gemini 3.1 Pro Preview is paid-only - and 3.1 Pro is the default for gemini_chat, gemini_deep_research (synthesis), analyze_image and generate_landing_page. On a free key those calls will fail unless you point them at a free model: set GEMINI_DEFAULT_MODEL=gemini-3-flash-preview (covers chat and landing pages), GEMINI_DEEP_RESEARCH_MODEL and GEMINI_IMAGE_ANALYSIS_MODEL likewise, or pass model per call. generate_svg already defaults to gemini-3-flash-preview. Check the pricing page for the other models before relying on them - Google changes these tiers.

Step 2: Add to your Claude Desktop config

Config file locations:

  • Windows: C:\Users\{username}\AppData\Roaming\Claude\claude_desktop_config.json

  • macOS: ~/Library/Application Support/Claude/claude_desktop_config.json

{
  "mcpServers": {
    "gemini": {
      "command": "npx",
      "args": ["@houtini/gemini-mcp"],
      "env": {
        "GEMINI_API_KEY": "your-api-key-here"
      }
    }
  }
}

Step 3: Restart Claude Desktop

That's it. Tools show up automatically. npx pulls the package on first run, so there's no separate install.

Local build instead

For development, or if you'd rather not rely on npx:

git clone https://github.com/houtini-ai/gemini-mcp
cd gemini-mcp
npm install --include=dev
npm run build

Then point your config at the local build:

{
  "mcpServers": {
    "gemini": {
      "command": "node",
      "args": ["C:/path/to/gemini-mcp/dist/index.js"],
      "env": {
        "GEMINI_API_KEY": "your-api-key-here"
      }
    }
  }
}

Claude Code (CLI)

Claude Code doesn't read claude_desktop_config.json. Use claude mcp add instead:

claude mcp add -e GEMINI_API_KEY=your-api-key-here -s user gemini -- npx -y @houtini/gemini-mcp

With an output directory for images and video:

claude mcp add \
  -e GEMINI_API_KEY=your-api-key-here \
  -e GEMINI_IMAGE_OUTPUT_DIR=/path/to/output \
  -s user \
  gemini -- npx -y @houtini/gemini-mcp

Check with claude mcp get gemini - you want to see Status: Connected.


What it does

Chat with Google Search grounding

Use gemini:gemini_chat to ask: "What changed in the MCP spec in the last month?"

Grounding is on by default. Gemini searches Google before it answers, so you get this month's information rather than a training-cutoff guess, and the sources come back as markdown links. For questions where you want pure reasoning ("explain this code", that sort of thing) set grounding: false.

Runs on gemini-3.1-pro-preview unless you say otherwise. Pass model: "gemini-3.8-flash" if you'd rather have speed than depth. thinking_level works on every Gemini 3.x model: high for the hard stuff, low to keep it snappy.

Deep research

Use gemini:gemini_deep_research with:
  research_question="What are the current approaches to AI agent memory management?"
  max_iterations=5

Runs grounded search passes and then writes them up as one report. The passes run on gemini-3.8-flash with low thinking (they're gathering facts, not reasoning about them), and the synthesis at the end runs on gemini-3.1-pro-preview with high thinking. Two passes plus a synthesis is the default and lands in two to three minutes.

That split matters more than it sounds. Earlier versions ran everything on Pro with full thinking, and one pass on a broad question could run past Claude Desktop's four-minute timeout on its own. Keep max_iterations at 2 or 3 in Claude Desktop; in an IDE or an agent framework, 5 to 7 produces noticeably better synthesis. focus_areas takes an array if you want to steer each pass.

Image generation with search grounding

Use gemini:generate_image with:
  prompt="Stock price chart showing Apple (AAPL) closing prices for the last 5 trading days"
  use_search=true
  aspectRatio="16:9"

Default model is gemini-3-pro-image, which is Nano Banana Pro now that it's out of preview. It renders legible text, does 4K, and it's the only one that supports conversational editing. gemini-3.1-flash-image (Nano Banana 2) is the fast option - not far off in quality, a fraction of the time and the cost. There's a Lite variant too if you're doing hundreds.

With use_search=true, Gemini looks things up before it draws. Weather, prices, sports results, that kind of data-driven image works reliably. The full-resolution file always goes to disk; the inline preview is resized to fit the MCP transport cap but the original is untouched.

Video generation with Veo 3.1

Use gemini:generate_video with:
  prompt="A close-up shot of a futuristic coffee machine brewing a glowing blue espresso, steam rising dramatically. Cinematic lighting."
  resolution="1080p"
  durationSeconds=8

Google's Veo 3.1. Four to eight second clips at up to 4K with native, synchronised audio. It's asynchronous on Google's side and takes two to five minutes; the tool polls until it's ready so you don't have to.

Options worth knowing about:

  • aspectRatio - 16:9 landscape or 9:16 for vertical

  • generateAudio - on by default; dialogue and effects that match the prompt

  • firstFrameImage - animate from a still (that's how the cockpit clip above was made)

  • referenceImages - up to three, for character or style consistency

  • sampleCount - up to four variations in one call

  • seed - deterministic output across runs

  • generateThumbnail - pulls a frame out with ffmpeg, if you've got it in PATH

  • model - veo-3.1-lite-generate-preview if you want cheaper and faster

SVG generation

This is the one people underestimate. The output isn't a picture of a diagram, it's the SVG markup itself - drop it into a codebase, a slide, a web page, and it scales without a single raster artefact.

Use gemini:generate_svg with:
  prompt="Architecture diagram showing a microservices system with API gateway, three services, and a shared database"
  style="technical"
  width=1000
  height=600

Four styles:

Style

Best for

technical

Architecture diagrams, flowcharts, system maps

artistic

Illustrations, decorative graphics, icons

minimal

Clean data visualisations, simple charts

data-viz

Charts, dashboards, infographics

You get real SVG code back. Edit it, animate it, embed it, commit it. No export step, no Figma.

Runs on gemini-3-flash-preview unless you pass model - it's quick, it handles SVG markup comfortably, and it has a free tier. Pass model: "gemini-3.1-pro-preview" for denser diagrams if you're on a paid key.

Image editing and analysis

Conversational editing. Nano Banana Pro keeps context between turns. Pass the thought signature from the previous call back in and it remembers what it was working on:

Use gemini:edit_image with:
  prompt="Change the colour scheme to blue and green"
  images=[{filePath: "C:/output/gemini-123.png", thoughtSignature: "fromPreviousCall"}]

Analysis, two tools for two jobs:

  • describe_image - quick general descriptions on gemini-3.8-flash

  • analyze_image - structured extraction and proper reasoning on gemini-3.1-pro-preview

Local files:

Use gemini:load_image_from_path with filePath="C:/screenshots/error.png"

Or just pass filePath straight into any image tool's images array and the server reads it itself - that skips the MCP transport limit entirely.

Media resolution control

Cut token usage by up to 75% where the task doesn't need the detail:

Level

Tokens

Savings

Best for

MEDIA_RESOLUTION_LOW

280

75%

Simple tasks, bulk operations

MEDIA_RESOLUTION_MEDIUM

560

50%

PDFs and documents (OCR saturates here)

MEDIA_RESOLUTION_HIGH

1120

default

Detailed analysis

MEDIA_RESOLUTION_ULTRA_HIGH

2000+

per-image only

Maximum detail

For PDF OCR, MEDIUM gives me identical text extraction to HIGH at half the tokens. I've not found a case where it didn't.

Landing page generation

Use gemini:generate_landing_page with:
  brief="A SaaS tool that helps developers monitor API latency"
  companyName="PingWatch"
  primaryColour="#6366F1"
  style="startup"
  sections=["hero", "features", "pricing", "cta"]

One self-contained HTML file: inline CSS, vanilla JS, no external dependencies. Styles are minimal, bold, corporate and startup. The PingWatch page in the gallery above is exactly this call.

Professional chart design systems

gemini_prompt_assistant carries nine chart design systems you can ask for by name:

System

Inspiration

Best for

storytelling

Cole Nussbaumer Knaflic

Executive presentations

financial

Financial Times

Editorial journalism - FT pink, serif titles

terminal

Bloomberg / fintech

High-density dark mode with neon

modernist

W.E.B. Du Bois

Bold geometric blocks, stark contrasts

professional

IBM Carbon / Tailwind

Enterprise dashboards

editorial

FiveThirtyEight / Economist

Data journalism

scientific

Nature / Science

Academic rigour

minimal

Edward Tufte

Maximum data-ink ratio

dark

Observable

Modern dark mode

Help system

Use gemini:gemini_help with topic="overview"

The full documentation without leaving Claude. Topics: overview, image_generation, image_editing, image_analysis, chat, deep_research, grounding, media_resolution, models, all.


Image output and storage

By default, images come back as inline previews rendered directly in Claude, and the full-size file is written next to the package. Set GEMINI_IMAGE_OUTPUT_DIR if you'd rather they all landed somewhere sensible:

"env": {
  "GEMINI_API_KEY": "your-api-key-here",
  "GEMINI_IMAGE_OUTPUT_DIR": "C:/Users/username/Pictures/gemini-output"
}

Two files per image:

File

What it is

Full-res

Saved to disk immediately, untouched

Preview

Resized JPEG for inline transport, sized to fit under the cap

Gemini returns 2 to 5 MB images. The resize measures the non-image overhead in each response, works out the binary budget left, and steps the preview down (800, 600, 400, 300, 200px) until it fits under the 1 MB MCP transport limit. The full image is always there on disk.

Inline viewers in Claude Desktop

Image, SVG, video and landing-page results each open in an MCP App viewer with zoom, the saved path and a copy button. If you're on Claude Desktop and the viewer sat on "Waiting for image..." forever in an older version, that was Claude Desktop stripping the structured data the viewer reads (ext-apps#696). Since 2.7.0 the viewer fetches it back from the server itself, so it renders either way.


Configuration reference

Variable

Required

Default

Description

GEMINI_API_KEY

Yes

-

Google AI API key from AI Studio

GEMINI_DEFAULT_MODEL

No

gemini-3.1-pro-preview

Model for gemini_chat

GEMINI_DEEP_RESEARCH_MODEL

No

gemini-3.1-pro-preview

Synthesis model for gemini_deep_research

GEMINI_DEEP_RESEARCH_SEARCH_MODEL

No

gemini-3.8-flash

Model for the grounded search passes in gemini_deep_research

GEMINI_IMAGE_ANALYSIS_MODEL

No

gemini-3.1-pro-preview

Model for analyze_image

GEMINI_IMAGE_DESCRIBE_MODEL

No

gemini-3.8-flash

Model for describe_image

GEMINI_IMAGE_GENERATION_MODEL

No

gemini-3-pro-image

Model for generate_image and edit_image

GEMINI_DEFAULT_GROUNDING

No

true

Set to false to turn Google Search grounding off by default

GEMINI_IMAGE_OUTPUT_DIR

No

-

Where generated images and videos are saved

GEMINI_ALLOW_EXPERIMENTAL

No

false

Include experimental and preview models in auto-discovery

GEMINI_REQUEST_TIMEOUT_MS

No

240000

Per-request timeout for chat and analysis calls, in milliseconds

GEMINI_MCP_RETRY_ATTEMPTS

No

3

Total attempts per Gemini API request. Transient network failures (fetch failed, ECONNRESET, proxy or VPN drops) are retried with backoff. 1 disables it

GEMINI_MCP_LOG_FILE

No

false

Write logs to ~/.gemini-mcp/logs/

DEBUG_MCP

No

false

Log to stderr for debugging tool calls

Tools reference

Tool

Description

gemini_chat

Chat with Gemini 3.1 Pro. Google Search grounding on by default. Supports thinking_level

gemini_deep_research

Grounded search passes on Flash, synthesised into a report by 3.1 Pro. Default 2 passes

gemini_list_models

Lists the models your API key can see, live

gemini_help

Documentation for every tool without leaving Claude

gemini_prompt_assistant

Expert guidance for image generation with nine chart design systems

generate_image

Image generation with optional search grounding. Full-res saved to disk

edit_image

Edit images with natural-language instructions. Multi-turn continuity via thought signatures

describe_image

Fast image descriptions on Gemini 3.8 Flash

analyze_image

Structured extraction and analysis on Gemini 3.1 Pro

load_image_from_path

Read a local image file and return base64 for any image tool

generate_video

Video generation with Veo 3.1: 4 to 8 seconds at up to 4K with native audio

generate_svg

Production-ready SVG: diagrams, illustrations, icons, data visualisations

generate_landing_page

Self-contained HTML landing pages with inline CSS and JS

gemini_viewer_payload

Internal. The inline viewers use it to fetch their display data; you'll never call it


Model reference

Checked against the live models API on 22 September 2026. gemini_list_models will tell you what your key can see today.

Model

Used by

Notes

gemini-3.1-pro-preview

gemini_chat, gemini_deep_research (synthesis), analyze_image, generate_landing_page

Default. Still the strongest reasoning model Google ships, preview label or not. Paid-only - no free tier

gemini-3.8-flash

describe_image, gemini_deep_research (search passes)

Default. Google's GA workhorse as of September 2026; a good gemini_chat choice when you want speed

gemini-3-flash-preview

generate_svg

Default. Has a free tier (Google pricing page, 24 September 2026)

gemini-3.7-flash, 3.6, 3.5, 3.5-flash-lite, 3.1-flash-lite

any text tool

Earlier GA Flash releases, all accepted

gemini-3-pro-image

generate_image, edit_image

Default. Nano Banana Pro, GA. 4K, real text, conversational editing

gemini-3.1-flash-image

generate_image, edit_image

Nano Banana 2, GA. Near-Pro quality, Flash speed and price

gemini-3.1-flash-lite-image

generate_image

Nano Banana 2 Lite. Fastest and cheapest

gemini-3-pro-image-preview, nano-banana-pro-preview, gemini-2.5-flash-image

image tools

Older IDs, still accepted so existing configs keep working

veo-3.1-generate-preview

generate_video

Default. Cinematic, native audio, up to 4K

veo-3.1-lite-generate-preview

generate_video

Cheaper and faster on the same API

Why the defaults are what they are. I went back and forth on this. Gemini 3.8 Flash is GA and newer, but 3.1 Pro is still what Google calls its strongest reasoning model, and reasoning is the point of gemini_chat and deep research - so Pro stays the default there, and Flash takes the lighter describe_image job. For images, Nano Banana Pro left preview and kept its quality lead, so it's the default; Nano Banana 2 is there when you'd rather trade a little quality for a lot of speed. Gemini's newer Omni video model uses a different API, so it's not wired in yet.

Gemini 3 notes: temperature is forced to 1.0 on every 3.x model (Google's requirement, lower values cause looping). thinking_level applies to gemini_chat.

Token budgets: max_tokens defaults to each model's full output ceiling as reported live by the models API (65,536 on current Gemini 3 text models; the 1M figure is input context). It's a cap, not consumption, so unused headroom costs nothing. Values below 4,096 are ignored because Gemini 3 thinking burns tiny budgets before you see any output, which looks exactly like a timeout, and values above the model's real limit are clamped.


Requirements

  • Node.js 18+

  • A Gemini API key from Google AI Studio

  • ffmpeg (optional, for video thumbnails)

Licence

Apache-2.0

Available Tools

13 tools
analyze_imageAnalyze ImageA

Analyze and extract information from one or more images using Gemini multimodal understanding. Returns a text analysis - no image is generated. Default model: gemini-3-pro-preview. DO NOT SET max_tokens - the server allocates the model's full output ceiling automatically; a small cap is spent on Gemini 3 thinking and returns empty output that looks like a timeout. [MCP_RECOMMENDED_TIMEOUT_MS: 300000]

ParametersJSON Schema
NameRequiredDescriptionDefault
modelNoOmit to use gemini-3-pro-preview. Other valid options: gemini-3.1-pro-preview, gemini-3-flash-preview. Do NOT pass gemini-1.5-* or gemini-pro-vision — those are out of support.
imagesYesOne or more images to analyze
promptYesWhat to analyze or extract from the image(s)
max_tokensNoOutput token budget INCLUDING Gemini 3 thinking tokens. OMIT THIS — the server allocates the model's full output ceiling (a cap, not consumption; unused headroom costs nothing). Values below 4096 are IGNORED; values above the model's real limit are clamped.
global_media_resolutionNoGlobal image quality for cost optimization. MEDIUM recommended for PDFs (50% savings).

Output Schema

ParametersJSON Schema
NameRequiredDescription
contentYes
successYes

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries full behavioral burden and does so excellently. It discloses output modality, the default model, the server-side max_tokens allocation behavior, the Gemini 3 thinking-token pitfall that can mimic a timeout, and a recommended timeout value. This is genuine transparency beyond what the schema provides.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with purpose, then provides output modality, model default, a critical warning, and a timeout recommendation in only four sentences. Every sentence earns its place; there is no filler or redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 5-parameter tool with a 100% documented schema and an output schema, the description is complete enough. It covers what the tool does, what it returns, the default model, the critical max_tokens constraint, and timeout expectations. The structured schema and output schema handle parameter details and return values, so the description does not need to repeat them.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% with detailed per-parameter descriptions, so the baseline is 3. The description adds value by explicitly warning against setting max_tokens and explaining why ('a small cap is spent on Gemini 3 thinking and returns empty output that looks like a timeout'), plus pointing out the default model. This elevates it above the schema-only baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Analyze and extract information from one or more images using Gemini multimodal understanding.' It also clarifies the output type ('Returns a text analysis - no image is generated'), which distinguishes it from image-generation siblings like generate_image. This is a clear, non-tautological statement of purpose.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear use context: analyze/extract information from images and produce text output, not an image. It also provides critical operational guidance such as omitting max_tokens, the default model, and a recommended timeout. However, it does not explicitly name alternative tools like describe_image or state when NOT to use this tool versus those siblings, so it falls just short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

describe_imageDescribe Image (Nano Banana Pro)A

Analyze and describe one or more images using Google Gemini image models (Nano Banana Pro). Returns a text description — no image is generated. Default model: gemini-3-flash-preview. [MCP_RECOMMENDED_TIMEOUT_MS: 180000]

ParametersJSON Schema
NameRequiredDescriptionDefault
modelNoOmit to use gemini-3-flash-preview. Other valid options: gemini-3-pro-preview, gemini-3.1-pro-preview. Do NOT pass gemini-1.5-* or gemini-pro-vision — those are out of support.
imagesYesOne or more images to describe/analyze
promptNoOptional custom analysis prompt (default: general description)
global_media_resolutionNoGlobal image quality for cost optimization. MEDIUM recommended for PDFs (50% savings).

Output Schema

ParametersJSON Schema
NameRequiredDescription
contentYes
successYes

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. It explicitly states the output is text and that no image is generated, and mentions the default model. The timeout recommendation is also included. However, it does not mention potential side effects, rate limits, cost implications, or any details about how the images are processed besides the model hint. This is adequate but not rich.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact — two sentences plus a timeout recommendation. The core purpose is front-loaded: 'Analyze and describe...' The rest is supplementary. Every part adds value: the output type, the no-generation assurance, and the default model. No fluff or redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's moderate complexity (4 params, nested images object, output schema present), the description covers the essential purpose, output, and default model. The schema handles parameter details, and the output schema clarifies return values. The timeout hint is a plus. However, it doesn't differentiate from the similar analyze_image sibling, which might be a slight gap. Overall, it's sufficiently complete for a description-focused tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% — every parameter (model, images, prompt, global_media_resolution) is well-described in the schema, including sub-fields like filePath, mimeType, and mediaResolution. The description itself adds minimal parameter info beyond the schema, but since the schema covers everything, the baseline of 3 applies. The description does note the default model, which matches the schema's instructions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's action: 'Analyze and describe one or more images' using Gemini models, and specifies the output: 'Returns a text description'. It also explicitly notes that no image is generated, which differentiates it from generation tools like generate_image. The verb+resource+output is specific and unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear context on what the tool does (describes images, returns text) and implicitly indicates it's for description rather than generation. However, it does not explicitly mention when to use this over alternatives like analyze_image or when not to use it. The 'no image is generated' hint suggests it's not for image generation, but no explicit alternatives or exclusions are given.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

edit_imageEdit Image with GeminiA

Edit one or more images using Google Gemini image models (Nano Banana Pro). Provide images and natural-language instructions for how to modify them. Returns edited image with inline preview and saves full-resolution to disk.

ParametersJSON Schema
NameRequiredDescriptionDefault
modelNoGemini image model to use (default: gemini-3-pro-image-preview)
imagesYesOne or more images to edit
promptYesInstructions for how to edit the image(s)
outputPathNoOptional file path to save the edited image (e.g., ./output/edited.png)
use_searchNoEnable Google Search grounding for data-driven editing
global_media_resolutionNoGlobal image quality setting (default: HIGH). See generate_image for details.

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description discloses that the tool returns an edited image with preview and saves to disk, and mentions the underlying model family. However, it omits details such as token costs, whether the original is altered, or any side effects beyond saving. With no annotations, these gaps leave the description only partially transparent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two concise sentences, front-loaded with purpose. Every word contributes, no redundancy or unnecessary detail.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has six parameters and no output schema, the description adequately covers the core workflow. It mentions return and disk saving, but does not elaborate on cost or how to choose between filePath and data for large images (though the schema covers that). Overall, it is reasonably complete for a typical use case.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has full (100%) coverage and detailed descriptions for every parameter, including resolution options and thoughtSignature. The tool description itself adds no new information about parameters beyond what the schema already states, so a baseline score of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool edits images using Google Gemini image models (Nano Banana Pro), explicitly distinguishing it from generation tools. It also notes the return of an edited image with preview and disk saving, making its purpose unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description instructs to provide images and natural-language instructions, which implies the primary use case. It does not explicitly compare with alternatives like generate_image or mention when not to use it, but the context is clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

gemini_chatGemini ChatA

Chat with Google Gemini models. Grounded in Google Search by default, on gemini-3.1-pro-preview. DO NOT SET max_tokens - the server allocates the model's full output ceiling automatically. It is a cap, not consumption, so unused headroom costs nothing; setting a small one makes Gemini 3 thinking burn the whole budget and return empty output that looks like a timeout. [MCP_RECOMMENDED_TIMEOUT_MS: 300000]

ParametersJSON Schema
NameRequiredDescriptionDefault
modelNoOmit to use the configured default (gemini-3.1-pro-preview). Other valid options: gemini-3-pro-preview, gemini-3-flash-preview. Do NOT pass gemini-1.5-* or gemini-pro — those are out of support.
messageYesThe message to send
groundingNoEnable Google Search grounding for real-time information
max_tokensNoOutput token budget INCLUDING Gemini 3 thinking tokens. OMIT THIS — the server allocates the model's full output ceiling (queried live, 65,536 on current Gemini 3 text models). It is a cap, not consumption — unused headroom costs nothing. Values below 4096 are IGNORED (thinking burns them before any visible output) and values above the model's real limit are clamped to it.
temperatureNoControls randomness (0.0 to 1.0). Ignored on Gemini 3+ (forced to 1.0 per Google docs).
system_promptNoOptional system instruction
thinking_levelNoThinking depth for Gemini 3 models only. "low" minimises latency for simple tasks. "high" (default for Gemini 3) maximises reasoning depth. "medium"/"minimal" available on Gemini 3 Flash only. Ignored for non-Gemini-3 models.

Output Schema

ParametersJSON Schema
NameRequiredDescription
contentYes
successYes

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It discloses critical behavioral traits: grounding by default, the max_tokens cap behavior (not consumption), the fact that values below 4096 are ignored, and that temperature is ignored on Gemini 3+ (forced to 1.0). It also explains the thinking_level parameter's scope and defaults. This is exemplary transparency for a complex tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with the core purpose and then dives into critical warnings. It is dense but not bloated; every sentence adds value. The only minor issue is that the max_tokens warning is repeated in both the description and the schema parameter description, which is slightly redundant but reinforces the critical point. Overall, it's well-structured and efficient.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (7 parameters, 1 required, 100% schema coverage, output schema present), the description is remarkably complete. It covers model selection, grounding, token budget behavior, temperature quirks, thinking levels, and timeout recommendations. The output schema exists, so return values don't need explanation. This is a model example of a complete tool description.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, but the description adds significant value beyond the schema. For max_tokens, it explains the cap-vs-consumption distinction, the 4096 threshold, and clamping behavior. For temperature, it notes the forced 1.0 on Gemini 3+. For thinking_level, it clarifies which models support which values. The description enriches every parameter with practical context.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: 'Chat with Google Gemini models.' It specifies the default model (gemini-3.1-pro-preview) and grounding behavior, distinguishing it from sibling tools like gemini_deep_research or gemini_prompt_assistant. The verb 'chat' plus the resource 'Google Gemini models' is specific and unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides explicit usage guidance: it warns against setting max_tokens, explains the server's automatic allocation, and clarifies that small values cause thinking tokens to burn the budget and return empty output. It also includes a recommended timeout (300000 ms) and the schema details valid model options and exclusions (e.g., 'Do NOT pass gemini-1.5-* or gemini-pro'). This is comprehensive and actionable.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

gemini_deep_researchGemini Deep ResearchA

Conduct deep research on complex topics using iterative multi-step analysis with Gemini. This performs multiple searches and synthesizes comprehensive research reports (takes several minutes). [MCP_RECOMMENDED_TIMEOUT_MS: 900000]

ParametersJSON Schema
NameRequiredDescriptionDefault
modelNoModel to use for deep research (defaults to latest available)
focus_areasNoOptional: specific areas to focus the research on
max_iterationsNoNumber of research iterations (1-10, default 1). Environment guidance: Claude Desktop: use 1-2 (4-min timeout). Agent SDK/IDEs (VSCode, Cursor, Windsurf)/AI platforms (Cline, Roo-Cline): can use 5-7 (longer timeout tolerance)
research_questionYesThe complex research question or topic to investigate deeply

Output Schema

ParametersJSON Schema
NameRequiredDescription
contentYes
successYes

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the burden of explaining behavior. It discloses that the tool performs multiple searches, synthesizes reports, and takes several minutes, including a recommended timeout. This is solid transparency for common behavioral concerns, even if it does not mention output format or potential failure modes.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences and front-loads the core purpose and approach. Every sentence adds value, and the recommended timeout is embedded in a compact, parseable tag rather than verbose prose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

This is a complex tool with an output schema, and the description covers the key behavioral side: iterative searches, report synthesis, and time requirement. It could say a bit more about what types of comprehensive reports are produced, but the output schema likely covers return-shape expectations.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema already provides 100% description coverage for all four parameters, so the baseline is 3. The description does not add extra parameter-level detail beyond what the schema supplies, but it reinforces that research_question should be complex and that the process is iterative.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly identifies the tool as conducting deep research on complex topics with iterative multi-step analysis, which uses a specific verb and resource. This distinguishes it from sibling tools like gemini_chat and gemini_prompt_assistant by emphasizing multiple searches and comprehensive research reports.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear context for when to use this tool: complex topics requiring deep research, multiple searches, and comprehensive reports. It does not explicitly name alternatives for simpler queries, but the emphasis on complex topics and multi-step iteration makes the intended use case reasonably clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

gemini_helpGemini MCP HelpB

Get comprehensive help about Gemini MCP features, settings, and best practices

ParametersJSON Schema
NameRequiredDescriptionDefault
topicNoHelp topic to displayoverview

TDQS

B3.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. The description is generic, saying 'Get comprehensive help' but does not specify what the tool returns (e.g., text output, how it behaves when topic is invalid, whether it requires network access). It doesn't disclose any potential side effects or limitations, leaving the agent with minimal behavioral context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is one concise sentence that clearly states the tool's purpose. It is front-loaded with the main verb 'Get' and includes the resource 'Gemini MCP' and the key aspects it covers, making it efficient and to the point.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a help tool with a simple schema and no output schema, the description is adequate but not outstanding. It tells the agent what the tool does but does not explain what a help response looks like, how the topic parameter affects the output, or provide any usage examples. The tool has 12 siblings, so a bit more context on which topics are most commonly needed could help, but the current description is minimally sufficient.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema describes the single parameter 'topic' with an enum of valid values and a default of 'overview', providing full coverage (100%). The description does not add additional meaning beyond the schema, but with 100% coverage, the baseline of 3 is appropriate. The description's mention of 'features, settings, and best practices' may hint at the content of the topics, but it's not explicit.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool provides comprehensive help about Gemini MCP features, settings, and best practices. It distinguishes it from sibling tools that perform specific actions like chat or image generation, making the purpose evident.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies use when users need help or information about the Gemini MCP suite, but it doesn't explicitly state when not to use it or compare to alternatives. Sibling tools handle specific tasks, so the need for guidance is there, but the purpose is clear enough that an agent might infer when to use it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

gemini_list_modelsList Gemini ModelsA

List available Gemini models and their descriptions

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
contentYes
successYes

TDQS

A3.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of disclosing behavioral traits. It simply states the tool lists models without mentioning return format, pagination, or whether it's read-only. For a simple list operation, this is a significant gap, though the operation is inherently safe.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, concise sentence that directly states the tool's function without fluff. It's front-loaded and efficient.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple list tool with no parameters and an output schema (not shown), the description adequately covers the core functionality. It doesn't elaborate on usage context, but given the simplicity, it's sufficient.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parametersтное schema is empty, so the description doesn't need to add parameter meaning. Baseline for 0 params is 4; the description correctly omits any param details.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'List available Gemini models and their descriptions' clearly states the tool's function with a specific verb ('list') and resource ('Gemini models'). It unambiguously distinguishes this tool from sibling tools like gemini_chat or gemini_prompt_assistant, which perform actions rather than enumeration.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives. While its purpose is self-evident, it lacks any explicit context or exclusions, such as whether it should be used before invoking chat tools or whether it lists all models or only certain ones.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

gemini_prompt_assistantImage & Chart Prompt AssistantA

Get expert prompt templates and guidance for Gemini image generation. Covers photography (portraits, products, cinematic), chart/diagram design (9 professional design systems including FT, Bloomberg, Tufte, Du Bois), lighting, colour grading, lens simulation, and style aesthetics. For charts: use chart_design with a color_scheme to get a full professional design system prompt.

ParametersJSON Schema
NameRequiredDescriptionDefault
emphasisNoDesign emphasis / audience priority
use_caseNoSpecific use case for template (photography request types)
chart_typeNoChart type for chart-specific guidelines
color_schemeNoChart colour scheme / design system (for chart_design, optimize_chart, get_palette)
request_typeYesType of assistance needed
current_promptNoCurrent prompt to optimize or troubleshoot
desired_outcomeNoDescription of what you want to achieve

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It clearly frames the tool as advice-giver ('Get expert prompt templates and guidance'), implying a read-only, non-destructive assistant. It does not disclose any hidden side effects or special behaviors, but for a guidance tool, this is minimal risk. No contradictions with annotations exist since none are present.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences, tightly packed with useful information. It front-loads the primary purpose, lists covered domains, and ends with a targeted usage note. Every sentence earns its place; no fluff or repetition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With 7 parameters (5 enums) and no output schema, the description covers the core functionality and highlights the chart_design+color_scheme combination. It doesn't elaborate on every request_type, but the enum in the schema provides that list, and the description gives sufficient context for the most common branching scenario. It could mention that it also handles prompt optimization/troubleshooting, but the schema implies it via the request_type enum.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so baseline is 3. The description adds a valuable cross-parameter hint: combining chart_design (request_type) with color_scheme yields a professional design system prompt. This enriches understanding beyond the schema's isolated parameter descriptions. It does not explain each parameter in depth, but the schema already does that adequately.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states it provides 'expert prompt templates and guidance for Gemini image generation' and lists specific areas (photography, charts, lighting, etc.). It distinguishes itself from siblings like generate_image by focusing on guidance rather than generation, and explicitly mentions chart_design for chart-specific needs.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description includes an explicit usage hint for charts: 'For charts: use chart_design with a color_scheme to get a full professional design system prompt.' It doesn't enumerate when not to use it or list alternatives, but this guidance covers a key use case. Context for other request types is implied through the enum in the schema.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

generate_imageGenerate Image with GeminiA

Generate an image using Google Gemini image models (Nano Banana Pro). Returns image with inline preview in Claude Desktop and saves full-resolution to disk. Default model: gemini-3-pro-image-preview.

ParametersJSON Schema
NameRequiredDescriptionDefault
modelNoGemini image model to use (default: gemini-3-pro-image-preview). Options: gemini-3-pro-image-preview, gemini-2.5-flash-image, nano-banana-pro-preview
imagesNoOptional reference images to guide generation
promptYesDescription of the image to generate
imageSizeNoResolution of the generated image (only for image-specific models)
outputPathNoOptional file path to save the generated image (e.g., ./output/image.png)
use_searchNoEnable Google Search grounding for data-driven image generation. Use for: weather forecasts, current events, stock prices, sports scores, statistics. The model will search the web for real-time data to inform image generation.
aspectRatioNoAspect ratio of the generated image1:1
global_media_resolutionNoGlobal image quality setting for cost optimization (default: HIGH). LOW (280 tokens, 75% savings) - Simple tasks, bulk operations. MEDIUM (560 tokens, 50% savings) - PDFs/documents (OCR saturates at medium). HIGH (1120 tokens) - Best quality, detailed analysis. Can be overridden per-image using mediaResolution in images array.

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It discloses key behaviors: returns inline preview in Claude Desktop, saves full-resolution to disk, and mentions the default model. It also explains the use_search grounding behavior. However, it does not mention potential side effects like file overwriting, cost implications, or rate limits, which would be more transparent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise: two sentences that state the core function, the model family, the output behavior, and the default model. It is front-loaded with the primary action and avoids unnecessary fluff. Every sentence adds value.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (8 parameters, no output schema, no annotations), the description is fairly complete. It covers the main purpose, default model, output behavior, and search grounding. However, it could mention that the tool can also edit images when provided with reference images (as implied by the images parameter and thoughtSignature), which would help agents understand its dual generation/editing capability.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all parameters. The description adds value by explaining the default model and the use_search use cases (weather, current events, etc.). It also clarifies the global_media_resolution cost trade-offs, which is beyond the schema. However, it doesn't elaborate on the thoughtSignature parameter's role in conversational editing, which is partially covered in the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: 'Generate an image using Google Gemini image models (Nano Banana Pro).' It specifies the action (generate), the resource (image), and the technology (Gemini image models). It also distinguishes from siblings like edit_image, generate_svg, and generate_video by focusing on image generation with Gemini.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear context on when to use the tool (for generating images with Gemini) and mentions the default model. It also hints at usage for data-driven generation via use_search. However, it does not explicitly state when NOT to use it or mention alternatives like edit_image for editing existing images, which would be a stronger guideline.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

generate_landing_pageGenerate Landing PageA

Generate a complete, self-contained HTML landing page using Gemini. Returns inline preview with responsive viewport controls. No external dependencies; inline CSS and vanilla JS only. [MCP_RECOMMENDED_TIMEOUT_MS: 300000]

ParametersJSON Schema
NameRequiredDescriptionDefault
briefYesDescription of the product/service and page goals
modelNoOmit to use gemini-3.1-pro-preview. Other valid options: gemini-3-pro-preview, gemini-3-flash-preview.
styleNoVisual design stylestartup
sectionsNoSections to include (e.g. ["hero", "features", "pricing", "cta"])
outputPathNoOptional file path to save the HTML (e.g. C:/dev/output/landing.html)
companyNameNoCompany or product name
primaryColourNoPrimary brand colour (e.g. #3B82F6 or "deep blue")

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Since no annotations are provided, the description carries the full behavioral burden. It discloses that the output is an inline preview with responsive viewport controls, that the HTML has no external dependencies, and it includes a recommended timeout of 300000 ms. It doesn't cover side effects of saving a file or model differences, but the core behavior is well communicated.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is only three sentences plus a timeout metadata tag. Each sentence earns its place: purpose, preview behavior, and the no-dependency constraint. It is front-loaded with the core action. The timeout metadata feels slightly out-of-place but is useful for the agent.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 7-parameter tool with no output schema, the description provides the most critical context: an HTML landing page is generated, a preview is returned, and no external dependencies are used. The inline preview covers the return behavior. It does not explain what outputPath does in detail, but the schema covers that parameter. Overall it is sufficiently complete for the agent to invoke successfully.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema already describes 100% of parameters, so the description does not need to add much parameter-level detail. The parameter descriptions cover meanings and defaults. The tool description itself doesn't add significant parameter semantics beyond the schema, so a baseline score of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool generates a complete, self-contained HTML landing page via Gemini. This is a specific verb and resource, and it distinguishes from sibling tools like generate_image, generate_svg, and generate_video. The addition of inline preview and responsive viewport controls further clarifies the output format.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear context: use this to create a self-contained HTML landing page and preview it in the IDE. It doesn't explicitly name alternatives or state when-not-to-use, but the purpose is unambiguous enough for an agent to choose this over sibling generation tools. The dependency-free constraint also helps narrow expectations.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

generate_svgGenerate SVG GraphicA

Generate scalable vector graphics (SVG) using Gemini. Creates clean, production-ready SVG code for diagrams, illustrations, icons, and data visualizations. Returns inline preview with SVG viewer. [MCP_RECOMMENDED_TIMEOUT_MS: 240000]

ParametersJSON Schema
NameRequiredDescriptionDefault
modelNoOmit to use gemini-3.1-pro-preview. Other valid options: gemini-3-pro-preview, gemini-3-flash-preview.
styleNoVisual style: technical (diagrams), artistic (illustrations), minimal (simple), data-viz (charts)technical
widthNoSVG width in pixels (default: 800)
heightNoSVG height in pixels (default: 600)
promptYesDescription of the SVG graphic to generate
outputPathNoOptional file path to save the SVG (e.g. C:/output/diagram.svg)

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Without annotations, the description carries the transparency burden. It mentions the inline preview and a recommended timeout, which gives operational insight. However, it does not disclose side effects like file writing when outputPath is provided, whether the operation is read-only, or any permission requirements. The generative nature implies creation, but this is not explicitly stated.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences plus a timeout note, with the primary action stated upfront. Every sentence adds value: the first defines purpose, the second describes output and preview, and the timeout note is practical. No filler or redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (6 parameters, no annotations, no output schema), the description is fairly complete. It covers what it generates, the output format, and a timeout recommendation. However, it omits details like error handling, whether outputPath is required for saving, and how to specify style constraints beyond schema enums. These gaps are minor for a generation tool but prevent a higher score.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so parameter descriptions already exist. The tool description adds contextual value by naming the types of graphics (diagrams, icons, etc.) but does not elaborate on parameter usage beyond what the schema provides. It meets the baseline for full schema coverage without adding significant new meaning.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool generates SVG graphics via Gemini, listing specific use cases (diagrams, illustrations, icons, data visualizations). This distinguishes it from sibling tools like generate_image (which likely produces raster images) and edit_image (which modifies existing images), making the purpose specific and unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage for vector graphic creation but does not explicitly state when to prefer this over alternatives such as generate_image or when not to use it. It provides context (SVG, data visualization) but lacks direct guidance on tool selection or exclusion criteria.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

generate_videoGenerate VideoA

Generate videos using Google Veo 3.1 AI model. Creates realistic 4-8 second videos from text prompts with optional first-frame image and reference images for character/style consistency. Supports native audio generation. Processing time: 2-5 minutes for 1080p videos. Returns video file path with optional thumbnail and HTML preview player. ⚠️ IMPORTANT: Video generation is ASYNC and takes 2-5 minutes. The tool will poll for completion automatically.

ParametersJSON Schema
NameRequiredDescriptionDefault
seedNoOptional seed for deterministic output. Use the same seed with the same prompt for consistent results.
modelNoVideo generation model (default: veo-3.1-generate-preview)veo-3.1-generate-preview
promptYesDetailed description of the video to generate. Be specific about actions, camera movements, lighting, and style. Example: "A close-up shot of a futuristic coffee machine brewing a glowing blue espresso, with steam rising dramatically. Cinematic lighting, 4K quality."
outputPathNoOptional custom output path for the video file (e.g., C:/videos/output.mp4). If not provided, saves to default output directory with timestamped filename.
resolutionNoVideo resolution. Higher resolutions take longer to generate and result in larger files.1080p
aspectRatioNoVideo aspect ratio: 16:9 (landscape) or 9:16 (portrait/vertical)16:9
sampleCountNoNumber of video samples to generate (1-4). Each sample is a separate generation.
generateAudioNoGenerate native synchronized audio effects and dialogue based on the prompt
durationSecondsNoVideo duration in seconds — 4, 6, or 8. Other values in range are rounded to the nearest of those.
firstFrameImageNoStarting frame image for image-to-video generation. Provide via filePath (local file) or data+mimeType (base64). The video will animate from this image. Supports JPEG, PNG, WebP.
referenceImagesNoUp to 3 reference images for character/style consistency. Each needs a referenceType ("asset" or "style") and an image.
generateThumbnailNoExtract thumbnail from video (requires ffmpeg installed). Thumbnail is saved alongside video.
generateHTMLPlayerNoGenerate interactive HTML video player with preview and download options

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden. It discloses the async nature, processing time, and that it polls automatically, as well as dependencies like ffmpeg for thumbnails. It also implies mutations (creating files) but without explicit warnings, but this is covered by the nature of generation. It adds value beyond the schema by explaining behavior like automatic polling.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is relatively compact for a complex tool, with key information front-loaded (model, duration, inputs, audio), followed by a warning about async processing. It uses bullet points effectively. It could be more concise, but for the complexity it is well-structured and not verbose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the high schema coverage and descriptive schema (nested objects, enums, bounds), the description completes the picture by explaining output artifacts (file path, thumbnail, HTML player) and processing time. It does not explain return structure, but there's no output schema, and the description hints at what is returned. For a generation tool with no output schema, this is adequate.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema already has 100% coverage with detailed descriptions for each parameter, including example prompts, resolution trade-offs, and image input alternatives. The description reinforces the overall purpose but does not add additional parameter-level semantics beyond what's in the schema. Since coverage is high, a baseline of 3 is appropriate, but the description's mention of output paths and thumbnail/HTML player adds slight context, hence a 4.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states that it generates videos using Google Veo 3.1, with specific details on duration, input types (text prompt, first-frame image, reference images), audio generation, and output artifacts (video file, thumbnail, HTML player). It distinguishes it from sibling image tools by emphasizing video generation and asynchronous processing.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explains key usage context: asynchronous with 2-5 minute processing time, automatic polling, and supports optional inputs like first-frame and reference images. It does not explicitly mention when NOT to use it or alternatives, but the context is clear enough for an agent to decide, especially with sibling tools like generate_image.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

load_image_from_pathLoad Image from File PathA

Read a local image file and return it as base64-encoded data ready to pass to generate_image, edit_image, describe_image, or analyze_image tools. Supports JPEG, PNG, GIF, WebP, BMP.

ParametersJSON Schema
NameRequiredDescriptionDefault
filePathYesAbsolute or relative path to the image file

Output Schema

ParametersJSON Schema
NameRequiredDescription
contentYes
successYes

TDQS

A4.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Since no annotations are provided, the description carries the full behavioral burden; it discloses the base64 return format and supported image types, which are useful facts. However, it doesn't mention error behavior, file size limits, or security/permission considerations, so transparency is adequate but not comprehensive.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences efficiently deliver the essential information: reads a local image, returns base64, names its consumers, and lists supported formats. There is no filler or redundant repetition of the schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter tool with an output schema, the description covers the key behavioral context: input, output, and supported formats. It could mention potential failure conditions, but these are not critical for a defensive agent's selection and invocation of such a straightforward tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema already fully describes filePath, so the baseline is 3. The description supplements it by emphasizing 'local image file' and listing supported formats, which adds mildly relevant context beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('Read a local image file') and the resource being handled, with an explicit output type. It also differentiates itself from sibling tools by explaining the result is ready to feed into generate_image, edit_image, describe_image, or analyze_image.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly names the downstream tools that accept its output, giving strong context for when to use this tool. It doesn't state an explicit 'do not use when' condition, but for a focused loading utility, this is a minor gap.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 13 tool updatesv2.6.2
    • First observedanalyze_image
    • First observeddescribe_image
    • First observededit_image
    • First observedgemini_chat
    • First observedgemini_deep_research
    • First observedgemini_help
    • First observedgemini_list_models
    • First observedgemini_prompt_assistant
    • First observedgenerate_image
    • First observedgenerate_landing_page
    • First observedgenerate_svg
    • First observedgenerate_video
    • First observedload_image_from_path

TDQS

A4/5.0

Scored across 13 tools

Disambiguation4/5

Most tools are clearly distinct (help, list_models, chat, research, image generation/editing, video generation). However, describe_image and analyze_image overlap significantly—both analyze images and return text—with only subtle differences in default model and phrasing, which could cause misselection. Also, gemini_help overlaps with what an agent might expect from general documentation but is distinct enough.

Naming Consistency4/5

Most tools follow a verb_noun pattern (gemini_list_models, generate_image, edit_image, generate_landing_page, generate_svg, generate_video), but some tools omit the 'gemini_' prefix (describe_image, analyze_image, load_image_from_path) creating minor inconsistency. The use of 'generate' for different output types is clear, but 'describe' vs 'analyze' could be more distinct. Patterns are mostly predictable.

Tool Count5/5

With 13 tools, this server is well-scoped for a multimodal AI assistant covering chat, research, image, video, and text generation. Each tool has a clear purpose and covers distinct capabilities (help, models, prompting, chat, deep research, image in/out, editing, landing page, SVG, video). The count is appropriate without being excessive.

Completeness5/5

The tool surface covers the key workflows: image generation (generate_image), image editing (edit_image), image analysis (describe/analyze_image), local image loading (load_image_from_path), video generation (generate_video), text generation (generate_svg/landing_page), and interactive use (chat, deep_research). A clear lifecycle exists for image tasks (load→analyze→generate/edit). Missing features like image manipulation beyond editing or direct video editing are minor and likely out of scope.

Maintenance

ActivityMaintained
ResponsivenessSlow

Related MCP Connectors

Related MCP Servers