gemini-relay
Gemini Relay is a local MCP server that lets your coding agent delegate large-context reading, analysis, planning, image generation, and brainstorming to Google Gemini, returning only concise answers instead of filling the agent's context window.
ask-gemini: Ask questions, review code, or request changes; include files/folders with
@, choose model/effort/mode, enforce JSON schema output, resume conversations, get token usage, and receive structured changeMode edits.gemini-plan: Generate architectural and implementation blueprints, dependency analysis, and risk assessments without modifying code.
gemini-image: Generate images from text descriptions, choose aspect ratios, and save them directly into your project workspace.
brainstorm: Structured ideation using frameworks like SCAMPER or Design Thinking, with constraints, domain context, and feasibility analysis.
gemini-models: List available Gemini models, defaults, reasoning capabilities, and backend status.
gemini-doctor: Diagnose CLI/agy installation, backend, login, version, and remaining quota.
fetch-chunk: Retrieve additional cached chunks from a prior changeMode response.
ping and Help: Verify server connectivity and get the active backend CLI's help.
Enables image generation through Google Imagen with selectable aspect ratios and saving generated assets to the workspace path.
Provides access to Google Gemini models via the Antigravity CLI, supporting configurable reasoning effort, planning mode, structured JSON schema outputs, token usage metrics, and workspace context inlining.
TL;DR — Gemini Relay is a bridge between your coding agent and Google Gemini. Claude asks, Gemini answers, and Claude keeps working. Image generation, a second pair of eyes on your code, plans, brainstorms, huge files, and models Claude cannot reach on its own. Ask in plain English. Point at files with
@, or let Gemini find them itself.
What is this?
Claude is a great coding agent. It is also one model, with one set of strengths and one quota.
Gemini Relay is a small MCP server that sits beside it. Your agent sends it a request, the request goes to Google Gemini through the Antigravity CLI, and the answer comes back as if Claude had done the work itself. No copy-pasting between two chats. No switching windows.
That turns Gemini into a set of tools Claude can pick up whenever they fit: draw an image, review a diff, design a refactor, read a folder that would never fit in its window, or ask a different model the same question and compare.
Gemini can read the project on its own, too. Point it at files with @ when you want exactly those. Say nothing and it goes looking.
Related MCP server: MCP Gemini Server
What Gemini brings
Things Claude gets by having Gemini next door.
Images. Gemini generates pictures. Claude does not. Ask for a hero image for the README, an icon, a diagram, a mock screenshot, and it lands in your repo as a file.
A second opinion. A review from a different model catches different bugs. Hand Gemini a diff or a folder and get back a report that Claude did not write, then let Claude act on it.
Plans and brainstorms. A dedicated planning tool for architecture and refactors, and a brainstorming tool with a choice of methods, both run on Gemini's deepest reasoning.
Big reading. Gemini takes the files, Claude takes only the answer. The 181 KB package-lock.json in this repo would cost Claude roughly 45,000 tokens to open. Sent through the relay it came back in 19 seconds as one line:
{ "count": 392 }
More models, more quota. Whatever Antigravity offers is available: Gemini Flash and Pro at every effort level, plus Claude and GPT-OSS models on a separate quota bucket. When one well runs dry, another is a parameter away.
Install
You need two things: this server, and Google's Antigravity CLI that it drives.
1. Install the CLI and sign in.
curl -fsSL https://antigravity.google/cli/install.sh | bash
agyRun agy once and it walks you through signing in. On Windows, use the official installer from https://goo.gle/gemini-cli-migration instead of the curl line.
2. Add the server.
Claude Code, one command:
claude mcp add gemini-relay -- npx -y gemini-relayClaude Desktop, in claude_desktop_config.json:
{
"mcpServers": {
"gemini-relay": {
"command": "npx",
"args": ["-y", "gemini-relay"]
}
}
}Cursor and Windsurf: add a server named gemini-relay with the command npx -y gemini-relay.
3. Check it. Ask your agent to run gemini-doctor. It reports whether the CLI was found, whether you are signed in, and how much quota is left. It costs nothing to run.
Node 18.19 or newer.
What you can ask for
Talk to your agent normally. These are the shapes that work.
You want to… | Say something like |
Make a picture | "Use gemini-image for a 16:9 dark hero image, save it to |
Get a second review of your code | "Have gemini review |
Ask another model the same question | "Ask gemini the same question, but with Claude Opus." |
Get a plan before you write anything | "Use gemini-plan to design retry with backoff for the upload queue." |
Kick ideas around | "Brainstorm ten ways to cut our cold-start time." |
Let Gemini go find the problem itself | "Ask gemini to find the riskiest code in this repo." |
Look at something too big to open | "Ask gemini what's in |
Get an answer your code can parse | "Ask gemini for the outdated deps as JSON." |
Pick up an earlier thread | "Run gemini-conversations and continue the one about the upload queue." |
Stop a run that is taking too long | "Run gemini-cancel." |
See what models you have | "Run gemini-models." |
Find out why it broke | "Run gemini-doctor." |
@ accepts a file, a folder, @. for the whole project, or a glob like @src/**/*.ts.
Eleven tools, every one prefixed gemini-. Every parameter, every default. A twelfth, gemini-timeout-test, appears only when GEMINI_MCP_TEST_TOOLS is set.
Tool | Parameter | Type · default | Notes |
|
| string, required | Supports |
| string | Any id | |
|
| Thinking depth. | |
|
|
| |
| object | string | Enforces structured JSON. Suppresses | |
| string[] | Extra directories agy may see. | |
| string | Resume a thread. A plain reply reports the id it created or continued. | |
| string | Run a custom | |
| boolean · | Let a | |
| boolean · |
| |
| boolean · | Appends tokens and timing. Ignored with | |
| boolean · | Forwarded, but agy does not isolate tool execution headless, and says so in a notice. The legacy | |
| boolean · | Gemini emits | |
| number | string | Which chunk (1-based). With | |
| string | Exactly 8 lowercase hex characters, or the call is refused. | |
|
| string, required | The thing to plan. |
| string | Constraints, or | |
| string · | Pinned unless you override it. | |
|
| ||
| string[] | ||
| boolean · | Never reports a conversation id. | |
|
| string, required | |
| enum · |
| |
|
| Omit to let Gemini pick. | |
| string | Relative workspace path. Escaping the root is refused. | |
|
| string, required | |
|
| ||
| string | ||
| integer · | ||
| boolean · | Never reports a conversation id. | |
|
| string, number — both required | Both reported by the initial |
|
| integer · | Recent conversations from agy's local store, newest first: id, last activity, cwd, first prompt. Max 100. No model turn. |
| — | Kills every CLI run in flight. Each cancelled call returns an error to its caller. | |
| — | Live catalogue from | |
| — | Binaries, versions, backend, plus login and quota via a free | |
|
| string · | Answered in process. Proves the transport is alive, not the CLI. |
| — | The backend CLI's own |
One rule for every flag. The relay sends a flag only when the installed agy advertised it in --help. If that probe finds nothing, no flags are sent at all and the run falls back to agy's defaults — an unknown flag makes agy exit non-zero and fails the whole request.
Models. Gemini 3.8 / 3.7 / 3.6 Flash in high, medium and low; Gemini 3.1 Pro in high and low; plus claude-sonnet-4-6, claude-opus-4-6-thinking and gpt-oss-120b-medium, which draw on a separate quota bucket.
What @ actually sends. A file inlines. A folder or @. inlines the text files beneath it. A glob inlines its matches. A token that resolves to nothing is left in the prompt verbatim. During folder and glob expansion node_modules, .git, dist and secret-looking files are skipped — name @.env directly and it is sent. Any file is dropped if it is binary, unreadable, or past the budget of 256 KB per file and 2 MB per prompt; a cut file carries TRUNCATED:, and dropped files are named in OMITTED: and UNREADABLE: footers. Nothing outside the project root is ever read.
Variable | Default | What it does |
| resolves by date |
|
| auto-detected | Full path to |
| auto-detected | Full path to the legacy |
|
| Wrapper timeout in minutes. Fractions accepted. Read once at load. |
| derived | Go duration forwarded to |
| unset |
|
| unset | Registers the test-only |
Going deeper
Everything technical lives here, so this page can stay short.
Every tool schema, recipes, and how to spend a context window well | |
What happens between your question and the answer | |
What | |
The catalogue, reasoning effort, and the quota buckets | |
Parameters and defaults, in long form | |
The errors you will actually see, and what to do | |
Why the backend moved, and what the code still guards against |
Good to know
Gemini reads your project on its own. Not only what you send with
@. It has file, search, web, memory and shell tools, and it uses them in whatever folder the server runs in.mode: "plan"keeps a run read-only.Headless runs are not sandboxed. Asking for
sandboxforwards the flag but does not isolate tool execution on theagybackend, and you get a notice saying so. Your ownagypermission settings are what hold.Secrets are skipped when a folder is expanded, not when you name one.
@.envsends the file.Quota is shared with your other agy use.
gemini-doctorshows what is left, free of charge. Claude and GPT-OSS models sit on their own bucket.Windows finds
agyat%LOCALAPPDATA%\agy\bin\agy.exe. If the server cannot see it, setAGY_CLI_PATHto the full path.Nothing leaves your machine except the prompt. Files are read locally and sent to Google as prompt text, the same as if you had pasted them.
Working on it
npm run doctor # is the environment sane
npm test # 150 unit + integration tests
npm run test:e2e # build, then drive the real CLI
npm run lint # type-check source and tests
npm run build # compile to dist/Support
If this saves you tokens or time, buy me a coffee.
License
MIT — see LICENSE.
Available Tools
11 toolsgemini-askA
Send one prompt to Google Gemini (or the Claude and GPT-OSS models the Antigravity CLI also offers) and get the answer back: questions, code review, analysis, edits, structured JSON. Reference project files with @path to send their contents along. Pass conversationId to continue a thread. For a phased implementation blueprint use gemini-plan; for idea generation use gemini-brainstorm; for everything else use this.
| Name | Required | Description | Default |
|---|---|---|---|
| mode | No | Agent execution mode: 'plan' for architectural analysis without modifying files, 'accept-edits' for direct edit application. | |
| agent | No | Optional custom agent to run instead of the default one (the name of an agent.md agent, e.g. a reviewer persona). Ignored by an agy build without --agent. | |
| model | No | Model to use — any id 'gemini-models' lists, or the aliases 'flash' and 'pro'. Leave it unset and no --model is sent at all, so agy answers on whatever model it is itself configured to use. | |
| effort | No | Reasoning effort ('low', 'medium', 'high') for Gemini 3.8 Flash, 3.7 Flash, and 3.1 Pro. Controls depth of thinking tokens. | |
| prompt | Yes | Analysis request. Use @ syntax to include files (e.g., '@largefile.js explain what this does') or ask general questions | |
| addDirs | No | Optional additional workspace directories to provide context to Gemini. | |
| sandbox | No | Ask the CLI to run in sandbox mode. The agy backend forwards it but does not isolate tool execution in headless runs, and says so in a notice; the legacy gemini backend does isolate. | |
| changeMode | No | Enable structured change mode - formats prompts to prevent tool errors and returns structured edit suggestions that Claude can apply directly | |
| chunkIndex | No | Which chunk of a changeMode response to return (1-based). With chunkCacheKey it reads the cached chunk; alone, it selects which chunk of a fresh changeMode result to return. | |
| jsonSchema | No | Optional JSON schema to enforce structured output from Gemini. | |
| includeUsage | No | Set to true to append token usage and timing metrics to the response. Ignored when jsonSchema is set, so the body stays valid JSON. | |
| chunkCacheKey | No | The 8-hex-character cache key a multi-chunk changeMode response reported, to fetch a later chunk without re-running the analysis. | |
| conversationId | No | Optional conversation ID to resume a previous session. A plain-text reply reports the ID of the conversation it created or continued; a jsonSchema or changeMode reply does not, because its body is parsed. | |
| skipPermissions | No | Set to true to run agy with --dangerously-skip-permissions. This genuinely changes behaviour: since agy 1.1.5 headless runs honour the persisted permission settings, so without it a tool call the settings do not allow is refused with nobody there to approve it. | |
| allowSlashCommands | No | Set to true to let a prompt starting with '/' expand as an agy command or skill (e.g. '/usage', '/skills'), which answers for free without a model turn. Default false: the prompt goes to the model verbatim. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries full burden. It discloses several beyond-schema behaviors: file inclusion via @path, conversation continuation, response-ID reporting caveats in plain-text vs jsonSchema/changeMode, agy backend's sandbox limitations, and skipPermissions behavior. Some details live in the schema, but the description adds A plain-text reply vs parsed-body distinction and the basic file-inclusion mechanism.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences: core action, feature highlights, and sibling routing. Front-loaded with what the tool does, no filler, and the routing closes the description. Every sentence earns its place; it stays compact despite handling a complex 15-parameter tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Richly covers a complex general-purpose tool: what it returns, how to reference files, how to continue threads, and when to use siblings. There is no output schema to document return values, but the description's core promise ('get the answer back') plus the conversationId/plain-text caveat suffice. No essential calling context is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema documents every parameter, making the baseline 3. The description adds meaning above that by encoding the @path file-inclusion convention directly ('Reference project files with @path to send their contents along') and by grouping the conversational state (conversationId to continue a thread). It also gives a quick semantic to mode via 'accept-edits' and 'plan' examples in the schema; the prose adds the general mental model, which lifts it slightly above baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description leads with a clear verb+resource ('Send one prompt to Google Gemini... and get the answer back'), and enumerates concrete uses: questions, code review, analysis, edits, structured JSON. It also distinguishes itself from gemini-plan and gemini-brainstorm by name, which is enough to tell it apart from siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says when to use this tool vs alternatives: 'For a phased implementation blueprint use gemini-plan; for idea generation use gemini-brainstorm; for everything else use this.' That is direct when/when-not routing. Thread-continuation guidance (conversationId) is also included.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
gemini-brainstormA
Generate a ranked list of ideas for an open-ended challenge using a named creativity method (SCAMPER, design thinking, lateral, divergent, convergent), with optional feasibility analysis. Ideation only: use gemini-ask for answers and gemini-plan for implementation blueprints.
| Name | Required | Description | Default |
|---|---|---|---|
| model | No | Model to use — any id 'gemini-models' lists, or the aliases 'flash' and 'pro'. Leave it unset and no --model is sent at all, so agy answers on whatever model it is itself configured to use. | |
| domain | No | Domain context for specialized brainstorming (e.g., 'software', 'business', 'creative', 'research', 'product', 'marketing') | |
| effort | No | Reasoning effort ('low', 'medium', 'high') to control depth of ideation and feasibility analysis. | |
| prompt | Yes | Primary brainstorming challenge or question to explore | |
| ideaCount | No | Target number of ideas to generate (default: 10-15) | |
| constraints | No | Known limitations, requirements, or boundaries (budget, time, technical, legal, etc.) | |
| methodology | No | Brainstorming framework: 'divergent' (generate many ideas), 'convergent' (refine existing), 'scamper' (systematic triggers), 'design-thinking' (human-centered), 'lateral' (unexpected connections), 'auto' (AI selects best) | auto |
| existingContext | No | Background information, previous attempts, or current state to build upon | |
| includeAnalysis | No | Include feasibility, impact, and implementation analysis for generated ideas |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description must carry the behavioral disclosure burden. It does this well by clarifying the tool only generates ideas, produces a ranked list, and optionally includes feasibility analysis. It could go further by describing the return structure or any model-related caveats, but the key behavioral boundary—ideation only, not execution—is explicit.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences, front-loaded with the core function, and contains no filler. The second sentence earns its place by preventing misuse through clear sibling routing.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 9-parameter tool with no output schema, the schema covers parameters thoroughly, and the description supplies the ideation boundary and sibling alternatives. It could be slightly more complete by mentioning what the output looks like beyond 'ranked list,' but the essential context for selecting and invoking the tool is present.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, and each parameter already has meaningful descriptions and enums. The tool description adds high-level context about ranked output and feasibility but does not add per-parameter semantics beyond what the schema already provides. The baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Generate a ranked list of ideas for an open-ended challenge.' It names concrete creativity methods and mentions optional feasibility analysis, making the tool's function unmistakable. It also differentiates itself from siblings by explicitly saying this is ideation only, not for answers or implementation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit routing guidance: 'Ideation only: use gemini-ask for answers and gemini-plan for implementation blueprints.' This tells an agent exactly when this tool is appropriate and when to choose a sibling, leaving no ambiguity.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
gemini-cancelA
Stop every Gemini CLI run the relay has in flight right now: a slow gemini-ask, gemini-plan, gemini-brainstorm or gemini-image. Each cancelled call returns an error to its caller. Nothing else is touched.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It explains the side effect precisely: each cancelled call returns an error to its caller, and nothing else is affected. This is strong transparency for a cancel operation, though it doesn't mention reversibility or authorization requirements.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tightly written sentences. The first states the action and scope, the second clarifies side effects and boundaries. Every phrase earns its place with no filler or repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a no-parameter cancellation tool with no output schema, the description is complete: it names the target, the scope, the effect on callers, and the exclusion of everything else. An agent has enough information to invoke it correctly and predict outcomes.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters and an empty schema, so there is no parameter meaning to convey. Per the baseline for 0-param tools, 4 is appropriate; the description correctly focuses on behavior rather than inventing parameter details.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Stop') and resource ('every Gemini CLI run the relay has in flight'), and explicitly names affected sibling tools (gemini-ask, gemini-plan, gemini-brainstorm, gemini-image). This clearly distinguishes it from the read/execute tools in the sibling list.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context for when to use the tool: when there are slow in-flight Gemini CLI runs that need to be aborted 'right now.' It also scopes what is and isn't touched ('Nothing else is touched'), effectively ruling out unintended effects. It doesn't explicitly name alternatives, but for a cancellation operation the context is sufficient.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
gemini-conversationsA
List recent Gemini conversations the relay can resume: id, last activity, working directory, first prompt. Pass an id to gemini-ask as conversationId to continue that thread. Reads agy's local store; no model turn.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | How many conversations to list, newest first (default 20, max 100). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the burden. It discloses that the tool reads 'agy's local store' and makes no model turn, signaling a read-only, low-cost operation. This is meaningful behavioral context beyond a generic list action.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, each adding distinct value: what is listed, how to use the result, and the operational side effect. The important action is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple list tool with one optional parameter and no output schema, the description covers the returned fields, the continuation workflow, and the read-only/no-model-turn behavior. Nothing essential is missing for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The only parameter, limit, is fully documented in the input schema with default, max, and ordering. The description does not need to restate it, and at 100% schema coverage the baseline is 3.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description leads with a specific verb and resource ('List recent Gemini conversations the relay can resume') and lists the returned fields. It also names gemini-ask as the downstream consumer, which helps an agent distinguish listing from asking.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly tells the agent to pass a returned id to gemini-ask as conversationId when continuing a thread, giving a clear use case. It does not state when to avoid this tool in favor of gemini-ask for a new conversation, but the context is clear enough.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
gemini-doctorA
Diagnose and verify Gemini CLI / Antigravity CLI (agy) installation, active backend, CLI version, and system readiness.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral disclosure burden. It lists what is checked (installation, backend, version, readiness) but does not explicitly state whether the tool is read-only, whether it modifies anything, what output format to expect, or whether it requires any prerequisites. 'Diagnose and verify' strongly implies non-mutating, but this is left implicit.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single sentence with no filler. The action verb is front-loaded, and every phrase adds a distinct diagnostic aspect: installation, backend, version, and readiness.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter diagnostic tool with no output schema and no annotations, the description covers the key dimensions an agent needs to understand its purpose. It does not state what the verification output looks like or how to interpret results, but the low complexity makes that a minor gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so there is no parameter documentation burden. Schema coverage is trivially complete, and the description does not need to explain parameters that do not exist.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses specific verbs ('diagnose', 'verify') and names exact resources: Gemini CLI / Antigravity CLI (agy) installation, active backend, CLI version, and system readiness. This clearly distinguishes it from sibling tools like ask-gemini or gemini-image, which are content-generation tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The diagnostic scope clearly implies use when checking environment health or CLI readiness, which is distinct from the content-generation siblings. However, it does not explicitly state when not to use it or name alternative tools for similar checks.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
gemini-fetch-chunkA
Retrieve one chunk of a chunked changeMode reply from gemini-ask, by the cacheKey and chunkIndex that reply reported. No model turn; the cache lives 10 minutes.
| Name | Required | Description | Default |
|---|---|---|---|
| cacheKey | Yes | The cache key provided in the initial changeMode response | |
| chunkIndex | Yes | Which chunk to retrieve (1-based index) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the disclosure burden. It discloses two non-obvious behaviors: no model turn is consumed and cached data expires after 10 minutes. For a simple retrieval tool this is adequate, though failure behavior on cache expiry is not stated.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with no filler; the verb and target are front-loaded and the behavioral caveats are packaged cleanly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter cached retrieval tool, the essentials are present: what to pass, where the values come from, and that the cache is short-lived. Minor omissions are the exact return shape and expiry/error behavior, but an agent has enough to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema covers both parameters 100%, so baseline is 3. The description adds provenance guidance—the cacheKey and chunkIndex come from the initial gemini-ask reply—which is more useful than either schema description alone.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a precise action: retrieve one chunk of a chunked changeMode reply, scoped to data produced by gemini-ask, identified by cacheKey and chunkIndex. This clearly distinguishes it from siblings like gemini-ask or gemini-cancel.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the trigger: use after gemini-ask returns a chunked changeMode reply and supplies cacheKey/chunkIndex. It identifies this as a retrieval step rather than a model call and gives a 10-minute validity window, but it does not explicitly name alternatives or exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
gemini-helpA
Print the active backend CLI's own --help text (agy or gemini), to see which flags the installed version supports.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the behavioral disclosure burden. It accurately states that the tool prints the CLI's --help text and mentions the two possible CLIs, 'agy or gemini', which is useful transparency. It does not discuss output size or failure modes, but for a zero-argument read-only help tool this is a minor omission.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence with no filler. It front-loads the action ('Print'), names the resource, and gives the purpose, all in one efficient, well-structured statement.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with no parameters and no output schema, the description is complete: it explains what is printed, which CLI's help text is shown, and why the agent would call it. Nothing essential is missing for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With zero parameters, the baseline is 4. The description adds meaning by clarifying that the tool needs no user input and instead introspects the 'active backend CLI', which helps the agent understand that invocation is automatic and requires no configuration.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb and resource: 'Print the active backend CLI's own --help text (agy or gemini)'. It also immediately explains the purpose, 'to see which flags the installed version supports', making the tool's role clear and distinct from the gemini-* sibling commands.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly implies when to use the tool: whenever the agent needs to discover which flags the installed backend CLI supports. It does not name an alternative or exclusion explicitly, but no sibling provides the same help/introspection function, so the usage context is reasonably clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
gemini-imageA
Generate images from text descriptions through the Antigravity CLI's image generation tool. Supports aspect ratio and output size selection, plus optional export directly into your project.
| Name | Required | Description | Default |
|---|---|---|---|
| size | No | Optional output resolution ('512', '1K', '2K' or '4K'). Omit to let Gemini pick. | |
| prompt | Yes | Detailed visual description of the image to create (subject, environment, lighting, artistic style, colors). | |
| outputPath | No | Optional path, relative to the project root, where the generated image should be copied (e.g., 'assets/hero.jpg'). A path escaping the project root is refused. | |
| aspectRatio | No | Aspect ratio of the generated image (default: '1:1'). | 1:1 |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It does disclose a material side effect: optional export directly into the project, and it places the tool in the Antigravity CLI context. However, it does not mention overwrite behavior, what happens when outputPath is omitted, or generation failure modes, leaving some behavioral gaps.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with no filler. The first sentence states the core purpose, and the second compactly lists the key options. Information is front-loaded and every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the 100% schema coverage and straightforward nature of image generation, the tool definition is largely complete: prompt, sizes, aspect ratios, and export path are all documented. The only notable gap is the lack of any statement about what the tool returns when outputPath is omitted, which matters more because there is no output schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description's second sentence ('Supports aspect ratio and output size selection, plus optional export') merely summarizes the schema's parameters without adding new meaning; the schema itself already documents each parameter precisely.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb phrase 'Generate images from text descriptions', naming the resource (image generation) and the input (text). It also names the context 'Antigravity CLI's image generation tool' and the supported options, clearly distinguishing it from the chat/planning sibling tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The intended use is clear from the first sentence: generate an image when you have a text description. The description also previews optional export, helping an agent infer workflow use. It does not explicitly name alternatives, but no sibling tool performs image generation, so the context is unambiguous.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
gemini-modelsA
Lists available Gemini models, default selections, reasoning capabilities, and active backend status.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. The verb 'Lists' signals a read-only operation, and 'active backend status' discloses that it queries live backend state. It does not mention rate limits or authentication, but for a zero-parameter discovery tool these are low-risk omissions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single efficient sentence that fronts the core action ('Lists available Gemini models') and then packs the important output dimensions into a compact list. Every phrase adds information without repetition or filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple no-parameter, no-output-schema tool, the description covers what the tool returns: models, defaults, reasoning capabilities, and backend status. It could add a bit more about output format or how status is represented, but nothing essential is missing for invoking it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters and the schema is an empty object with 100% coverage, so there is nothing for the description to add about parameter meaning. The baseline of 4 for no-parameter tools is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Lists') and identifies a distinct resource ('available Gemini models'), then enumerates the exact content categories returned: default selections, reasoning capabilities, and active backend status. This clearly separates it from action-oriented siblings like ask-gemini, gemini-plan, and gemini-image.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description makes the use context clear: an agent should call this when it needs to know which Gemini models are available, which are defaults, what reasoning capabilities exist, or whether the backend is active. It does not explicitly name alternatives or exclusions, but no sibling tool appears to cover model discovery, so the context is sufficient.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
gemini-pingA
Liveness check: echoes the prompt back from the relay itself, without touching the CLI or any model. Use gemini-doctor to check the CLI.
| Name | Required | Description | Default |
|---|---|---|---|
| prompt | No | Message to echo |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden. It meaningfully discloses that the tool avoids the CLI and any model and only echoes the prompt from the relay, strongly implying a safe, side-effect-free operation. It does not describe return format or error cases, but for a simple ping tool this is adequate.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two short sentences with no filler. The primary purpose is front-loaded, and the alternative routing is stated in the second sentence without redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a low-complexity, optional-parameter ping tool, the description fully enables correct selection and invocation. The absence of an output schema is not a significant gap because the echo behavior makes the return value self-evident.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema already fully describes the single 'prompt' parameter as 'Message to echo', so the description adds little beyond confirming the echo behavior. This matches the baseline expectation for full schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses specific language ('liveness check', 'echoes the prompt back') and precisely identifies the resource/scope: the relay itself, not the CLI or any model. It clearly differentiates this tool from gemini-doctor and other sibling tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly states when to use this tool for a relay liveness check and explicitly directs agents to gemini-doctor for CLI checks. This gives clear selection guidance and distinguishes the tool from its closest alternative.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
gemini-planA
Produce a read-only, phased implementation blueprint for one task — steps, dependencies, risks — on Gemini's highest reasoning effort. Never edits files and never answers general questions; use gemini-ask for those.
| Name | Required | Description | Default |
|---|---|---|---|
| task | Yes | The architectural task, complex feature, or refactoring goal to plan out. | |
| model | No | Model to plan with. Unlike gemini-ask, this tool pins 'gemini-3.8-flash-high' when you name none. 'gemini-3.1-pro-high' is the usual step up. | |
| effort | No | Reasoning effort level (default: 'high'). Allocates deep thinking tokens for comprehensive plan design. | high |
| addDirs | No | Additional directories to add to workspace context. | |
| context | No | Additional context, constraints, or reference files (supports @file syntax). | |
| includeUsage | No | Include token metrics in the response. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It discloses key behavioral traits: read-only nature, no file editing, restriction to a single task, and use of highest reasoning effort. It does not mention potential long-running behavior or response format, but the most important side-effect and scope constraints are clearly stated.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with zero waste. The core behavior and scope are front-loaded in the first sentence, and the exclusion/alternative is packed into the second. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only planning tool with 6 parameters and no output schema, the description covers purpose, scope, exclusions, and a routing pointer. The main gap is the absence of behavior around response structure or asynchronous execution, but the schema handles parameters and the description gives sufficient context for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Input schema has 100% description coverage for all 6 parameters, so the baseline is 3. The description adds no parameter-specific detail beyond the schema; it only reinforces that the tool plans a task, which is already clear from the 'task' parameter. This is acceptable since the schema already documents each parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description opens with a specific verb ('Produce') and identifies the deliverable: 'a read-only, phased implementation blueprint' with steps, dependencies, and risks. It clearly differentiates itself from gemini-ask by stating 'never answers general questions; use gemini-ask for those', so an agent can distinguish between the two without opening schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly states when to use this tool ('for one task' needing a blueprint) and when not to ('Never edits files and never answers general questions'), and names the alternative tool for general questions. This gives agents clear routing instructions with minimal inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
14 tool updates
v2.0.0- Removed
ask-gemini - Removed
brainstorm - Removed
fetch-chunk - Added
gemini-ask - Added
gemini-brainstorm - Added
gemini-cancel - Added
gemini-conversations - Added
gemini-fetch-chunk - Added
gemini-help - Changed
gemini-image3 fields changed- changed
Input schema / properties / aspectRatio / enumPrevious value: -[ - "1:1", - "16:9", - "9:16", - "4:3", - "3:4", - "3:2", - "2:3" -]New value: +[ + "1:1", + "16:9", + "9:16", + "4:3", + "3:4", + "3:2", + "2:3", + "5:4", + "4:5", + "21:9", + "4:1", + "1:4", + "8:1", + "1:8" +] - changed
Input schema / properties / outputPath / descriptionPrevious value: -"Optional relative file path in your project workspace where the generated image file should be copied (e.g., 'assets/hero.jpg')."New value: +"Optional path, relative to the project root, where the generated image should be copied (e.g., 'assets/hero.jpg'). A path escaping the project root is refused." - added
Input schema / properties / sizeAdded value: +{ + "description": "Optional output resolution ('512', '1K', '2K' or '4K'). Omit to let Gemini pick.", + "enum": [ + "512", + "1K", + "2K", + "4K" + ], + "type": "string" +}
- Added
gemini-ping - Changed
gemini-plan1 field changed- changed
Input schema / properties / model / descriptionPrevious value: -"Gemini model to use (default: 'gemini-3.8-flash-high' or 'gemini-3.1-pro-high')."New value: +"Model to plan with. Unlike gemini-ask, this tool pins 'gemini-3.8-flash-high' when you name none. 'gemini-3.1-pro-high' is the usual step up."
- Removed
Help - Removed
ping
9 tool updates
v1.2.0- First observed
ask-gemini - First observed
brainstorm - First observed
fetch-chunk - First observed
gemini-doctor - First observed
gemini-image - First observed
gemini-models - First observed
gemini-plan - First observed
Help - First observed
ping
TDQS
Scored across 11 tools
Each tool targets a distinct mode: gemini-ask for general answers, gemini-plan for blueprints, gemini-brainstorm for ideation, and gemini-image for generation. The descriptions explicitly cross-reference which tool to use, so there is little risk of selecting the wrong one.
All tools share the gemini- prefix and kebab-case style, making the family instantly recognizable. However, gemini-models and gemini-conversations are noun-style subcommands while others like gemini-ask and gemini-fetch-chunk are verb-style, a minor deviation from a strict verb_noun pattern.
11 tools is well within the ideal range and each tool earns its place: four generation modes, conversation and chunk retrieval support, cancellation, and diagnostics. There is no obvious redundancy or padding.
The surface covers generation, planning, ideation, images, conversation resumption, chunked-response retrieval, cancellation, and environment diagnostics—no dead ends for the relay's stated purpose. Optional capabilities like conversation deletion are not implied by the server's scope.
Maintenance
Related MCP Connectors
MCP server unifying ERPs, CRMs, APIs and knowledge base for Claude, ChatGPT and Gemini.
A comprehensive Model Context Protocol (MCP) server that enables AI assistants to interact with yo…
One MCP endpoint for Claude, GPT & Gemini: 100+ tools + no-code connectors + agent workers.
MCP server for building and testing AI agents with multi-model experimentation and insights.
Related MCP Servers
- AlicenseAqualityFmaintenanceModel Context Protocol (MCP) server implementation that enables Claude Desktop to interact with Google's Gemini AI models.659 npm261MIT
- FlicenseNot gradedqualityDmaintenanceA server implementing the Model Context Protocol that enables AI assistants like Claude to interact with Google's Gemini API for text generation, text analysis, and chat conversations.-
- AlicenseNot gradedqualityNot gradedmaintenanceAn MCP server implementation that allows using Google's Gemini AI models (specifically Gemini 1.5 Pro) through Claude or other MCP clients via the Model Context Protocol.1MIT
- AlicenseNot gradedqualityDmaintenanceA Model Context Protocol (MCP) server implementation for the Google Gemini language model. This server allows Claude Desktop users to access the powerful reasoning capabilities of Gemini-2.0-flash-thinking-exp-01-21 model.1MIT