Skip to main content
Glama
irrumi

antigravity-plugin-codex

by irrumi

antigravity-plugin-codex

Use Antigravity from a local Codex chat to get a second opinion, review a diff, or delegate a coding task. This plugin connects Codex to the Antigravity CLI (agy) and returns the result for you to inspect.

Русский · Compatibility evidence · Architecture · Security

Quick start

Requires Node.js 22+, Git, a local Codex CLI with plugin support, and an installed, signed-in Antigravity CLI. This project is installed from source; it has no npm release or separate API key.

git clone https://github.com/irrumi/antigravity-plugin-codex.git
cd antigravity-plugin-codex
codex plugin marketplace add .
codex plugin add antigravity-plugin-codex@antigravity-local
node plugins/antigravity-plugin-codex/src/cli.mjs doctor

Open a new local Codex chat and ask:

Ask Antigravity to explain Git worktrees in one sentence. Do not run commands or change files.

Codex starts the local Antigravity task and brings its answer back to the chat. In a verified smoke run, the MCP result reported state: "succeeded", exitCode: 0, and the requested answer AGY_CODEX_E2E_OK. The doctor command checks the installed CLI and supported flags; it does not verify sign-in unless you add --probe-auth, which sends a real request.

In Windows PowerShell, use codex.cmd if execution policy blocks its .ps1 shim. Sign in with an interactive agy session if needed. Installation uses Codex's marketplace commands and does not install or reconfigure Antigravity. See update and uninstall for later changes.

Related MCP server: sub-antigravity

Why use it?

Switching between coding agents manually means copying context out, running a second CLI, and reconciling its answer or edits. This bridge lets Codex request a second opinion and track the local task while you stay in the same chat. Codex remains the coordinator and asks you to review any returned code changes.

  • Ask or review: Send a focused question, selected files, or a staged, unstaged, or revision-based diff.

  • Delegate code work: Give Antigravity an independent clone of committed HEAD; inspect its returned diff before applying anything.

  • Track and stop tasks: Get status and results from the MCP server, with time and output limits and cancellation support.

  • See a Codex worker: When native subagents are available, the skill runs Antigravity through a visible Codex subagent. The external CLI itself remains a separate process.

Codex chat → Codex skill/worker → local MCP adapter → agy → answer or diff

Common requests

Ask Antigravity for a second opinion on src/parser.ts in /absolute/path/to/project.

Have Antigravity review the staged diff in /absolute/path/to/project and return concrete findings.

Delegate empty-input handling in /absolute/path/to/project to Antigravity; report its diff and test results.

The review request needs your explicit acceptance of a writable disposable workspace because this CLI has no verified fully read-only mode. A delegated task runs against committed HEAD; your uncommitted files are not copied. See review safety and tool behavior.

Installation notes and supported environments

This repository supplies a Codex plugin and skill, a Node.js adapter with no runtime dependencies, and a local MCP stdio server. The adapter discovers the standard user-local agy executable before trying PATH; an Antigravity IDE launcher is not enough. In the tested Codex version, the plugin's MCP working directory resolves to the installed plugin root.

The project has no published npm package or GitHub release. Local Windows use was verified with Codex CLI 0.155.1, Antigravity CLI 1.2.13, and Node 26.5.0. Ubuntu CI runs with Node 22 and 24; Windows and macOS CI jobs are paused after a known test path-comparison failure. Linux/macOS desktop use is not verified. See compatibility evidence for the exact test scope.

Approve the intended MCP calls under your normal Codex policy. The install does not replace your Codex config; testing used a separate CODEX_HOME. Keep your existing CODEX_HOME for normal use because it also affects authentication. Natural-language requests and the installed antigravity skill are the supported entry points; there are no /agy:* slash commands.

Tools

MCP tool

Purpose

antigravity_doctor

Version/help/capability checks. probeAuth: true makes a small real request.

antigravity_ask

Start a prompt in a scratch directory; optional selected context.

antigravity_delegate

Start a code task in an independent clone of committed HEAD.

antigravity_review

Select staged, unstaged, or base-to-working-tree diff.

antigravity_status

Poll task state.

antigravity_result

Retrieve answer, stderr, exit code, duration, diff and changed files.

antigravity_cancel

Request owned process-tree termination; then poll for completion.

antigravity_forget

Release a completed result from memory, preserving workspace files.

Every task requires an explicit absolute cwd. Starts immediately return a task ID; later calls use taskId. IDs and the concurrency limit belong to one MCP server process. Closing/restarting Codex loses task metadata; workspaces remain. No detached daemon or automatic task resumption exists.

Example tool input:

{
  "cwd": "/absolute/path/to/project",
  "prompt": "Fix empty-input handling and run the relevant tests. Report what changed.",
  "contextFiles": ["src/parser.ts"],
  "timeoutMs": 120000
}

On Windows, use C:/path/to/project. Selected context is text sent to the provider, not a file overlay. Delegate copies committed HEAD only: dirty/staged/untracked changes in the source are preserved and not copied. Return patches are never automatically applied, merged, pushed or published. New untracked file names are listed separately; inspect their contents in the returned workspace. Check reports are external, unverified tool reports.

Review safety

No verified full read-only mechanism is available in the tested CLI. A nonempty review fails with READ_ONLY_UNAVAILABLE by default. --mode plan is not treated as a filesystem guarantee.

If you explicitly accept review in a writable disposable directory, pass:

{"cwd":"/absolute/project","mode":"staged","allowSnapshotWrites":true}

The diff is supplied as text; the source repo is not the process working directory. This avoids ordinary edit collisions but is not OS containment against malicious tools. Empty diffs return no_changes without launching Antigravity. For comparison to a revision use mode: "base", base: "main"; this is git diff main, not a merge-base diff.

Visible subagent workflow

The bundled skill requests this workflow by default:

Codex coordinator → built-in Codex worker → Antigravity MCP/CLI → result → independent checks

Ask: “Use Antigravity through a Codex subagent to review these two files. Wait for its result and verify the findings.” The worker appears in the native subagent activity list on clients that support it. That entry represents the Codex intermediary; the external agy process does not become a native Codex agent or a selectable Codex model.

The coordinator creates one worker and passes the task, absolute project path, selected context, permissions and limits. The worker loads the skill, runs Antigravity, waits for a terminal result and checks its claims. It does not create another intermediary. It also owns MCP task IDs in its session, so the coordinator requests cancellation through the worker rather than assuming those IDs work across sessions. For large reviews, use focused sections and report which files were covered.

If subagents are unavailable or you request direct execution, the same MCP/CLI task runs in the current agent without a subagent card. Direct MCP or CLI calls alone do not create native subagents. This routing is skill guidance, not a new MCP endpoint; the existing isolation, permission and review-consent requirements still apply. Reinstall the updated plugin and start a new chat to load the revised skill; see Update and uninstall.

The workflow was exercised in a Windows Codex desktop chat: a built-in worker called the bundled CLI for a focused review, received succeeded / exit 0, and independently reproduced a finding. See evidence and remaining limits.

Configuration and direct CLI

Defaults: 5-minute task deadline, 1 MiB input, 1 MiB combined stdout/stderr, 2 concurrent tasks, 100 retained tasks. Optional config example:

{"timeoutMs": 120000, "maxConcurrent": 1}

Save it as JSON and pass its absolute path through AGY_CODEX_CONFIG in the Codex host environment, or use --config FILE with the direct CLI. Unspecified fields keep their defaults; see the full example. Do not put secrets in the config. executableArgs is a trusted host setting for wrappers, never a task argument. Windows .cmd/.bat/.ps1 launchers are rejected; use the native .exe or node with an absolute JS entrypoint.

model is passed verbatim through the locally confirmed --model flag. No alias table and no global model-setting edits. Unsupported capabilities fail closed. An unset model uses the CLI's normal selection. Headless requests use stdin, not shell quoting or long command-line arguments.

node plugins/antigravity-plugin-codex/src/cli.mjs doctor --probe-auth
node plugins/antigravity-plugin-codex/src/cli.mjs ask --request request.json
node plugins/antigravity-plugin-codex/src/cli.mjs delegate --request request.json
node plugins/antigravity-plugin-codex/src/cli.mjs review --request review.json

The direct CLI is synchronous; use MCP for start/status/result/cancel. A successful process launch is not a successful task: the bridge requires a final SUCCESS result and exit 0, and reports permission soft denials separately. Authentication is unknown unless a live request succeeds; a settings directory is never evidence of authentication. The real CLI can still try interactive authentication in headless mode, so the bridge detects its prompts and stops.

MCP-only fallback

For Codex versions without plugin installation, use the officially supported local stdio MCP registration:

codex mcp add antigravity -- node /absolute/repo/plugins/antigravity-plugin-codex/src/cli.mjs serve

For custom config append --config /absolute/config.json after serve. Do not register this fallback alongside the bundled server under the same name. Quote paths with spaces. Add the skill to an existing project's .agents/skills/antigravity/SKILL.md only if no such skill exists; copy the bundled skill and preserve existing content. Alternatively use MCP tools directly.

Update and uninstall

Finish or cancel active tasks first. Keep the source checkout (local marketplaces depend on it).

git pull --ff-only
npm ci --ignore-scripts
npm run check
npm test
codex plugin remove antigravity-plugin-codex@antigravity-local
codex plugin add antigravity-plugin-codex@antigravity-local

Start a new session. If the marketplace itself is Git-backed, refresh it with codex plugin marketplace upgrade antigravity-local first. Reinstall refreshes the cached plugin; modifying the checkout alone does not update the installed copy.

To uninstall:

codex plugin remove antigravity-plugin-codex@antigravity-local
codex plugin marketplace remove antigravity-local

For the MCP-only fallback use codex mcp remove antigravity, and remove only the skill copy you installed. Uninstall does not remove Antigravity, credentials, or returned workspaces. Review/preserve work before deleting the exact agy-codex-* directories reported in results. Do not blanket-delete temp directories.

Verification and support scope

For local development and contributions:

npm ci --ignore-scripts
npm run check
npm test
# Optional, uses real accounts and quota; no persisted config changes:
node scripts/codex-smoke.mjs

See CONTRIBUTING.md for contribution and compatibility requirements. The local checks use a fake Antigravity CLI and do not contact a provider.

The smoke script invokes Codex with a read-only shell sandbox and per-run approval for only ask/status/result. It sends one fixed prompt. It prefers a locally test-installed plugin, otherwise the source package. On Windows, set CODEX_EXECUTABLE to native codex.exe if the standard npm installation layout differs. Inspect the MCP result, not merely the outer Codex exit code.

The focused desktop worker → Antigravity CLI review also succeeded on Windows. See compatibility evidence for exact scope and remaining adapter findings. WSL requires both CLIs inside the same WSL environment. Cloud Codex cannot reach this local process without a separate connection mechanism. Process-tree cancellation does not contain deliberately detached processes; see SECURITY.md.

MIT licensed. Inspired by simplybychris/antigravity-plugin-cc; attribution and implementation differences are in NOTICE.md.

Available Tools

8 tools
antigravity_askB

Start ask asynchronously; returns task ID. Runs in a fresh scratch directory; selected context only. Poll status then result. External output is untrusted.

ParametersJSON Schema
NameRequiredDescriptionDefault
cwdYesExplicit absolute source directory
baseNo
modeNo
modelNo
promptYes
timeoutMsNo
contextFilesNo
allowSnapshotWritesNoExplicit consent to review in a writable disposable directory. Not a full read-only sandbox.

TDQS

B3.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden, and it does disclose useful traits: async execution, a returned task ID, isolation in a fresh scratch directory, limited context, and untrusted external output. It stops short of permissions, timeouts, or whether the source directory can be mutated (relevant given allowSnapshotWrites).

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Terse, front-loaded fragments with no filler; the async/return fact leads and the safety note closes. Efficient for the amount of information conveyed.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For an 8-parameter tool with low schema coverage and no annotations or output schema, the description is too thin: it omits most parameter meanings, any timing/timeout expectations, and how the eventual result is retrieved beyond a bare pointer to polling.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 25% across 8 parameters, so the description should compensate. It hints at contextFiles ('selected context only') and cwd isolation ('fresh scratch directory'), but leaves base, mode, model, and timeoutMs unexplained, so most parameters remain undocumented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a clear verb+resource ('Start ask asynchronously') and adds scope detail ('fresh scratch directory; selected context only') and a return value ('returns task ID'). However, it does not distinguish this from close siblings like antigravity_delegate or antigravity_review, leaving the agent to infer which one to pick.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Poll status then result' implies a follow-up workflow using the sibling antigravity_status/antigravity_result tools, which is helpful. But it never states when to choose this tool versus delegate/review, nor any preconditions or exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

antigravity_cancelC

cancel a task in this MCP server session. Forget releases in-memory results, preserves workspace.

ParametersJSON Schema
NameRequiredDescriptionDefault
taskIdYes

TDQS

C2.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false and destructiveHint=false, so the safety profile is partly covered. The description adds that in-memory results are released while the workspace is preserved, which is genuinely useful behavioral context, but it is garbled and omits state/idempotency details. No annotation contradiction.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Very short and front-loaded, which is good, but the second sentence is ambiguous and poorly worded ('Forget releases in-memory results, preserves workspace'), so not every sentence cleanly earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutating, single-parameter tool with no output schema, the description leaves the agent without taskId semantics, the resulting task state, or idempotency behavior. It is under-specified relative to the tool's complexity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There is one required parameter (taskId) with 0% schema description coverage, and the description says nothing about what taskId is or where it comes from. The description fails to compensate for the schema gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('cancel a task') scoped to 'this MCP server session', which is clear enough to distinguish from most siblings. The sideline reference to 'Forget' muddies it slightly, but the core purpose is legible.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No explicit when-to-use or when-not-to-use guidance, and no condition that selects this over antigravity_forget or antigravity_status. The second sentence hints at a relationship with forget but never states when one should be preferred.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

antigravity_delegateA

Start delegate asynchronously; returns task ID. Edits an independent clone of committed HEAD. No merge/push. Poll status then result. External output is untrusted.

ParametersJSON Schema
NameRequiredDescriptionDefault
cwdYesExplicit absolute source directory
baseNo
modeNo
modelNo
promptYes
timeoutMsNo
contextFilesNo
allowSnapshotWritesNoExplicit consent to review in a writable disposable directory. Not a full read-only sandbox.

TDQS

A3.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does well: async execution, task-ID return, edits confined to an independent clone of committed HEAD, no merge or push, and a warning that external output is untrusted. It omits permission requirements, timeout/expiry behavior, and what happens if the clone diverges, which keeps it short of a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Five terse clauses, front-loaded with the action and return value, then the isolation guarantee and the security caveat. No filler sentences; every statement carries operational weight.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For an 8-parameter mutation-adjacent tool with no annotations and no output schema, the description supplies the critical async/clone/untrusted-output context but leaves the majority of parameters (mode, base, model, timeoutMs, contextFiles) unexplained. Adequate for invoking it, incomplete for tuning or predicting behavior.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Eight parameters with only 25% schema description coverage, and the description adds no per-parameter meaning at all. Schema-undocumented params like base, mode, model, timeoutMs, and contextFiles (including its 64-item cap) get no explanation or example from the description, so the coverage gap is not compensated.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a concrete action (start an async delegate task) and states its key output (task ID), plus the workspace model: it edits an independent clone of committed HEAD with no merge/push. That distinguishes it from read-only siblings like ask/review. The phrase 'Start delegate' is slightly awkward, but the resource and effect are unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Poll status then result' implies the follow-up workflow using the antigravity_status and antigravity_result siblings, giving implied routing. However, it never explicitly names those tools, nor states when to prefer antigravity_ask or antigravity_review instead of delegating a mutation.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

antigravity_doctorA

Inspect installed agent CLI/version/help. probeAuth sends a small real prompt and may incur usage. Never infer auth from config directories.

ParametersJSON Schema
NameRequiredDescriptionDefault
probeAuthNo

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full disclosure burden, and it delivers two meaningful traits: probeAuth performs a real call that 'may incur usage' (a cost/side-effect warning), and the explicit anti-pattern 'Never infer auth from config directories'. It still omits whether the inspection itself is read-only or what it returns, so it is not fully complete.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences with zero filler; the primary purpose is front-loaded, followed by the parameter cost caveat and the methodology warning in priority order.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-parameter diagnostic tool with no output schema, the description covers what it does and the main behavioral risk. It could add whether the command is non-destructive and roughly what it reports, but nothing critical to correct invocation is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must explain the single parameter, and it does: probeAuth 'sends a small real prompt and may incur usage'. That conveys the semantics and the cost implication the bare boolean type could not. Minor gap: no mention of the default when omitted.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Inspect) and resource (installed agent CLI/version/help), which is a diagnostic intent clearly distinct from the ask/delegate/review/status siblings. It does not explicitly name an alternative, so it falls short of the 5 bar.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied (run this to inspect the installed CLI) but there is no explicit when-to-use/when-not framing or named sibling for comparison. The probeAuth caveat gives conditional guidance for that parameter, not for choosing the tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

antigravity_forgetB

forget a task in this MCP server session. Forget releases in-memory results, preserves workspace.

ParametersJSON Schema
NameRequiredDescriptionDefault
taskIdYes

TDQS

B3.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false and destructiveHint=false, and the description adds real context beyond them: what is released (in-memory results) versus what survives (workspace). This tells the agent the operation mutates session state without destroying persisted work, which is exactly the kind of detail annotations cannot convey. It still omits reversibility of the forget action.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences with the core action front-loaded and the effect detail second; there is no filler. The phrasing is slightly clipped and the second sentence could be tighter, but nothing is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-parameter session-mutation tool with no output schema, the description covers the main effect but omits how to obtain taskId, whether forgetting is reversible, and what happens on an unknown id. Minimum viable but with clear gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% for the single taskId parameter, and the description never explains where taskId comes from or what form it takes (presumably returned by antigravity_delegate/antigravity_ask). A single undocumented parameter leaves the agent guessing about how to supply the required argument.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb-resource pair (forget a task) scoped to the MCP server session, which is concrete and actionable. It does not distinguish itself from the closest sibling, antigravity_cancel, so an agent cannot tell from the description alone which of the two is appropriate.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit when-to-use or when-not-to-use guidance, and no mention of alternatives such as antigravity_cancel. The only usage signal is implicit in the effect statement 'releases in-memory results, preserves workspace,' which hints at intent but does not route the agent.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

antigravity_resultC
Read-only

result a task in this MCP server session. Forget releases in-memory results, preserves workspace.

ParametersJSON Schema
NameRequiredDescriptionDefault
taskIdYes

TDQS

C2.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false, setting a low bar. The phrase "preserves workspace" adds a modest behavioral claim beyond the annotations, but the surrounding sentence is incoherent and it is unclear what state is touched or returned.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Brief, but the size comes from under-specification rather than economy. The first sentence is not front-loaded with a clear action and the second sentence's meaning is ambiguous, so neither sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema and no annotations covering return values, the description should explain what a "result" actually is, yet it does not. For a session-scoped tool sitting among seven siblings, this leaves an agent unable to know when or how to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There is one parameter (taskId) with 0% schema description coverage, so the description must compensate and it does not explain what a taskId is, where to obtain it, or whether the task must have been started in the current session. The parameter is effectively undocumented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

"result a task in this MCP server session" is grammatically broken and never states a clear verb+resource (retrieve? submit? finalize?). An agent cannot confidently tell whether this fetches a task's output or produces a result, and it is not clearly distinguished from siblings like antigravity_status or antigravity_forget.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The second sentence gestures at a distinction from a forget-style operation ("Forget releases in-memory results"), but it is garbled enough that no reliable when-to-use rule can be extracted. No explicit conditions, prerequisites, or named alternatives are given.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

antigravity_reviewB

Start review asynchronously; returns task ID. Read-only unavailable: nonempty diff requires explicit allowSnapshotWrites consent. Poll status then result. External output is untrusted.

ParametersJSON Schema
NameRequiredDescriptionDefault
cwdYesExplicit absolute source directory
baseNo
modeNo
modelNo
promptNo
timeoutMsNo
contextFilesNo
allowSnapshotWritesNoExplicit consent to review in a writable disposable directory. Not a full read-only sandbox.

TDQS

B3.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and discloses several non-obvious traits: the operation is asynchronous and returns a task ID, it is not truly read-only (nonempty diff needs allowSnapshotWrites consent), and external output is untrusted. It omits auth/permission requirements and failure behavior, which keeps it short of a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four dense, front-loaded fragments with no filler and the async/return behavior stated first. The telegraphic phrasing trades some readability for compactness, but nothing is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For an 8-parameter, annotation-free tool with no output schema, the description covers the async workflow, the consent gating, and the untrusted-output warning, which is the important behavioral context. It is still incomplete on several parameter meanings and on what the eventual result contains.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 25% across 8 parameters, so the description must compensate but largely does not. It clarifies the consent condition tied to allowSnapshotWrites (which the schema already partly documents) and implies mode=base, but base, model, prompt, timeoutMs, contextFiles, and cwd remain undocumented in both places.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

"Start review asynchronously; returns task ID" gives a specific verb, resource, and outcome, and the reference to polling status and result differentiates it from the status/result/cancel siblings. It is clear what the tool does, though it does not explicitly contrast itself with antigravity_ask or antigravity_delegate.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

"Poll status then result" implies the follow-up workflow and the consent condition hints at a usage prerequisite, but there is no explicit statement of when to choose this over antigravity_ask or antigravity_delegate. Usage is implied rather than stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

antigravity_statusD
Read-only

status a task in this MCP server session. Forget releases in-memory results, preserves workspace.

ParametersJSON Schema
NameRequiredDescriptionDefault
taskIdYes

TDQS

D1.8/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=true and destructiveHint=false, yet the description states that it "[f]orgets/releases in-memory results" — a state mutation — which conflicts with a read-only profile. The definition also never says what is returned (task status payload, terminal vs pending states, etc.), so an agent cannot rely on it for behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is short, but the second sentence appears to be unrelated/leftover text about forgetting and releasing results, which wastes space and confuses rather than front-loading the core action. Two sentences exist but one of them does not earn its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With one required parameter, no output schema, and annotations that are contradicted by the text, the description leaves the agent without the information needed to call the tool correctly or interpret it. It should at minimum state what status is returned and how taskId is obtained.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

taskId has 0% schema description coverage and no inline description, so the description must compensate; it only implies the parameter identifies "a task" in the session without stating the format, where the ID comes from, or how workspace scope interacts with it.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The phrase "status a task in this MCP server session" gestures at a verb (status) and resource (task) but never says concretely whether it fetches, sets, or reports task status, and it is not distinguished from the sibling antigravity_result. The second sentence then describes forgetting/releasing behavior, which muddies the purpose further.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit when-to-use or when-not-to-use guidance, and no alternative sibling is named. The clause "preserves workspace" vaguely gestures at a contrast with antigravity_forget, but the agent is not told the condition under which this tool should be selected over result, review, or cancel.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 8 tool updatesv0.1.0
    • First observedantigravity_ask
    • First observedantigravity_cancel
    • First observedantigravity_delegate
    • First observedantigravity_doctor
    • First observedantigravity_forget
    • First observedantigravity_result
    • First observedantigravity_review
    • First observedantigravity_status

TDQS

B3/5.0

Scored across 8 tools

Disambiguation4/5

Each tool targets a distinct action in the async task lifecycle: doctor inspects, ask/delegate/review start different work modes, and status/result/cancel/forget manage existing tasks. However, status and result have nearly identical descriptions, and the four task-session tools share boilerplate wording, which could cause minor hesitation.

Naming Consistency4/5

All names use the antigravity_ prefix and snake_case, forming a predictable scheme. The suffixes mix action verbs (ask, delegate, review, cancel, forget) with noun-like states (doctor, status, result), which is a minor deviation rather than a serious inconsistency.

Tool Count5/5

Eight tools are well-scoped for a plugin that exposes agent CLI operations and task management. The count supports distinct workflows without obvious redundancy or bloat.

Completeness4/5

The set covers setup/diagnosis, three work modes, and the full async task lifecycle (status/result/cancel/forget). It lacks a task-listing or workspace-cleanup operation, but the core ask/delegate/review workflows are supported with viable workarounds.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    A
    quality
    C
    maintenance
    Enables AI harnesses to delegate code analysis, modification, testing, and long-running tasks to the locally installed Antigravity CLI via stdio MCP, with job status tracking and conversation continuity.
    3
    MIT
  • F
    license
    Not graded
    quality
    C
    maintenance
    Acts as a local stdio MCP service that lets Codex delegate implementation tasks to Antigravity CLI while conserving tokens; it returns compact review packets with changed files, diff stats, and run history for Codex to review and iterate.
    -
  • A
    license
    A
    quality
    C
    maintenance
    Enables OpenAI Codex to delegate tasks to Google Antigravity CLI, running them headlessly and polling for results.
    4
    MIT