Codex Antigravity Subagent MCP
该 MCP 服务器封装 Antigravity/Gemini 子代理的调用、后台任务管理和可视化看板,支持同步/流式执行、HITL 交互和状态查询。
delegate_to_gemini:同步委派任务,默认模型 gemini-3.8-flash-high,支持 mode/agent/model/effort/project/new_project/session_mode/working_directory 等参数。
start_gemini_task:后台启动长任务并立即返回 job_id,适合 /teamwork-preview 等耗时操作。
get_gemini_task:按 job_id 查询后台任务状态和最终结果。
cancel_gemini_task:按 job_id 取消运行中的后台任务。
interact_gemini_task:向 stream 模式任务注入交互式指令或人工输入,并可选择 end_session 优雅结束会话。
finish_gemini_task:主动结束 stream 模式任务,关闭 stdin 管道并等待任务收敛为 success 终态。
antigravity_status:检查 AGY CLI 版本、可用模型以及默认模型 gemini-3.8-flash-high 是否存在。
open_dashboard:打开 Web 看板,实时展示子代理的工作流、动作流与物理日志。
Delegates tasks to Google Antigravity CLI (agy), leveraging Google's Gemini models for AI-driven command-line execution.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Codex Antigravity Subagent MCPRun a /teamwork-preview to analyze the codebase and suggest improvements"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Codex Antigravity Subagent MCP
项目状态:Final GA(v1.5.3 状态完全可信与看板视图稳健修复版)
本项目已全面解决用户复测报告抓出的 4 项残留真问题(stream 模式深层子代理变动穿透 SSE 指纹广播job_updated、前端快照防覆盖用户自定义筛选/分页视图、按持久化唯一标识selectedSubagentId锚定子代理与 DAG 节点彻底根除换人失焦、BUT_NOT_FINISHED转折分句未闭环一票否决规则),离线 11 套回归套件与 23 项专项反例 100% 成功通过,生产目录 105 个文件零污染。
⚡ 5 分钟快速上手 (Quick Start)
通过本指南,你可以在 5 分钟内将 Google Antigravity 的高并发级联 Agent 能力无缝集成到你的 Codex 工作流中。
Codex (LLM / IDE)
│
├─ stdio (MCP Protocol)
│
Codex Antigravity Subagent MCP (node src/server.mjs)
│
├─ Dual-Track Process Supervisor (Print / Stream NDJSON)
├─ Real-time God's-eye Dashboard (HTTP :3721 / SSE)
│
Google Antigravity CLI (agy)
│
Gemini 3.8 Flash High (Multi-Agent Swarm / Teamwork)第一步:准备基础环境
在使用本 MCP 前,请确保本机已具备以下环境:
Node.js:
>= 20.0.0(推荐 Node 20 LTS 或 Node 22)Git:用于拉取代码
Codex:正常安装与配置
Google Antigravity CLI (
agy):Google 官方代理终端
第二步:安装与授权 Google Antigravity CLI
根据你的操作系统,使用官方一键脚本安装 agy:
Windows (PowerShell):
irm https://antigravity.google/cli/install.ps1 | iexWindows 默认安装路径:
%LOCALAPPDATA%\agy\bin(即C:\Users\<用户名>\AppData\Local\agy\bin)。
macOS / Linux (Bash/Zsh):
curl -fsSL https://antigravity.google/cli/install.sh | bash首次登录与信任验证:
安装完成后,在终端运行一次 agy 完成浏览器登录并信任工作区:
# 验证版本
agy --version
# 首次运行并完成浏览器登录授权
agy第三步:克隆 MCP 仓库与离线自测
将本项目克隆到本地目录(以 C:\Tools\antigravity-subagent-mcp 或 ~/Tools/... 为例):
# 1. 克隆仓库
git clone https://github.com/zamatewi-cell/antigravity-subagent-mcp.git
cd antigravity-subagent-mcp
# 2. 安装依赖 (严格基于 package-lock.json 对齐)
npm ci
# 3. 运行全量离线自动化测试套件 (11 门单测聚合)
npm run test:offline若看到 11 套测试全部绿色通过(
All Tests Passed),说明 Node.js 依赖、MCP 本地组件及离线回归测试正常(AGY CLI 凭据与模型连通性由后续步骤验证)。
第四步:配置接入 Codex
你可以选择命令行快速添加或编辑配置文件(强烈推荐)。
方式 A:命令行快速添加(CLI)
在终端中执行:
# Windows 示例(请替换为你本机的实际仓库路径)
codex mcp add antigravity-subagent -- node "C:\Tools\antigravity-subagent-mcp\src\server.mjs"
# macOS / Linux 示例
codex mcp add antigravity-subagent -- node "/path/to/antigravity-subagent-mcp/src/server.mjs"方式 B:直接编辑 config.toml(推荐,最稳妥方案)
在 Codex 配置文件 ~/.codex/config.toml 中追加如下配置。明确指定 AGY_CLI_PATH 可以彻底杜绝因环境变量 PATH 缺失导致的寻址失败:
[mcp_servers.antigravity-subagent]
command = "node"
args = ["C:\\Tools\\antigravity-subagent-mcp\\src\\server.mjs"]
[mcp_servers.antigravity-subagent.env]
# 显式指定 agy.exe 绝对路径(Windows 官方默认安装路径示例如下)
AGY_CLI_PATH = "C:\\Users\\YOUR_USERNAME\\AppData\\Local\\agy\\bin\\agy.exe"
# 默认使用 auto-approve 模式,确保无头执行时工具调用免人工交互阻塞
ANTIGRAVITY_PERMISSION_MODE = "auto-approve"
# 默认模型(可选,默认 gemini-3.8-flash-high)
ANTIGRAVITY_DEFAULT_MODEL = "gemini-3.8-flash-high"第五步:验证连通性
打开 Codex,输入 /mcp 即可查看已配置的 MCP 工具列表。随后向 Codex 发送指令进行连通性自检:
提示词示范:
“检查 antigravity-subagent MCP 是否可用,调用antigravity_status,告诉我当前 AGY CLI 版本、默认模型以及运行状态。”
如果配置正常,Codex 会调用 antigravity_status 并返回真实的就绪结构:
{
"status": "READY",
"cli_path": "C:\\Users\\...\\AppData\\Local\\agy\\bin\\agy.exe",
"cli_version": "1.x.x",
"default_model": "gemini-3.8-flash-high",
"default_model_available": true,
"default_permission_mode": "auto-approve",
"models": [
"gemini-3.8-flash-high",
"gemini-3.8-pro",
"..."
],
"agents": ["..."],
"errors": []
}第六步:核心使用场景实战
现在你可以随心所欲地让 Codex 调度 Gemini 多代理团队协同工作:
场景 1:普通单次委派任务(代码审查 / 独立分析)
你对 Codex 说:
“把当前仓库代码交给 Gemini 做一次深度代码审查,重点排查潜在并发漏洞与未处理异常,只出具报告不要修改源码。”
(Codex 将自动调用delegate_to_gemini并实时返回执行进度)
场景 2:启动大型长任务或 Teamwork 多代理集群
你对 Codex 说:
“在后台拉起 Gemini 任务,使用/teamwork-preview对当前项目做一次完整架构、安全与测试审计。”
(Codex 将调用start_gemini_task生成任务 ID,由 Project Sentinel、Top Orchestrator 与多个 Worker 级联并行推进)
场景 3:双向交互流传输(Interactive HITL 多轮会话)
你对 Codex 说:
“以 stream 模式启动一个 Gemini 交互任务,先给出重构方案,等我确认后再继续写代码。”
(走start_gemini_task(session_mode="stream")→ 收到第一轮回复 →interact_gemini_task注入第二轮提示 →finish_gemini_task优雅结束并收敛为 success)
场景 4:一键唤起上帝视角监控看板
你对 Codex 说:
“打开 Antigravity 监控看板。”
(Codex 将调用open_dashboard并在默认浏览器弹出http://localhost:3721)
Related MCP server: sub-antigravity
🖥️ 实时可视化监控看板 (Visual Dashboard)
本项目内置纯原生零外部依赖(Zero External Dependencies)的暗黑极客风 Web 监控看板:

看板核心特性:
微观上帝视角网格:主编排器(Orchestrator)、子代理矩阵(Workers)、审计员(Auditor)全员卡片网格,实时展示各 Agent 当前步数、微观动作描述与激活工具标签(如
[run_command]、[write_to_file])。原生 SVG Agent DAG 拓扑连线:纯原生 HTML5/SVG 渲染父子调用拓扑,搭载硬件加速的 CSS 霓虹流光动画(
wireFlow/wireBreath),直观呈现 Agent 编排层级。SSE 增量差异广播:利用服务端事件流(
/api/stream)推送细粒度变更(job_created,job_updated,job_removed),前端内存字典局部打补丁,彻底杜绝高频全量 JSON 广播带来的网络浪费与重绘卡顿。历史任务分页与检索:支持按任务状态(
running、success、error等)精确过滤,支持关键词和 Job ID 毫秒级检索。人机软介入(HITL)控制台:在网页端直接向运行中子进程 stdin 注入按键或指令(快捷键
[Y]、[N]、[Enter]),支持一键安全终止与优雅结束。
启动方式:
MCP 联动模式:设置环境变量
ANTIGRAVITY_ENABLE_DASHBOARD=1,或通过 MCP 工具open_dashboard。独立进程模式:直接运行
npm run dashboard(支持--port <port>与--open)。
🛠️ MCP 工具清单与接口契约
Codex 连接本服务后,将获得以下 8 个高内聚工具:
工具名称 | 读写属性 | 核心功能与使用说明 |
| 写入/执行 | 同步/准实时委托:阻塞执行并定期(每 2s)通过 |
| 写入/执行 | 异步后台启动:立即返回全局唯一 |
| 只读查询 | 进度轮询与探查:获取任务最新状态、详细进度对象 |
| 写入/交互 | HITL 人机交互介入:向运行中的长会话注入第二轮输入,支持 |
| 写入/控制 | 会话优雅收官:关闭 stream 任务的输入管道,等待子进程退出并稳定收敛至 |
| 写入/控制 | 统一任务取消:原子取消 |
| 只读查询 | 环境探查诊断:快速检测本机 |
| 只读辅助 | 浏览器唤起:确保看板服务在线并自动调用系统默认浏览器打开 |
🛡️ 架构设计与系统硬化亮点
┌──────────────────────────────────────────────────────────────────────────────┐
│ Codex Antigravity Subagent MCP Engine │
├──────────────────────────────────┬───────────────────────────────────────────┤
│ 1. 双轨执行架构 (Dual-Track) │ 2. 动态双层解析架构 (Two-Tier Transcript) │
│ - Batch Print (--print=...) │ - parseLocalTranscript: 静态解析真 LRU │
│ - Stream NDJSON (bidirect) │ - composeAgentTree: 实时动态拓扑装配 │
├──────────────────────────────────┼───────────────────────────────────────────┤
│ 3. 生命周期统一状态机 │ 4. 纵深防御与安全控制 │
│ - 首发终止原因胜出保护 │ - 独立看板 HTTP 409 防篡改只读拦截 │
│ - 严防 completed 误判与漂移 │ - Host 白名单与非本地 Origin 403 阻断 │
│ - 同目录串行互斥锁 (AsyncLock)│ - 物理日志凭据(Bearer/Token)脱敏 │
└──────────────────────────────────┴───────────────────────────────────────────┘双轨执行架构(Dual-Track):
自动化/流水线委托保持极稳的
--output-format json --print=<prompt>轨道;交互长会话无缝接入官方
--input-format stream-json --output-format stream-json轨道,使用StreamLineParser解决 TCP 粘包与断行恢复。
两层 Transcript 解析架构:
底层
parseLocalTranscript负责单文件语法静态解析与真 LRU 缓存管理;上层
composeAgentTree动态递归遍历子代理目录结构,彻底解决“父 transcript 不变导致子 Agent 状态冻结”的历史顽疾。
工作区目录串行锁(
withDirectoryLock):同一工作目录并发委托自动排队互斥,杜绝并发 Antigravity 实例引发的文件读写冲突;不同工作目录完全并发。
严格的只读与安全防护机制:
独立运行的看板进程(无内存句柄时)严禁通过 REST API 伪造取消或注入输入(一律拦截并返回
HTTP 409 Conflict);严格防御 DNS Rebinding 与 CSRF 攻击,仅放行受信任本地请求(
localhost/127.0.0.1)。
❓ 常见问题排错 (Troubleshooting)
Q1: antigravity_status 报告 ready: false 或找不到 CLI 路径?
原因:
agy.exe未在系统PATH环境变量中,或安装在非标准路径。解决:在
~/.codex/config.toml中显式配置AGY_CLI_PATH为绝对路径,例如:AGY_CLI_PATH = "C:\\Users\\YOUR_NAME\\AppData\\Local\\agy\\bin\\agy.exe"。
Q2: 任务执行时提示工具权限被拒绝(TOOL_CONFIRMATION_DENIED)?
原因:Antigravity 默认在需要执行终端或写入文件时弹出确认提示,无头(Headless)运行中无法交互确认。
解决:确保环境变量
ANTIGRAVITY_PERMISSION_MODE="auto-approve"(本项目默认值),该模式会在底层自动注入--dangerously-skip-permissions。
Q3: 首次调用时提示认证失败或阻塞?
原因:Google Antigravity CLI 尚未完成初始授权或工作区信任。
解决:打开系统原生命令行,直接输入
agy回车,按照浏览器指引完成一次授权,并对当前工作目录执行/trust。
🧪 离线与端到端自动化测试验证
本项目包含完善的自动化测试基础设施,涵盖 11 套单测与端到端套件:
# 执行全部 11 门离线回归套件(沙箱隔离,零生产数据污染):
npm run test:offline
# 独立单测执行:
node test/hitl-stream-transport.test.mjs # 交互式流协议与优雅结束单测 (8/8 PASS)
node test/hitl-real-agy.e2e.test.mjs # 真实 AGY 进程物理 E2E 验证 (PASS)
node test/pagination.test.mjs # 历史任务分页过滤单测 (12/12 PASS)
node test/sse-delta.test.mjs # SSE 细粒度增量广播单测 (7/7 PASS)
node test/dag-topology.test.mjs # 原生 SVG DAG 拓扑单测 (4/4 PASS)
node test/hitl.test.mjs # 软介入与 EPIPE 防崩安全单测 (9/9 PASS)
node test/cancel-consistency.test.mjs # 统一取消与防漂移单测 (6/6 PASS)
node test/counterexamples.test.mjs # 深度反例与状态防伪单测 (18/18 PASS)
node test/dashboard.test.mjs # Web 看板 REST/SSE 接口单测 (7/7 PASS)
node test/lifecycle.test.mjs # 生命周期离线验证单测 (PASS)
node test/progress.test.mjs # 进度解析与持久化单测 (PASS)
node test/diagnostics.test.mjs # 诊断日志脱敏与错误分析单测 (PASS)📜 详细版本变更历史 (Changelog)
Version 1.5.1 (Feature Freeze Final Release):
Graceful Stream Session Termination (
finish_gemini_task):Added official MCP tool
finish_gemini_task(job_id)and Dashboard endpointPOST /api/jobs/:id/finish, allowing clients and operators to gracefully terminate long-running stream sessions by safely invokingstdin.end()and awaiting exit completion.Enhanced
interact_gemini_taskwithend_session: boolean (default: false)to send a final instruction and immediately close the input stream in a single step.
Stream Exit State Machine Decoupling:
Decoupled process close handling in stream mode from single-turn
parseAgyJson. Directly leverages stream-parsedlastTurnResult, preventing NDJSON{"event": "result"}wrappers from corrupting task status and ensuring clean convergence tostate: "success".
Production MCP Real-AGY E2E Verification:
test/hitl-real-agy.e2e.test.mjsrewritten to connect directly as a live MCP client throughsrc/server.mjs, validating the complete lifecycle (start_gemini_task->interact_gemini_task->finish_gemini_task) with actual AGY binaries and exit code 0.
Version 1.5.0:
Interactive Stream Transport & Native Multi-Turn HITL Closure:
Dual-Track Execution Architecture: Retains the rock-solid one-off delegation track (
--output-format json --print=<prompt>) for batch automation, while introducing the official Headless Streaming channel (--input-format stream-json --output-format stream-json --dangerously-skip-permissions) activated viasession_mode: "stream".NDJSON Stream Protocol Parser & Framing: Implemented
StreamLineParserinsrc/stream-transport.mjsto reliably handle chunk fragmentation, line-boundary recovery, and TCP socket packet coalescing. Standardized input packaging viaencodeStreamUserMessageadhering strictly to official Headless NDJSON event schema ({"event": "user", "message": {"content": "..."}}).Live Handshake & Conversation Tracking: Intercepts
event: initto immediately extract the genuine backendconversation_id(job.conversationId), continuously digestsstep_updateto stream text deltas, and handlesresultto increment multi-turn counters (numTurns).Persistent Session HITL Pipeline: Upgraded
sendInputToJoband DashboardPOST /api/jobs/:id/interactto directly inject structured NDJSON into running child processes, enabling persistent multi-turn conversations in the same child process without restarting CLI or creating disjoint sessions.New MCP Tool
interact_gemini_task: Directly exposed to MCP clients for orchestrating interactive conversations with long-running subagents.Real AGY End-to-End Test Suite: Validated full physical round-trip on actual local hardware with
D:\Antigravity\agy\bin\agy.exe(test/hitl-real-agy.e2e.test.mjs), confirming prompt -> response -> Web interaction injection -> turn 2 response -> graceful exit (Exit Code: 0).
Version 1.4.0:
Historical Jobs Pagination & Multidimensional Filtering (R1):
Added query parameter parsing to
GET /api/jobssupportingpage(default: 1),limit(default: 20, clamped to 1-100),stateexact filtering, andsearchsubstring search matching prompt and job ID.Standardized paginated response structure
{ data: Job[], pagination: { total, page, limit, totalPages, hasMore } }with transparent backward-compatibility for unparameterized requests.
Single-Job Replacement SSE Event Streaming & Lightweight Snapshots (R2):
Re-architected
GET /api/streamto dispatch initial lightweightsnapshot(all active jobs + recent 20 terminal jobs, offloading deep history to REST pagination) with incrementing cursorseq.Upgraded differential change detection from full-list broadcast to single-job entity replacement (
job_created,job_updated,job_removed), sending lightweight:keep-alivecomments during idle intervals.Hardened
getJobFingerprintto deeply serialize subagent status, step, current action, last tool, and topology tree, ensuring subagent-only micro-activities immediately trigger live streaming updates.
Native Zero-Dependency SVG Agent DAG Topology (R3):
Augmented recursive agent tree assembly with directional topological metadata (
parentId,depth,childrenIds,nodeType).Rendered compact hierarchical DAG visualizer using a dual-layer architecture: underlying SVG cubic Bézier flowing wires with pure CSS hardware-accelerated neon breathing keyframes (
wireFlow,wireBreath), and top-layer absolute HTML interactive cards.
Experimental Human-in-the-loop (HITL) Soft Intervention Pipeline (R4):
Configured child process spawn with
stdio: ["pipe", "pipe", "pipe"]as an experimental bidirectional stream channel designed for stream protocol interactions (--input-format stream-json).Implemented asynchronous
sendInputToJobwith physical write confirmation callbacks, newline auto-completion, and robust EPIPE / pipe-destruction crash guards.
Version 1.3.3:
Two-tier transcript parser architecture resolving worker cache freezing.
Decoupled
src/directory-lock.mjsmodule.Memory bounding with genuine LRU eviction on transcript and disk caches.
Multi-OS GitHub Actions CI workflow (
.github/workflows/ci.yml).
Version 1.3.2:
Strict standalone monitor readonly protection (HTTP 409 Conflict).
Subagent state priority realignment (
killedstatus permanence).CSRF & DNS Rebinding protection (Host header whitelisting & Origin 403).
Version 1.3.1:
DOM-based XSS elimination via HTML entity escaping.
Tightened CORS headers, physical log tail sanitization, earliest termination reason protection.
Version 1.3.0:
Standalone real-time visual monitor dashboard with God's-eye View Grid.
Micro-activity timeline and zero-polling SSE push.
New MCP tool
open_dashboard.
Version 1.2.2:
Complete negation and in-progress defense in lifecycle evaluation.
Queued task cancellation and restart convergence.
Version 1.2.1:
Strict evidence-based lifecycle evaluation.
Working MCP progress notification stream via
ctx.mcpReq.notify.Asynchronous per-workspace serialization directory lock.
Version 1.2.0:
Deep cascade subagent transcript penetration for spawned Teamwork workers.
Version 1.1.0:
Real-time progress and activity tracking with lightweight job persistence.
Version 1.0.3:
Credential redaction in diagnostic logs and cross-platform process tree cleanup.
📄 开源许可证
本项目基于 MIT License 开源。欢迎提交 Issue 与改进建议!
Available Tools
8 toolsantigravity_status检查 Antigravity 子代理状态ARead-only
检查 AGY CLI 版本、可用模型,以及默认模型 gemini-3.8-flash-high 是否存在。
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the read-only, non-destructive, open-world safety profile. The description adds useful behavioral scope by stating exactly what is inspected: CLI version, available models, and the presence of the default model. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single concise sentence with the action and checked resources front-loaded. It contains no filler, tautology, or unnecessary background.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter, read-only status tool with supporting annotations, the description names all relevant checks and is sufficient for an agent to select and invoke it correctly. The lack of an output schema is not a significant gap given the low invocation complexity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters and the schema contains an empty properties object, so there is no parameter detail to document. The description appropriately focuses on what the status check reports rather than inputs.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('检查') and concrete resources: AGY CLI version, available models, and whether the default model gemini-3.8-flash-high exists. This clearly distinguishes it from sibling task-management tools like delegate_to_gemini or cancel_gemini_task.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied from the tool's nature as a status/health-check command, but the description does not explicitly say when to use it versus alternatives or provide any exclusions. No comparable status-checking sibling is named, so the guidance is implicit rather than explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
cancel_gemini_task取消 Gemini 子代理任务ADestructive
终止当前 MCP 进程中仍在运行的 Gemini/Teamwork 后台任务。
| Name | Required | Description | Default |
|---|---|---|---|
| job_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already state destructiveHint=true and readOnlyHint=false, matching the terminate semantics. The description adds useful context by scoping cancellation to tasks 'still running in the current MCP process', which clarifies that tasks in other processes or already-completed tasks are not affected. It does not describe side effects like partial work loss, but the destructive annotation covers the core safety signal.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single concise sentence, front-loaded with the action verb, and contains no filler. It efficiently communicates scope, target resource, and the action.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple one-parameter cancellation tool, the description, combined with destructiveHint and the 'current MCP process' scoping, is mostly complete. The main gap is not stating that job_id should come from a starting/delegating tool, but this is inferable from siblings and not critical for invoking the tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description never mentions job_id or explains where it comes from. The parameter name and UUID format are self-explanatory to some degree, so the tool is callable, but the description adds no meaning beyond what the schema already provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb '终止' (terminate) with a clear resource: Gemini/Teamwork background tasks still running in the current MCP process. This clearly distinguishes the tool from siblings like start_gemini_task and get_gemini_task, which have different actions.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase '仍在运行' implies the tool should be used for running tasks, which gives some usage context. However, there is no explicit when/when-not guidance, no mention of alternatives such as get_gemini_task for finding the job_id, and no note about what to do if the task is already finished.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
delegate_to_gemini委派任务给 Gemini 子代理ADestructive
通过 Antigravity CLI 调用 Gemini,默认模型为 gemini-3.8-flash-high。保留 prompt 并补充绝对工作目录,启用斜杠命令展开,可直接调用 /teamwork-preview <任务> 等系统命令。调用会等待结果;预计超过数分钟的 Teamwork 任务请使用 start_gemini_task。
| Name | Required | Description | Default |
|---|---|---|---|
| mode | No | Antigravity 执行模式。plan 只规划;accept-edits 允许提出和执行修改。 | |
| agent | No | 可选的 Antigravity agent 名称。 | |
| model | No | Antigravity 模型名;默认 gemini-3.8-flash-high。 | |
| effort | No | 推理强度。 | |
| prompt | Yes | 交给 Gemini 的完整提示词。保留开头的系统斜杠指令,并在末尾补充绝对工作目录,例如:/teamwork-preview <完整任务>。 | |
| project | No | 可选的 Antigravity project ID 或名称。 | |
| new_project | No | 为本次调用新建 Antigravity project。 | |
| session_mode | No | AGY 运行轨道模式:print 为常规单次委托执行(自动化/批量),stream 为交互式流传输长会话(支持多轮人机交互)。 | |
| add_directories | No | 额外加入 AGY 工作区的目录。 | |
| continue_latest | No | 继续最近一次 Antigravity 会话。 | |
| conversation_id | No | 继续指定 Antigravity 会话;值来自上一次结果的 conversation_id。 | |
| permission_mode | No | Antigravity 权限模式;默认 auto-approve(auto-approve 确保在无头 headless 模式下工具调用免人工交互确认)。 | auto-approve |
| timeout_seconds | No | 单次 AGY 运行超时,30 到 21600 秒。Teamwork 长任务应提高该值或使用后台任务工具。 | |
| working_directory | No | Gemini 工作目录的绝对路径。涉及当前仓库时必须传入当前 Codex 工作目录。 |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
注释已声明 destructiveHint=true,描述没有矛盾。描述额外揭示了调用会等待结果、会保留并补充 prompt、启用斜杠命令展开等行为,这些超出了注释本身能提供的信息。只是没有明确说明可能产生的破坏性副作用,但已有 destructiveHint 覆盖。
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
描述仅三句话,第一句直接说明工具作用和默认模型,第二句说明 prompt 处理和斜杠命令,第三句给出等待语义和与兄弟工具的区分。信息密度高且没有冗余。
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
对于一个有 14 个参数、无输出 schema 的复杂工具,描述涵盖了核心调用语义、参数补充规则、超时建议和备选工具。虽然没有明确返回值,但 schema 也完全没有输出定义。基本足够,但可以再补充说明返回结果的结构。
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
schema 覆盖率为 100%,每个参数都有描述。描述对 prompt 参数特别补充了'保留开头的系统斜杠指令并在末尾补充绝对工作目录'的行为,并给出具体示例,这对 agent 构造正确调用至关重要。working_directory 和 permission_mode 的描述也提供了 schema 之外的调用上下文。
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
描述明确说明通过 Antigravity CLI 调用 Gemini,并指出默认模型,还点名了与 start_gemini_task 的区别(等待结果 vs 后台任务)。资源、动作和范围都清晰,能与同级工具区分。
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
描述明确指出预计超过数分钟的 Teamwork 任务应使用 start_gemini_task,timeout_seconds 参数也补充了长任务应提高值或使用后台任务工具。这提供了明确的替代工具路由和适用场景。
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
finish_gemini_task优雅结束 Gemini 子代理长会话 (HITL)A
主动结束处于 stream 模式的运行中任务,关闭 stdin 管道并等待任务优雅收敛进入 success 终态。
| Name | Required | Description | Default |
|---|---|---|---|
| job_id | Yes | 目标 stream 运行中任务的 job_id |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description discloses meaningful mechanics—closing the stdin pipe and waiting for graceful convergence—beyond what the annotations state. Annotations already indicate the operation is not read-only and not destructive, and the description aligns with those traits. It does not cover failure or timeout behavior, but that is a minor gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One dense sentence conveys action, condition, mechanism, and outcome with no filler. It is front-loaded with the primary purpose and remains easy to parse.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter tool with strong schema coverage and supporting annotations, the description is largely complete. The only notable omission is what happens if the task fails to converge or times out, along with any return-value expectations, but neither is critical for selecting and invoking this tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema provides 100% coverage for the single job_id parameter with a precise description, so the tool description adds no extra parameter-level meaning. Baseline 3 applies because the schema carries the burden.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific operation—proactively ending a running stream-mode task—and states the goal of graceful convergence into the success terminal state. This clearly distinguishes it from the sibling cancel_gemini_task, which implies abrupt termination.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives a clear precondition: use it on running tasks that are in stream mode. This tells an agent when to invoke the tool. It does not explicitly state when not to use it or contrast it with cancel_gemini_task, so a full exclusion is missing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_gemini_task查询 Gemini 子代理任务ARead-only
按 job_id 查询后台 Gemini/Teamwork 任务状态和最终结果。
| Name | Required | Description | Default |
|---|---|---|---|
| job_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and destructiveHint=false, covering the safety profile. The description adds that the tool returns status and final result, which is useful context, but it doesn't disclose potential error states, whether the task must be completed, or any behavior around unfinished tasks.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence with no wasted words. It immediately states the action, the resource, and the parameter without redundant phrasing.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read-only lookup with one parameter and annotations covering safety, the description is mostly complete. It names the high-level return content (status and final result), though it lacks output shape details and relationship to sibling tools. These are minor gaps given the low complexity and available annotations.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, but the single parameter job_id is self-explanatory and is further constrained by uuid format and pattern. The description confirms job_id is the lookup key, yet it doesn't explain where the id comes from or what values are valid beyond the schema pattern, so compensation for missing schema descriptions is partial.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb (查询/query), a resource (后台 Gemini/Teamwork 任务状态和最终结果), and the key parameter (job_id). It is clear and distinct from mutation siblings like start_gemini_task and cancel_gemini_task, though it doesn't explicitly differentiate from antigravity_status.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage when you need the status or final result of a Gemini task by job_id, but it provides no explicit guidance on when not to use this tool or which sibling to choose instead. It doesn't mention that job_id would come from start_gemini_task or delegate_to_gemini.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
interact_gemini_task与运行中的 Gemini 子代理交互 (HITL)A
向正在运行的长会话子代理(尤其处于 stream 模式的任务)注入交互式指令或人工输入,支持可选在完成当轮后优雅结束会话。
| Name | Required | Description | Default |
|---|---|---|---|
| input | Yes | 需要向子代理输入的指令或交互内容 | |
| job_id | Yes | 目标运行中任务的 job_id | |
| end_session | No | 是否在本次交互发送后自动关闭输入流,触发任务在当前轮次完成后收敛为 success 终态 |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare the tool is not read-only and not destructive. The description adds meaningful behavioral detail beyond annotations: it injects input into a running stream and, when end_session is set, closes the input stream after the current turn to trigger a graceful success terminal state. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single dense sentence front-loads the core purpose and then adds the optional end-session behavior. Every phrase earns its place, with no filler or repetition of schema details.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The definition covers the target state (running long-session, stream mode), the action (inject input), and the optional termination semantics. Annotations handle the safety profile and the schema covers parameters, so the main missing piece is return-value or failure behavior, which is not critical for invoking this tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all three parameters. The description mirrors the input and end_session semantics but adds no new format, constraints, or usage details beyond what the schema provides, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description names a specific verb-resource pair: inject interactive instructions/human input into a running long-session sub-agent, with special emphasis on stream-mode tasks. It clearly conveys what the tool does, though it doesn't explicitly name a sibling to distinguish itself from, so it stops short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context for when to use it: for running long-session sub-agents, especially those in stream mode. It does not list exclusions or mention alternatives like finish_gemini_task or cancel_gemini_task, so it lacks explicit when-not-to-use guidance, but the stated context is specific and actionable.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
open_dashboard打开 Antigravity 子代理实时可视化看板ARead-only
在系统默认浏览器中打开独立的可视化监控看板(Web UI),实时展示全员子代理微观工作流、动作流与物理日志。
| Name | Required | Description | Default |
|---|---|---|---|
| auto_open | No | 是否自动在默认浏览器中弹出看板窗口 |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and destructiveHint=false. The description adds valuable behavioral context by noting that the tool opens a browser window as a side effect and that the dashboard shows real-time data. This goes beyond the annotations and helps the agent set user expectations, though it does not address possible failures or return behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, compact sentence that front-loads the main action and purpose. There is no redundant wording or repetition of schema information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is simple, has no required parameters, and annotations cover its read-only, non-destructive nature. The description adequately explains what the dashboard shows and that it opens in the default browser, though it does not mention what the tool returns after opening, which would be slightly more complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the single auto_open parameter already includes a clear description and default value. The tool description adds no further parameter-specific detail, but the schema carries the semantic weight, so the baseline score of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb (open), a clear resource (standalone visualization dashboard / Web UI), and its purpose (real-time display of subagent workflows, action streams, and physical logs). It is clearly distinguishable from the sibling task-management tools like start_gemini_task or cancel_gemini_task.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage context: an agent would call this when the user wants to visually monitor subagent activity. However, it does not explicitly state when to prefer it over alternatives or when not to use it, and it names no sibling tools as alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
start_gemini_task后台启动 Gemini 子代理ADestructive
后台启动 Antigravity/Gemini 任务并立即返回 job_id,适合 /teamwork-preview 等长任务。之后用 get_gemini_task 查询;后台任务依赖当前 MCP 服务进程保持运行。
| Name | Required | Description | Default |
|---|---|---|---|
| mode | No | Antigravity 执行模式。plan 只规划;accept-edits 允许提出和执行修改。 | |
| agent | No | 可选的 Antigravity agent 名称。 | |
| model | No | Antigravity 模型名;默认 gemini-3.8-flash-high。 | |
| effort | No | 推理强度。 | |
| prompt | Yes | 交给 Gemini 的完整提示词。保留开头的系统斜杠指令,并在末尾补充绝对工作目录,例如:/teamwork-preview <完整任务>。 | |
| project | No | 可选的 Antigravity project ID 或名称。 | |
| new_project | No | 为本次调用新建 Antigravity project。 | |
| session_mode | No | AGY 运行轨道模式:print 为常规单次委托执行(自动化/批量),stream 为交互式流传输长会话(支持多轮人机交互)。 | |
| add_directories | No | 额外加入 AGY 工作区的目录。 | |
| continue_latest | No | 继续最近一次 Antigravity 会话。 | |
| conversation_id | No | 继续指定 Antigravity 会话;值来自上一次结果的 conversation_id。 | |
| permission_mode | No | Antigravity 权限模式;默认 auto-approve(auto-approve 确保在无头 headless 模式下工具调用免人工交互确认)。 | auto-approve |
| timeout_seconds | No | 单次 AGY 运行超时,30 到 21600 秒。Teamwork 长任务应提高该值或使用后台任务工具。 | |
| working_directory | No | Gemini 工作目录的绝对路径。涉及当前仓库时必须传入当前 Codex 工作目录。 |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
标注已声明 readOnlyHint=false、destructiveHint=true,描述没有与标注矛盾。描述补充了标注之外的有用行为信息:任务依赖 MCP 进程存活、异步返回 job_id 的工作流。但鉴于标注已标记破坏性操作,描述并未就破坏性影响、权限需求或安全性提供额外语境,对高破坏性的工具来说披露不够充分。
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
仅两句描述,核心异步语义前置,随后的进程依赖警告和查询指引都是必要信息,无冗余。结构与篇幅对该工具恰到好处。
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
工具复杂度高(14 参数、4 个枚举、无输出 schema),描述虽交代了异步工作流(返回 job_id、用 get_gemini_task 查询)和进程依赖,但未说明返回值结构或其他操作细节。由于无输出 schema,描述本应承担更多返回信息说明职责;params 已由 schema 100% 覆盖,整体基本够用但偏薄。
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema 描述覆盖率达 100%,所有 14 个参数(含 prompt 的斜杠指令与工作目录示例、permission_mode 的 headless 说明、timeout 的 teamwork 建议)均已由 schema 文档化。描述本身未添加参数级信息,仅提及返回 job_id。在高覆盖场景下取基线 3 分,schema 已承担主要说明职责。
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
描述明确表述了具体动词(后台启动)、资源(Antigravity/Gemini 任务)和模式(立即返回 job_id 的异步行为),并指出适用于 /teamwork-preview 等长任务。通过'后台'和'立即返回 job_id'的表述与 get_gemini_task(查询)以及 delegate_to_gemini(很可能为同步)等兄弟工具清晰区分。虽然未显式点名 delegate_to_gemini,但异步语义本身已足够区分。
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
描述给出了明确的使用场景——'适合 /teamwork-preview 等长任务',并指出后续用 get_gemini_task 查询,还警示后台任务依赖 MCP 服务进程保持运行,这是重要的使用前提。但未明确说明何时不该使用(如短任务或需要交互的场景),也未点名 delegate_to_gemini 作为同步替代方案,缺少明确的排除性指引。
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
4 tool updates
v1.5.3- Changed
delegate_to_gemini1 field changed- added
Input schema / properties / session_modeAdded value: +{ + "default": "print", + "description": "AGY 运行轨道模式:print 为常规单次委托执行(自动化/批量),stream 为交互式流传输长会话(支持多轮人机交互)。", + "enum": [ + "print", + "stream" + ], + "type": "string" +}
- Added
finish_gemini_task - Added
interact_gemini_task - Changed
start_gemini_task1 field changed- added
Input schema / properties / session_modeAdded value: +{ + "default": "print", + "description": "AGY 运行轨道模式:print 为常规单次委托执行(自动化/批量),stream 为交互式流传输长会话(支持多轮人机交互)。", + "enum": [ + "print", + "stream" + ], + "type": "string" +}
6 tool updates
v1.3.0- First observed
antigravity_status - First observed
cancel_gemini_task - First observed
delegate_to_gemini - First observed
get_gemini_task - First observed
open_dashboard - First observed
start_gemini_task
TDQS
Scored across 8 tools
Most tools target distinct actions: starting background tasks, starting synchronous tasks, querying, canceling, finishing, interacting, checking status, and opening the dashboard. The only mild ambiguity is between start/delegate and cancel/finish, but the descriptions clearly separate async vs sync and graceful vs hard termination.
The naming convention is largely consistent snake_case with verb_object patterns like start_gemini_task, get_gemini_task, and cancel_gemini_task. Minor deviations are delegate_to_gemini and antigravity_status, but they remain readable and predictable.
Eight tools is well-scoped for a task/subagent management server. Each tool covers a meaningful lifecycle or support function without unnecessary duplication or bloat.
The core lifecycle is covered: start sync/async, query status/results, cancel, finish, interact, and verify environment. A minor gap is the lack of a way to list all jobs or past tasks without knowing a job_id, though the dashboard partially compensates.
Maintenance
Related MCP Connectors
- DazbenchOAuthapp.dazbench
Task management your AI agents can actually run. One line becomes a context-ready task over MCP.
Develop, manage, and debug Railway projects, services, and deployments from within agents.
Project memory, tasks and Telegram notifications for your coding agent. Chip account required.
Source-checked CLI guides and model-aware planning for Claude Code, Codex, and Grok Build.
Related MCP Servers
- AlicenseBqualityAmaintenanceAn MCP bridge that lets Codex delegate long-running agent work to the Antigravity CLI, providing observable and resumable tool-based execution with project scoping.1473 PyPI12MIT
- AlicenseAqualityCmaintenanceEnables AI harnesses to delegate code analysis, modification, testing, and long-running tasks to the locally installed Antigravity CLI via stdio MCP, with job status tracking and conversation continuity.3MIT
- AlicenseAqualityCmaintenanceEnables OpenAI Codex to delegate tasks to Google Antigravity CLI, running them headlessly and polling for results.4MIT
- AlicenseAqualityBmaintenanceEnables Codex to delegate bounded code analysis, review, isolated edits, multimodal research, and native image generation to supervised Antigravity workers, with queued execution, retries, and inspectable patches.232MIT