Skip to main content
Glama
HackerWilson

agy-mcp-server

by HackerWilson

agy-mcp-server

一个零第三方 Python 运行依赖的 MCP stdio 服务,用于调用本机 agy CLI。支持同步执行、 超时自动转后台任务、会话续跑,以及任务状态和结果查询。

要求

  • Python 3.9+

  • macOS 或 Linux

  • 已安装并能独立运行的 agy

服务不会读取或复制 OAuth token;认证由 agy 自身完成。

Related MCP server: mcp-job-queue

安装

直接使用源码:

git clone https://github.com/HackerWilson/agy-mcp-server.git
cd agy-mcp-server
python3 server.py

也可以安装命令行入口:

pipx install git+https://github.com/HackerWilson/agy-mcp-server.git
# 或
uv tool install git+https://github.com/HackerWilson/agy-mcp-server.git

配置

使用源码:

{
  "mcpServers": {
    "antigravity": {
      "command": "python3",
      "args": ["/absolute/path/to/agy-mcp-server/server.py"]
    }
  }
}

使用 pipx 或 uv tool 安装后:

{
  "mcpServers": {
    "antigravity": {
      "command": "agy-mcp-server",
      "args": []
    }
  }
}

MCP stdio 服务由客户端按需启动,并通过标准输入输出通信。直接执行后没有终端提示、一直等待 输入属于正常现象;客户端关闭连接后,服务进程也会退出。

如果客户端的环境变量中找不到 agy,可以显式指定路径:

{
  "env": {
    "ANTIGRAVITY_MCP_AGY_PATH": "/absolute/path/to/agy"
  }
}

工具

  • run_antigravity:同步执行新任务,达到等待阈值后返回后台 Job。

  • antigravity_continue:续跑已有会话,同样支持超时降级。

  • antigravity_start:直接启动异步任务。

  • antigravity_job_status、antigravity_job_result、antigravity_job_cancel:管理后台任务。

  • antigravity_sessions、antigravity_transcript:读取本机历史会话和执行轨迹。

  • antigravity_doctor、antigravity_models:环境诊断和模型列表。

read_only=true 会启用 agy --sandbox,但这不是严格的文件系统只读保证。 yolo=true 会传递 --dangerously-skip-permissions,仅应在可信工作目录中使用。 timeout_seconds 只控制同步工具等待多久后返回后台 Job;print_timeout_seconds 控制 agy 任务本身的最长执行时间,默认为 3600 秒,可设置为 60–86400 秒。 续跑时可以传入完整会话 ID,也可以传入至少 8 位的十六进制前缀;短前缀只有在本机会话 数据库中唯一匹配时才会展开。完整 UUID 在本机数据库中不存在时会拒绝启动,避免 agy 忽略 --conversation 后静默创建新会话。任务完成后,状态和结果会分别显示请求/实际会话 ID、续跑判定和 turn 数。 agy 要求 prompt 作为 -p 的命令行参数;运行期间,本机同一用户的进程查看工具可能看到 prompt,因此不要在任务文本中直接放置密码或 token。单次 prompt 上限为 120 KiB。

Job 状态默认保存在 ~/.gemini/antigravity-mcp/jobs。目录权限为 0700,文件权限为 0600;可通过 ANTIGRAVITY_MCP_JOBS_DIR 修改位置。

更新与卸载

# 源码安装
git -C /path/to/agy-mcp-server pull --ff-only

# pipx
pipx upgrade agy-mcp-server
pipx uninstall agy-mcp-server

# uv
uv tool upgrade agy-mcp-server
uv tool uninstall agy-mcp-server

源码更新不会热替换已经运行的 stdio MCP 进程;更新后需要让 MCP Client 断开并重新创建该 进程。若客户端会缓存失败或已断开的 MCP 通道,则需要重启客户端应用。

测试

python3 -m unittest -v

测试使用临时 fake agy,不会访问网络、读取认证信息或消耗模型额度。

参考

Available Tools

10 tools
antigravity_continueA

CONTINUE an existing Antigravity conversation by ID. Synchronous execution with a bounded timeout guard (auto-converts to a background job if still running). Supports YOLO autonomous execution or lightweight terminal sandboxing.

ParametersJSON Schema
NameRequiredDescriptionDefault
yoloNoIf true (default), runs with --dangerously-skip-permissions for auto-approving tools.
modelNoOptional model identifier override.
promptYesFollow-up instructions or questions.
work_dirNoOptional working directory path.
read_onlyNoIf true, enables Antigravity --sandbox; workspace writes may still be allowed. Defaults to false.
conversation_idYesThe Conversation ID to continue (for example, '123e4567-e89b-12d3-a456-426614174000').
timeout_secondsNoSeconds to wait before returning the background job ID. Range 1-40; defaults to 40.
print_timeout_secondsNoMaximum agy execution time. Range 60-86400; defaults to 3600 seconds.

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations present, the description carries the full behavioral burden, and it does disclose meaningful behavioral traits: synchronous execution, a bounded timeout guard, automatic conversion to a background job if still running, and support for two execution modes (YOLO autonomous execution vs terminal sandboxing. It does not mention return payload shape, error behavior, or permissions, but it offers far more transparency than a typical terse mutation tool description.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is three compact sentences with no filler. The main action is front-loaded, and the behavioral caveats are packed efficiently into the subsequent clauses.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity and the absence of an output schema, the description covers the key execution modes, timeout guard, and background-job fallback behavior, and the schema fully documents all parameters. What is missing is an explicit mention of how to retrieve results or how the background-job conversion connects to sibling job tools. That gap is real but moderate, and the schema already hints at background-job ID returns.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is already 3; the description adds value beyond the schema by framing yolo and read_only as two execution modes ('YOLO autonomous execution' vs 'lightweight terminal sandboxing') and by explaining the timeout behavior as an auto-conversion to a background job. It does not elaborate every parameter, but the high schema coverage makes that unnecessary.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the core purpose: 'CONTINUE an existing Antigravity conversation by ID', which identifies both the action and the specific resource. It is distinct from starting a new conversation, but it does not explicitly name sibling tools like antigravity_start or antigravity_start to reinforce the contrast. That differentiation is implied rather than stated.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear context for when to use the tool: you have an existing conversation ID and want to continue it. However, it does not provide explicit when-not-to-use guidance or point to siblings like antigravity_job_status/result for background follow-up. The usage case is inferred from 'by ID' rather than directly contrasted with alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

antigravity_doctorA

Diagnose Antigravity CLI health, authentication token status, YOLO permission mode, and storage metrics.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A3.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. 'Diagnose' implies a read-only operation, but it does not explicitly state non-destructiveness, authentication requirements, side effects, or whether it performs network calls. It also does not mention what happens if the CLI is unhealthy or if storage metrics are unavailable. This is a minimal disclosure for a tool with behavioral risk.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single sentence with no filler, front-loaded with the verb 'Diagnose'. Every part of the sentence adds meaning, listing the four diagnostic dimensions. It is appropriately sized for a parameterless tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is simple (0 params), but there is no output schema, so the description should compensate by explaining what the agent will receive. It does not mention return format, whether it returns a success/failure status, metrics, or logs. It also doesn't clarify prerequisite conditions (e.g., CLI must be installed). For a diagnostic tool, the output is critical, so the description is incomplete on that front, though the rest is adequate for a zero-parameter tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters and the input schema is empty, so there is no parameter semantics to add. The description lists the diagnostic areas, which is all that is needed. Baseline 4 for 0 parameters is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb ('Diagnose') and a clear resource ('Antigravity CLI health, authentication token status, YOLO permission mode, and storage metrics'). It is immediately distinct from the sibling tools, which focus on running commands, jobs, sessions, transcripts, and models. An agent can tell this is the health/diagnostic tool without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied: it is the diagnostic tool among siblings. However, it gives no explicit guidance on when to use it versus alternatives (e.g., 'use antigravity_sessions for session info, antigravity_doctor for health'). No exclusions or preferred scenarios are stated, leaving the agent to infer context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

antigravity_job_cancelA

Cancel or kill an ongoing background Antigravity job by ID.

ParametersJSON Schema
NameRequiredDescriptionDefault
job_idYesThe Job ID to cancel.

TDQS

A3.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It mentions the action is destructive ('cancel or kill') but does not elaborate on consequences such as irreversibility, what happens to the job's results, or any permission requirements. For a mutation tool, this is a significant gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence with zero waste. It states the action, the resource, and the parameter in a compact format, which is ideal for quick parsing by an agent.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is simple (one parameter, no output schema), and the description covers the core purpose and parameter. However, for a destructive operation, it lacks details on error conditions (e.g., job not found, already completed) and post-cancellation behavior. Given the absence of annotations, a bit more context would improve completeness.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema covers the single parameter with 100% description coverage ('The Job ID to cancel.'). The tool description adds nothing beyond that – it merely says 'by ID', which is already implied. Since the schema does the heavy lifting, a baseline of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action (cancel/kill) on a specific resource (ongoing background Antigravity job) and the identifier (by ID). It is distinct from sibling tools like antigravity_job_status (status) and antigravity_job_result (retrieving results), making its purpose unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies the intended use case: cancelling jobs that are still ongoing. It explicitly conditions the action on 'ongoing' jobs, which helps an agent avoid using it on completed jobs. However, it does not explicitly name alternatives or state when not to use it, but the context is clear enough.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

antigravity_job_resultA

Fetch the full execution report and final output of a background job once completed. Returns status, duration, full step trajectory, and final answer.

ParametersJSON Schema
NameRequiredDescriptionDefault
job_idYesThe Job ID to fetch results for.

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden. It does disclose that the tool returns status, duration, trajectory, and final answer, and 'Fetch' implies a read-only operation. However, it does not describe behavior when the job is not yet complete, authentication requirements, or side-effect guarantees beyond the implied read.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence with no filler. It immediately names the action and result, then briefly lists return contents. Every word earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with one parameter, no output schema, and no annotations, the description covers purpose, precondition, and key returned fields. It does not discuss error cases or whether to poll with a status tool first, but it is essentially complete for a simple result-fetch tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema already documents job_id at 100% coverage, so the description is not required to add much. It adds no extra detail about the ID's format, where to get it, or lifecycle, but for a single straightforward identifier parameter this is acceptable. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific action ('Fetch the full execution report and final output') and a specific resource ('background job'), making the purpose clear. The 'once completed' qualifier also hints at the distinction from status-checking siblings, though it never names them or contrasts the two explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives clear temporal guidance: use this tool after a background job has completed. It doesn't explicitly mention alternatives such as antigravity_job_status for checking progress, nor does it state 'do not call before completion,' but the precondition is inferable and not misleading.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

antigravity_job_statusA

Check the real-time progress and live step status of a background job. Fast local query (returns in < 20ms). Shows elapsed time, process alive status, and latest tool execution actions.

ParametersJSON Schema
NameRequiredDescriptionDefault
job_idYesThe Job ID returned by antigravity_start or graceful timeout fallback.

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It reveals that the tool is fast (<20ms), shows elapsed time, process alive status, and latest tool execution actions, which gives a clear picture of its output and non-destructive nature. It does not mention error conditions or side effects, but for a read-only status check, this is adequate.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and front-loaded. The first sentence states the purpose, and the second adds performance and output details without any redundant phrasing. Every word serves a function, and the structure is easily scannable.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple tool with one fully documented parameter and no output schema, the description covers the essential information: what it does, how fast it responds, and what it returns. It does not describe error handling or edge cases, but given the low complexity and the presence of sibling tools, the agent has enough context to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema already provides a complete description of job_id, including its origin from antigravity_start or graceful timeout fallback, achieving 100% schema coverage. The tool description adds no additional parameter meaning beyond what the schema states, so the baseline of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb 'check' and the resource 'real-time progress and live step status of a background job', which is distinct from siblings like antigravity_job_result (final result) and antigravity_start (initiation). It leaves no ambiguity about the tool's core function.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage for monitoring an in-progress job by highlighting 'real-time' and 'live step status', and the performance note 'returns in < 20ms' suggests it is suitable for frequent polling. However, it does not explicitly contrast with alternatives like antigravity_job_result or state when to stop using it, so it stops short of full guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

antigravity_modelsA

List available models supported by this Antigravity installation.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the burden of behavioral disclosure. It clearly indicates a read-only listing operation, which is useful. However, it does not disclose details such as whether the list is cached, whether it requires authentication, or what the response format looks like. The description is accurate but minimal.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, clear sentence that front-loads the action ('List') and the resource ('available models'). Every word earns its place, and there is no redundancy or unnecessary detail.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple list tool with no parameters and no output schema, the description is mostly complete. However, it does not mention what the output will look like (e.g., model IDs, names, metadata) or whether any filtering or pagination is available. Given the tool's simplicity, this is a minor gap, but the description could be slightly more informative.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, so the schema provides no parameter semantics. The description correctly indicates that the tool lists available models, which is sufficient for a no-parameter tool. Baseline 4 is appropriate because there are no parameters to document.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('List') and resource ('available models supported by this Antigravity installation'), which clearly identifies the tool's function. It doesn't explicitly differentiate from siblings, but the resource is distinct enough that an agent can infer its purpose among the sibling tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage context: it is for listing available models, which is a discovery/read operation. However, it does not explicitly state when to use this tool versus alternatives like run_antigravity or antigravity_sessions, nor does it mention any exclusions or prerequisites.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

antigravity_sessionsA

List or search historical Antigravity sessions from local database (fast, zero model cost). Returns session IDs, summaries, step counts, and last modified times. Use this to find a conversation ID to resume.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMaximum number of sessions to return. Defaults to 20.
queryNoOptional search keyword to filter session titles and summaries.

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations provided, so the description carries the burden. It discloses performance ('fast, zero model cost') and data source ('local database'), and the verb 'list or search' implies read-only behavior. However, it doesn't explicitly confirm that no modifications occur, nor does it mention potential errors or limitations. Adequate but not thorough.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with the core action and value proposition, followed by the return fields and usage hint. Zero waste, all sentences earn their place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a straightforward list/search tool, the description covers purpose, returns, and usage context. The limit parameter is in the schema, so pagination is handled. No output schema exists, but the listed return fields suffice. Missing only minor details like error handling or sort order, which are not critical for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and both parameters are well-described in the schema (limit with default, query for filtering titles/summaries). The tool description adds minimal extra meaning beyond restating the search capability. Baseline 3 is appropriate since the schema already handles parameter documentation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description states a specific verb ('List or search'), resource ('historical Antigravity sessions'), and key differentiators ('local database, fast, zero model cost'). It clearly distinguishes from siblings like run_antigravity or antigravity_start, which execute sessions, by focusing on historical lookup.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly states 'Use this to find a conversation ID to resume', which is clear usage guidance. It doesn't enumerate alternatives, but the context makes it obvious that this is for retrieval, not execution. Could be stronger by mentioning when not to use, but the sentence provides actionable direction.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

antigravity_startA

Submit an asynchronous task to Antigravity and return a job ID without waiting for model completion. Use this for long-running tasks (e.g. multi-minute analyses, large refactors). Supports both new tasks and continuing existing sessions (if conversation_id is provided).

ParametersJSON Schema
NameRequiredDescriptionDefault
yoloNoIf true (default), runs with --dangerously-skip-permissions for auto-approvals.
modelNoOptional model identifier override.
promptYesTask instructions or continuation prompt.
work_dirNoOptional working directory path.
read_onlyNoIf true, enables Antigravity --sandbox; workspace writes may still be allowed. Defaults to false.
conversation_idNoOptional Conversation ID to continue an existing session.
print_timeout_secondsNoMaximum agy execution time. Range 60-86400; defaults to 3600 seconds.

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Because there are no annotations, the description carries the behavioral disclosure burden. It does disclose the key behavior clearly: this tool is asynchronous, returns a job ID, and has a continuation path. However, it does not disclose tool-level side effects such as the dangerous default yolo/skip-permissions behavior or what happens to workspace writes, which are meaningful transparency gaps for a no-annotation tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences with no filler. The core behavior is front-loaded first, followed by the main use case and the conversation_id edge case. Every sentence earns its place and the structure makes it easy for an agent to quickly select or reject the tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For an asynchronous submission tool with no output schema, the description covers the essential contract: no waiting, a job ID is returned, long-running tasks are appropriate, and existing conversations can be continued via conversation_id. It would be more complete if it pointed to the companion status/result tools or called out the default yolo safety behavior, but what is needed to invoke the tool correctly is mostly present.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema description coverage is 100%, so the baseline is 3 even if the description adds little parameter-level detail. The description only restates the conversation_id continuation concept, which the schema already covers. It does not introduce new meaning about defaults, formats, ranges, or interactions among the seven parameters, so there is no reason to score above baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the exact action: 'Submit an asynchronous task to Antigravity and return a job ID without waiting for model completion.' This names the verb, resource, and response behavior, and it clearly differentiates itself from the synchronous-looking sibling run_antigravity. The continuation support via conversation_id is also explicitly called out, adding specificity without ambiguity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit context with 'Use this for long-running tasks (e.g. multi-minute analyses, large refactors),' which tells an agent when this tool is appropriate. It also clarifies the continuation case via conversation_id. It does not explicitly name siblings like antigravity_continue as alternatives or state when not to use them, so it stops short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

antigravity_transcriptA

Read the detailed execution transcript of any past or current Antigravity session by ID. Fast local read (zero token cost). Returns the step-by-step history of user prompts, tools executed, command outputs, and conclusions.

ParametersJSON Schema
NameRequiredDescriptionDefault
max_stepsNoMaximum number of recent steps to return. Defaults to 40.
only_finalNoIf true, returns only the final summary and conclusion. Defaults to false.
conversation_idYesThe Conversation ID to inspect.

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It discloses that the operation is a fast local read with zero token cost, and specifies the return content (user prompts, tools executed, command outputs, conclusions). It does not mention auth or edge cases, but for a read-only operation this is adequate.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with purpose, and the second adds relevant behavioral context. No wasted words or repetition of schema details.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with three well-documented parameters and no output schema, the description adequately explains what is returned and the nature of the operation. It lacks caveats about missing conversations or pagination beyond max_steps, but these are minor given the simplicity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description mentions 'by ID' for conversation_id and 'recent steps' for max_steps, but these add no substantive meaning beyond the schema's own descriptions. The performance note is not parameter-specific.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb (Read), resource (execution transcript), and scope (by ID), clearly distinguishing it from siblings like antigravity_sessions (list) and antigravity_job_status (status). The phrase 'detailed execution transcript' leaves no ambiguity about what is retrieved.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description clearly implies the use case (reading step-by-step history of a session) but does not explicitly name alternatives or exclusion conditions. It is obvious when to use it vs. listing sessions, but no direct comparison is given.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

run_antigravityA

Start a NEW task using Google Antigravity AI Agent (agy). Synchronous execution with a bounded timeout guard (auto-converts to a background job if still running). Supports YOLO autonomous execution or lightweight terminal sandboxing.

ParametersJSON Schema
NameRequiredDescriptionDefault
yoloNoIf true (default), runs with --dangerously-skip-permissions for full autonomous non-interactive execution.
modelNoOptional model identifier (e.g. 'gemini-3.8-flash-high', 'gemini-3.1-pro-high', 'claude-sonnet-4-6').
promptYesDetailed task instructions and requirements.
work_dirNoOptional working directory path. Defaults to the current workspace root.
read_onlyNoIf true, enables Antigravity --sandbox. This restricts terminal access but may still allow workspace writes. Defaults to false.
timeout_secondsNoSeconds to wait before returning the background job ID. Range 1-40; defaults to 40.
print_timeout_secondsNoMaximum agy execution time. Range 60-86400; defaults to 3600 seconds.

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the behavioral burden and does meaningful work: it discloses synchronous execution, a bounded timeout, automatic conversion to a background job, YOLO autonomous execution, and terminal sandboxing. It does not cover side effects like workspace writes or permission implications, but the key runtime behaviors are disclosed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two focused sentences with no filler. The first sentence states the core purpose, and the second covers the most important behavioral distinctions. Every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 7-parameter tool with no output schema and no annotations, the description covers the essential runtime contract: new task, synchronous wait, timeout guard, and background-job fallback. The schema fills in parameter details, and sibling tools cover follow-up actions, so only minor gaps remain (e.g., exact return shape).

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all 7 parameters individually. The description adds high-level context by mentioning YOLO and sandboxing, but it does not add meaning beyond what the parameter descriptions already provide. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action: 'Start a NEW task using Google Antigravity AI Agent (agy).' It also adds distinguishing behavioral details such as synchronous execution and timeout-guard conversion to a background job, which separates it from continuation and job-management siblings. However, it does not explicitly differentiate itself from the closely named sibling antigravity_start.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description conveys useful context: it is for new tasks, runs synchronously up to a timeout, and automatically converts to background execution. But it does not explicitly say when not to use this tool or name alternatives like antigravity_start for long-running/background workflows, nor does it direct the agent to antigravity_job_status/result after conversion.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 10 tool updatesv1.0.4
    • First observedantigravity_continue
    • First observedantigravity_doctor
    • First observedantigravity_job_cancel
    • First observedantigravity_job_result
    • First observedantigravity_job_status
    • First observedantigravity_models
    • First observedantigravity_sessions
    • First observedantigravity_start
    • First observedantigravity_transcript
    • First observedrun_antigravity

TDQS

A4.1/5.0

Scored across 10 tools

Disambiguation5/5

Each tool has a distinct purpose: starting new tasks, continuing sessions, submitting async jobs, checking status, fetching results, canceling, listing sessions, diagnosing, reading transcripts, and listing models. There is no ambiguity between them.

Naming Consistency4/5

The majority follow the 'antigravity_' prefix pattern with verb suffixes (e.g., antigravity_continue, antigravity_start). The one exception is 'run_antigravity', which inverts the order but remains clear. Overall, the naming is predictable and readable.

Tool Count5/5

With 10 tools, the server is well-scoped for managing Antigravity agent tasks. Each tool serves a necessary function, covering execution, monitoring, history, diagnostics, and model discovery without bloat.

Completeness5/5

The toolset covers the full lifecycle of agent tasks: starting (sync/async), continuing, checking status, retrieving results, canceling, listing sessions, reading transcripts, and diagnosing the environment. No critical gaps are apparent for the intended domain.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    A
    quality
    D
    maintenance
    Enables MCP clients to submit long-running jobs that are executed safely in isolated child processes with a durable SQLite queue, configurable timeouts, retries with backoff, and backpressure.
    5
    MIT
  • A
    license
    A
    quality
    A
    maintenance
    Durable MCP server for managing long-running jobs locally, over SSH, or on Slurm clusters. Jobs survive client disconnects and return exit codes, bounded logs, and JSON artifacts.
    11
    41 PyPI
    1
    MIT