Skip to main content
Glama
Zythenth

Antigravity MCP Bridge

by Zythenth

Antigravity MCP Bridge

Servidor MCP local para delegar tarefas de programação ao Google Antigravity por meio do CLI oficial agy. O Codex pode iniciar tarefas, acompanhar eventos, consultar resultados e cancelar processos. O projeto é independente e não é afiliado ao Google ou à OpenAI.

Codex -- MCP stdio --> bridge -- subprocesso --> agy oficial --> Antigravity
Codex <-- eventos e resultado estruturados <-- bridge <-- stdout/stderr

O bridge não acessa endpoints privados, cookies ou arquivos de autenticação. A comunicação com o serviço é feita pelo próprio agy.

Requisitos

O bridge foi testado com agy 1.2.16 e @modelcontextprotocol/sdk 1.30.1. Ele exige que o CLI anuncie --sandbox e stream-json; versões futuras podem exigir adaptação. Consulte abaixo a limitação observada na execução de testes no Windows.

Related MCP server: codex-antigravity-bridge

Versões disponíveis

A 0.5.1 está publicada no npm e inclui perfis, espera com progresso, transferência entre papéis, comparação de modelos e papéis personalizados. Os exemplos npx abaixo usam essa versão.

O código-fonte e o plugin deste repositório preparam a 0.5.2, com correção do ambiente de execução Windows, diagnósticos específicos do sandbox e recusa de recibos com bypass. A publicação dessa versão no npm está pendente.

A 0.5.1 amplia a margem de observação dos testes para suportar a preparação de cópias em runners mais lentos, preservando as verificações dos timeouts de execução.

Instalação

Servidor MCP via npm/npx

Para iniciar o servidor pelo pacote npm 0.5.1, configure seu cliente MCP com:

{
  "mcpServers": {
    "antigravity": {
      "command": "npx",
      "args": ["--yes", "antigravity-mcp-bridge@0.5.1"]
    }
  }
}

No Windows, clientes que exigem o nome completo do comando podem usar npx.cmd. O executável inicia o servidor em stdio; não é um comando interativo. O pacote contém o servidor compilado e os avisos das dependências incluídas. Node.js 24, Git e o agy autenticado continuam necessários. Para validar uma versão ainda não publicada, use npm run build:plugin, npm pack e o tarball local com npx --yes --package <caminho-do-tarball> antigravity-mcp-bridge.

Instalação pelo código-fonte

Os exemplos PowerShell usam npm.cmd, o lançador do npm para Windows, que funciona mesmo quando a política de execução bloqueia npm.ps1. Em outros shells, use npm.

git clone https://github.com/Zythenth/antigravity-mcp-bridge.git
cd antigravity-mcp-bridge
npm.cmd ci
npm.cmd run build:plugin
npm.cmd test

npm run build:plugin compila o servidor e gera plugin/server.mjs. Esse arquivo também acompanha o repositório para que o plugin possa ser instalado sem executar o build. npm start inicia o servidor MCP em stdio; a saída padrão fica reservada para JSON-RPC.

Instalar como plugin do Codex

O repositório contém um catálogo em .agents/plugins/marketplace.json e o plugin em plugin/. Instale o catálogo e o plugin:

codex plugin marketplace add Zythenth/antigravity-mcp-bridge
codex plugin add antigravity@antigravity-mcp-bridge

Para atualizar, execute codex plugin marketplace upgrade antigravity-mcp-bridge e codex plugin add antigravity@antigravity-mcp-bridge. Abra uma nova conversa depois da instalação ou atualização. Peça, por exemplo: “Use $antigravity para implementar esta mudança e revisar o resultado.” A skill orienta a escolha de modelos, o acompanhamento da tarefa e a revisão final. Sessões já abertas não recarregam as ferramentas do plugin. Se o aplicativo não encontrar agy, configure AGY_PATH no ambiente em que o Codex é iniciado.

O guia oficial de plugins explica o formato do catálogo e outras opções de instalação.

Registrar somente o servidor MCP

Para usar o bridge sem a skill do plugin, compile o projeto e registre o servidor:

$server = (Resolve-Path .\dist\src\index.js).Path
codex mcp add antigravity -- node $server

Há também um exemplo de configuração TOML. Use uma forma de registro por vez para evitar ferramentas duplicadas.

Perfis de ferramentas

Defina BRIDGE_TOOL_PROFILE no ambiente do servidor e reinicie a conexão MCP:

Valor

Ferramentas na 0.5.1

Catálogo e execução

full (padrão)

29

Todas as ferramentas; preserva a configuração existente

query

20

Consulta, modelos, sessões, handoff e comparação; tarefas somente em leitura

review

23

Consulta mais prévia, leitura de patches e verificação; tarefas somente em leitura

implementation

29

Fluxo completo, incluindo testes, integração confirmada e descarte

O perfil é informado em antigravity_health.toolProfile. Ferramentas fora do perfil não são registradas e chamadas diretas são recusadas. query e review também recusam mode: "write"; omitir o modo seleciona leitura. Perfis reduzem o catálogo e restringem essas tarefas; o sandbox e a confirmação de integração continuam necessários. Valores desconhecidos impedem a inicialização.

Espera com progresso

Prefira antigravity_wait com taskId, after e timeoutSeconds (1–60, padrão 30). Clientes que enviam _meta.progressToken recebem notificações notifications/progress com a sequência e o tipo de eventos reais, incluindo ferramentas, recibos de testes e conclusão. O número é um cursor de eventos, sem total ou porcentagem. Nenhum texto de arquivo, prompt ou saída do terminal é enviado nessas notificações.

A resposta inclui ready, timedOut, estado, uso de tokens e até 1.000 eventos. Continue com nextCursor; truncated informa eventos antigos perdidos. Timeout da espera e cancelamento da chamada MCP preservam a execução. Para parar a tarefa, use antigravity_cancel. Sem suporte a progresso, a resposta final continua disponível. A espera consulta também o estado persistido para observar tarefas de outro processo do bridge.

Transferência entre papéis

Use antigravity_context após uma tarefa concluída para inspecionar critérios, relatórios e treeSha256. Em seguida, chame antigravity_handoff com sourceTaskId, esse hash, role, prompt e, opcionalmente, model e decisions.

A nova tarefa recebe uma cópia independente dos arquivos atuais, incluindo alterações ainda não integradas. Mantém o baseline do projeto original, critérios, até oito relatórios de planejamento/revisão e a origem dos testes. Decisões são marcadas como relatos do cliente; relatórios e recibos mantêm sua origem e não autorizam integração. O pacote de contexto inteiro deve caber no limite do prompt; não há corte silencioso.

Planejamento → implementação → revisão pode mudar de papel e modelo sem reusar a conversa do papel anterior. O contexto é dado a conferir, não uma instrução de prioridade superior. Hashes são conferidos ao aceitar a tarefa e ao copiar, também quando ela aguardou na fila. A revisão verifica que sua cópia permaneceu intacta, incluindo arquivos ignorados. A cópia da implementação permanece disponível para testes, verificação e integração confirmada. Uma implementação iniciada a partir da revisão preserva o patch acumulado contra o original.

A cópia continua respeitando os ignores do projeto e limites de arquivos/bytes. Se a seleção mudar por novas regras de ignore, o handoff falha; inspecione o contexto novamente. Não inclua segredos em decisões. antigravity_resume conserva papel e cópia; antigravity_handoff cria outro papel em outra cópia e sessão.

Comparação entre modelos

Após antigravity_context, chame antigravity_compare com sourceTaskId, expectedContextSha256, prompt e models contendo 2 a 4 IDs distintos devolvidos por agy models. Cada modelo recebe uma cópia independente da mesma versão e executa uma revisão em leitura. A comparação consome a quota de cada tarefa; respeita MAX_CONCURRENT_TASKS, fila, retenção e limites de cópia. As cópias são preparadas antes de iniciar o grupo, evitando que revisores disputem a cópia de origem.

Guarde comparisonId e acompanhe cada taskId com antigravity_wait. antigravity_comparison reúne pareceres, erros, uso de tokens, modelos ausentes e os achados por arquivo/linha/citação. identical significa achados literalmente iguais; different, interpretações diferentes no mesmo trecho; not-reported-by-all, um trecho não relatado por todos. Ausência de achados não prova concordância nem correção.

complete exige todos os pareceres concluídos e suas cópias ainda correspondentes ao conteúdo comparado. contextStale sinaliza que a origem ou alguma cópia mudou, ficou indisponível ou não pôde ser conferida; confira contextMatches por parecer. Erros de início ficam em startErrors; falhas ou tarefas removidas não são ocultadas. O Codex deve conferir as fontes e sintetizar recomendações, divergências e limites de cada parecer. O agrupamento não faz votação semântica e não autoriza integração. As tarefas de comparação podem ser canceladas e descartadas individualmente.

Limites de interação e contagem prévia

antigravity_health.bridgeLimitations informa duas capacidades indisponíveis:

  • interactiveReplies.available: false: o protocolo headless verificado recusa control_request e control_response. Mensagens de texto em novos turnos não respondem a solicitações pendentes de permissão. O bridge encerra o stdin após seu prompt; para continuar a conversa concluída, use antigravity_resume. Examine pedidos negados nos eventos e erros, sem contornar o sandbox. A confirmação MCP da integração continua sendo uma operação separada do bridge.

  • preflightTokenCount.available: false, exactTokens: null: não há comando do agy verificado para contar tokens antes do envio. O uso informado pelo CLI é observado após execução. Tamanho em caracteres não é uma contagem exata de tokens nem um orçamento de quota. O bridge permanece no CLI oficial, sem API adicional, novas credenciais ou estimador apresentado como contagem exata.

Essas limitações não são resolvidas por manter o processo aberto ou inventar mensagens do protocolo. Um suporte futuro exige verificar a versão e o contrato oferecido pelo CLI antes de adicionar a operação.

Papéis personalizados

Defina BRIDGE_CUSTOM_ROLES como um array JSON no ambiente do servidor e reinicie a conexão. Exemplo PowerShell:

$env:BRIDGE_CUSTOM_ROLES = '[{"name":"security-review","baseRole":"reviewer","description":"Revisão de controles de acesso","instruction":"Examine os controles de acesso do escopo solicitado e cite evidências reais."}]'

antigravity_roles lista os nomes disponíveis, descrições, base e tamanho das instruções. Selecione o nome em role de antigravity_run ou antigravity_handoff; os schemas MCP anunciam os nomes configurados. Cada papel tem nome de até 32 caracteres em letras minúsculas, números e hífens, baseRole e instruções de até 8.000 caracteres. A descrição é opcional, até 500 caracteres. Há até 20 papéis; nomes duplicados, substituição dos três nomes nativos, bases desconhecidas ou instruções vazias impedem a inicialização.

As bases planner e reviewer conservam modo de leitura, contratos JSON e conferência de citações. implementer conserva o fluxo de escrita na cópia e as mesmas exigências de verificação e confirmação para integração. Instruções personalizadas não dão permissões extras e entram no limite total do prompt. A definição usada é salva com a tarefa; a retomada mantém instruções e base originais, mesmo após mudar a configuração. Para trocar de papel, crie uma nova tarefa ou faça handoff.

Ferramentas

Todas as ferramentas publicam outputSchema com campos e tipos de suas respostas estruturadas. O contrato contempla sucesso e error: { code, message }. O SDK confere os campos obrigatórios antes de entregar respostas de sucesso; clientes também podem validar o JSON recebido. Dados brutos do CLI continuam com tipo aberto porque seu formato pertence ao provedor. O contrato não transforma uma alegação do modelo em prova de execução.

Ferramenta

Função

antigravity_health

Verifica executável, versão, autenticação aparente e capacidades

antigravity_list_models

Lista os IDs devolvidos por agy models

antigravity_get_model / antigravity_set_model

Consulta ou persiste o modelo padrão; null seleciona Auto

antigravity_context / antigravity_handoff

Inspeciona e transfere plano, decisões, critérios e evidências para outro papel em cópia independente

antigravity_compare / antigravity_comparison

Solicita pareceres de 2 a 4 modelos e reúne achados, divergências, falhas e uso observado

antigravity_roles

Lista os papéis nativos e personalizados configurados

antigravity_run

Inicia uma tarefa e retorna o taskId

antigravity_list_project_files

Lista os arquivos elegíveis para a cópia

antigravity_preview

Mostra A/M/D, totais de linhas, estatísticas por arquivo, patch, hash e testes relatados

antigravity_record_test

Registra comando, saída e exit code relatados pelo cliente, vinculados ao hash

antigravity_test

Executa testes pelo terminal do agy com sandbox nativo e captura recibos reais

antigravity_verify

Confere critérios da tarefa e evidências de revisão contra arquivos reais

antigravity_read_patch

Lê o patch por arquivo ou em trechos vinculados ao hash completo

antigravity_read_result

Lê o JSON final do CLI em trechos com hash de conteúdo

antigravity_usage

Consolida tokens observados por tarefa, sessão e modelo

antigravity_integrate

Solicita confirmação via MCP e aplica o patch revisado ao original

antigravity_tasks

Recupera IDs e metadados de tarefas persistidas localmente

antigravity_status

Consulta estado, processo, sessão, uso e snapshots Git

antigravity_wait

Espera até 60 segundos com notificações MCP dos eventos observados; timeout não cancela a tarefa

antigravity_events

Lê eventos após um cursor after

antigravity_result

Consulta o resultado ou informa ready: false

antigravity_discard

Remove a cópia e o baseline de uma tarefa finalizada

antigravity_cleanup

Remove cópias finalizadas cujo prazo de retenção expirou

antigravity_cancel

Cancela tarefa na fila ou encerra o processo local

antigravity_sessions

Lista sessões conhecidas no estado local

antigravity_resume

Retoma uma conversa conhecida pelo sessionId

antigravity_run recebe prompt, workingDirectory absoluto na raiz de um repositório Git e, opcionalmente, model, timeoutSeconds, mode e includePaths (arquivos ou pastas relativos à raiz). Sem includePaths, copia todos os arquivos rastreados e não rastreados que não correspondam a .gitignore, .git/info/exclude ou às outras regras de ignore do Git. O filtro também exclui arquivos rastreados que passaram a ser ignorados. includePaths apenas reduz essa seleção; não permite incluir arquivos ignorados. Links simbólicos e caminhos fora da raiz são recusados. Consulte antigravity_list_models antes de selecionar um modelo.

Para revisão, diagnóstico ou segunda opinião, passe mode: "read-only". O bridge exige suporte a agy --mode plan, confere ao final que nenhum arquivo foi criado, modificado ou removido (inclusive arquivos novos ignorados pelo Git) e recusa integração dessas tarefas. Esse modo é uma restrição do CLI com verificação posterior, sem garantia de bloqueio físico de escrita. O padrão mode: "write" preserva o fluxo de implementação na cópia. Uma sessão retomada mantém seu modo original.

Fluxo típico:

  1. Consulte antigravity_health para conferir CLI, perfil e confirmação disponível. Use antigravity_list_models e antigravity_roles para selecionar IDs e papéis anunciados pelo servidor.

  2. Liste os arquivos elegíveis com antigravity_list_project_files, selecione includePaths quando necessário e defina acceptanceCriteria cobrindo os requisitos antes de iniciar a implementação com antigravity_run. Guarde o taskId.

  3. Acompanhe com antigravity_wait, preservando nextCursor e repetindo a espera quando ready for falso. timedOut encerra apenas a espera. Use antigravity_events para consultar os eventos detalhados.

  4. Após o término, confira o status e os erros em antigravity_result com includeResult: false. Leia o resultado com antigravity_read_result e o patch com antigravity_preview (includePatch: false) e antigravity_read_patch. Se houver falha, confira a causa antes de retomar; completed não comprova os requisitos.

  5. Execute os testes pertinentes com antigravity_test, usando o hash da prévia. A ferramenta devolve outro taskId: acompanhe e revise esse ID, confira recibos, exit codes e evidências desatualizadas e obtenha a prévia atual novamente. antigravity_record_test registra testes executados pelo cliente, como relatos; não substitui os recibos observados do executor.

  6. Confira os arquivos reais contra cada critério e envie evidências de revisão a antigravity_verify com o hash atual. Com verificação aprovada e atual, confira também se os testes observados continuam válidos e chame antigravity_integrate na tarefa de implementação mais recente dessa cópia para solicitar a confirmação humana. Alterações no patch, nas evidências ou nos arquivos afetados do original exigem nova conferência.

  7. Informe o uso observado com antigravity_usage e, quando o trabalho puder ser removido, descarte a cópia com antigravity_discard.

Esse fluxo de implementação requer full ou implementation. Para planejamento, revisão ou comparação, use os papéis de leitura e os fluxos específicos acima. Uma revisão por handoff recebe outra cópia; ela não altera qual é a tarefa mais recente da cópia de implementação.

A integração exige suporte do cliente a MCP form elicitation. antigravity_health informa integrationApproval.available. O formulário mostra origem, tarefa, hash, arquivos e contagens de linhas; só accept com confirm: true permite aplicar. Recusa, cancelamento, timeout ou falta de suporte preservam o original. O hash identifica o patch e a confirmação vem de uma resposta separada do cliente; nenhum argumento approved é aceito como autorização. Após a resposta, o bridge confere novamente hash e origem. A confirmação depende de um cliente confiável que apresente a decisão ao usuário.

Verificação dos resultados

Consumo de tokens

antigravity_status e antigravity_result incluem task.tokenUsage. antigravity_usage consolida as tarefas retidas e aceita filtros por taskId, sessionId e model. A resposta contém byTask, bySession, byModel e os contadores inputTokens, outputTokens, totalTokens, thinkingTokens e cacheReadTokens.

O resultado final do agy informa uso cumulativo da sessão. O bridge salva os contadores anteriores ao retomar e usa a diferença para a tarefa seguinte, inclusive quando o modelo muda. Assim, duas respostas cumulativas de 120 e 170 tokens representam 170 tokens na sessão e 50 na segunda tarefa. observedCumulative preserva o último total de sessão informado pelo CLI, separado do consumo das tarefas retidas.

Contadores ausentes, inválidos, reiniciados ou sem baseline conhecido ficam null; available, partial, source e warnings indicam a qualidade dos dados. Não há estimativa de tokens nem substituição silenciosa por zero. O consumo de tarefas falhas também é incluído quando o CLI devolve os contadores finais. Antes desse resultado, o total da tarefa pode estar indisponível. Tarefas antigas removidas pela retenção deixam de compor o consolidado local. Modelo null significa que não foi informado um ID; não se presume um modelo padrão.

Esses números são relatos do CLI, sem cálculo de cobrança ou acesso à quota global da conta. Cache e raciocínio são dimensões separadas e não devem ser somados novamente a totalTokens. A skill orienta o Codex a informar o consumo disponível ao concluir ou relatar falhas.

Papéis de trabalho

antigravity_run aceita os papéis nativos "implementer" (padrão), "planner" e "reviewer", além dos nomes personalizados anunciados por antigravity_roles. As bases de planejamento e revisão usam obrigatoriamente mode: "read-only" e agy --mode plan; selecionar escrita nesses papéis é recusado. A retomada mantém o papel e a definição originais.

Os papéis de consulta exigem suporte a agy --json-schema. O planejamento devolve summary, steps com arquivos e verificações observáveis, e unverified. A revisão devolve summary, reviewedFiles, findings e unverified; cada achado contém gravidade P0–P3, caminho, linha, citação literal, mensagem, impacto e sugestão.

O relatório validado aparece em task.report na resposta completa. Com resposta compacta, use antigravity_read_result para ler o structured_output original do CLI. Relatórios são identificados como agy-reported; na revisão, citationsChecked: true significa que arquivos e citações foram conferidos, sem atestar a interpretação ou provar ausência de defeitos. Formato inválido, arquivos inexistentes ou citações inventadas impedem a conclusão normal da tarefa. Uma resposta de planejamento não significa que suas etapas foram executadas.

Ler respostas grandes em partes

Use antigravity_preview com includePatch: false para obter arquivos, estatísticas, hash e patchLength sem enviar o diff inteiro. Depois chame antigravity_read_patch com expectedSha256 e, opcionalmente, um path devolvido na prévia. A seleção é feita pelo Git, incluindo arquivos binários e caminhos com espaços; não depende de interpretar cabeçalhos do patch. Qualquer alteração do patch completo invalida a leitura, mesmo quando você seleciona apenas um arquivo.

antigravity_result aceita includeResult: false para omitir o resultado bruto, o prompt, relatórios transferidos, instruções do papel e a lista completa de arquivos da cópia. Quando ready: true, consulte antigravity_read_result para ler o resultado serializado como JSON. Guarde contentSha256 e envie-o como expectedContentSha256 nas páginas seguintes para detectar mudanças.

Os leitores recebem offset (padrão 0) e limit (padrão 10.000, entre 2 e 50.000). Retornam text, nextOffset, hasMore, totalLength e contentSha256. Concatene text até hasMore: false; use sempre o nextOffset devolvido. Os offsets usam unidades UTF-16 e o leitor preserva caracteres representados por pares substitutos, como emojis. As opções antigas continuam devolvendo o conteúdo inteiro quando os campos de omissão não são usados.

Defina acceptanceCriteria antes de iniciar uma tarefa. Cada critério tem id, description e, opcionalmente, check com kind, path e text. As verificações disponíveis são file-exists, file-absent, file-contains e file-not-contains; as duas últimas exigem text. Critérios são preservados em retomadas e não podem ser substituídos depois da execução.

Após revisar o patch, chame antigravity_verify com o hash atual e reviews. Para cada critério, informe criterionId, verdict (passed, failed ou unverified), path, line, quote e explanation. O bridge lê os arquivos da cópia, executa as verificações e confere se a citação corresponde exatamente à linha indicada. Uma alegação do Gemini, um resultado SUCCESS ou um registro de teste do cliente não substitui essas evidências.

A integração exige todos os critérios aprovados e uma revisão atual. Critérios ausentes ou pendentes geram VERIFICATION_REQUIRED. Alterações no patch ou nos arquivos usados como evidência invalidam a verificação, inclusive alterações em arquivos ignorados que não aparecem no diff. O servidor verifica novamente após a confirmação humana. A prévia inclui verification e stale.

As verificações automáticas demonstram apenas as condições declaradas; a correção funcional mais ampla depende dos testes pertinentes e da revisão do Codex. O parecer permanece identificado como client-reported: conferir uma citação não demonstra que sua interpretação está correta. Esse fluxo aplica a distinção entre alegação e estado final e combina verificações determinísticas com revisão, conforme a orientação sobre avaliações de agentes.

Eventos, sessões e cancelamento

O bridge exige suporte aos formatos stream-json e a --sandbox, conferidos na descoberta do CLI. A versão verificada neste projeto é agy 1.2.16. Eventos estruturados chegam como NDJSON; diagnósticos de stderr permanecem separados. Linhas inválidas são expostas como stream.unparsed. O EventStore mantém um buffer limitado: truncated: true indica perda de eventos antigos. Registros de tarefas, sessões, eventos disponíveis, preferência de modelo e referências às cópias são persistidos por escrita atômica em ~/.antigravity-mcp-bridge (ou BRIDGE_STATE_DIRECTORY). O estado contém prompts e resultados: mantenha esse diretório privado, fora dos projetos versionados e de pastas compartilhadas.

Modelo padrão

Defina BRIDGE_DEFAULT_MODEL com um ID devolvido por agy models para o padrão inicial. antigravity_set_model grava a preferência no estado privado, compartilhada por servidores que usam o mesmo diretório. A prioridade é: model da tarefa, preferência salva, variável de ambiente e padrão do agy. A seleção é conferida contra a lista atual antes de executar; IDs indisponíveis causam MODEL_NOT_AVAILABLE.

Passe model: null para Auto: por tarefa, ignora os padrões do bridge; em antigravity_set_model, persiste o padrão do agy mesmo quando a variável está configurada. Omitir model preserva a preferência vigente. Para voltar ao padrão inicial do ambiente, pare o servidor e remova somente model-selection.json do diretório privado de estado.

antigravity_resume usa o conversation_id de uma tarefa concluída, inclusive após reinício, e reutiliza sua cópia isolada. Use antigravity_tasks para recuperar IDs e antigravity_sessions para consultar sessões persistidas. Execuções interrompidas não são repetidas automaticamente: recebem SERVER_RESTARTED quando o processo anterior já terminou. Se o PID registrado ainda estiver vivo, a cópia fica bloqueada com ORPHAN_PROCESS_RUNNING; o bridge não encerra processos recuperados apenas por PID. Tarefas de outro servidor ativo podem ser acompanhadas, mas devem ser canceladas no servidor que as iniciou. Locks locais impedem uso simultâneo da mesma cópia. O cancelamento encerra o subprocesso local; alterações parciais na cópia podem permanecer e devem ser revisadas.

Cópia, revisão e integração

Antes de chamar agy, o bridge cria uma cópia temporária dos arquivos elegíveis e mantém um baseline Git separado da cópia. O CLI recebe a cópia como diretório de trabalho e a opção --sandbox. O projeto original só muda por antigravity_integrate, depois da revisão do patch. O bridge não faz commit, merge nem push.

Quando o CLI anuncia --add-dir e --new-project, o bridge declara a cópia como workspace e cria um projeto CLI separado na primeira execução. A retomada preserva o projeto da conversa. Essas opções não equivalem a uma comprovação de todas as fronteiras do sandbox; permissões negadas continuam sendo respeitadas.

O resultado informa copyDirectory e includedFiles. Cópias temporárias permanecem para revisão por 7 dias após a última tarefa finalizada. O servidor limpa cópias expiradas na inicialização e a cada minuto; antigravity_cleanup permite antecipar a verificação. Use antigravity_discard para remover imediatamente uma cópia pelo MCP. Tarefas retomadas compartilham a mesma cópia; todas perdem acesso após descarte. Cópias em uso são preservadas. A expulsão do último registro pelo limite de retenção também remove sua cópia. isolateWorktree: true é aceito apenas por compatibilidade e usa o mesmo fluxo de cópia; false é recusado.

antigravity_preview preserva files, patch e sha256 e acrescenta summary (totais A/M/D, linhas e arquivos binários), fileSummaries (inserções/remoções por caminho) e tests. Binários usam null nas contagens de linhas. Registros de antigravity_record_test têm source: "client-reported": são relatos cujo comando o bridge não executou. Registros capturados por antigravity_test têm source: "agy-tool" e incluem o recibo observado do terminal. Cada registro leva hash, data, exit code e até 4.000 caracteres de saída. Resultados antigos recebem stale: true quando a evidência já não corresponde aos arquivos atuais. Falhas são preservadas e devem ser apresentadas na revisão.

A cópia aceita até 10.000 arquivos e 256 MiB por padrão. A seleção é medida antes da criação e os bytes efetivamente copiados são conferidos novamente para detectar crescimento da origem. includePaths pode reduzir a seleção. Mais de 100 arquivos alterados gera CHANGE_LIMIT_EXCEEDED ao finalizar, revisar ou integrar. O original permanece intacto; reduza a tarefa ou ajuste os limites explicitamente no ambiente do servidor.

Executar testes sem Docker

Chame antigravity_test com taskId, expectedSha256 e command: { executable, args }. A ferramenta inicia uma continuação na mesma cópia e devolve outro taskId. Acompanhe esse ID pelos eventos e pelo resultado; revise e integre a tarefa mais recente. retries vale 0 por padrão e aceita até 3 tentativas de correção adicionais, solicitadas explicitamente. timeoutSeconds vale 600 por padrão.

O executor usa agy --sandbox e a ferramenta nativa run_command. Não exige Docker nem executa o comando diretamente no host como alternativa. Um runner temporário, conferido por SHA-256 antes da execução, captura saída, exit code e fingerprints dos arquivos elegíveis antes/depois do teste. O bridge aceita o recibo apenas no evento da chamada exata do terminal; uma mensagem do Gemini dizendo que o teste passou não conta. Os registros têm source: "agy-tool" e sandbox: "agy-native-requested". A saída é limitada a 4.000 caracteres, com indicação de truncamento.

Quando não há recibo válido e a causa não foi identificada, o resultado é TEST_EXECUTION_UNVERIFIED. Os diagnósticos específicos abaixo preservam falhas conhecidas do runtime. Um comando não zero produz TEST_FAILED; mudanças nos arquivos durante/depois do comando produzem TEST_CHANGED_PATCH. Testes observados precisam continuar atuais para a integração; um relato manual não substitui um teste observado falho. Após TEST_FAILED, é possível retomar a conversa para corrigir ou executar os testes novamente. A orientação de correção mantém o comando original e proíbe enfraquecer os testes; tentativas observadas além do limite encerram a tarefa.

Código

Evidência e ação

AGY_SANDBOX_ACCESS_DENIED

O runtime relatou falha ao conceder acesso ao alvo. Confira o caminho e as ACLs; esse erro sozinho não demonstra que o setup inicial está ausente.

AGY_SANDBOX_SETUP_REQUIRED

A solicitação administrativa de setup não foi concluída, inclusive quando denied_actions informa escalate_admin. Confira a disponibilidade do broker do runtime atual.

AGY_SANDBOX_BYPASS_DENIED

O runtime relatou uma recusa de execução fora do sandbox. Preserve essa recusa.

AGY_SANDBOX_BYPASS_REQUESTED

A chamada declarou bypass ou um valor incompatível para essa opção. Seu recibo não verifica execução isolada.

No Windows, o bridge preenche PATHEXT ausente no subprocesso com .COM;.EXE;.BAT;.CMD, para que o PowerShell encontre os executáveis instalados. Valores definidos pelo cliente, inclusive um valor vazio, são preservados. O SDK pode omitir essa variável no ambiente herdado, fazendo node parecer ausente mesmo quando o arquivo está instalado. A correção vale para as sondagens e para todas as tarefas, sem configuração por projeto.

No Windows, use executáveis nativos como node.exe e python.exe; para npm, use npm.cmd. Arquivos .cmd/.bat aceitam argumentos comuns, mas metacaracteres de shell são recusados. A execução depende das permissões do sandbox nativo do CLI. Uma restrição de terminal não demonstra isolamento de todas as ferramentas do agente nem permite afirmar proteção completa do sistema de arquivos. Consulte a configuração oficial do sandbox e os eventos de ferramentas no modo headless.

Na configuração inicial do sandbox Windows, o CLI pode pedir uma elevação UAC. Quando o runtime indicar setup administrativo pendente, abra agy --sandbox em uma pasta descartável e, no prompt do próprio Antigravity, peça Execute node.exe --version. O cartão de configuração apresenta Yes, elevate; confira o aplicativo solicitante e confirme o diálogo do Windows. Uma tela de sandbox bypass é uma solicitação distinta. Depois, verifique o executor pelo recibo de antigravity_test. Um comando executado no PowerShell após sair do agy e antigravity_health.capabilities.sandbox: true não comprovam a preparação do sandbox. Falhas de ACL precisam de diagnóstico do alvo; repetir o setup para cada projeto não é uma correção demonstrada.

A versão 0.4.1 foi validada com execução real no Windows após essa configuração: o recibo capturou exit code 0, a origem permaneceu intacta e um marcador artificial fora da cópia teve leitura e escrita negadas. O snapshot Git do runner fica temporariamente dentro da cópia montada no sandbox, excluído do fingerprint e removido ao terminar. Esse teste comprova os cenários observados; não atesta todas as ferramentas do agente nem todas as fronteiras do sistema de arquivos. A suíte padrão continua usando um CLI simulado.

No agy 1.2.16, os testes reais passaram enquanto o broker administrativo da sessão de configuração estava ativo. Depois de encerrar essa sessão e seu filho, uma nova execução headless pediu escalate_admin e terminou sem recibo. Portanto, a confirmação anterior de setup não comprova disponibilidade após o encerramento do processo que mantinha o broker. O bridge ainda não gerencia esse ciclo de vida; esse limite precisa ser resolvido antes de afirmar funcionamento independente no Windows.

Para testar com a conta real em um projeto descartável, execute npm run build e node tests/native-integration.mjs. Esse teste usa a conta do agy e verifica a captura de uma execução real; a suíte padrão usa o CLI simulado e não consome quota.

Configuração e segurança

Variável

Padrão

Uso

BRIDGE_CUSTOM_ROLES

[]

Até 20 papéis em JSON com nome, base e instruções

BRIDGE_TOOL_PROFILE

full

Catálogo: full, query, review ou implementation

AGY_PATH

agy

Caminho do CLI oficial

MAX_CONCURRENT_TASKS

1

Processos simultâneos

MAX_QUEUED_TASKS

20

Tarefas aguardando

MAX_RETAINED_TASKS

100

Tarefas persistidas retidas

DEFAULT_TIMEOUT_SECONDS

1800

Prazo máximo por execução

EVENT_BUFFER_SIZE

2000

Eventos mantidos em memória

BRIDGE_STATE_DIRECTORY

~/.antigravity-mcp-bridge

Diretório privado de tarefas e sessões

BRIDGE_DEFAULT_MODEL

vazio

Modelo inicial, usado quando não há preferência salva

MAX_COPY_FILES

10000

Máximo de arquivos selecionados para a cópia

MAX_COPY_BYTES

268435456

Máximo de bytes copiados (256 MiB)

MAX_CHANGED_FILES

100

Máximo de arquivos alterados para revisão e integração

COPY_RETENTION_HOURS

168

Prazo de retenção das cópias finalizadas

MAX_PROMPT_CHARS

50000

Tamanho máximo do prompt enviado, incluindo critérios, contexto transferido e instruções do bridge e do papel

FORBIDDEN_DIRECTORIES

vazio

Diretórios bloqueados, separados por ; no Windows

Entradas e diretórios são validados. O processo é iniciado com spawn sem shell e exige --sandbox; não passa --dangerously-skip-permissions. Não inclua credenciais ou documentos privados nos prompts. O sandbox do CLI restringe comandos de terminal, mas não constitui garantia de isolamento completo do sistema de arquivos no Windows. Mantenha arquivos sensíveis fora da cópia por regras de ignore e selecione apenas os caminhos necessários com includePaths. Consulte a documentação do sandbox e do modo headless.

Erros comuns incluem AGY_NOT_FOUND, AGY_AUTH_REQUIRED, MODEL_NOT_AVAILABLE, INVALID_WORKING_DIRECTORY, QUEUE_FULL e AGY_PROCESS_FAILED. Se houver AGY_AUTH_REQUIRED, faça login no agy interativo. Se as ferramentas não aparecerem no Codex, confirme a instalação com codex plugin list --json e abra uma conversa nova.

Testes

npm.cmd run typecheck
npm.cmd run lint
npm.cmd test
npm.cmd run build:plugin
npm.cmd run test:package

npm test usa um mock do agy e não consome quota. A integração real é opcional: npm run test:integration cria um repositório descartável, executa uma tarefa pelo cliente MCP, confere que o original permanece intacto até a integração e remove o repositório. Esse teste simula a resposta de confirmação somente para seu projeto descartável; a confirmação humana da interface deve ser usada nos projetos reais. Execute-a apenas com agy autenticado e quando quiser usar a conta real.

Licença e políticas

O projeto usa a licença MIT. Consulte a política de segurança, a política de privacidade, o histórico de versões e os avisos das dependências distribuídas. npm run build:plugin atualiza o bundle e os avisos a partir das dependências efetivamente incluídas.

Available Tools

23 tools
antigravity_cancelCancel Antigravity taskA

Cancel a queued task or terminate its local agy process.

ParametersJSON Schema
NameRequiredDescriptionDefault
taskIdYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
taskNo
errorNo

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare destructiveHint=false and readOnlyHint=false, so the agent already knows this is a state-changing but non-destructive operation. The description adds the meaningful distinction between cancelling a queued item and killing a live process, but it does not say what happens on failure, whether cancellation is reversible, or what state the task ends in. Terminating a process sits in mild tension with destructiveHint=false, though this is a hint-scope ambiguity rather than a clear contradiction.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single well-formed sentence with the two operating modes front-loaded and zero filler. Nothing to trim and nothing buried.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained. The description covers both cancel paths but leaves open the practical questions an agent faces: behavior on an already-finished task, required permissions, and effects on subsequent antigravity_result/antigravity_status calls.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There is a single required taskId parameter and 0% schema description coverage, so the description would ideally clarify which id it expects. It names no parameter at all; the schema's uuid format and pattern do the documenting instead, making the description neutral rather than additive.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb pair (cancel/terminate) and the exact targets (a queued task, its local agy process), which is more informative than the title alone. It does not, however, distinguish itself from nearby siblings like antigravity_discard or antigravity_cleanup, so the agent must infer the boundary.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Phrasing implies the tool applies to a queued task or a running local process, which is a useful contextual hint. There is no explicit when-to-use/when-not guidance and no mention of the alternatives (discard, cleanup) for stopping or removing work.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

antigravity_cleanupClean expired isolated copiesA
Destructive

Remove finished copies older than COPY_RETENTION_HOURS. Active copies are preserved.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
errorNo
discardedTaskIdsNo

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, readOnlyHint=false, and closed-world, so the safety profile is covered. The description adds genuinely useful behavioral context beyond that by naming exactly what is destroyed (finished copies past retention) and what is protected (active copies), which is precisely the context that matters for a destructive tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short, front-loaded sentences with zero waste. The destructive scope leads and the preservation guarantee follows, which is the right ordering for an agent deciding whether to call it.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained, and there are no parameters to document. The description covers scope and safeguards adequately for a zero-argument cleanup; it omits only peripheral details such as whether COPY_RETENTION_HOURS is configurable or what permissions are needed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so there is nothing for the description to disambiguate and the baseline of 4 applies. No parameter semantics gaps exist.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Remove) and resource (finished copies older than COPY_RETENTION_HOURS) plus a scoping rule (active copies preserved). It does not explicitly contrast itself with similar siblings like antigravity_discard or antigravity_cancel, so an agent must infer that this is the bulk/expiry-based cleanup rather than the single-item removal.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is only implied: the retention rule signals when it applies, but there is no explicit when-to-use/when-not guidance and no named alternative for targeted removal. An agent can infer the intent but is not steered away from misuse.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

antigravity_discardDiscard an isolated copyA
Destructive

Delete the copy and baseline of a finished task, including resumed tasks sharing that copy. Active copies are refused. The source project is preserved.

ParametersJSON Schema
NameRequiredDescriptionDefault
taskIdYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
taskNo
errorNo

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true and readOnlyHint=false, so the bar is lower, yet the description adds real value: what is destroyed (copy AND baseline, plus resumed tasks sharing that copy), a refusal rule for active copies, and the reassuring fact that the source project is preserved. It stops short of noting irreversibility or permission requirements.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, each earning its place, with the core action front-loaded and the constraints following. No filler or repetition of the title.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained, and annotations cover the safety profile. The description supplies the destructive scope, refusal condition, and preservation guarantee, leaving only minor gaps such as permissions or irreversibility wording.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the single taskId parameter carries no schema text, so the description is the only source of context. It establishes that the operation targets a task, which implicitly scopes the ID, but adds no detail on the taskId itself (format, whether it must reference a finished task, etc.).

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (delete/discard) and resource (the copy and baseline of a finished task), including the resumed-tasks-sharing-copy case. It is clear what the tool does, but it never differentiates itself from closely related siblings like antigravity_cancel or antigravity_cleanup.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Implies usage ('of a finished task') and states a precondition ('active copies are refused'), which helps the agent understand when it will fail. However, it names no alternatives or when-not conditions relative to cancel/cleanup siblings, so the routing guidance is only implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

antigravity_eventsRead Antigravity eventsB
Read-only

Read live normalized agy events after a sequence cursor. Includes original agy event payloads.

ParametersJSON Schema
NameRequiredDescriptionDefault
afterNo
limitNo
taskIdYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
errorNo
eventsNo
truncatedNo
nextCursorNo
oldestAvailableNo

TDQS

B3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly/non-destructive/non-openWorld, so safety is covered. The description adds real behavioral context beyond that: events are 'live', 'normalized', returned 'after a sequence cursor', and include original payloads. It omits pagination/limit behavior and termination semantics, keeping it below a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, front-loaded with the action and scope. The second sentence is return-value detail that borders on redundant given the output schema exists, but it is not wasteful enough to drop below a 4.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Output schema exists so return values need little explanation, and annotations cover safety. However, with a required taskId and a limit parameter at 0% schema coverage, and no guidance on cursor advancement or loop termination, the definition is only minimally sufficient.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description carries the burden. It explains the `after` cursor implicitly, but says nothing about `limit` (max 1000) or the required `taskId`, leaving two of three parameters undocumented anywhere.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a clear verb ('Read') and resource ('events') plus a cursor-based retrieval mode, but 'agy' is undefined jargon and nothing distinguishes it from siblings like antigravity_status or antigravity_result, which an agent could easily confuse with an event stream.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'after a sequence cursor' hints at incremental polling, but there is no explicit when-to-use, no when-not-to-use, and no named alternative (e.g., status vs. events). The agent must infer the polling pattern.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

antigravity_get_modelGet selected modelA
Read-only

Return the default model selected for subsequent bridge tasks. Null means agy chooses its own default.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
errorNo
modelNo

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare the safe read-only, non-destructive, closed-world profile, so the bar is lower. The description adds genuinely useful behavior by explaining that a null result means agy falls back to its own default, which is not derivable from the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, zero filler, with the core purpose front-loaded and the null-handling caveat immediately after. Every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return-value structure need not be documented, and the description still covers the one ambiguous case (null). For a zero-parameter read-only getter, nothing needed to invoke it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so the baseline is 4. There is no parameter syntax to clarify, and the sentence about null concerns the return value rather than inputs.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Return the default model selected') and clarifies scope with 'for subsequent bridge tasks'. It is distinguishable from antigravity_set_model and antigravity_list_models by the word 'selected', though it does not explicitly name those siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is only implied: an agent can infer this reads the current selection before calling antigravity_set_model, but the description never states when to use it versus antigravity_list_models or antigravity_status. No prerequisites or exclusions are given.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

antigravity_healthCheck Antigravity CLIB
Read-only

Inspect installed agy version, authentication and supported capabilities.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
pathNo
errorNo
versionNo
installedNo
capabilitiesNo
authenticatedNo
integrationApprovalNo

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, destructiveHint=false, and openWorldHint=false, so the safety profile is covered. The description adds which facets are checked (version, auth, capabilities), but says nothing about cost, latency, or side effects like cache refresh — adequate but not rich.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single tight sentence that front-loads the verb and lists the inspected facets. No filler, though it is terse enough that it could have afforded one clause of usage context.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With zero parameters and an output schema present, the description need not explain return values, so the coverage of what is inspected is sufficient for an agent to call it correctly. Only usage routing to sibling diagnostics is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so per the rubric the baseline is 4. There is nothing to disambiguate, and the description adds no contradictory parameter claims.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Inspect') and enumerates the resource facets: installed agy version, authentication, and supported capabilities. This clearly separates it from sibling tools like antigravity_status or antigravity_usage, though it never names those siblings explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance is given. It is only implied by the word 'Inspect' that this is a diagnostic/health probe, and there is no indication of when to prefer it over antigravity_status or the other read-oriented siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

antigravity_integrateIntegrate reviewed changesA
Destructive

Request human confirmation through MCP form elicitation, then apply the reviewed patch. Requires the SHA-256 from antigravity_preview. Clients without form elicitation cannot integrate; no tool argument substitutes for user confirmation.

ParametersJSON Schema
NameRequiredDescriptionDefault
taskIdYes
expectedSha256Yes

Output Schema

ParametersJSON Schema
NameRequiredDescription
errorNo
filesNo
patchNo
sha256No
summaryNo
copyDirectoryNo
fileSummariesNo
sourceDirectoryNo

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true and openWorldHint=true, but the description adds non-obvious behavior: a mandatory human-confirmation flow via form elicitation that precedes the mutation, and a client-capability gate. This is exactly the kind of context annotations cannot convey, and it does not contradict them.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences: operation, prerequisite, and constraint, ordered by importance. No filler and no repetition of schema or annotation content.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Covers the mutation, its prerequisite, and the confirmation gate, and an output schema exists so return values need no explanation. The only gap is the undefined taskId, minor given its self-evident naming and the 0% schema description coverage.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must carry parameter meaning. It partially does for expectedSha256 by tying it to the antigravity_preview SHA-256, but taskId is never explained and no format/source detail is added beyond the regex and pattern in the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific operation ('apply the reviewed patch') preceded by a distinct step ('request human confirmation through MCP form elicitation'), which is more precise than the title alone. The mention of antigravity_preview as the source of the SHA distinguishes it from siblings like antigravity_preview, antigravity_verify, and antigravity_discard.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly states the prerequisite ('Requires the SHA-256 from antigravity_preview') and the disqualifying condition ('Clients without form elicitation cannot integrate'), and rules out workarounds ('no tool argument substitutes for user confirmation'). The agent knows both when to call it and when it will fail.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

antigravity_list_modelsList Antigravity modelsB
Read-only

List model IDs actually returned by agy models for this account.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
errorNo
modelsNo

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, destructiveHint=false and openWorldHint=false, so the safety profile is covered. The description adds one genuinely useful behavioral fact — the list is dynamic, reflecting what the backing CLI actually returns for this account rather than a static catalog. It does not mention caching, refresh, or empty-result behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler; the account qualifier is attached directly to the resource. It is arguably too terse to carry usage routing, but nothing in the text is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need no elaboration, and a 0-parameter read-only list tool has a very small surface to document. The account-specific, runtime-derived nature of the result is stated, which is the main thing an agent could get wrong. Only the sibling relationship is unaddressed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Zero parameters, so there is nothing for the description to disambiguate and the baseline of 4 applies. The account scoping statement is the only relevant context and it is present.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (List) and resource (model IDs) and scopes it to the current account, which distinguishes it from antigravity_set_model and antigravity_get_model. The phrase 'actually returned by agy models' is slightly cryptic (it presumes knowledge of an external CLI command) but the intent — enumerating this account's available models — is recoverable.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance and no reference to the obvious sibling antigravity_get_model, which presumably fetches a single model. The agent must infer that this is the discovery step before set_model, with nothing in the text to confirm or exclude that.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

antigravity_list_project_filesList files eligible for a project copyB
Read-only

List tracked and untracked files excluding Git ignore and local exclude matches.

ParametersJSON Schema
NameRequiredDescriptionDefault
workingDirectoryYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
errorNo
filesNo

TDQS

B3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, destructiveHint=false, and openWorldHint=false, so the safety profile is covered. The description adds useful scope context by specifying that gitignored and locally excluded files are filtered out. However, it says nothing about authentication, rate limits, or how results are ordered/paginated.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single tight sentence that front-loads the filtering semantics with no wasted words. It is perhaps too terse, leaving gaps, but the structure itself is efficient.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Low complexity (one read-only parameter) and an output schema exists, so return values need not be explained. Still, the meaning of 'eligible for a project copy' is never clarified and the sole parameter is undocumented, leaving the definition minimally adequate.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the single required parameter workingDirectory is never mentioned in the description, so the agent gets no explanation of what path to supply (repo root? any directory?). The description implies a repository context but does not compensate for the documentation gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (List) and resource (files) with meaningful scope qualifiers: tracked, untracked, excluding gitignore/local excludes. This is clearly distinguishable from sibling list tools like list_models, tasks, or sessions, though the description never explicitly contrasts itself with them.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance, no preconditions, and no alternatives named. The title's 'eligible for a project copy' hints at intent, but the description itself never tells the agent when this tool is appropriate versus other antigravity list/status tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

antigravity_previewPreview isolated changesC
Read-only

Return A/M/D files, per-file line statistics, totals, binary markers, patch, SHA-256 and client-reported test evidence with stale markers. Requires a finished task.

ParametersJSON Schema
NameRequiredDescriptionDefault
taskIdYes
includePatchNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
errorNo
filesNo
patchNo
testsNo
sha256No
summaryNo
patchLengthNo
verificationNo
copyDirectoryNo
fileSummariesNo
sourceDirectoryNo

TDQS

C2.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare this a safe, closed-world, non-destructive read, so the safety profile is covered. The description adds the finished-task precondition and indicates the payload includes 'client-reported test evidence with stale markers', which hints at freshness semantics. It does not describe error behaviour, size limits, or what happens when includePatch is omitted.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences, front-loaded with the payload contents and ending on the precondition. It is compact, but the enumerated return fields largely restate what the output schema already provides, so much of the sentence does not earn its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be enumerated, leaving the description thin on the things that matter: neither input parameter is explained and no sibling differentiation is offered. For a read tool with two undocumented parameters and many similar siblings, it is under-specified.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description carries the full burden for taskId and includePatch, yet it explains neither. The mention of 'patch' refers to an output field, not the includePatch input, so it may even muddy which parameter controls patch inclusion.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description outlines a diff/preview payload (A/M/D files, stats, patch, SHA-256, test evidence), which conveys what the tool returns and, indirectly, what it does. However it frames the tool as a return-value list rather than a verb+resource, and it names no sibling such as antigravity_read_patch or antigravity_read_result, so the agent cannot distinguish it from those on the description alone.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

"Requires a finished task" is a genuine state precondition that tells the agent when the call will be valid, which is real usage guidance. There is no routing against alternatives (read_patch, read_result, result) and no statement of when NOT to use it, so usage is only partially covered.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

antigravity_read_patchRead patch by file or chunkA
Read-only

Read at most 50000 UTF-16 units of the current patch, optionally selecting a changed path. Bind every read to the full preview hash. Follow nextOffset until hasMore is false. Use preview with includePatch false for metadata.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathNo
limitNo
offsetNo
taskIdYes
expectedSha256Yes

Output Schema

ParametersJSON Schema
NameRequiredDescription
pathNo
textNo
errorNo
lengthNo
offsetNo
sha256No
taskIdNo
hasMoreNo
nextOffsetNo
offsetUnitNo
totalLengthNo
contentSha256No

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover the safety profile (readOnlyHint=true, destructiveHint=false), so the description adds value with the 50000-unit read cap, the 'follow nextOffset until hasMore is false' loop, and the requirement to bind each read to the full preview hash. The hash-binding constraint is a meaningful behavioral detail beyond annotations, though failure behavior on mismatch is left implicit.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four short, front-loaded sentences with no filler; the core read action leads, followed by path selection, hash binding, pagination, and the alternative. Every sentence adds distinct operational value.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, return values need not be described, and the description covers pagination, size cap, hash binding, and sibling routing. The only notable omission is the meaning of taskId, but otherwise it is complete for a read tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description carries the burden and largely succeeds: it explains limit (max 50000 UTF-16 units), path (optional changed path), offset (via nextOffset), and expectedSha256 ('bind every read to the full preview hash'). Only taskId is left unexplained, keeping it short of a full 5.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Read ... the current patch') and clarifies scope (up to 50000 UTF-16 units, optional path selection). It also distinguishes itself from the sibling antigravity_preview by routing metadata lookups there, so an agent can tell the two apart without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear operating context: select a 'changed path' optionally, 'Follow nextOffset until hasMore is false' for pagination, and use 'preview with includePatch false for metadata' as the alternative. It names an alternative tool but stops short of explicit when-not-to-use conditions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

antigravity_read_resultRead result in chunksA
Read-only

Read the final CLI result serialized as JSON in bounded chunks. Returns ready false while active. Keep contentSha256 for subsequent requests and reconstruct the JSON by concatenating text.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
offsetNo
taskIdYes
expectedContentSha256No

Output Schema

ParametersJSON Schema
NameRequiredDescription
textNo
errorNo
readyNo
lengthNo
offsetNo
statusNo
taskIdNo
hasMoreNo
nextOffsetNo
offsetUnitNo
totalLengthNo
contentSha256No

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare a safe read profile (readOnlyHint=true, destructiveHint=false). Beyond that, the description discloses meaningful behavior: a 'ready' flag that is false while active, chunked/bounded retrieval, and a contentSha256 that must be carried across requests and used to reconstruct the JSON by concatenation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, zero filler, and the core chunked-read purpose is front-loaded. The two operative hints (ready-false polling, sha retention) follow in priority order.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return-shape explanation is not required. The description covers the essential chunked-read workflow, but leaves the sha-mismatch/error behavior and the meaning of a non-ready response underspecified for a stateful paging tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across 4 parameters, so the description carries the burden. It conveys chunking (limit/offset) and the role of contentSha256, but omits the 2-50000 bounds, offset semantics, and what happens when expectedContentSha256 mismatches — leaving key semantics unexplained.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Read the final CLI result') and clarifies its distinguishing trait: 'serialized as JSON in bounded chunks.' This differentiates it semantically from the full-result sibling antigravity_result, though it never names that sibling explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Returns ready false while active' implies a polling loop and 'keep contentSha256 for subsequent requests' implies repeated paged calls, so usage is inferable. However, there is no explicit statement of when to prefer this over antigravity_result or antigravity_status, and no exclusion guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

antigravity_record_testRecord review test evidenceA

Record a test already executed by the client in the isolated copy. Bind command, exit code and output to the reviewed patch. These are client-reported results; the bridge does not execute or independently verify the command. Do not include secrets in output.

ParametersJSON Schema
NameRequiredDescriptionDefault
outputNo
taskIdYes
commandYes
exitCodeYes
expectedSha256Yes

Output Schema

ParametersJSON Schema
NameRequiredDescription
errorNo
outputNo
sha256No
sourceNo
attemptNo
commandNo
sandboxNo
exitCodeNo
truncatedNo
recordedAtNo
testTaskIdNo
treeSha256No
executionErrorNo
beforeTreeSha256No

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false and destructiveHint=false, so the agent knows this writes state non-destructively. The description adds valuable context beyond annotations: it clarifies the tool records client-reported results and that the bridge does not execute or independently verify the command. It also warns against including secrets. It does not describe failure modes or what happens on duplicate recordings.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four sentences, each front-loaded with essential information: action, binding scope, verification stance, security constraint. No filler or redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values needn't be explained. The description covers the key behavioral caveats (client-reported, no verification, no secrets) for a write tool with annotations. It is nearly complete, missing only usage alternatives and some parameter meaning (taskId, expectedSha256).

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It does by naming what gets bound (command, exit code, output) and the security constraint on output. However, it does not explain taskId or expectedSha256 semantics, leaving two of five parameters undocumented in both schema and description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Specific verb (record) and resource (client-executed test evidence) with the binding scope stated (command, exit code, output tied to the reviewed patch). It is distinguishable from siblings like antigravity_test (which presumably runs tests) by clarifying that the bridge does not execute the command. Sibling differentiation is implicit rather than naming an alternative explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage: record a test the client already ran in the isolated copy. However, it does not state when to use this versus antigravity_test or antigravity_verify, nor any preconditions (e.g., patch must be reviewed first). Usage context is implied, not spelled out.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

antigravity_resultGet Antigravity resultA
Read-only

Return terminal result, usage and error once the task finishes. Use antigravity_preview for the patch.

ParametersJSON Schema
NameRequiredDescriptionDefault
taskIdYes
includeResultNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
taskNo
errorNo
readyNo
reportAvailableNo
resultAvailableNo
includedFileCountNo

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, destructiveHint=false and openWorldHint=false, so the safety profile is covered. The description adds the terminal-state precondition, but says nothing about what happens if the task is still running (error, empty, or block), which is important for a result-retrieval call.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences with no waste, and the core purpose is front-loaded ahead of the sibling redirect.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be spelled out. Still missing: the meaning of includeResult and clarification against antigravity_read_result/antigravity_status, which leaves real ambiguity for an agent choosing among result-oriented siblings.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% for two parameters. The description never mentions taskId's UUID format requirement or, more importantly, what includeResult toggles, so the description fails to compensate for the undocumented schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a clear verb+resource: return terminal result, usage and error once the task finishes. It distinguishes itself from antigravity_preview by scoping that sibling to the patch. However, it does not differentiate from the nearby sibling antigravity_read_result, leaving an overlap an agent must resolve on its own.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It states the condition that selects this tool ('once the task finishes') and explicitly routes patch retrieval to antigravity_preview. It does not exclude or address antigravity_read_result or antigravity_status, which appear to cover adjacent ground.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

antigravity_resumeResume Antigravity conversationC
Destructive

Continue a completed session in its existing isolated copy.

ParametersJSON Schema
NameRequiredDescriptionDefault
modeNo
roleNo
modelNo
promptYes
sessionIdYes
includePathsNo
timeoutSecondsNo
isolateWorktreeNo
workingDirectoryYes
acceptanceCriteriaNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
taskNo
errorNo

TDQS

C2.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, destructiveHint=true, and openWorldHint=true. The description adds only that the session continues in an existing isolated copy; it does not explain what 'continue' entails, what may be modified, or any permission or rate-limit context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single short sentence with no wasted words, but it is severely under-specified for a 10-parameter, destructive, open-world tool. Conciseness here reflects omission rather than efficient communication.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists and annotations cover the safety profile, so the description need not explain return values. However, for a complex tool with 10 parameters, no schema descriptions, and destructive behavior, the description leaves usage, parameter semantics, and operation details almost entirely unaddressed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across 10 parameters, and the description mentions none of them. It provides no meaning for prompt, sessionId, workingDirectory, mode, role, model, includePaths, timeoutSeconds, isolateWorktree, or acceptanceCriteria.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Continue') and resource ('a completed session'), with the scope 'in its existing isolated copy.' It is clear what the tool does, but it does not distinguish itself from siblings like antigravity_run or antigravity_sessions.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'Continue a completed session' implies usage when a session is already completed, but it gives no explicit when-to-use guidance, prerequisites, or alternatives to sibling tools such as antigravity_run.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

antigravity_runRun Antigravity taskB
Destructive

Copy non-ignored project files to a temporary directory and run agy --sandbox there. includePaths narrows copied files or folders. Returns a taskId; source is unchanged.

ParametersJSON Schema
NameRequiredDescriptionDefault
modeNo
roleNo
modelNo
promptYes
sessionIdNo
includePathsNo
timeoutSecondsNo
isolateWorktreeNo
workingDirectoryYes
acceptanceCriteriaNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
taskNo
errorNo

TDQS

B3.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, destructiveHint=true, and openWorldHint=true, so the safety profile is covered. The description still adds real value beyond them by disclosing the sandboxing mechanism (copies into a temp dir, runs agy --sandbox), the includePaths narrowing behavior, and the taskId return token. The 'source is unchanged' claim sits in mild tension with destructiveHint=true but is not a direct contradiction, since the destruction happens inside the sandbox/task state rather than the source tree.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences, front-loaded with the mechanism (copy to temp dir and run in sandbox), then the includePaths qualifier, then the return/immutability note. No filler and no repetition of the title.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained, but for a 10-parameter tool with 0% schema description coverage the description covers only one parameter and no workflow positioning. An agent can call it with the two required params but has no basis for choosing mode, role, isolateWorktree, timeoutSeconds, or how to build acceptanceCriteria.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and there are 10 parameters, so the schema provides names and types but no semantics. The description only explains includePaths ('narrows copied files or folders'), leaving mode, role, model, sessionId, timeoutSeconds, isolateWorktree, workingDirectory, and the whole acceptanceCriteria structure undocumented in prose. That is far too thin to compensate for a 0% coverage schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a concrete verb+resource ('Copy non-ignored project files to a temporary directory and run agy --sandbox there') and states the returned handle (taskId). It clearly separates this from read-only siblings like antigravity_preview/antigravity_status, though it never names a specific sibling to disambiguate from antigravity_resume or antigravity_integrate.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit when-to-use guidance: nothing says this starts a new task versus sitting alongside antigravity_resume, antigravity_cancel, or antigravity_verify. The phrase 'Returns a taskId' hints that this is the entry point, but the agent must infer that. No exclusions or prerequisites are given.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

antigravity_sessionsList known Antigravity sessionsA
Read-only

List conversation IDs recovered from local persisted tasks. agy does not advertise a session-list command.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
errorNo
scopeNo
sessionsNo

TDQS

A3.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, destructiveHint=false, and openWorldHint=false, so the safety profile is covered. The description adds useful provenance context (IDs are recovered from locally persisted tasks rather than fetched from a live session API), but says nothing about freshness, ordering, or whether empty results mean no sessions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, no filler, with the core behavior front-loaded ahead of the clarifying note about the missing CLI command.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained, and the description supplies the important provenance caveat. Coverage is solid for a zero-parameter read tool, with only minor gaps around result ordering and empty-state meaning.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so there is nothing for the description to disambiguate and the baseline of 4 applies. The description correctly implies no inputs are needed.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource (list conversation IDs) and the provenance of the data ('recovered from local persisted tasks'), which an agent can distinguish from most siblings. It does not, however, explicitly differentiate from the closest siblings like antigravity_tasks or antigravity_resume.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is only implied: an agent infers it should call this when it needs available session/conversation IDs. The note that 'agy does not advertise a session-list command' explains why the tool exists but does not state when to prefer this over antigravity_tasks, antigravity_resume, or antigravity_status.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

antigravity_set_modelSelect Antigravity modelA

Persist an exact model ID from agy models as the bridge default. Null persists Auto (agy default), overriding BRIDGE_DEFAULT_MODEL. Does not alter agy global settings.

ParametersJSON Schema
NameRequiredDescriptionDefault
modelYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
errorNo
modelNo

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=false, destructiveHint=false and openWorldHint=false, covering the safety profile. The description adds real behavioral context beyond that: the write target (bridge default), the env var it overrides (BRIDGE_DEFAULT_MODEL), the null-to-Auto semantic, and the explicit scope limitation that agy global settings are untouched. Authorization needs and error behavior remain unstated.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences, front-loaded with the action and followed by the null case and the scope boundary. No filler, every clause carries information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need no explanation, and the description covers the single parameter plus the side-effect boundary. What is missing is validation/error behavior for an invalid model ID and whether the value is verified against the agy model list at call time.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and there is one required parameter, so the description carries the burden. It does so reasonably: the value must be an 'exact model ID from agy models' and null persists Auto rather than being an error. The lack of an enum or concrete ID example keeps it from being a 5.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (persist) and resource (an exact model ID as the bridge default), which cleanly separates it from the sibling readers antigravity_get_model and antigravity_list_models. It stops short of naming an alternative explicitly, but the read/write split is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage by describing the null-persists-Auto behavior, so an agent can infer that passing a value selects a model and passing null restores the default. However there is no explicit when-to-use guidance, no statement of prerequisites, and no instruction that the ID must first be obtained from 'agy models'.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

antigravity_statusGet Antigravity task statusB
Read-only

Return task metadata, status, process ID and isolated copy path when available.

ParametersJSON Schema
NameRequiredDescriptionDefault
taskIdYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
taskNo
errorNo

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already establish the safety profile (readOnlyHint=true, destructiveHint=false, openWorldHint=false), so the bar is lower. The description adds useful context on what fields come back and the 'when available' caveat signaling optional/absent fields, which is more than the annotations provide.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence that leads with the verb and enumerates the returned fields. No filler or redundancy, though it is lean enough to leave gaps rather than being fully economical.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be re-explained, and annotations carry the safety profile. The remaining gap is sibling routing: with 22 peers including antigravity_tasks and antigravity_result, the definition does not say when this status call wins.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description should compensate for the undocumented taskId parameter. It implies the input is a task but adds no format or constraint detail; the self-explanatory name plus the uuid format in the schema keep this at a baseline 3.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Return') and resource ('task metadata, status, process ID and isolated copy path'), making the tool's output scope clear. However, it does not differentiate itself from close siblings like antigravity_tasks or antigravity_result, leaving the agent to infer which retrieval surface to use.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no when-to-use guidance and names no alternatives, despite a crowded sibling set (antigravity_tasks, antigravity_result, antigravity_events, antigravity_sessions) where routing matters. Nothing tells the agent whether this is for a single task ID versus a listing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

antigravity_tasksList persisted tasksA
Read-only

Recover task IDs and metadata from local bridge state, including tasks from earlier server processes.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
errorNo
tasksNo

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, openWorldHint=false, destructiveHint=false, so the safety profile is covered. The description adds genuine context beyond that: the data comes from local bridge state and survives server process restarts, telling the agent results persist across sessions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One tight sentence that front-loads the action (Recover task IDs and metadata) and then the distinguishing scope. No filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained. For a no-param, read-only list tool the description covers purpose and persistence scope adequately; only explicit sibling routing is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Zero-parameter tool, so there is nothing for the description to disambiguate; baseline 4 applies. No parameter confusion is possible.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource (recover task IDs and metadata) and scopes the source (local bridge state, persisting across server processes). Clear enough to distinguish from antigravity_status/antigravity_result, though no sibling is named explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'tasks from earlier server processes' implies the recovery scenario (after restart/process loss), which is useful implied usage. There is no explicit when-to-use or when-not-to-use versus siblings like antigravity_result or antigravity_status.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

antigravity_testRun tests in the native agy sandboxA
Destructive

Start an asynchronous continuation that executes an exact command through agy run_command. Captures actual output, exit status and file fingerprints. Optionally permits up to three repair retries. No Docker or host execution fallback. Use the returned taskId for events, result, review and integration.

ParametersJSON Schema
NameRequiredDescriptionDefault
taskIdYes
commandYes
retriesNo
expectedSha256Yes
timeoutSecondsNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
taskNo
errorNo

TDQS

A3.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already flag destructive/openWorld/non-readOnly, and the description adds substantive context beyond them: asynchronous execution, captured exit status and file fingerprints, optional up to three repair retries, and the absence of Docker/host fallback. That said, it doesn't explain what gets mutated or how expectedSha256 gates execution.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four tight sentences, front-loaded with the core action and followed by capability, constraint, and routing information. No filler, though the last sentence is somewhat list-like.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return-value explanation isn't required, and the description covers the async workflow and downstream taskId usage. However, for a destructive tool with 0% schema coverage on a mandatory hash parameter, the description leaves meaningful gaps about how taskId and expectedSha256 are obtained and what timeoutSeconds controls.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across 5 parameters including a nested command object, so the description must compensate and largely does not. It mentions retry behavior (matching the retries max of 3) but says nothing about taskId, expectedSha256, timeoutSeconds, or the command's executable/args shape.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource (start an asynchronous continuation that executes a command) plus what it captures, which separates it from siblings like antigravity_run and antigravity_verify. It never explicitly names a sibling to contrast against, so it stops short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives real routing guidance ('Use the returned taskId for events, result, review and integration') and a negative constraint ('No Docker or host execution fallback'), telling the agent when this path is and isn't appropriate. It does not say when to prefer antigravity_run or antigravity_record_test, so no explicit alternative comparison.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

antigravity_usageRead observed token usageA
Read-only

Consolidate final CLI usage by retained task, session and requested model. Resumed task counters are session deltas, not repeated cumulative totals. Missing counters stay null. This does not report account quota or billing; disclose partial or unavailable usage to the user.

ParametersJSON Schema
NameRequiredDescriptionDefault
modelNo
taskIdNo
sessionIdNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
errorNo
scopeNo
byTaskNo
byModelNo
countersNo
bySessionNo
taskCountNo
measuredTaskCountNo

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false, so the safety profile is covered. The description adds real behavioral context beyond that: resumed task counters are session deltas rather than cumulative totals, missing counters remain null, and partial/unavailable usage should be surfaced to the user.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four tight sentences, front-loaded with the core action before the caveats. Each sentence carries information, though the phrasing is somewhat telegraphic in the opening clause.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values need not be explained, and the description covers delta semantics, null handling, and scope boundaries. For a read-only, zero-required-param tool this is largely complete, missing only explicit alternative routing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the schema provides no parameter meaning. The description partially compensates by naming the three filter dimensions (task, session, requested model), which map to taskId, sessionId and model, but gives no format or exclusivity guidance.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (consolidate) and resource (CLI token usage) scoped by task, session and model. It distinguishes itself by clarifying it is not account quota or billing, though it does not name a sibling alternative directly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description clarifies the boundary ('does not report account quota or billing') which is useful when-not guidance, but it never states when to prefer this over siblings like antigravity_status or antigravity_result. Usage is implied rather than routed.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

antigravity_verifyVerify task acceptance criteriaB

Check actual artifacts and ground Codex review quotes in file lines. Requires criteria defined before the task. A CLI SUCCESS or unsupported claim is not verification; client review remains client-reported.

ParametersJSON Schema
NameRequiredDescriptionDefault
taskIdYes
reviewsNo
expectedSha256Yes

Output Schema

ParametersJSON Schema
NameRequiredDescription
errorNo
checksNo
reviewNo
sha256No
statusNo
checkedAtNo
fileHashesNo

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=false, openWorldHint=false, and destructiveHint=false, so the agent knows this is a non-destructive but state-changing call. The description adds the meaningful caveat that CLI success is not verification and that client review is client-reported, but it doesn't explain what state is mutated, whether prior verification is replaced, or what permissions are needed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three compact clauses with the core action front-loaded and the caveat trailing. No filler, though the sentence rhythm reads as a list of assertions rather than a structured statement of prerequisites and behavior.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need no explanation, and the annotations cover the safety profile. What remains missing for a required-hash, criteria-based verification call is the meaning of expectedSha256 and taskId, so the definition is adequate but not complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across three parameters, so the description carries the full burden. It gestures at the reviews structure ('ground Codex review quotes in file lines', 'criteria defined before the task') but says nothing about taskId or the required expectedSha256 pin, leaving the two required parameters undocumented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific action (check artifacts, ground review quotes in file lines) against a named resource (task acceptance criteria), so an agent understands this tool validates evidence rather than producing it. It stops short of differentiating explicitly from siblings like antigravity_test or antigravity_record_test, leaving some overlap to inference.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives one prerequisite ('Requires criteria defined before the task') and a negative boundary ('A CLI SUCCESS or unsupported claim is not verification'), which is useful context. However, it never says when to call this versus antigravity_test, antigravity_record_test, or antigravity_read_result, so the routing decision is only implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 23 tool updatesv0.4.0
    • First observedantigravity_cancel
    • First observedantigravity_cleanup
    • First observedantigravity_discard
    • First observedantigravity_events
    • First observedantigravity_get_model
    • First observedantigravity_health
    • First observedantigravity_integrate
    • First observedantigravity_list_models
    • First observedantigravity_list_project_files
    • First observedantigravity_preview
    • First observedantigravity_read_patch
    • First observedantigravity_read_result
    • First observedantigravity_record_test
    • First observedantigravity_result
    • First observedantigravity_resume
    • First observedantigravity_run
    • First observedantigravity_sessions
    • First observedantigravity_set_model
    • First observedantigravity_status
    • First observedantigravity_tasks
    • First observedantigravity_test
    • First observedantigravity_usage
    • First observedantigravity_verify

TDQS

B3.4/5.0

Scored across 23 tools

Disambiguation4/5

Tools generally target distinct lifecycle stages (run/resume/cancel, preview/read_patch/read_result, test/verify/integrate), but several reading/status tools have adjacent purposes. The descriptions help separate result vs read_result and preview vs read_patch, though an agent could still confuse them.

Naming Consistency4/5

All tools consistently use the antigravity_ prefix with snake_case, which makes the set predictable. Most names are verb_noun or verb_noun-phrase, but some are noun-only query tools (usage, status, tasks, events, result, sessions), a minor deviation from a strict pattern.

Tool Count3/5

23 tools is heavy for a single MCP server and sits in the borderline 16-25 range. The domain is complex, but several output-reading and status tools could be consolidated, making the surface feel somewhat over-scoped.

Completeness4/5

The surface covers model selection, task lifecycle, file listing, patch/result reading, testing, verification, integration, and cleanup. Minor gaps exist, such as no explicit authentication/login tool or direct source-file read outside task artifacts, but core workflows are covered.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers