Skip to main content
Glama
JamesZor

Antigravity MCP Server

by JamesZor
README.md
# Antigravity MCP Server

**Run Google's Antigravity/Gemini CLI (`agy`) as an MCP server — a multi-model conductor/executor for AI agents.**

[![License: MIT](https://img.shields.io/badge/License-MIT-yellow.svg)](LICENSE)
[![Python 3.12+](https://img.shields.io/badge/python-3.12+-blue.svg)](https://www.python.org/downloads/)
[![MCP](https://img.shields.io/badge/MCP-server-purple.svg)](https://modelcontextprotocol.io/)

## What & why

A [Model Context Protocol](https://modelcontextprotocol.io/) server that exposes the
[Antigravity CLI](https://antigravity.google/docs/cli-using) (`agy`, Google's Gemini agent) as a set of
tools usable from Claude Code, Claude Desktop, Cursor, and Windsurf.

The animating idea is **cost discipline through model tiering**: a frontier model (Claude) acts as the
**conductor**, and cheaper Gemini (`agy`) is the **executor** it offloads bulky, token-heavy work to —
web research, codebase indexing, cross-model review, commit messages. The heavy output stays on disk and
in the cheap model's context; the conductor ingests only short digests, so its own context stays lean and
its bill stays low. Tiers (`flash` → `pro` → cross-family `sonnet`/`opus`/`gpt-oss`) let you dial
cost against quality per task.

## Architecture

```
        ┌─────────────────────────────┐
        │   Conductor (Claude)        │   plans, verifies, keeps context lean
        │   via any MCP client        │
        └──────────────┬──────────────┘
                       │  MCP tool calls (stdio)
        ┌──────────────▼──────────────┐
        │   Antigravity MCP server    │   antigravity_mcp/  (this repo)
        │   FastMCP · 13 tools/3 prompts
        └──────────────┬──────────────┘
                       │  subprocess  (prompt via stdin / tempfile path)
        ┌──────────────▼──────────────┐
        │   agy CLI  →  Gemini        │   Executor: web search, file reads, generation
        └──────────────┬──────────────┘
                       │  detached workers write here
        ┌──────────────▼──────────────┐
        │  $ANTIGRAVITY_JOBS (~/.antigravity-jobs)
        │  per-job dirs: out · err · rc · subreport.md · manifest.json
        └─────────────────────────────┘
```

- **Single FastMCP server.** Every tool is a `@mcp.tool()`-decorated function; every prompt is
  `@mcp.prompt()`. All tools shell out to `agy` via `subprocess`, guarding on `shutil.which("agy")` first.
- **Model tiering.** `model_for_tier()` maps a semantic tier to a concrete Gemini/cross-family model — the
  one place to update when Antigravity renames models.
- **Filesystem-backed background jobs.** Long jobs run as detached processes that redirect to `out`/`err`
  and write their exit code to `rc`. State is reconstructed purely from those files, so **jobs survive the
  MCP server restarting**.
- **Parallel fan-out pipelines.** A shared `_start_batch` primitive launches one worker per sub-question
  (research) or per aspect (review). Workers write full reports to disk and print only a short digest;
  batch-generic collectors gather them.
- **Large inputs never hit the command line.** Prompts pipe via stdin; big diffs/files are written to a
  tempfile and only the *path* is passed, dodging OS argument-length limits.

A fuller, auto-generated breakdown lives in [`ARCHITECTURE.md`](ARCHITECTURE.md) — itself produced by this
server's own `index_code` tool (see [Dogfooding](#dogfooding) below).

## Quickstart

### Prerequisites
- [Antigravity CLI](https://antigravity.google/docs/cli-using) (`agy`) installed **and authenticated**
  (run `agy` once interactively to sign in).
- Python ≥ 3.12
- [`uv`](https://docs.astral.sh/uv/)

### Add to Claude Code
```bash
mcp add antigravity-server uv run --directory /absolute/path/to/this/repo main.py
```

### Add to Claude Desktop
Add to your `claude_desktop_config.json`:
```json
{
  "mcpServers": {
    "antigravity": {
      "command": "uv",
      "args": ["run", "--directory", "/absolute/path/to/this/repo", "main.py"]
    }
  }
}
```

### Run the server directly
```bash
uv run main.py          # stdio transport
```

### Try the pipeline (copy-paste)
[`examples/deep_research_example.py`](examples/deep_research_example.py) drives the fan-out research tools
end-to-end (fan out → poll → collect digests) against the cheap `flash-lo` tier:
```bash
uv run python examples/deep_research_example.py
```
> Uses web search, so it spends real Antigravity quota.

## Tools

18 tools across five groups, plus 4 orchestration prompts.

| Tool | Group | What it does |
| --- | --- | --- |
| `delegate_to_antigravity` | Delegation | Run a synchronous task on `agy`; targeted quota/auth/timeout error hints. |
| `start_background_job` | Delegation | Dispatch a long-running task to a detached process; returns a `job_id`. |
| `check_job_status` | Delegation | Reconstruct a background job's status/output from its on-disk files. |
| `propose_research_questions` | Deep research | Cheap pre-flight: draft clarifying questions + sub-questions to sharpen a brief before spending quota. |
| `research_fanout` | Deep research | Launch one parallel grounded-research worker per sub-question (web search → report on disk + digest). |
| `research_status` | Deep research \* | Aggregate progress of every worker in a batch. |
| `collect_digests` | Deep research \* | Gather workers' short digests plus on-disk report paths, keeping context lean. |
| `propose_design_questions` | Architect/build | The *grill*: draft a requirements interview + candidate requirements for a build/improvement. |
| `review_fanout` | Architect/build | Launch one parallel code-review worker per aspect (architecture, security, tests, …). |
| `draft_design_doc` | Architect/build | Have `agy` draft a full design doc (with work packages) from on-disk batches + a verified brief. |
| `list_agy_skills` | Science skills | Catalog the skills installed in `agy` (name, plugin, one-liner). Zero quota — reads the filesystem. |
| `check_science_credentials` | Science skills | Report which API keys each skill wants and which are present in `~/.env`. Never prints a value. |
| `propose_science_plan` | Science skills | Cheap pre-flight: draft sub-questions already mapped to the right databases, plus a clarifying interview. |
| `delegate_with_skills` | Science skills | Run ONE task pinned to named skills, in an isolated workspace. |
| `science_fanout` | Science skills \* | Launch one parallel worker per sub-question, each pinned to real database CLIs. |
| `cross_model_review` | Code & git | Independent diff review — use `tier='gpt-oss'`/`'sonnet'` for a different model family than the author. |
| `auto_git_commit` | Code & git | Stage, generate a conventional commit message, commit, and optionally push. |
| `index_code` | Code & git | Distill directories/files into an architectural index without pulling raw code into the conductor's context. |

\* `research_status` and `collect_digests` are **batch-generic** — they read any batch's
`sub_NN/subreport.md` + `manifest.json`, so the review *and* science pipelines reuse them unchanged.

**Prompts:** `antigravity_research_recipe` (deep-research recipe), `antigravity_build_recipe`
(Spec-Driven Requirements → Design → Tasks → Implement loop), `antigravity_science_recipe`
(primary-source science research), and `antigravity_workflow` (the core conductor/executor
cost-discipline rules).

### Model tiers
`flash` · `flash-med` · `flash-lo` · `pro` (default for research/review) · `pro-lo` · and cross-family
`sonnet` · `opus` · `gpt-oss`. Tier→model resolution is centralised in `model_for_tier()`; run
`agy models` for the live list.

## Three pipelines

- **Deep research** — `research_fanout` → `research_status` → `collect_digests`. Claude plans and
  adversarially verifies; `agy` does the grounded web legwork in parallel.
- **Architect/build** — `propose_design_questions` → `review_fanout`/`research_fanout` →
  `draft_design_doc`. A Spec-Driven loop for creating or improving codebases; `agy` drafts, Claude refines
  and drives implementation.
- **Science (primary sources)** — `propose_science_plan` → `science_fanout` → `research_status` →
  `collect_digests`. See below.

## Science skills (primary-source research)

`agy` can load **agent skills**, notably Google DeepMind's
[science-skills](https://github.com/google-deepmind/science-skills) bundle — ~39 skills wrapping
arXiv, OpenAlex, PubMed, UniProt, PDB, ChEMBL, ClinVar, gnomAD, AlphaFold, AlphaGenome,
ClinicalTrials.gov and more, each a rate-limited CLI over the real API.

This matters because **`research_fanout` is web search, and web search will hand you a
plausible-looking DOI that does not exist.** These skills call the actual databases and are
forbidden to fabricate identifiers. `science_fanout` is the primary-source counterpart — same cost
profile (agy works, Claude reads digests), but every ID it returns came from a real API call.

**Despite the name, this is not biomedical-only.** Two families live here:

- **All-discipline literature** — `literature-search-arxiv` covers every arXiv category
  (statistics, maths, CS, physics, economics, quant-finance) and `literature-search-openalex`
  indexes *all* scholarly work in every field, with real DOIs and citation counts. A literature
  scan on Bayesian count models or transformer architectures is squarely in scope.
- **Domain databases** — PubMed, UniProt, PDB, ChEMBL, ClinVar, gnomAD, AlphaFold,
  ClinicalTrials.gov, for primary biomedical and chemical records.

Rule of thumb: `/science-research` when you need citations you can **trust**; `/deep-research` when
you need **breadth** across the open web (news, blogs, docs, market scans). For a pure derivation,
neither — just do the maths.

**Setup:** install the bundle in Antigravity (Settings → Customizations → Build with Google Plugins →
Science). No API key is needed to start — most skills work keyless at lower rate limits, and
`check_science_credentials()` tells you exactly which keys would help and how to add them safely.

Two design notes worth knowing, since they're not obvious:

- `agy` has **no `--skill` flag**. It auto-loads every installed skill's description and triggers on
  prompt *content*, so pinning a worker to a skill means naming it in the prompt and handing over its
  absolute path (`skill_preamble()` in `agy.py`).
- The bundled `credentials` skill tells an agent to **halt and prompt the user** when a key is
  missing. A detached worker has no user, so it would stall until timeout. `skill_preamble()`
  resolves credentials up front and explicitly overrides that protocol; a genuinely *required* key
  (only AlphaGenome has one) refuses the launch instead.

Each worker's `DATA_STATUS: OK | PARTIAL | BLOCKED` line is the health signal that catches the one
failure mode that looks like success — a worker quietly answering from memory instead of querying.

## Companion Claude Code skills

`/deep-research`, `/grill-me-research` and `/architect` orchestrate the first two pipelines. They live
in the user's `~/.claude/skills/` and are **not shipped here** — the server and its `@mcp.prompt()`
recipes are self-contained without them.

The science pipeline's skill **is** shipped, in [`skills/science-research/`](skills/science-research/).
Install it with:

```bash
cp -r skills/science-research ~/.claude/skills/
```

## Dogfooding

This repo was tidied up for release **using its own tools** — a nice end-to-end proof that they work:
- `index_code` distilled the package into [`ARCHITECTURE.md`](ARCHITECTURE.md).
- `cross_model_review` gave an independent second-model pass over the release diff.

## Testing

```bash
uv run python test_offline.py        # fast, offline, no `agy`, no quota (pure-Python helper tests)
uv run python smoke_test_manual.py   # manual end-to-end smoke test — invokes real `agy`, spends quota
uv run python smoke_test_science.py  # manual science-pipeline smoke test — needs the science plugin, spends quota
```

## Caveats

- This is a thin **wrapper around the external `agy` CLI**: it doesn't call Gemini directly, so it inherits
  `agy`'s auth and quota. Tools return human-readable error strings (never exceptions) with remediation tips.
- It's a **personal project**, not an official Google or Anthropic product.
- Several tools (`cross_model_review`, `auto_git_commit`, `index_code`, the fan-out workers) run `agy` with
  `--dangerously-skip-permissions` because they need autonomous file/web access. Point them at code you
  trust.
- Model tier names track Antigravity's current model lineup and may drift as Google renames models — update
  `model_for_tier()` when they do.

## License

[MIT](LICENSE) © James Zoryk

TDQS

A4/5.0

Scored across 13 tools

Disambiguation4/5

Most tools have distinct purposes, but there is some overlap between delegate_to_antigravity and start_background_job (both run agy tasks), and between check_job_status and research_status (both poll for status). However, descriptions differentiate them well and the contexts are separate.

Naming Consistency4/5

Tools follow a consistent snake_case verb_noun pattern (e.g., check_job_status, collect_digests, propose_design_questions). Minor inconsistencies: auto_git_commit uses a prefix, and some verbs are compound (cross_model_review). Overall, the pattern is clear and predictable.

Tool Count5/5

With 13 tools, the server covers a broad domain (git, research, design, code review, background jobs) without being excessive. Each tool earns its place, and the count is appropriate for the advertised capabilities.

Completeness5/5

The tool surface covers the full lifecycle of research and design pipelines: question proposal, fanout, status polling, collecting digests, and drafting documents. Additionally, git operations, code indexing, and cross-model review are included. No obvious gaps for the intended use case.

Maintenance

ActivityStale
ResponsivenessNo issues