Antigravity MCP Server
# Antigravity MCP Server
**Run Google's Antigravity/Gemini CLI (`agy`) as an MCP server — a multi-model conductor/executor for AI agents.**
[](LICENSE)
[](https://www.python.org/downloads/)
[](https://modelcontextprotocol.io/)
## What & why
A [Model Context Protocol](https://modelcontextprotocol.io/) server that exposes the
[Antigravity CLI](https://antigravity.google/docs/cli-using) (`agy`, Google's Gemini agent) as a set of
tools usable from Claude Code, Claude Desktop, Cursor, and Windsurf.
The animating idea is **cost discipline through model tiering**: a frontier model (Claude) acts as the
**conductor**, and cheaper Gemini (`agy`) is the **executor** it offloads bulky, token-heavy work to —
web research, codebase indexing, cross-model review, commit messages. The heavy output stays on disk and
in the cheap model's context; the conductor ingests only short digests, so its own context stays lean and
its bill stays low. Tiers (`flash` → `pro` → cross-family `sonnet`/`opus`/`gpt-oss`) let you dial
cost against quality per task.
## Architecture
```
┌─────────────────────────────┐
│ Conductor (Claude) │ plans, verifies, keeps context lean
│ via any MCP client │
└──────────────┬──────────────┘
│ MCP tool calls (stdio)
┌──────────────▼──────────────┐
│ Antigravity MCP server │ antigravity_mcp/ (this repo)
│ FastMCP · 13 tools/3 prompts
└──────────────┬──────────────┘
│ subprocess (prompt via stdin / tempfile path)
┌──────────────▼──────────────┐
│ agy CLI → Gemini │ Executor: web search, file reads, generation
└──────────────┬──────────────┘
│ detached workers write here
┌──────────────▼──────────────┐
│ $ANTIGRAVITY_JOBS (~/.antigravity-jobs)
│ per-job dirs: out · err · rc · subreport.md · manifest.json
└─────────────────────────────┘
```
- **Single FastMCP server.** Every tool is a `@mcp.tool()`-decorated function; every prompt is
`@mcp.prompt()`. All tools shell out to `agy` via `subprocess`, guarding on `shutil.which("agy")` first.
- **Model tiering.** `model_for_tier()` maps a semantic tier to a concrete Gemini/cross-family model — the
one place to update when Antigravity renames models.
- **Filesystem-backed background jobs.** Long jobs run as detached processes that redirect to `out`/`err`
and write their exit code to `rc`. State is reconstructed purely from those files, so **jobs survive the
MCP server restarting**.
- **Parallel fan-out pipelines.** A shared `_start_batch` primitive launches one worker per sub-question
(research) or per aspect (review). Workers write full reports to disk and print only a short digest;
batch-generic collectors gather them.
- **Large inputs never hit the command line.** Prompts pipe via stdin; big diffs/files are written to a
tempfile and only the *path* is passed, dodging OS argument-length limits.
A fuller, auto-generated breakdown lives in [`ARCHITECTURE.md`](ARCHITECTURE.md) — itself produced by this
server's own `index_code` tool (see [Dogfooding](#dogfooding) below).
## Quickstart
### Prerequisites
- [Antigravity CLI](https://antigravity.google/docs/cli-using) (`agy`) installed **and authenticated**
(run `agy` once interactively to sign in).
- Python ≥ 3.12
- [`uv`](https://docs.astral.sh/uv/)
### Add to Claude Code
```bash
mcp add antigravity-server uv run --directory /absolute/path/to/this/repo main.py
```
### Add to Claude Desktop
Add to your `claude_desktop_config.json`:
```json
{
"mcpServers": {
"antigravity": {
"command": "uv",
"args": ["run", "--directory", "/absolute/path/to/this/repo", "main.py"]
}
}
}
```
### Run the server directly
```bash
uv run main.py # stdio transport
```
### Try the pipeline (copy-paste)
[`examples/deep_research_example.py`](examples/deep_research_example.py) drives the fan-out research tools
end-to-end (fan out → poll → collect digests) against the cheap `flash-lo` tier:
```bash
uv run python examples/deep_research_example.py
```
> Uses web search, so it spends real Antigravity quota.
## Tools
18 tools across five groups, plus 4 orchestration prompts.
| Tool | Group | What it does |
| --- | --- | --- |
| `delegate_to_antigravity` | Delegation | Run a synchronous task on `agy`; targeted quota/auth/timeout error hints. |
| `start_background_job` | Delegation | Dispatch a long-running task to a detached process; returns a `job_id`. |
| `check_job_status` | Delegation | Reconstruct a background job's status/output from its on-disk files. |
| `propose_research_questions` | Deep research | Cheap pre-flight: draft clarifying questions + sub-questions to sharpen a brief before spending quota. |
| `research_fanout` | Deep research | Launch one parallel grounded-research worker per sub-question (web search → report on disk + digest). |
| `research_status` | Deep research \* | Aggregate progress of every worker in a batch. |
| `collect_digests` | Deep research \* | Gather workers' short digests plus on-disk report paths, keeping context lean. |
| `propose_design_questions` | Architect/build | The *grill*: draft a requirements interview + candidate requirements for a build/improvement. |
| `review_fanout` | Architect/build | Launch one parallel code-review worker per aspect (architecture, security, tests, …). |
| `draft_design_doc` | Architect/build | Have `agy` draft a full design doc (with work packages) from on-disk batches + a verified brief. |
| `list_agy_skills` | Science skills | Catalog the skills installed in `agy` (name, plugin, one-liner). Zero quota — reads the filesystem. |
| `check_science_credentials` | Science skills | Report which API keys each skill wants and which are present in `~/.env`. Never prints a value. |
| `propose_science_plan` | Science skills | Cheap pre-flight: draft sub-questions already mapped to the right databases, plus a clarifying interview. |
| `delegate_with_skills` | Science skills | Run ONE task pinned to named skills, in an isolated workspace. |
| `science_fanout` | Science skills \* | Launch one parallel worker per sub-question, each pinned to real database CLIs. |
| `cross_model_review` | Code & git | Independent diff review — use `tier='gpt-oss'`/`'sonnet'` for a different model family than the author. |
| `auto_git_commit` | Code & git | Stage, generate a conventional commit message, commit, and optionally push. |
| `index_code` | Code & git | Distill directories/files into an architectural index without pulling raw code into the conductor's context. |
\* `research_status` and `collect_digests` are **batch-generic** — they read any batch's
`sub_NN/subreport.md` + `manifest.json`, so the review *and* science pipelines reuse them unchanged.
**Prompts:** `antigravity_research_recipe` (deep-research recipe), `antigravity_build_recipe`
(Spec-Driven Requirements → Design → Tasks → Implement loop), `antigravity_science_recipe`
(primary-source science research), and `antigravity_workflow` (the core conductor/executor
cost-discipline rules).
### Model tiers
`flash` · `flash-med` · `flash-lo` · `pro` (default for research/review) · `pro-lo` · and cross-family
`sonnet` · `opus` · `gpt-oss`. Tier→model resolution is centralised in `model_for_tier()`; run
`agy models` for the live list.
## Three pipelines
- **Deep research** — `research_fanout` → `research_status` → `collect_digests`. Claude plans and
adversarially verifies; `agy` does the grounded web legwork in parallel.
- **Architect/build** — `propose_design_questions` → `review_fanout`/`research_fanout` →
`draft_design_doc`. A Spec-Driven loop for creating or improving codebases; `agy` drafts, Claude refines
and drives implementation.
- **Science (primary sources)** — `propose_science_plan` → `science_fanout` → `research_status` →
`collect_digests`. See below.
## Science skills (primary-source research)
`agy` can load **agent skills**, notably Google DeepMind's
[science-skills](https://github.com/google-deepmind/science-skills) bundle — ~39 skills wrapping
arXiv, OpenAlex, PubMed, UniProt, PDB, ChEMBL, ClinVar, gnomAD, AlphaFold, AlphaGenome,
ClinicalTrials.gov and more, each a rate-limited CLI over the real API.
This matters because **`research_fanout` is web search, and web search will hand you a
plausible-looking DOI that does not exist.** These skills call the actual databases and are
forbidden to fabricate identifiers. `science_fanout` is the primary-source counterpart — same cost
profile (agy works, Claude reads digests), but every ID it returns came from a real API call.
**Despite the name, this is not biomedical-only.** Two families live here:
- **All-discipline literature** — `literature-search-arxiv` covers every arXiv category
(statistics, maths, CS, physics, economics, quant-finance) and `literature-search-openalex`
indexes *all* scholarly work in every field, with real DOIs and citation counts. A literature
scan on Bayesian count models or transformer architectures is squarely in scope.
- **Domain databases** — PubMed, UniProt, PDB, ChEMBL, ClinVar, gnomAD, AlphaFold,
ClinicalTrials.gov, for primary biomedical and chemical records.
Rule of thumb: `/science-research` when you need citations you can **trust**; `/deep-research` when
you need **breadth** across the open web (news, blogs, docs, market scans). For a pure derivation,
neither — just do the maths.
**Setup:** install the bundle in Antigravity (Settings → Customizations → Build with Google Plugins →
Science). No API key is needed to start — most skills work keyless at lower rate limits, and
`check_science_credentials()` tells you exactly which keys would help and how to add them safely.
Two design notes worth knowing, since they're not obvious:
- `agy` has **no `--skill` flag**. It auto-loads every installed skill's description and triggers on
prompt *content*, so pinning a worker to a skill means naming it in the prompt and handing over its
absolute path (`skill_preamble()` in `agy.py`).
- The bundled `credentials` skill tells an agent to **halt and prompt the user** when a key is
missing. A detached worker has no user, so it would stall until timeout. `skill_preamble()`
resolves credentials up front and explicitly overrides that protocol; a genuinely *required* key
(only AlphaGenome has one) refuses the launch instead.
Each worker's `DATA_STATUS: OK | PARTIAL | BLOCKED` line is the health signal that catches the one
failure mode that looks like success — a worker quietly answering from memory instead of querying.
## Companion Claude Code skills
`/deep-research`, `/grill-me-research` and `/architect` orchestrate the first two pipelines. They live
in the user's `~/.claude/skills/` and are **not shipped here** — the server and its `@mcp.prompt()`
recipes are self-contained without them.
The science pipeline's skill **is** shipped, in [`skills/science-research/`](skills/science-research/).
Install it with:
```bash
cp -r skills/science-research ~/.claude/skills/
```
## Dogfooding
This repo was tidied up for release **using its own tools** — a nice end-to-end proof that they work:
- `index_code` distilled the package into [`ARCHITECTURE.md`](ARCHITECTURE.md).
- `cross_model_review` gave an independent second-model pass over the release diff.
## Testing
```bash
uv run python test_offline.py # fast, offline, no `agy`, no quota (pure-Python helper tests)
uv run python smoke_test_manual.py # manual end-to-end smoke test — invokes real `agy`, spends quota
uv run python smoke_test_science.py # manual science-pipeline smoke test — needs the science plugin, spends quota
```
## Caveats
- This is a thin **wrapper around the external `agy` CLI**: it doesn't call Gemini directly, so it inherits
`agy`'s auth and quota. Tools return human-readable error strings (never exceptions) with remediation tips.
- It's a **personal project**, not an official Google or Anthropic product.
- Several tools (`cross_model_review`, `auto_git_commit`, `index_code`, the fan-out workers) run `agy` with
`--dangerously-skip-permissions` because they need autonomous file/web access. Point them at code you
trust.
- Model tier names track Antigravity's current model lineup and may drift as Google renames models — update
`model_for_tier()` when they do.
## License
[MIT](LICENSE) © James Zoryk
TDQS
Scored across 13 tools
Most tools have distinct purposes, but there is some overlap between delegate_to_antigravity and start_background_job (both run agy tasks), and between check_job_status and research_status (both poll for status). However, descriptions differentiate them well and the contexts are separate.
Tools follow a consistent snake_case verb_noun pattern (e.g., check_job_status, collect_digests, propose_design_questions). Minor inconsistencies: auto_git_commit uses a prefix, and some verbs are compound (cross_model_review). Overall, the pattern is clear and predictable.
With 13 tools, the server covers a broad domain (git, research, design, code review, background jobs) without being excessive. Each tool earns its place, and the count is appropriate for the advertised capabilities.
The tool surface covers the full lifecycle of research and design pipelines: question proposal, fanout, status polling, collecting digests, and drafting documents. Additionally, git operations, code indexing, and cross-model review are included. No obvious gaps for the intended use case.