AGY Visual Witness MCP
# AGY Visual Witness MCP
A provider-specific visual reader for text-only agents. It uses Google's
Antigravity (`agy`) CLI and returns structured **visual evidence**, not an
acceptance verdict.
Three operations are intentionally separate:
- `WHOLE_READ`: scene-level structured description through pinned ModLens.
- `TARGETED_WITNESS`: one image plus one visible-evidence question.
- `COMPARE_PREFILTER`: two images plus a list of visible difference candidates.
Every response declares:
```text
control_surface = QA_ONLY
epistemic_label = INFERRED
authority = evidence_only_no_gate_no_acceptance_no_promotion
reproducibility_class = NOT_REPRODUCIBLE
```
It must not replace deterministic size/binding/AOV checks, pixel diffs,
histograms, IoU measurements, human identity/art-direction judgment, or final
promotion authority.
## Relationship to ModLens
`WHOLE_READ` invokes
[`@liustack/modlens`](https://github.com/liustack/modlens) as an external CLI.
The targeted and comparison operations are separate clean-room adapters that
invoke `agy` structured output directly. No ModLens source is vendored here.
## Install
1. Install and sign in to the official `agy` CLI.
2. Install Node.js/npm if you want `WHOLE_READ`.
3. Install this package:
```bash
pipx install .
# or
uv tool install .
```
## CLI
```bash
agy-visual doctor
agy-visual read --root /path/to/images --image /path/to/images/scene.png \
--operation WHOLE_READ
agy-visual read --root /path/to/images --image /path/to/images/hand.png \
--operation TARGETED_WITNESS \
--prompt "How many fingers are visibly countable? Mark occluded digits unknown."
agy-visual read --root /path/to/images \
--image /path/to/images/before.png --image /path/to/images/after.png \
--operation COMPARE_PREFILTER --prompt "List visible geometry differences only."
```
Use `--dry-run` to validate paths, hashes, operation contract, and runtime
identity without sending an image to a provider.
## MCP
Codex `config.toml` example:
```toml
[mcp_servers.agy-visual-witness]
command = "agy-visual-witness-mcp"
env = { AGY_VISUAL_ROOTS = "/absolute/allowed/image/root" }
```
Tools:
- `describe_image_gemini`
- `agy_visual_doctor`
## Boundaries
- Local files only; URLs are rejected.
- Images must stay under `AGY_VISUAL_ROOTS` or the roots passed by the caller.
- Supported formats: PNG, JPEG, WebP, GIF, BMP; maximum 25 MiB each.
- Provider execution receives staged copies in an otherwise empty temporary
directory. Image text is treated as untrusted data, never as instructions.
- Input evidence includes SHA-256 and byte size.
- `TARGETED_WITNESS` requires exactly one image and a non-empty question.
- `COMPARE_PREFILTER` requires exactly two images.
## Runtime identity for comparison
Multi-image behavior has changed across `agy` releases. Comparison therefore
fails closed unless the binary identity is known:
- Windows `agy` 1.1.9 is recognized by its tested SHA-256.
- Other builds can be pinned with `AGY_VISUAL_EXPECTED_SHA256` after independent
verification.
- `AGY_VISUAL_ALLOW_UNVERIFIED_COMPARE=1` is an explicit escape hatch. The
output still reports `OBSERVED_NOT_PINNED`; do not treat it as reproducible.
Version probes run with AGY auto-update disabled and the executable is hashed
again after the probe. If a configured/known hash does not match, or the binary
changes during probing, **every** provider operation is blocked before egress.
The unverified-compare escape hatch never overrides an explicit identity
mismatch.
`AGY_BIN` selects an explicit executable. `AGY_VISUAL_MODEL` overrides the
default `gemini-3.6-flash-low`. Provider model availability and quota can change.
## Security
Do not attach secret-bearing screenshots unless the selected external provider
is authorized to receive them. Never promote model output to a deterministic
gate merely because the call returned successfully.
See [SECURITY.md](SECURITY.md).
Observed pre-release checks and their evidence limits are recorded in
[docs/VALIDATION.md](docs/VALIDATION.md).
## License
MIT. This project is independent and is not endorsed by Google, ModLens, or
OpenAI. ModLens remains under its own MIT license and is used only through its
documented CLI interface.
TDQS
Scored across 2 tools
The two tools have clearly distinct purposes: one analyzes images via Gemini, the other checks runtime presence and identity evidence. There is no overlap or ambiguity between them.
The naming conventions are inconsistent: 'describe_image_gemini' follows a verb_noun_model pattern, while 'agy_visual_doctor' uses a product-prefixed noun phrase with no verb. This makes the set feel disjointed.
With only two tools, the server is minimal but still coherent with its stated 'Visual Witness' purpose. It feels slightly thin but not unreasonable for a specialized utility.
The tools cover image description and runtime diagnostics, but the 'Visual Witness' domain might benefit from additional operations like image comparison or evidence storage. There is no obvious dead end, but the surface is narrow.