VisionCLI
<p align="center"><img src="docs/logo.png" width="128" alt="VisionCLI logo"></p>
<h1 align="center">VisionCLI</h1>
<p align="center">
Let Claude see your screen. Ask about the file, error, UI, database table or web page you're looking at,<br>
by typing in <b>Claude Code or Gemini CLI</b> or <b>by voice</b>, with the answer floating above your mouse pointer.
</p>
<p align="center">
<a href="https://github.com/thethirdsourcers/visioncli/releases/latest"><b>⬇︎ Download VisionCLI for macOS</b></a>
</p>
---
- **"What is this file for?"** Claude looks at the window under your mouse, finds the real file in your project and answers with `file:line` references.
- **Hold ⌥ Space and ask out loud.** A loader rides on your cursor, and the answer appears in a bubble above it, read aloud.
- **Works with your Claude Code session.** It types into the `claude` already running for that project, or opens one, so you can type and talk in the same conversation.
- **Points at things.** Claude draws glowing boxes on your screen around what it's talking about.
- **Fast browser control.** "Open the Stripe tab", "click sign in", "pause the video" run in under a second (via TypeSafe Jev), with a visible AI cursor.
- **Knows your data.** Read-only database questions ("how are payments and orders related?") using your project's `.env`.
## Install
### Requirements
- macOS 13 or later (Apple Silicon or Intel)
- [Claude Code](https://claude.com/claude-code) (`claude` CLI), signed in, **and/or** [Gemini CLI](https://github.com/google-gemini/gemini-cli) (`gemini`) with a Gemini API key (see [Gemini CLI](#gemini-cli))
- [Node.js](https://nodejs.org) 18 or later
### 1. Download the app
1. Download **`VisionCLI-<version>-macos.dmg`** from the [latest release](https://github.com/thethirdsourcers/visioncli/releases/latest).
2. Open it and drag **VisionCLI** into **Applications**.
3. The app isn't notarized by Apple yet, so the first time:
- **Right-click VisionCLI → Open → Open**, or
- run `xattr -dr com.apple.quarantine /Applications/VisionCLI.app` in Terminal.
### 2. First launch
The 👁 icon appears in the menu bar and a welcome screen offers to set up the CLIs it finds. That registers VisionCLI's tools for every project: `claude mcp add visioncli --scope user …` for Claude Code, and `gemini mcp add -s user --trust visioncli …` for Gemini CLI. You can also do it later from 👁 → Settings → *Set Up Claude Code* / *Set Up Gemini CLI*.
macOS will ask for three permissions. VisionCLI needs all of them:
| Permission | Why |
|---|---|
| **Microphone** and **Speech Recognition** | to hear your question while you hold ⌥ Space |
| **Screen & System Audio Recording** for **VisionCLI** | to see the window you're asking about in voice mode. After allowing it, quit and reopen VisionCLI. |
| **Screen & System Audio Recording** for **your terminal app** (iTerm, Terminal, Warp…) | when you *type* to Claude Code, the screenshot is taken by the terminal running `claude`. Without it, Claude only sees the wallpaper. Restart the terminal after allowing it. |
| **Automation** (iTerm, Chrome/Brave) | asked the first time it types into your terminal or controls the browser |
If something doesn't work, check System Settings → Privacy & Security, and 👁 → Settings → *Open Log*.
### 3. Optional extras (auto-detected)
| Add | Get | How |
|---|---|---|
| **Whisper** | Better transcription of names and jargon (~0.7s) | `brew install whisper-cpp`, then put a model such as [`ggml-small.en.bin`](https://huggingface.co/ggerganov/whisper.cpp) in `~/Library/Application Support/visioncli/models/` |
| **Kokoro** | Natural voices for spoken answers | install [`uv`](https://docs.astral.sh/uv/), then put [`kokoro-v1.0.onnx` and `voices-v1.0.bin`](https://github.com/thewh1teagle/kokoro-onnx/releases) in the same `models/` folder |
| **TypeSafe Jev** | Instant browser commands | add your key via 👁 → Settings → *API Keys*, or set `TYPESAFE_API_KEY` |
| **Browser control** | Click, scroll and read pages | in Chrome or Brave: **View → Developer → Allow JavaScript from Apple Events** |
## Gemini CLI
VisionCLI works with Gemini CLI as well as Claude Code:
- **Typing in `gemini`:** after *Set Up Gemini CLI*, the visioncli tools (look, browser, database, device screenshots…) are available in every Gemini CLI project.
- **Voice with Gemini:** choose 👁 → **Assistant → Gemini CLI**.
- *Gemini Terminal* mode types your question into the `gemini` running for the project (or opens one), attaches screenshots as `@file`, and shows the answer in both the terminal and the bubble.
- *Bubble Only* mode runs `gemini -p … -o stream-json`, resuming the same session each time.
- **Pointing at a terminal** that runs `claude` or `gemini` uses that CLI, whatever the Assistant setting.
**Sign-in:** Google no longer serves Gemini CLI on a *personal Google sign-in* (the free "Code Assist for individuals" tier says to move to Antigravity). Use a **Gemini API key** from [aistudio.google.com/apikey](https://aistudio.google.com/apikey): put it in 👁 → Settings → *API Keys* (`"gemini": {"apiKey": "..."}`) or set `GEMINI_API_KEY` before `visioncli voice`. VisionCLI then runs `gemini` with API-key auth for its own sessions (a per-run settings override), without changing your Gemini settings. Paid Code Assist (Standard/Enterprise) sign-ins keep working as-is.
**Limits:** in Bubble Only mode Gemini can't ask for approval, so it can only read (code, screen, browser, database); for edits and commands use Gemini Terminal mode. Highlight boxes, history, speech and Jev work the same with either assistant.
## Use it
### Voice (anywhere)
| Keys | What happens |
|---|---|
| **Hold ⌥ Space**, speak, let go | Ask about what's under the mouse (or tap to start, tap again to send) |
| **Hold ⌥ ⇧ Space** | Ask, and drag a box around exactly what you mean |
| **⌥ P** | Pin the answer: scroll, select text, click `file:line` links, toolbar (copy, read aloud, history) |
| **⌥ O** | Open the first `file:line` in your editor (JetBrains IDEs, Xcode, VS Code, Cursor) |
| **⌥ ↩ / esc** | Approve or deny an edit or command (background mode) |
| **esc** | Dismiss, or stop the answer |
**Which project?** By default VisionCLI follows the window under your mouse. PhpStorm `myshop – payments` maps to `~/PhpstormProjects/myshop`, and a terminal running `claude` maps to its folder. With several projects open, pick one in **👁 → Project**. The menu bar shows the current project, with 📌 when you've chosen one.
**Which session?** In **Terminal** mode (the default), voice questions are typed into the `claude` (or `gemini`, see [Gemini CLI](#gemini-cli)) already running for that project in iTerm, or a new one is opened. You can type there too; it's one conversation, and answers show in both the terminal and the bubble. **Bubble Only** mode uses a hidden, leaner session instead. Switch in 👁 → *Answer In*.
### In Claude Code (typing)
After *Set Up Claude Code*, just talk about what's on screen in any project:
- *"what is this file for?"*, *"why is this view misaligned in the simulator?"*
- *"let me select an area"*, *"read the page I have open and summarize it"*
- *"how are the orders and payments tables related?"*
Slash commands: `/mcp__visioncli__see <question>`, `/mcp__visioncli__point <question>`.
<details>
<summary><b>All MCP tools</b></summary>
| Tool | |
|---|---|
| `look` | Screenshot the window behind the terminal, the one under the mouse, a named app, or a display; reports the pointer position |
| `list_windows`, `select_region`, `clipboard_image` | Pick a window, drag a box, or use the clipboard image |
| `device_screenshot` | Clean iOS Simulator (needs Xcode) or Android (`adb`) screenshot |
| `db_schema`, `db_query` | Tables, columns, foreign keys; single read-only SELECT/SHOW/DESCRIBE/EXPLAIN/WITH queries in a read-only transaction (MySQL, Postgres, SQLite from `.env`) |
| `browser_tabs`, `browser_open`, `browser_switch_tab`, `browser_navigate`, `browser_read`, `browser_elements`, `browser_click`, `browser_type`, `browser_scroll` | Fast Chrome/Brave control via AppleScript |
| `start_voice` | Start the voice overlay |
</details>
## Menu
**👁 Project** (Auto, or pick from running `claude`/`gemini` sessions, open editors and recent projects) · **Assistant** (Claude Code or Gemini CLI) · **Answer In** (terminal, or bubble only) · **Model** (Sonnet or Opus, for Claude) · **Voice & Audio** (speak answers, speed up to 2×, voice, speech recognition, microphone, speaker) · History · Copy Last Answer · New Conversation · **Settings** (Set Up Claude Code / Gemini CLI, Start at Login, instant browser commands, AI cursor, API keys, logs).
## API keys
👁 → Settings → *API Keys* opens `~/Library/Application Support/visioncli/integrations.json` (readable only by you):
```json
{
"typesafe": { "apiKey": "..." },
"gemini": { "apiKey": "..." }
}
```
- `typesafe`: instant browser commands with [Jev](https://typesafe.ai). Without it, every request goes to the assistant.
- `gemini`: lets VisionCLI run Gemini CLI with API-key auth.
`TYPESAFE_API_KEY` / `GEMINI_API_KEY` in your shell are copied here by `visioncli voice` if the file has none.
## Privacy
- Screenshots and questions go to Claude through your own Claude Code login.
- With Jev enabled, quick commands send your spoken words, open tab titles and URLs, and visible link/button text to api.typesafe.ai. Turn it off in Settings.
- Whisper and Kokoro run locally. History is stored in `~/Library/Application Support/visioncli/`.
## Build from source
```bash
git clone https://github.com/thethirdsourcers/visioncli.git && cd visioncli
npm install # builds the MCP server (dist/)
node dist/index.js install # register with Claude Code (and Gemini CLI, if installed) for all projects
npm run build # build VisionCLI.app and install it into /Applications
node dist/index.js voice # start voice mode (under launchd: restarts after a crash)
```
- `npm run setup-signing` (optional, once) creates a local code-signing identity so macOS keeps permissions across rebuilds.
- `npm run release` builds the downloadable universal `.dmg` and `.zip` into `release/`. Pushing a `v*` tag does the same on GitHub Actions and attaches them to a release.
**Layout:** `src/` is the MCP server and CLI (TypeScript). `overlay/Sources/` is the menu-bar app (Swift). `scripts/` has the build, release and signing scripts.
## Troubleshooting
| Problem | Fix |
|---|---|
| "VisionCLI can't be opened" | Right-click → Open, or `xattr -dr com.apple.quarantine /Applications/VisionCLI.app` |
| Claude sees only the wallpaper | Voice: enable Screen Recording for VisionCLI. Typing in Claude Code: enable it for your terminal app. Then quit and reopen that app. |
| Tools missing in Claude Code / Gemini CLI | 👁 → Settings → *Set Up Claude Code* / *Set Up Gemini CLI* (again after moving the app or switching Node versions), then restart `claude` / `gemini` |
| Gemini: "no longer supported…" or "API key not valid" | Add a valid key from aistudio.google.com/apikey in 👁 → Settings → *API Keys* |
| Answers only in the terminal, or none at all | 👁 → Settings → *Open Log*; check `claude` is installed and signed in |
| Microphone errors | Pick a specific mic in 👁 → Voice & Audio → Microphone |
| It crashed | It restarts automatically; details are in 👁 → Settings → *Open Crash Log* |
## License
[MIT](LICENSE) © The Third Sourcers
TDQS
Scored across 17 tools
Tools cluster into clear domains (browser control, screen/device capture, database, voice), and the capture tools each target a distinct source (screen, simulator/device, clipboard, user selection). Minor overlap between browser_open (new tab) and browser_navigate (active tab), and between browser_switch_tab and browser_tabs, but descriptions disambiguate them.
Browser tools share a consistent browser_ prefix and db tools a db_ prefix, mostly verb_noun. The vision tools deviate slightly with a bare verb ('look') alongside verb_noun forms (list_windows, select_region, device_screenshot), but all remain snake_case and readable.
17 tools is slightly heavy but justified: browser control legitimately needs ~8 tools, capture needs several sources, database needs schema+query, plus voice. No tool feels redundant or out of scope.
Covers browser navigation, page reading, element interaction, multiple capture sources, DB introspection, and voice. Gaps are minor and mostly by design (read-only DB, no explicit back/forward history navigation), so agents can work around them.