Skip to main content
Glama
README.md
<p align="center"><img src="docs/logo.png" width="128" alt="VisionCLI logo"></p>

<h1 align="center">VisionCLI</h1>

<p align="center">
Let Claude see your screen. Ask about the file, error, UI, database table or web page you're looking at,<br>
by typing in <b>Claude Code or Gemini CLI</b> or <b>by voice</b>, with the answer floating above your mouse pointer.
</p>

<p align="center">
<a href="https://github.com/thethirdsourcers/visioncli/releases/latest"><b>⬇︎ Download VisionCLI for macOS</b></a>
</p>

---

- **"What is this file for?"** Claude looks at the window under your mouse, finds the real file in your project and answers with `file:line` references.
- **Hold ⌥ Space and ask out loud.** A loader rides on your cursor, and the answer appears in a bubble above it, read aloud.
- **Works with your Claude Code session.** It types into the `claude` already running for that project, or opens one, so you can type and talk in the same conversation.
- **Points at things.** Claude draws glowing boxes on your screen around what it's talking about.
- **Fast browser control.** "Open the Stripe tab", "click sign in", "pause the video" run in under a second (via TypeSafe Jev), with a visible AI cursor.
- **Knows your data.** Read-only database questions ("how are payments and orders related?") using your project's `.env`.

## Install

### Requirements

- macOS 13 or later (Apple Silicon or Intel)
- [Claude Code](https://claude.com/claude-code) (`claude` CLI), signed in, **and/or** [Gemini CLI](https://github.com/google-gemini/gemini-cli) (`gemini`) with a Gemini API key (see [Gemini CLI](#gemini-cli))
- [Node.js](https://nodejs.org) 18 or later

### 1. Download the app

1. Download **`VisionCLI-<version>-macos.dmg`** from the [latest release](https://github.com/thethirdsourcers/visioncli/releases/latest).
2. Open it and drag **VisionCLI** into **Applications**.
3. The app isn't notarized by Apple yet, so the first time:
   - **Right-click VisionCLI → Open → Open**, or
   - run `xattr -dr com.apple.quarantine /Applications/VisionCLI.app` in Terminal.

### 2. First launch

The 👁 icon appears in the menu bar and a welcome screen offers to set up the CLIs it finds. That registers VisionCLI's tools for every project: `claude mcp add visioncli --scope user …` for Claude Code, and `gemini mcp add -s user --trust visioncli …` for Gemini CLI. You can also do it later from 👁 → Settings → *Set Up Claude Code* / *Set Up Gemini CLI*.

macOS will ask for three permissions. VisionCLI needs all of them:

| Permission | Why |
|---|---|
| **Microphone** and **Speech Recognition** | to hear your question while you hold ⌥ Space |
| **Screen & System Audio Recording** for **VisionCLI** | to see the window you're asking about in voice mode. After allowing it, quit and reopen VisionCLI. |
| **Screen & System Audio Recording** for **your terminal app** (iTerm, Terminal, Warp…) | when you *type* to Claude Code, the screenshot is taken by the terminal running `claude`. Without it, Claude only sees the wallpaper. Restart the terminal after allowing it. |
| **Automation** (iTerm, Chrome/Brave) | asked the first time it types into your terminal or controls the browser |

If something doesn't work, check System Settings → Privacy & Security, and 👁 → Settings → *Open Log*.

### 3. Optional extras (auto-detected)

| Add | Get | How |
|---|---|---|
| **Whisper** | Better transcription of names and jargon (~0.7s) | `brew install whisper-cpp`, then put a model such as [`ggml-small.en.bin`](https://huggingface.co/ggerganov/whisper.cpp) in `~/Library/Application Support/visioncli/models/` |
| **Kokoro** | Natural voices for spoken answers | install [`uv`](https://docs.astral.sh/uv/), then put [`kokoro-v1.0.onnx` and `voices-v1.0.bin`](https://github.com/thewh1teagle/kokoro-onnx/releases) in the same `models/` folder |
| **TypeSafe Jev** | Instant browser commands | add your key via 👁 → Settings → *API Keys*, or set `TYPESAFE_API_KEY` |
| **Browser control** | Click, scroll and read pages | in Chrome or Brave: **View → Developer → Allow JavaScript from Apple Events** |

## Gemini CLI

VisionCLI works with Gemini CLI as well as Claude Code:

- **Typing in `gemini`:** after *Set Up Gemini CLI*, the visioncli tools (look, browser, database, device screenshots…) are available in every Gemini CLI project.
- **Voice with Gemini:** choose 👁 → **Assistant → Gemini CLI**.
  - *Gemini Terminal* mode types your question into the `gemini` running for the project (or opens one), attaches screenshots as `@file`, and shows the answer in both the terminal and the bubble.
  - *Bubble Only* mode runs `gemini -p … -o stream-json`, resuming the same session each time.
- **Pointing at a terminal** that runs `claude` or `gemini` uses that CLI, whatever the Assistant setting.

**Sign-in:** Google no longer serves Gemini CLI on a *personal Google sign-in* (the free "Code Assist for individuals" tier says to move to Antigravity). Use a **Gemini API key** from [aistudio.google.com/apikey](https://aistudio.google.com/apikey): put it in 👁 → Settings → *API Keys* (`"gemini": {"apiKey": "..."}`) or set `GEMINI_API_KEY` before `visioncli voice`. VisionCLI then runs `gemini` with API-key auth for its own sessions (a per-run settings override), without changing your Gemini settings. Paid Code Assist (Standard/Enterprise) sign-ins keep working as-is.

**Limits:** in Bubble Only mode Gemini can't ask for approval, so it can only read (code, screen, browser, database); for edits and commands use Gemini Terminal mode. Highlight boxes, history, speech and Jev work the same with either assistant.

## Use it

### Voice (anywhere)

| Keys | What happens |
|---|---|
| **Hold ⌥ Space**, speak, let go | Ask about what's under the mouse (or tap to start, tap again to send) |
| **Hold ⌥ ⇧ Space** | Ask, and drag a box around exactly what you mean |
| **⌥ P** | Pin the answer: scroll, select text, click `file:line` links, toolbar (copy, read aloud, history) |
| **⌥ O** | Open the first `file:line` in your editor (JetBrains IDEs, Xcode, VS Code, Cursor) |
| **⌥ ↩ / esc** | Approve or deny an edit or command (background mode) |
| **esc** | Dismiss, or stop the answer |

**Which project?** By default VisionCLI follows the window under your mouse. PhpStorm `myshop – payments` maps to `~/PhpstormProjects/myshop`, and a terminal running `claude` maps to its folder. With several projects open, pick one in **👁 → Project**. The menu bar shows the current project, with 📌 when you've chosen one.

**Which session?** In **Terminal** mode (the default), voice questions are typed into the `claude` (or `gemini`, see [Gemini CLI](#gemini-cli)) already running for that project in iTerm, or a new one is opened. You can type there too; it's one conversation, and answers show in both the terminal and the bubble. **Bubble Only** mode uses a hidden, leaner session instead. Switch in 👁 → *Answer In*.

### In Claude Code (typing)

After *Set Up Claude Code*, just talk about what's on screen in any project:

- *"what is this file for?"*, *"why is this view misaligned in the simulator?"*
- *"let me select an area"*, *"read the page I have open and summarize it"*
- *"how are the orders and payments tables related?"*

Slash commands: `/mcp__visioncli__see <question>`, `/mcp__visioncli__point <question>`.

<details>
<summary><b>All MCP tools</b></summary>

| Tool | |
|---|---|
| `look` | Screenshot the window behind the terminal, the one under the mouse, a named app, or a display; reports the pointer position |
| `list_windows`, `select_region`, `clipboard_image` | Pick a window, drag a box, or use the clipboard image |
| `device_screenshot` | Clean iOS Simulator (needs Xcode) or Android (`adb`) screenshot |
| `db_schema`, `db_query` | Tables, columns, foreign keys; single read-only SELECT/SHOW/DESCRIBE/EXPLAIN/WITH queries in a read-only transaction (MySQL, Postgres, SQLite from `.env`) |
| `browser_tabs`, `browser_open`, `browser_switch_tab`, `browser_navigate`, `browser_read`, `browser_elements`, `browser_click`, `browser_type`, `browser_scroll` | Fast Chrome/Brave control via AppleScript |
| `start_voice` | Start the voice overlay |
</details>

## Menu

**👁 Project** (Auto, or pick from running `claude`/`gemini` sessions, open editors and recent projects) · **Assistant** (Claude Code or Gemini CLI) · **Answer In** (terminal, or bubble only) · **Model** (Sonnet or Opus, for Claude) · **Voice & Audio** (speak answers, speed up to 2×, voice, speech recognition, microphone, speaker) · History · Copy Last Answer · New Conversation · **Settings** (Set Up Claude Code / Gemini CLI, Start at Login, instant browser commands, AI cursor, API keys, logs).

## API keys

👁 → Settings → *API Keys* opens `~/Library/Application Support/visioncli/integrations.json` (readable only by you):

```json
{
  "typesafe": { "apiKey": "..." },
  "gemini": { "apiKey": "..." }
}
```

- `typesafe`: instant browser commands with [Jev](https://typesafe.ai). Without it, every request goes to the assistant.
- `gemini`: lets VisionCLI run Gemini CLI with API-key auth.

`TYPESAFE_API_KEY` / `GEMINI_API_KEY` in your shell are copied here by `visioncli voice` if the file has none.

## Privacy

- Screenshots and questions go to Claude through your own Claude Code login.
- With Jev enabled, quick commands send your spoken words, open tab titles and URLs, and visible link/button text to api.typesafe.ai. Turn it off in Settings.
- Whisper and Kokoro run locally. History is stored in `~/Library/Application Support/visioncli/`.

## Build from source

```bash
git clone https://github.com/thethirdsourcers/visioncli.git && cd visioncli
npm install                       # builds the MCP server (dist/)
node dist/index.js install        # register with Claude Code (and Gemini CLI, if installed) for all projects
npm run build                     # build VisionCLI.app and install it into /Applications
node dist/index.js voice          # start voice mode (under launchd: restarts after a crash)
```

- `npm run setup-signing` (optional, once) creates a local code-signing identity so macOS keeps permissions across rebuilds.
- `npm run release` builds the downloadable universal `.dmg` and `.zip` into `release/`. Pushing a `v*` tag does the same on GitHub Actions and attaches them to a release.

**Layout:** `src/` is the MCP server and CLI (TypeScript). `overlay/Sources/` is the menu-bar app (Swift). `scripts/` has the build, release and signing scripts.

## Troubleshooting

| Problem | Fix |
|---|---|
| "VisionCLI can't be opened" | Right-click → Open, or `xattr -dr com.apple.quarantine /Applications/VisionCLI.app` |
| Claude sees only the wallpaper | Voice: enable Screen Recording for VisionCLI. Typing in Claude Code: enable it for your terminal app. Then quit and reopen that app. |
| Tools missing in Claude Code / Gemini CLI | 👁 → Settings → *Set Up Claude Code* / *Set Up Gemini CLI* (again after moving the app or switching Node versions), then restart `claude` / `gemini` |
| Gemini: "no longer supported…" or "API key not valid" | Add a valid key from aistudio.google.com/apikey in 👁 → Settings → *API Keys* |
| Answers only in the terminal, or none at all | 👁 → Settings → *Open Log*; check `claude` is installed and signed in |
| Microphone errors | Pick a specific mic in 👁 → Voice & Audio → Microphone |
| It crashed | It restarts automatically; details are in 👁 → Settings → *Open Crash Log* |

## License

[MIT](LICENSE) © The Third Sourcers

TDQS

B3.4/5.0

Scored across 17 tools

Disambiguation4/5

Tools cluster into clear domains (browser control, screen/device capture, database, voice), and the capture tools each target a distinct source (screen, simulator/device, clipboard, user selection). Minor overlap between browser_open (new tab) and browser_navigate (active tab), and between browser_switch_tab and browser_tabs, but descriptions disambiguate them.

Naming Consistency4/5

Browser tools share a consistent browser_ prefix and db tools a db_ prefix, mostly verb_noun. The vision tools deviate slightly with a bare verb ('look') alongside verb_noun forms (list_windows, select_region, device_screenshot), but all remain snake_case and readable.

Tool Count4/5

17 tools is slightly heavy but justified: browser control legitimately needs ~8 tools, capture needs several sources, database needs schema+query, plus voice. No tool feels redundant or out of scope.

Completeness4/5

Covers browser navigation, page reading, element interaction, multiple capture sources, DB introspection, and voice. Gaps are minor and mostly by design (read-only DB, no explicit back/forward history navigation), so agents can work around them.

Maintenance

ActivityMaintained
ResponsivenessNo issues