Argus
Allows installing Argus from GitHub and links to its demo and CI workflows.
Argus sends captured images to the AI tool's model provider, including Google's models, when using Gemini CLI.
Argus sends captured images to the AI tool's model provider, including OpenAI's models, when using Codex CLI.
Distributes the argus-screens package for installation via uv tool.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@ArgusLook at my screen and tell me what that error says"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Argus
Eyes for your AI tools across every monitor. Argus is an MCP server (plus a CLI and a tray app) for Windows that lets Claude Code, Claude Desktop, VS Code, Codex, Gemini CLI, Antigravity and other MCP clients see your screens. That covers the window you're looking at, any monitor, any window or any region, at full resolution when it matters.

Built from real Argus output on demo windows and a real Claude reply (how).
Named after Argus Panoptes, the hundred-eyed giant: one eye per monitor.
What you can say
You say | The AI calls |
"Look at this", "what's wrong on my screen?", "fix what I'm looking at" |
|
"Check my right monitor", "screenshot the laptop screen", "show me all screens" |
|
"Look at the Chrome window with localhost:3000" |
|
"Look at my snip" (after Win+Shift+S) |
|
(press Ctrl+Alt+S over a window) "look at what I marked" |
|
"What does that error say exactly?" |
|
"Where's the Submit button?" |
|
"Wait until the build finishes, then check it" |
|
"Did my CSS change break the page?" |
|
Related MCP server: desktop-touch-mcp
Install
You need Windows 10 (1903+) or 11, uv, and at least one MCP-capable AI tool.
uv tool install argus-screens --managed-python --python 3.13
argus install # registers Argus with the AI tools it finds; starts the tray app with Windows
argus doctor # checks monitors, OCR, the tray and each registrationTo run the latest unreleased code, install from GitHub instead:
uv tool install git+https://github.com/sahildayal/argus --managed-python --python 3.13.
Then open a new session in your AI tool and say "look at what I'm looking at". Running sessions don't pick up new MCP servers, and Claude Desktop needs a quit and reopen.
argus install configures whichever of these it finds, using each tool's own
mcp add command where it has one:
Tool | How it's registered |
Claude Code |
|
Claude Desktop |
|
VS Code (Copilot agent mode, MCP extensions) |
|
Codex CLI |
|
Gemini CLI |
|
Antigravity |
|
JSON files are backed up first (*.bak-argus-<time>). Your AI tools will ask before
each capture by default. argus install --trust pre-approves Argus in Claude Code
and Gemini CLI instead. Other options: --clients claude-code,vscode,
--no-startup and --dry-run. argus uninstall reverses everything.
Other MCP clients (Cursor, Windsurf, Zed, JetBrains, Cline, ...): add a stdio
server that runs argus-mcp. Use the full path, which where.exe argus-mcp prints:
{ "mcpServers": { "argus": { "command": "C:\\Users\\YOU\\.local\\bin\\argus-mcp.exe" } } }Updating: uv tool upgrade argus-screens (for a GitHub install, re-run the install command with --reinstall).
Use a uv-managed Python (the
--managed-pythonflag), not the Microsoft Store one: Store Python silently redirects writes underAppDatainto a private sandbox, so edits to other apps' configs would never reach them.
How "the window I'm looking at" works
When you type into a terminal or IDE, that window has focus, so it can't be
the answer. Argus keeps a per-session history of which windows had focus and works
out which program launched it (Windows Terminal, VS Code, Claude Desktop, ...). look then picks:
a hotkey mark you made in the last 2 minutes, if any;
otherwise the active window, if it isn't the chat itself (you switched after typing);
otherwise the window you used just before switching to the chat;
otherwise the front-most other window.
It always says why it chose that window and lists alternatives (for example the window under your mouse), so the AI can correct course.
Monitors
Monitors are numbered left to right (main row first, then any monitor above
or below), and each also answers to a position word worked out from your actual
layout: left, center, right, center-left, top, bottom, top-left...
Also: primary, laptop, cursor (where the mouse is), active (where the
focused window is), part of the model name (dell), or all. Mixed scaling and
monitors left of/above the primary (negative coordinates) are handled; the
layout is re-read on every call, so plugging and unplugging is fine.
Tools (MCP)
Tool | What it does |
| What you're looking at (see above). |
| A monitor, a window (covered windows render themselves), or a region; nothing = every monitor. |
| Full-resolution crop of an earlier capture by |
| Monitors and windows front to back, with handles and focus history. Text only, cheap. |
| Your newest Win+Shift+S / Snipping Tool / PrtScn snips, or a copied image. |
| Your newest hotkey mark (window or its monitor). |
| OCR text with positions (Windows' built-in engine, offline). |
| Where text is on screen, with a zoomed image of the best match. |
| Visual regression: numbered red boxes around every change. |
| Wait for |
Images are sent at up to 2000 px on the long side. A 1080p monitor goes
through untouched, and long sessions full of screenshots stay inside the
API's limits for requests with many images. Every capture is also saved as a
full-resolution PNG (path in the reply) with its metadata embedded, so zoom
can reopen it later by path.
Tray app and hotkeys
argus-tray runs in the notification area (and starts with Windows after argus install).
Ctrl+Alt+S: mark the window under the mouse (falls back to the focused window). A cyan outline flashes to show what was captured; the next
lookuses it.Ctrl+Alt+P: pause or resume Argus. While paused, every tool in every AI app refuses to capture (and the icon turns grey with a red slash).
Menu: pause for 15 minutes, open captures/marks folders, edit settings.
Change the hotkeys in ~/.argus/config.toml, then restart the tray.
Privacy and security
Argus gives AI tools sight of your screens, so it's built to make that a deliberate choice:
Blacked out before anything leaves Argus: password managers, WhatsApp, Signal, Telegram, Messenger, Phone Link, Windows notification toasts, and browser tabs whose titles match banking/payments/messaging patterns. The AI sees a dark box saying "Hidden by Argus" and never the title. The defaults are a starting point, so add your own bank, sites, apps or keywords under
[privacy]in~/.argus/config.toml. Chrome doesn't put "Incognito" in its window title, so private windows can't be detected automatically.Pause with Ctrl+Alt+P, the tray menu, or
argus pause [--minutes N].Retention: saved captures and marks are deleted after 7 days by default. Baselines are kept until you delete them.
Where images go: only to the AI tool that asked for them, and so to that tool's model provider (Anthropic, OpenAI, Google, ...). Argus itself never uploads anything, has no telemetry, and makes no network requests.
Prompt injection: anything on screen is input to the agent. A web page or document could try to steer an agent into capturing something. Keep per-capture prompts on (the default) unless you trust your setup, and pause Argus around sensitive work.
Everything personal lives in ~/.argus/ (settings, captures, marks, baselines and
logs), never in this repository. Set ARGUS_HOME to move it.
CLI
The CLI is useful for scripts and tests, and for any agent that can run shell
commands but doesn't speak MCP. It prints the saved PNG path for each capture;
add --json for machine-readable output.
argus list # monitors + windows
argus look # the window behind this terminal
argus shot -m right --grid # a monitor, with a labeled grid
argus shot -w "localhost:3000" -o page.png # a window, copied to page.png
argus zoom $env:USERPROFILE\.argus\captures\<date>\<file>.png --box 100,100,600,400
argus ocr -w "Windows Terminal" # exact text
argus find "Submit" -m cursor
argus wait text --text "Compiled successfully" -w "Terminal" -t 120 # exit 0 = seen, 1 = timeout
argus baseline save login -w "localhost:3000"; argus baseline compare login
argus pause --minutes 30; argus resume; argus statusSettings: ~/.argus/config.toml
Created with comments on first run. You can set the image size sent to the AI,
the grid default, the mouse marker, how fresh a mark must be for look, the
hotkeys, retention, and the privacy lists. Changes apply on the next capture;
hotkey changes need a tray restart.
Troubleshooting
Start with
argus doctor. It checks DPI awareness, monitors, OCR, the tray and each registration.The AI doesn't see Argus. Start a new session. In Claude Code,
claude mcp get argusshould say Connected.A hotkey does nothing. Another app may own it, in which case the tray shows a warning when it starts. Pick another in
config.tomland restartargus-tray.Small text looks blurry. Ask the AI to
zoomorread_textinstead of guessing.Logs:
~/.argus/logs/server.logand~/.argus/logs/tray.log.
Uninstall
argus uninstall [--purge] # unregisters everywhere, removes the startup entry, stops the tray
uv tool uninstall argus-screensContributing
Issues and pull requests are welcome. See CONTRIBUTING.md for setup, the two test suites, and the ground rules (tests never capture the real screen; privacy defaults stay conservative).
License
This server cannot be deployed
Maintenance
Related MCP Connectors
Eyes and hands on real Windows PCs — observe, click, type via Glasswarp API.
OCR, transcription, file extraction, and image generation for AI agents via MCP.
Clean PNG/JPEG screenshots via REST or MCP, with goal-driven multi-step navigation.
- mcpOAuthcom.screenshotink
Screenshot, diff, audit and sitemap-capture any web page — 5 MCP tools for AI agents.
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceMCP server for capturing screenshots of desktop windows on Windows. Allows AI assistants to see what's on screen for UI development, debugging, and iterating on designs.MIT
- AlicenseAqualityCmaintenanceAllows AI clients to see and control Windows 10/11 desktops via MCP, with screenshots, UI Automation, Chrome CDP, keyboard/mouse, and terminal using semantic element targeting.30541 npmMIT
- AlicenseNot gradedqualityBmaintenanceEnables AI agents to select Windows desktop UI elements, windows, or screen regions via mouse hover, then obtain context through UI Automation, screenshots, and local OCR for MCP-compatible clients.2MIT
- AlicenseNot gradedqualityAmaintenanceEnables MCP-compatible AI platforms to see the screen and operate any desktop software through real mouse clicks, text input, key presses, and scripted scenario execution.MIT