MCP Screenshot Server
Universal Screenshot MCP
Ein MCP (Model Context Protocol)-Server, der KI-Assistenten Screenshot-Funktionen bereitstellt – sowohl für die Aufnahme von Webseiten über Puppeteer als auch für plattformübergreifende System-Screenshots mithilfe nativer Betriebssystem-Tools.
Funktionen
Webseiten-Screenshots — Erfassen Sie jede öffentliche URL mit einem Headless-Chromium-Browser
Plattformübergreifende System-Screenshots — Vollbild-, Fenster- oder Bereichsaufnahmen mit nativen OS-Tools (macOS
screencapture, Linuxmaim/scrot/gnome-screenshot/etc., Windows PowerShell+.NET)Sicherheitsorientiertes Design — SSRF-Prävention, Schutz vor Pfad-Traversal, Schutz vor DNS-Rebinding, Verhinderung von Befehlsinjektionen und DoS-Begrenzung
MCP-nativ — Integriert sich direkt in Claude Desktop, Cursor und jeden MCP-kompatiblen Client
Related MCP server: Chrome DevTools MCP
Anforderungen
Node.js >= 18.0.0
Chromium wird beim ersten Start automatisch von Puppeteer heruntergeladen
Plattformspezifische Anforderungen für take_system_screenshot
Plattform | Erforderliche Tools | Hinweise |
macOS |
| Keine zusätzliche Installation erforderlich |
Linux | Eines der folgenden: |
|
Windows |
| Verwendet .NET |
Linux-Installationsbeispiele
# Ubuntu/Debian (recommended)
sudo apt install maim xdotool
# Fedora
sudo dnf install maim xdotool
# Arch Linux
sudo pacman -S maim xdotool
# Wayland (Sway, etc.)
sudo apt install grimNach der Installation können Sie Ihr Setup wie folgt überprüfen:
npx universal-screenshot-mcp --doctorDies untersucht den Host und gibt kopierbare Installationsbefehle für fehlende Tools aus, angepasst an Ihre erkannte Distribution.
Schnelleinstieg
Installation über npm
npm install -g universal-screenshot-mcpOder führen Sie es direkt mit npx aus:
npx universal-screenshot-mcpInstallation aus dem Quellcode
git clone https://github.com/sethbang/mcp-screenshot-server.git
cd mcp-screenshot-server
npm install
npm run buildKonfigurieren Sie Ihren MCP-Client
Fügen Sie den Server zur Konfiguration Ihres MCP-Clients hinzu. Für Claude Desktop bearbeiten Sie ~/Library/Application Support/Claude/claude_desktop_config.json:
{
"mcpServers": {
"screenshot-server": {
"command": "npx",
"args": ["-y", "universal-screenshot-mcp"]
}
}
}Oder bei Installation aus dem Quellcode:
{
"mcpServers": {
"screenshot-server": {
"command": "node",
"args": ["/absolute/path/to/mcp-screenshot-server/build/index.js"]
}
}
}Für Claude Code registrieren Sie den Server mit dem Befehl claude mcp add:
# Project scope (current directory only)
claude mcp add screenshot-server -- npx -y universal-screenshot-mcp
# User scope (available across all projects)
claude mcp add --scope user screenshot-server -- npx -y universal-screenshot-mcpOder bei Installation aus dem Quellcode:
claude mcp add screenshot-server -- node /absolute/path/to/mcp-screenshot-server/build/index.jsÜberprüfen Sie die Registrierung des Servers mit claude mcp list oder prüfen Sie den Live-Status innerhalb einer Sitzung mit /mcp.
Für Cursor oder andere MCP-Clients konsultieren Sie deren Dokumentation für die entsprechende Konfiguration.
Tools
Der Server stellt zwei MCP-Tools bereit:
take_screenshot
Erfasst eine Webseite (oder ein spezifisches Element) über einen Headless-Puppeteer-Browser.
Parameter | Typ | Erforderlich | Beschreibung |
| string | ✅ | Zu erfassende URL (nur http/https) |
| number | — | Viewport-Breite (1–3840) |
| number | — | Viewport-Höhe (1–2160) |
| boolean | — | Erfasst die gesamte scrollbare Seite |
| string | — | CSS-Selektor zur Erfassung eines spezifischen Elements |
| string | — | Vor der Aufnahme auf diesen Selektor warten |
| number | — | Verzögerung in Millisekunden (0–30000) |
| string | — | Ausgabedateipfad (Standard: |
Beispiel-Prompt:
Mache einen Screenshot von https://example.com mit 1920x1080
take_system_screenshot
Erfasst den Desktop, ein spezifisches Anwendungsfenster oder einen Bildschirmbereich mithilfe nativer OS-Tools. Funktioniert unter macOS, Linux und Windows.
Parameter | Typ | Erforderlich | Beschreibung |
| enum | ✅ |
|
| number | — | Fenster-ID für den Fenstermodus |
| string | — | App-Name (z. B. |
| object | — |
|
| number | — | Display-Nummer für Multi-Monitor-Setups |
| boolean | — | Mauszeiger in die Aufnahme einbeziehen |
| enum | — |
|
| number | — | Aufnahmeverzögerung in Sekunden (0–10) |
| string | — | Ausgabedateipfad (Standard: |
Plattformübergreifende Funktionsunterstützung
Funktion | macOS | Linux | Windows |
Vollbild | ✅ | ✅ | ✅ |
Bereich | ✅ | ✅ (maim, scrot, grim, import) | ✅ |
Fenster nach Name | ✅ | ⚠️ X11 + xdotool | ⚠️ best-effort |
Fenster nach ID | ✅ | ✅ nur X11 | ⚠️ HWND |
Multi-Display | ✅ | ⚠️ tool-abhängig | ✅ |
Cursor einbeziehen | ✅ | ⚠️ tool-abhängig | ⚠️ |
Verzögerung | ✅ | ✅ | ✅ |
Beispiel-Prompt:
Mache einen System-Screenshot des Safari-Fensters
Konfiguration
Umgebungsvariablen
Variable | Standard | Beschreibung |
|
| Standard-Ausgabeverzeichnis relativ zu |
|
| Auf |
Ausgabeverzeichnisse
Screenshots werden standardmäßig unter ~/Documents/screenshots gespeichert (konfigurierbar über SCREENSHOT_OUTPUT_DIR). Benutzerdefinierte Ausgabepfade müssen in eines dieser erlaubten Verzeichnisse aufgelöst werden:
Verzeichnis | Beschreibung |
| Standard-Ausgabeort (konfigurierbar) |
| Ursprünglicher Standardort |
| Benutzer-Download-Ordner |
| Benutzer-Dokumente-Ordner |
| System-Temporärverzeichnis |
Sicherheit
Dieser Server implementiert mehrere Ebenen der Sicherheitsabsicherung:
ID | Bedrohung | Gegenmaßnahme |
SEC-001 | SSRF / DNS-Rebinding | URLs werden gegen blockierte IP-Bereiche validiert; DNS wird vor der Anfrage aufgelöst mit IP-Pinning via |
SEC-003 | Befehlsinjektion | Alle Subprozesse verwenden |
SEC-004 | Pfad-Traversal | Ausgabepfade werden mit |
SEC-005 | Denial of Service | Gleichzeitige Puppeteer-Instanzen auf 3 via Semaphor begrenzt |
Für vollständige Details siehe docs/security.md.
Entwicklung
Skripte
Befehl | Beschreibung |
| Kompiliert TypeScript nach |
| Rekompiliert bei Dateiänderungen |
| Unit-Tests (schnell, vollständig gemockt) |
| Integrationstests (echtes DNS/Dateisystem) |
| E2E-Tests (echtes Puppeteer/native Tools) |
| Alle Testebenen zusammen |
| Linux E2E via Docker (erfordert Docker) |
| Tests im Watch-Modus ausführen |
| Tests mit Coverage-Bericht ausführen |
| Quellcode mit ESLint prüfen |
| Startet MCP Inspector zum Debuggen |
Projektstruktur
src/
├── index.ts # Entry point — stdio transport
├── server.ts # MCP server factory
├── config/
│ ├── index.ts # Static constants (limits, allowed dirs)
│ └── runtime.ts # Singleton semaphore, default directory
├── tools/
│ ├── take-screenshot.ts # Web page capture tool
│ └── take-system-screenshot.ts # macOS system capture tool
├── types/
│ └── index.ts # Shared TypeScript interfaces
├── utils/
│ ├── helpers.ts # Response builders, file utilities
│ ├── screenshot-provider.ts # Cross-platform provider interface + factory
│ ├── macos-provider.ts # macOS: screencapture wrapper
│ ├── linux-provider.ts # Linux: maim/scrot/gnome-screenshot/etc.
│ ├── windows-provider.ts # Windows: PowerShell + .NET System.Drawing
│ ├── macos.ts # Window ID lookup via CoreGraphics
│ └── semaphore.ts # Async concurrency limiter
└── validators/
├── path.ts # Output path validation (SEC-004)
└── url.ts # URL/SSRF validation (SEC-001)Testen
Tests verwenden Vitest in drei Ebenen:
Unit (
npm test) — Vollständige Dependency Injection, kein echtes I/O. Schnelle Feedbackschleife.Integration (
npm run test:integration) — Echte DNS-Auflösung, echtes Dateisystem mit temporären Verzeichnissen, echtes Puppeteer gegen einen lokalen HTTP-Server.E2E (
npm run test:e2e) — Echte native Screenshot-Tools. macOS-Tests laufen nativ; Linux-Tests laufen in Docker vianpm run test:linux.
npm test # Unit tests (~300ms)
npm run test:linux # Linux provider tests in Docker
npm run test:all # EverythingDebuggen mit dem MCP Inspector
npm run inspectorDies startet den MCP Inspector, der mit Ihrem gebauten Server verbunden ist, sodass Sie Tools interaktiv aufrufen können.
Lizenz
Apache-2.0 — Copyright 2026 Seth Bang
Available Tools
2 toolstake_screenshotB
Capture web page or element via headless browser. Saves to ~/Documents/screenshots by default (configurable via SCREENSHOT_OUTPUT_DIR env var).
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | URL to capture | |
| width | No | Viewport width | |
| height | No | Viewport height | |
| fullPage | No | Capture full page | |
| selector | No | CSS selector for element | |
| waitForSelector | No | Wait for selector | |
| waitForTimeout | No | Delay in ms | |
| outputPath | No | Absolute path, or relative to home dir |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries full burden but only mentions method and save location; it omits critical behavioral traits like destructiveness, permission needs, or return value.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two concise sentences: first states purpose, second details save behavior. Front-loaded and no wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite 8 parameters and no output schema, the description omits return value, error handling, and other essential context, leaving significant gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema covers all 8 parameters with descriptions, so baseline is 3. The description adds the default save path and env var override, providing marginal extra context.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses specific verb 'capture' and resource 'web page or element' with method 'via headless browser', clearly distinguishing it from sibling 'take_system_screenshot' which captures system screens.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for web screenshots but does not explicitly state when to use versus alternatives, nor does it mention prerequisites or exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
take_system_screenshotA
Capture desktop, window, or region screenshot. Cross-platform: macOS (screencapture), Linux (maim/scrot/gnome-screenshot/etc.), Windows (PowerShell+.NET). Saves to ~/Documents/screenshots by default (configurable via SCREENSHOT_OUTPUT_DIR env var). For window mode, provide windowName (app name like "Safari") or windowId.
| Name | Required | Description | Default |
|---|---|---|---|
| mode | Yes | fullscreen=entire screen, window=specific app (requires windowName or windowId), region=coordinates | |
| windowId | No | Window ID (for window mode) | |
| windowName | No | App name like "Safari", "Firefox" (for window mode) | |
| region | No | Region {x,y,width,height} | |
| display | No | Display number | |
| includeCursor | No | Include cursor | |
| format | No | Image format (png or jpg) | |
| delay | No | Delay seconds | |
| outputPath | No | Absolute path, or relative to home dir |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description must carry the full burden. It discloses the default save location and cross-platform tool dependencies but does not mention return behavior (e.g., file path or binary), permission requirements, or failure modes.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise with 4 sentences, front-loading the main action. It effectively uses bullet-like information for cross-platform details and mode instructions, though it could be slightly more structured.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given 9 parameters including a nested object and no output schema, the description covers default directory and platform support but omits return value, error handling, and comparison with the sibling tool. It is adequate but not fully complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description adds minor value by specifying default output directory and that windowName is an app name, but the schema already describes all parameters adequately.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states 'Capture desktop, window, or region screenshot' with specific verb and resource. It distinguishes from the sibling tool 'take_screenshot' by specifying cross-platform support and modes.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context for when to use each mode (fullscreen, window, region) and gives examples for window mode. However, it lacks guidance on when not to use this tool versus the sibling 'take_screenshot', missing explicit exclusion or alternative differentiation.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
2 tool updates
v1.2.0- Changed
take_screenshot10 fields changed- added
Input schema / $schemaAdded value: +"http://json-schema.org/draft-07/schema#" - added
Input schema / additionalPropertiesAdded value: +false - changed
Input schema / properties / fullPage / descriptionPrevious value: -"Capture full scrollable page"New value: +"Capture full page" - changed
Input schema / properties / height / descriptionPrevious value: -"Viewport height in pixels"New value: +"Viewport height" - changed
Input schema / properties / outputPath / descriptionPrevious value: -"Custom output path (optional)"New value: +"Absolute path, or relative to home dir" - added
Input schema / properties / selectorAdded value: +{ + "description": "CSS selector for element", + "type": "string" +} - changed
Input schema / properties / url / descriptionPrevious value: -"URL to capture (can be http://, https://, or file:///)"New value: +"URL to capture" - added
Input schema / properties / waitForSelectorAdded value: +{ + "description": "Wait for selector", + "type": "string" +} - added
Input schema / properties / waitForTimeoutAdded value: +{ + "description": "Delay in ms", + "maximum": 30000, + "minimum": 0, + "type": "number" +} - changed
Input schema / properties / width / descriptionPrevious value: -"Viewport width in pixels"New value: +"Viewport width"
- Added
take_system_screenshot
1 tool update
- First observed
take_screenshot
TDQS
Scored across 2 tools
The two tools are clearly distinct: one captures web pages/elements via headless browser, the other captures desktop/system screenshots. There is no overlap in functionality.
Both tools follow a consistent verb_noun pattern with snake_case (take_screenshot, take_system_screenshot), using the same 'take_' prefix.
Only 2 tools for a screenshot server is minimal but still covers the core use cases. More tools might be expected for management or configuration, but the count is not severely inadequate.
The tool set covers the primary screenshot domains (web and system). Minor gaps exist, such as listing or deleting screenshots, but core functionality is well-covered.
Maintenance
Related MCP Connectors
Browserless MCP — wraps the Browserless headless-Chromium REST API (browserless.io)
Screenshot any URL/HTML as PNG/JPEG/WebP, or read it as clean Markdown/text for LLMs.
Clean PNG/JPEG screenshots via REST or MCP, with goal-driven multi-step navigation.
Screenshot, PDF, OG-image, and page extraction (markdown/JSON) over MCP. Bearer key or x402.
Related MCP Servers
- AlicenseAqualityDmaintenanceCaptures screenshots of web pages using Puppeteer, allowing AI agents to visually verify web applications and see their progress when generating web apps.557MIT
- AlicenseNot gradedqualityAmaintenanceEnables AI coding assistants to control and inspect a live Chrome browser through Chrome DevTools. Provides browser automation, performance analysis, debugging capabilities, and network request monitoring.2,487,841 npm52,860Apache 2.0
- AlicenseCqualityDmaintenanceEnables taking screenshots of web pages with support for multiple devices (desktop, mobile, tablet), custom dimensions, full-page capture, and various image formats. Built with Playwright for reliable web page rendering and screenshot generation.1MIT
- AlicenseNot gradedqualityBmaintenanceEnables AI agents to capture any public URL as PNG, JPEG, or PDF via REST API or MCP tools, including screenshot capture, page description, and PDF rendering.16 npmMIT