MCP Screenshot Server
Универсальный скриншот-сервер MCP
Сервер MCP (Model Context Protocol), предоставляющий ИИ-ассистентам возможности создания скриншотов — как веб-страниц через Puppeteer, так и системных скриншотов на разных платформах с использованием нативных инструментов ОС.
Возможности
Скриншоты веб-страниц — захват любого публичного URL с помощью headless-браузера Chromium
Кроссплатформенные системные скриншоты — захват всего экрана, окна или области с помощью нативных инструментов ОС (macOS
screencapture, Linuxmaim/scrot/gnome-screenshot/и др., Windows PowerShell+.NET)Безопасность прежде всего — защита от SSRF, обхода путей (path traversal), перепривязки DNS, инъекций команд и ограничение DoS
Нативная поддержка MCP — прямая интеграция с Claude Desktop, Cursor и любым клиентом, совместимым с MCP
Related MCP server: Chrome DevTools MCP
Требования
Node.js >= 18.0.0
Chromium скачивается автоматически Puppeteer при первом запуске
Специфические требования для take_system_screenshot
Платформа | Необходимые инструменты | Примечания |
macOS |
| Дополнительная установка не требуется |
Linux | Один из: | Рекомендуются |
Windows |
| Использует .NET |
Примеры установки для Linux
# Ubuntu/Debian (recommended)
sudo apt install maim xdotool
# Fedora
sudo dnf install maim xdotool
# Arch Linux
sudo pacman -S maim xdotool
# Wayland (Sway, etc.)
sudo apt install grimПосле установки вы можете проверить настройки командой:
npx universal-screenshot-mcp --doctorЭта команда сканирует хост и выводит готовые команды для установки недостающих инструментов, адаптированные под ваш дистрибутив.
Быстрый старт
Установка из npm
npm install -g universal-screenshot-mcpИли запустите напрямую через npx:
npx universal-screenshot-mcpУстановка из исходного кода
git clone https://github.com/sethbang/mcp-screenshot-server.git
cd mcp-screenshot-server
npm install
npm run buildНастройка вашего MCP-клиента
Добавьте сервер в конфигурацию вашего MCP-клиента. Для Claude Desktop отредактируйте ~/Library/Application Support/Claude/claude_desktop_config.json:
{
"mcpServers": {
"screenshot-server": {
"command": "npx",
"args": ["-y", "universal-screenshot-mcp"]
}
}
}Или, если установка произведена из исходного кода:
{
"mcpServers": {
"screenshot-server": {
"command": "node",
"args": ["/absolute/path/to/mcp-screenshot-server/build/index.js"]
}
}
}Для Claude Code зарегистрируйте сервер командой claude mcp add:
# Project scope (current directory only)
claude mcp add screenshot-server -- npx -y universal-screenshot-mcp
# User scope (available across all projects)
claude mcp add --scope user screenshot-server -- npx -y universal-screenshot-mcpИли, если установка произведена из исходного кода:
claude mcp add screenshot-server -- node /absolute/path/to/mcp-screenshot-server/build/index.jsПроверьте регистрацию сервера командой claude mcp list или проверьте статус в реальном времени внутри сессии с помощью /mcp.
Для Cursor или других MCP-клиентов обратитесь к их документации для получения информации о соответствующей конфигурации.
Инструменты
Сервер предоставляет два инструмента MCP:
take_screenshot
Делает скриншот веб-страницы (или конкретного элемента) через headless-браузер Puppeteer.
Параметр | Тип | Обязательный | Описание |
| string | ✅ | URL для захвата (только http/https) |
| number | — | Ширина области просмотра (1–3840) |
| number | — | Высота области просмотра (1–2160) |
| boolean | — | Захват всей прокручиваемой страницы |
| string | — | CSS-селектор для захвата конкретного элемента |
| string | — | Ожидание селектора перед захватом |
| number | — | Задержка в миллисекундах (0–30000) |
| string | — | Путь к выходному файлу (по умолчанию: |
Пример запроса:
Сделай скриншот https://example.com с размером 1920x1080
take_system_screenshot
Делает скриншот рабочего стола, конкретного окна приложения или области экрана с помощью нативных инструментов ОС. Работает на macOS, Linux и Windows.
Параметр | Тип | Обязательный | Описание |
| enum | ✅ |
|
| number | — | ID окна для режима окна |
| string | — | Имя приложения (например, |
| object | — |
|
| number | — | Номер дисплея для конфигураций с несколькими мониторами |
| boolean | — | Включить курсор мыши в захват |
| enum | — |
|
| number | — | Задержка захвата в секундах (0–10) |
| string | — | Путь к выходному файлу (по умолчанию: |
Поддержка функций на разных платформах
Функция | macOS | Linux | Windows |
Весь экран | ✅ | ✅ | ✅ |
Область | ✅ | ✅ (maim, scrot, grim, import) | ✅ |
Окно по имени | ✅ | ⚠️ X11 + xdotool | ⚠️ best-effort |
Окно по ID | ✅ | ✅ только X11 | ⚠️ HWND |
Несколько дисплеев | ✅ | ⚠️ зависит от инструмента | ✅ |
Включить курсор | ✅ | ⚠️ зависит от инструмента | ⚠️ |
Задержка | ✅ | ✅ | ✅ |
Пример запроса:
Сделай системный скриншот окна Safari
Конфигурация
Переменные окружения
Переменная | По умолчанию | Описание |
|
| Директория вывода по умолчанию относительно |
|
| Установите |
Директории вывода
Скриншоты по умолчанию сохраняются в ~/Documents/screenshots (настраивается через SCREENSHOT_OUTPUT_DIR). Пользовательские пути вывода должны указывать на одну из разрешенных директорий:
Директория | Описание |
| Расположение по умолчанию (настраиваемое) |
| Исходное расположение по умолчанию |
| Папка загрузок пользователя |
| Папка документов пользователя |
| Системная временная директория |
Безопасность
Этот сервер реализует несколько уровней защиты:
ID | Угроза | Смягчение |
SEC-001 | SSRF / DNS rebinding | URL проверяются по заблокированным диапазонам IP; DNS разрешается до запроса с привязкой IP через |
SEC-003 | Инъекция команд | Все подпроцессы используют |
SEC-004 | Обход путей | Пути вывода проверяются через |
SEC-005 | Отказ в обслуживании | Параллельные экземпляры Puppeteer ограничены 3 через семафор |
Для получения полной информации см. docs/security.md.
Разработка
Скрипты
Команда | Описание |
| Компиляция TypeScript в |
| Перекомпиляция при изменениях файлов |
| Модульные тесты (быстрые, полностью замоканные) |
| Интеграционные тесты (реальный DNS/файловая система) |
| E2E-тесты (реальный Puppeteer/нативные инструменты) |
| Все уровни тестов вместе |
| Linux e2e через Docker (требуется Docker) |
| Запуск тестов в режиме наблюдения |
| Запуск тестов с отчетом о покрытии |
| Линтинг исходного кода с помощью ESLint |
| Запуск MCP Inspector для отладки |
Структура проекта
src/
├── index.ts # Entry point — stdio transport
├── server.ts # MCP server factory
├── config/
│ ├── index.ts # Static constants (limits, allowed dirs)
│ └── runtime.ts # Singleton semaphore, default directory
├── tools/
│ ├── take-screenshot.ts # Web page capture tool
│ └── take-system-screenshot.ts # macOS system capture tool
├── types/
│ └── index.ts # Shared TypeScript interfaces
├── utils/
│ ├── helpers.ts # Response builders, file utilities
│ ├── screenshot-provider.ts # Cross-platform provider interface + factory
│ ├── macos-provider.ts # macOS: screencapture wrapper
│ ├── linux-provider.ts # Linux: maim/scrot/gnome-screenshot/etc.
│ ├── windows-provider.ts # Windows: PowerShell + .NET System.Drawing
│ ├── macos.ts # Window ID lookup via CoreGraphics
│ └── semaphore.ts # Async concurrency limiter
└── validators/
├── path.ts # Output path validation (SEC-004)
└── url.ts # URL/SSRF validation (SEC-001)Тестирование
Тесты используют Vitest на трех уровнях:
Модульные (
npm test) — Полная инъекция зависимостей, никакого реального ввода-вывода. Быстрый цикл обратной связи.Интеграционные (
npm run test:integration) — Реальное разрешение DNS, реальная файловая система с временными директориями, реальный Puppeteer против локального HTTP-сервера.E2E (
npm run test:e2e) — Реальные нативные инструменты для скриншотов. Тесты macOS запускаются нативно; тесты Linux запускаются в Docker черезnpm run test:linux.
npm test # Unit tests (~300ms)
npm run test:linux # Linux provider tests in Docker
npm run test:all # EverythingОтладка с помощью MCP Inspector
npm run inspectorЭто запускает MCP Inspector, подключенный к вашему собранному серверу, позволяя вызывать инструменты в интерактивном режиме.
Лицензия
Apache-2.0 — Copyright 2026 Seth Bang
Available Tools
2 toolstake_screenshotB
Capture web page or element via headless browser. Saves to ~/Documents/screenshots by default (configurable via SCREENSHOT_OUTPUT_DIR env var).
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | URL to capture | |
| width | No | Viewport width | |
| height | No | Viewport height | |
| fullPage | No | Capture full page | |
| selector | No | CSS selector for element | |
| waitForSelector | No | Wait for selector | |
| waitForTimeout | No | Delay in ms | |
| outputPath | No | Absolute path, or relative to home dir |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries full burden but only mentions method and save location; it omits critical behavioral traits like destructiveness, permission needs, or return value.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two concise sentences: first states purpose, second details save behavior. Front-loaded and no wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite 8 parameters and no output schema, the description omits return value, error handling, and other essential context, leaving significant gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema covers all 8 parameters with descriptions, so baseline is 3. The description adds the default save path and env var override, providing marginal extra context.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses specific verb 'capture' and resource 'web page or element' with method 'via headless browser', clearly distinguishing it from sibling 'take_system_screenshot' which captures system screens.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for web screenshots but does not explicitly state when to use versus alternatives, nor does it mention prerequisites or exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
take_system_screenshotA
Capture desktop, window, or region screenshot. Cross-platform: macOS (screencapture), Linux (maim/scrot/gnome-screenshot/etc.), Windows (PowerShell+.NET). Saves to ~/Documents/screenshots by default (configurable via SCREENSHOT_OUTPUT_DIR env var). For window mode, provide windowName (app name like "Safari") or windowId.
| Name | Required | Description | Default |
|---|---|---|---|
| mode | Yes | fullscreen=entire screen, window=specific app (requires windowName or windowId), region=coordinates | |
| windowId | No | Window ID (for window mode) | |
| windowName | No | App name like "Safari", "Firefox" (for window mode) | |
| region | No | Region {x,y,width,height} | |
| display | No | Display number | |
| includeCursor | No | Include cursor | |
| format | No | Image format (png or jpg) | |
| delay | No | Delay seconds | |
| outputPath | No | Absolute path, or relative to home dir |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description must carry the full burden. It discloses the default save location and cross-platform tool dependencies but does not mention return behavior (e.g., file path or binary), permission requirements, or failure modes.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise with 4 sentences, front-loading the main action. It effectively uses bullet-like information for cross-platform details and mode instructions, though it could be slightly more structured.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given 9 parameters including a nested object and no output schema, the description covers default directory and platform support but omits return value, error handling, and comparison with the sibling tool. It is adequate but not fully complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description adds minor value by specifying default output directory and that windowName is an app name, but the schema already describes all parameters adequately.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states 'Capture desktop, window, or region screenshot' with specific verb and resource. It distinguishes from the sibling tool 'take_screenshot' by specifying cross-platform support and modes.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context for when to use each mode (fullscreen, window, region) and gives examples for window mode. However, it lacks guidance on when not to use this tool versus the sibling 'take_screenshot', missing explicit exclusion or alternative differentiation.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
2 tool updates
v1.2.0- Changed
take_screenshot10 fields changed- added
Input schema / $schemaAdded value: +"http://json-schema.org/draft-07/schema#" - added
Input schema / additionalPropertiesAdded value: +false - changed
Input schema / properties / fullPage / descriptionPrevious value: -"Capture full scrollable page"New value: +"Capture full page" - changed
Input schema / properties / height / descriptionPrevious value: -"Viewport height in pixels"New value: +"Viewport height" - changed
Input schema / properties / outputPath / descriptionPrevious value: -"Custom output path (optional)"New value: +"Absolute path, or relative to home dir" - added
Input schema / properties / selectorAdded value: +{ + "description": "CSS selector for element", + "type": "string" +} - changed
Input schema / properties / url / descriptionPrevious value: -"URL to capture (can be http://, https://, or file:///)"New value: +"URL to capture" - added
Input schema / properties / waitForSelectorAdded value: +{ + "description": "Wait for selector", + "type": "string" +} - added
Input schema / properties / waitForTimeoutAdded value: +{ + "description": "Delay in ms", + "maximum": 30000, + "minimum": 0, + "type": "number" +} - changed
Input schema / properties / width / descriptionPrevious value: -"Viewport width in pixels"New value: +"Viewport width"
- Added
take_system_screenshot
1 tool update
- First observed
take_screenshot
TDQS
Scored across 2 tools
The two tools are clearly distinct: one captures web pages/elements via headless browser, the other captures desktop/system screenshots. There is no overlap in functionality.
Both tools follow a consistent verb_noun pattern with snake_case (take_screenshot, take_system_screenshot), using the same 'take_' prefix.
Only 2 tools for a screenshot server is minimal but still covers the core use cases. More tools might be expected for management or configuration, but the count is not severely inadequate.
The tool set covers the primary screenshot domains (web and system). Minor gaps exist, such as listing or deleting screenshots, but core functionality is well-covered.
Maintenance
Related MCP Connectors
Browserless MCP — wraps the Browserless headless-Chromium REST API (browserless.io)
Screenshot any URL/HTML as PNG/JPEG/WebP, or read it as clean Markdown/text for LLMs.
Clean PNG/JPEG screenshots via REST or MCP, with goal-driven multi-step navigation.
Screenshot, PDF, OG-image, and page extraction (markdown/JSON) over MCP. Bearer key or x402.
Related MCP Servers
- AlicenseAqualityDmaintenanceCaptures screenshots of web pages using Puppeteer, allowing AI agents to visually verify web applications and see their progress when generating web apps.557MIT
- AlicenseNot gradedqualityAmaintenanceEnables AI coding assistants to control and inspect a live Chrome browser through Chrome DevTools. Provides browser automation, performance analysis, debugging capabilities, and network request monitoring.2,487,841 npm52,860Apache 2.0
- AlicenseCqualityDmaintenanceEnables taking screenshots of web pages with support for multiple devices (desktop, mobile, tablet), custom dimensions, full-page capture, and various image formats. Built with Playwright for reliable web page rendering and screenshot generation.1MIT
- AlicenseNot gradedqualityBmaintenanceEnables AI agents to capture any public URL as PNG, JPEG, or PDF via REST API or MCP tools, including screenshot capture, page description, and PDF rendering.16 npmMIT