Screenshot Website Fast
@just-every/mcp-screenshot-website-fast
CLIコーディングツール向けに最適化された、高速で効率的なウェブページスクリーンショットキャプチャツールです。フルページを自動的に1072x1072のチャンクにタイル分割し、最適な処理を実現します。
概要
AIビジョンワークフロー専用に構築されたこのツールは、Claude Vision APIやその他のAIモデルによる処理に最適な、自動解像度制限とタイル分割機能を備えた高品質なスクリーンショットをキャプチャします。互換性を最大化するため、スクリーンショットは1072x1072ピクセル(1.15メガピクセル)に完璧にサイズ調整されます。
Related MCP server: Webshot MCP
特徴
📸 高速スクリーンショットキャプチャ - Puppeteerヘッドレスブラウザを使用
🎯 Claude Vision最適化 - 自動解像度制限(最適な1.15メガピクセルを実現する1072x1072)
🔲 自動タイル分割 - フルページを自動的に1072x1072のタイルに分割
🎬 スクリーンキャストキャプチャ - 設定可能な間隔で一連のスクリーンショットを記録
🔄 常に最新のコンテンツ - キャッシュを使用しないため、常に最新のスクリーンショットを取得
📱 設定可能なビューポート - レスポンシブテストに対応
⏱️ 待機戦略 - 動的コンテンツに対応(networkidle、カスタム遅延)
📄 フルページキャプチャ - デフォルトでページ全体をキャプチャ
🎥 アニメーションWebPエクスポート - スクリーンキャストを高画質なアニメーションWebPファイルとして保存
💉 JavaScriptインジェクション - スクリーンキャストキャプチャ前にカスタムJSを実行
📦 最小限の依存関係 - 高速なnpmインストール
🔌 MCP統合 - シームレスなAIワークフローを実現
🪟 Windows対応ランチャー - npmインストールされたMCP利用に対応
🔋 リソース効率 - 60秒間非アクティブな場合にブラウザを自動クリーンアップ
🧹 メモリ管理 - リークを防ぐため、各スクリーンショット後にページを閉じる
インストール
Claude Code
claude mcp add screenshot-website-fast -s user -- npx -y @just-every/mcp-screenshot-website-fastVS Code
code --add-mcp '{"name":"screenshot-website-fast","command":"npx","args":["-y","@just-every/mcp-screenshot-website-fast"]}'Cursor
cursor://anysphere.cursor-deeplink/mcp/install?name=screenshot-website-fast&config=eyJzY3JlZW5zaG90LXdlYnNpdGUtZmFzdCI6eyJjb21tYW5kIjoibnB4IiwiYXJncyI6WyIteSIsIkBqdXN0LWV2ZXJ5L21jcC1zY3JlZW5zaG90LXdlYnNpdGUtZmFzdCJdfX0=JetBrains IDEs
設定 → ツール → AI Assistant → Model Context Protocol (MCP) → 追加
「JSONとして」を選択し、以下を貼り付けます:
{"command":"npx","args":["-y","@just-every/mcp-screenshot-website-fast"]}Raw JSON (任意のMCPクライアントで動作)
{
"mcpServers": {
"screenshot-website-fast": {
"command": "npx",
"args": ["-y", "@just-every/mcp-screenshot-website-fast"]
}
}
}これをクライアントのmcp.json(例: .vscode/mcp.json, ~/.cursor/mcp.json, またはClaude用の.mcp.json)に記述してください。
前提条件
Node.js 20.x以上
npm または npx
Chrome/Chromium (Puppeteerによって自動的にダウンロードされます)
クイックスタート
MCPサーバーの使用
IDEにインストールすると、以下のツールが利用可能になります:
利用可能なツール
take_screenshot- ウェブページの高品質なスクリーンショットをキャプチャしますパラメータ:
url(必須): キャプチャするHTTP/HTTPS URLwidth(任意): ビューポートの幅(ピクセル単位、最大1072、デフォルト: 1072)height(任意): ビューポートの高さ(ピクセル単位、最大1072、デフォルト: 1072)fullPage(任意): タイル分割を使用してフルページのスクリーンショットをキャプチャ(デフォルト: true)waitUntil(任意): 待機イベント: load, domcontentloaded, networkidle0, networkidle2(デフォルト: domcontentloaded)waitFor(任意): 追加の待機時間(ミリ秒単位)directory(任意): スクリーンショットを保存するディレクトリ - base64画像の代わりにファイルパスを返します
capture_selector- CSSセレクタに一致する特定のDOM要素のスクリーンショットをキャプチャしますパラメータ:
url(必須): キャプチャするHTTP/HTTPS URLselector(必須): キャプチャする要素のCSSセレクタwidth(任意): ビューポートの幅(ピクセル単位、最大1072、デフォルト: 1072)height(任意): ビューポートの高さ(ピクセル単位、最大1072、デフォルト: 1072)waitUntil(任意): 待機イベント: load, domcontentloaded, networkidle0, networkidle2(デフォルト: domcontentloaded)waitForMS(任意): 追加の待機時間(ミリ秒単位)selectorTimeoutMS(任意): セレクタが表示されるまで待機する時間(デフォルト: 5000)
使用例
デフォルトの使用方法(base64画像を返します):
take_screenshot(url="https://example.com")ディレクトリに保存(ファイルパスを返します):
take_screenshot(url="https://example.com", directory="/path/to/screenshots")特定の要素をキャプチャ:
capture_selector(url="https://example.com", selector="#main")directoryパラメータを使用する場合:
スクリーンショットはタイムスタンプ付きのPNGファイルとして保存されます
base64データの代わりにファイルパスが返されます
タイル分割されたスクリーンショットの場合、各タイルは個別のファイルとして保存されます
ディレクトリが存在しない場合は自動的に作成されます
take_screencast
スクリーンキャストを作成するために、一定期間にわたって一連のスクリーンショットをキャプチャします。ビューポートの最上部タイル(1072x1072)のみをキャプチャします。
パラメータ
url(必須): キャプチャするURLduration(任意): 合計時間(秒単位、デフォルト: 10)interval(任意): スクリーンショット間の間隔(秒単位、デフォルト: 2)jsEvaluate(任意): 開始時に実行するJavaScriptコードwaitUntil(任意): 待機戦略: 'load', 'domcontentloaded', 'networkidle0', 'networkidle2'waitForMS(任意): 開始前の追加待機時間directory(任意): アニメーションWebPとしてディレクトリに保存(1秒ごとにキャプチャ)
使用例
基本的なスクリーンキャスト(10秒間で5フレーム):
take_screencast(url="https://example.com")カスタムタイミング:
take_screencast(url="https://example.com", duration=15, interval=3)JavaScript実行を伴う場合:
take_screencast(
url="https://example.com",
jsEvaluate="document.body.style.backgroundColor = 'red';"
)アニメーションWebPとして保存:
take_screencast(url="https://example.com", directory="/path/to/output")directoryパラメータを使用する場合:
1秒間隔でアニメーションWebPが作成されます
個々のフレームもPNGファイルとして保存されます
アニメーションはデフォルトで無限ループします
WebPは優れた品質を提供します:
フルカラーサポート(256色の制限なし)
ウェブアニメーション向けの効率的な圧縮
グラデーション背景や滑らかなアニメーションに最適
GIFと比較して高品質かつ小さなファイルサイズ
開発での使用
インストール
npm install
npm run buildスクリーンショットのキャプチャ
# Full page with automatic tiling (default)
npm run dev capture https://example.com -o screenshot.png
# Viewport-only screenshot
npm run dev capture https://example.com --no-full-page -o screenshot.png
# Wait for specific conditions
npm run dev capture https://example.com --wait-until networkidle0 --wait-for 2000 -o screenshot.pngCLIオプション
-w, --width <pixels>- ビューポートの幅(最大1072、デフォルト: 1072)-h, --height <pixels>- ビューポートの高さ(最大1072、デフォルト: 1072)--no-full-page- フルページキャプチャとタイル分割を無効化--wait-until <event>- 待機イベント: load, domcontentloaded, networkidle0, networkidle2--wait-for <ms>- 追加の待機時間(ミリ秒単位)-o, --output <path>- 出力ファイルパス(タイル出力には必須)
自動再起動機能
MCPサーバーには、信頼性を向上させるためにデフォルトで自動再起動機能が含まれています:
クラッシュ時にサーバーを自動的に再起動
未処理の例外やPromiseの拒否を処理
指数バックオフを実装(1分間に最大10回試行)
監視のためにすべての再起動試行をログに記録
シャットダウン信号(SIGINT, SIGTERM)を適切に処理
自動再起動なしで開発/デバッグを行う場合:
# Run directly without restart wrapper
npm run serve:devアーキテクチャ
mcp-screenshot-website-fast/
├── src/
│ ├── internal/ # Core screenshot capture logic
│ ├── utils/ # Logger and utilities
│ ├── index.ts # CLI entry point
│ ├── serve.ts # MCP server entry point
│ └── serve-restart.ts # Auto-restart wrapper開発
# Run in development mode
npm run dev capture https://example.com -o screenshot.png
# Build for production
npm run build
# Run tests
npm test
# Type checking
npm run typecheck
# Linting
npm run lintなぜこのツールなのか?
AIビジョンワークフロー専用に構築されています:
Claude Vision API向けに最適化 - 1072x1072ピクセル(1.15メガピクセル)への自動解像度制限
自動タイル分割 - AI処理に最適なチャンクにフルページを分割
常に新鮮 - キャッシュなしで最新のコンテンツを確実に取得
MCPネイティブ - AI開発ツールとのファーストクラスの統合
シンプルなAPI - スクリーンショットをキャプチャするためのクリーンで直感的なインターフェース
貢献
貢献を歓迎します!以下の手順に従ってください:
リポジトリをフォークする
フィーチャーブランチを作成する
新機能のテストを追加する
プルリクエストを送信する
トラブルシューティング
Puppeteerの問題
Chrome/Chromiumがダウンロード可能であることを確認してください
ファイアウォールの設定を確認してください
PUPPETEER_SKIP_CHROMIUM_DOWNLOAD=trueを設定し、カスタム実行可能ファイルを提供してみてください
スクリーンショットの品質
ビューポートの寸法を調整してください
適切な待機戦略を使用してください
サイトが認証を必要とするか確認してください
タイムアウトエラー
--wait-forフラグで待機時間を増やしてください別の
--wait-until戦略を使用してくださいサイトにアクセス可能か確認してください
ライセンス
MIT
Available Tools
3 toolscapture_consoleARead-only
Capture console output from a web page. Accepts a URL, optional JS command to run, and duration to wait (default 4 seconds). Returns all console messages during that time.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | HTTP/HTTPS URL to capture console from | |
| jsCommand | No | Optional JavaScript command to execute on the page | |
| duration | No | Duration to capture console output in seconds | |
| waitUntil | No | Wait until event: load, domcontentloaded, networkidle0, networkidle2 | domcontentloaded |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already indicate read-only, non-destructive, and open-world traits, but the description adds valuable behavioral context: it specifies that the tool captures console messages over a duration, returns all messages during that time, and includes defaults (e.g., 4 seconds). This enhances understanding beyond annotations without contradiction.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the core purpose in the first sentence, followed by key parameters and return behavior in a second sentence. Every sentence adds value without redundancy, making it efficient and well-structured for quick comprehension.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's moderate complexity, rich annotations, and full schema coverage, the description is mostly complete. It covers purpose, key parameters, and output behavior, though it lacks details on error handling or specific use cases. With no output schema, it adequately explains returns, but could be slightly enhanced for full completeness.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 100% schema description coverage, the schema fully documents all parameters. The description adds minimal semantics by mentioning the URL, optional JS command, and duration with default, but does not provide additional meaning beyond what the schema already covers, such as explaining the waitUntil parameter or JS command usage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the specific action ('capture console output'), target resource ('from a web page'), and distinguishes from siblings by focusing on console messages rather than visual captures like take_screencast or take_screenshot. It uses precise language that defines the tool's unique function.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for capturing console output from web pages but does not explicitly state when to use this tool versus alternatives like take_screencast or take_screenshot. It provides some context with parameters but lacks explicit guidance on scenarios or exclusions, leaving usage somewhat inferred.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
take_screencastARead-only
Capture a series of screenshots of a web page over time, producing a screencast. Uses adaptive frame rates: 100ms intervals for ≤5s, 200ms for 5-10s, 500ms for >10s. PNG format: individual frames. WebP format: animated WebP with 4-second pause at end for looping.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | HTTP/HTTPS URL to capture | |
| duration | No | Total duration of screencast in seconds | |
| width | No | Viewport width in pixels (max 1072) | |
| height | No | Viewport height in pixels (max 1072) | |
| jsEvaluate | No | JavaScript code to execute. String: single instruction after first screenshot. Array: takes screenshot before each instruction, then continues capturing until duration ends. | |
| waitUntil | No | Wait until event: load, domcontentloaded, networkidle0, networkidle2 | domcontentloaded |
| directory | No | Save screencast to directory. Specify format with "format" parameter. | |
| format | No | Output format when using directory: "png" for individual PNG files, "webp" for animated WebP (default) | webp |
| quality | No | WebP quality level (only applies when format is "webp"): low (50), medium (75), high (90) | medium |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations provide readOnlyHint=true and destructiveHint=false, indicating safe operation. The description adds valuable behavioral context beyond annotations: adaptive frame rates (100ms, 200ms, 500ms intervals), output formats (PNG as individual frames, WebP as animated with 4-second pause), and format-specific details. It does not contradict annotations, as 'capture' aligns with read-only behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is appropriately sized and front-loaded, starting with the core purpose. It efficiently covers key behavioral traits in two sentences without redundancy. However, it could be slightly more structured by separating format details into distinct points for clarity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (9 parameters, no output schema) and rich annotations, the description is mostly complete. It explains adaptive frame rates and format behaviors, which are critical for usage. However, it does not cover all contextual aspects like error handling or performance implications, leaving minor gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema fully documents all 9 parameters. The description adds minimal parameter semantics, mentioning PNG and WebP formats and adaptive frame rates, which relate to 'format' and 'duration' parameters but do not provide significant additional meaning beyond the schema. Baseline 3 is appropriate given high schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: 'Capture a series of screenshots of a web page over time, producing a screencast.' It specifies the verb ('capture'), resource ('web page'), and output ('screencast'), distinguishing it from sibling tools like 'take_screenshot' (single screenshot) and 'capture_console' (different resource).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage through details like adaptive frame rates and format options, suggesting when to use it for time-based captures. However, it lacks explicit guidance on when to choose this tool over alternatives like 'take_screenshot' or 'capture_console', and does not mention prerequisites or exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
take_screenshotARead-only
Fast, efficient screenshot capture of web pages - optimized for CLI coding tools. Use this after performing updates to web pages to ensure your changes are displayed correctly. Automatically tiles full pages into 1072x1072 chunks for optimal processing.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | HTTP/HTTPS URL to capture | |
| width | No | Viewport width in pixels (max 1072) | |
| fullPage | No | Capture full page screenshot with tiling. If false, only the viewport is captured. | |
| waitUntil | No | Wait until event: load, domcontentloaded, networkidle0, networkidle2 | domcontentloaded |
| waitForMS | No | Additional wait time in milliseconds | |
| directory | No | Save tiled screenshots to a local directory (returns file paths instead of base64) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and destructiveHint=false, indicating a safe read operation. The description adds valuable behavioral context beyond annotations by specifying tiling behavior (1072x1072 chunks), optimization for CLI tools, and the purpose of verifying web page updates, though it doesn't cover rate limits or auth needs.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the core purpose, followed by usage guidelines and technical details, all in three concise sentences with zero wasted words, making it easy to scan and understand quickly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's moderate complexity, rich annotations, and full schema coverage, the description is mostly complete. It lacks details on output format (e.g., base64 vs. file paths) since there's no output schema, but otherwise covers purpose, usage, and key behaviors adequately.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema fully documents all 6 parameters. The description adds minimal parameter semantics by mentioning tiling and optimization, but doesn't provide additional details beyond what the schema already covers, meeting the baseline for high schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose with specific verbs ('capture', 'tiles') and resources ('web pages'), distinguishing it from sibling tools like capture_console and take_screencast by focusing on static screenshot functionality rather than console logs or video recordings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly states when to use this tool ('after performing updates to web pages to ensure your changes are displayed correctly') and provides context about its optimization for CLI coding tools, giving clear guidance without mentioning alternatives directly but implying its niche use case.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
3 tool updates
- First observed
capture_console - First observed
take_screencast - First observed
take_screenshot
TDQS
Scored across 3 tools
Each tool has a clearly distinct purpose: capture_console focuses on console output, take_screencast produces animated sequences, and take_screenshot captures static images. There is no overlap in functionality, making it easy for an agent to select the right tool based on the desired outcome.
All tool names follow a consistent verb_noun pattern (capture_console, take_screencast, take_screenshot) with clear, descriptive verbs. The naming is uniform and predictable, enhancing usability and reducing confusion.
With 3 tools, the server is well-scoped for its purpose of capturing different aspects of web pages (console output, screencasts, screenshots). Each tool earns its place by covering a distinct capture method, avoiding bloat or insufficiency.
The tool set provides comprehensive coverage for web page capture: console output, animated screencasts, and static screenshots. There are no obvious gaps, as these tools cover the main use cases for capturing web content in various formats and contexts.
Maintenance
Related MCP Connectors
Screenshot any URL/HTML as PNG/JPEG/WebP, or read it as clean Markdown/text for LLMs.
Screenshot, PDF and HTML-to-image rendering API so Claude and Cursor can see any web page.
Screenshot, PDF and HTML-to-image rendering API so Claude and Cursor can see any web page.
- GrabbitOAuthlive.grabbit
Screenshot any URL as a hosted image. No local browser; handles bot walls and full-page captures.
Related MCP Servers
- AlicenseAqualityDmaintenanceEnables AI assistants to capture screenshots of web pages using automated browser sessions. Supports full-page and element-specific screenshots, device simulation, and JavaScript execution for comprehensive web testing and monitoring.68 npmMIT
- AlicenseCqualityDmaintenanceEnables taking screenshots of web pages with support for multiple devices (desktop, mobile, tablet), custom dimensions, full-page capture, and various image formats. Built with Playwright for reliable web page rendering and screenshot generation.1MIT
- AlicenseNot gradedqualityDmaintenanceCaptures high-quality screenshots and screencasts of web pages, automatically tiling full pages into 1072x1072 chunks optimized for Claude Vision API and other AI vision models.362 npm26MIT
- AlicenseNot gradedqualityCmaintenanceCaptures comprehensive webpage screenshots with intelligent scrolling, text extraction, and HTML analysis, enabling AI tools to visually inspect and understand web content through the Model Context Protocol.3MIT