site-shot-mcp
OfficialSite-Shot MCP サーバー
Claude、Cursor、および他のAIエージェントに任意のWebページを見る能力を提供します — Site-Shot を Model Context Protocol 経由でウェブサイトのスクリーンショットを撮ります。
実際のChromiumレンダリング・全ページキャプチャ・国別プロキシ・自動広告・Cookieバナー除去(よりクリーンな画像、より少ないビジョントークン)。
クイックスタート (Claude Desktop)
Site-Shot APIキーを https://www.site-shot.com/start/ で取得します。
これをClaude Desktopの設定(
claude_desktop_config.json)に追加します:
{
"mcpServers": {
"site-shot": {
"command": "npx",
"args": ["-y", "site-shot-mcp"],
"env": { "SITESHOT_API_KEY": "YOUR_API_KEY" }
}
}
}Claude Desktopを再起動します。"https://news.ycombinator.com の全ページスクリーンショットを撮って" と依頼すると、サーバーを呼び出して画像を表示します。
他のMCPクライアント(Cursor、Cline、VS Code、LangChain、CrewAI)でも同様に動作します — 環境変数に SITESHOT_API_KEY を設定し、クライアントを npx -y site-shot-mcp に向けます。
Related MCP server: Webpage Screenshot MCP Server
ツール
capture_screenshot
Webページのスクリーンショットを撮ります(デフォルトはビューポート)。
パラメータ | 型 | デフォルト | 備考 |
| string (必須) | — | キャプチャするページ |
| boolean |
| スクロール可能なページ全体をキャプチャ |
| number | APIデフォルト | ビューポート/デバイスサイズ |
|
|
| 画像形式 |
| boolean |
| 広告を削除 |
| boolean |
| Cookie同意ポップアップを削除 |
| string | — | 2文字の ISO 3166-1 alpha-2 コードによるプロキシ国、例: |
| boolean |
| 国にプロキシがない場合、米国にフォールバックせずエラーにする |
| string | — | 手動オーバーライド |
| number | APIデフォルト | キャプチャ前の追加待機時間(SPA/アニメーション) |
| number | 20000(全ページ) | キャプチャする高さの上限 |
スクリーンショットをMCP画像として返します。
「APIデフォルト」はこのパッケージが指定できる数値ではありません。
width、height、wait_msは渡した場合のみ転送されるため、渡さなかった場合に適用される値はSite-Shot APIが決定し、ここでのリリースなしに変更される可能性があります。1.1.0までのバージョンでは、APIが使用しないwidth/heightのピクセルサイズを出力していました — それらを省略して「デフォルト」を取得したエージェントは、返された画像にそれを示すものがないまま、異なるビューポートを取得していました。サイズが重要になる場合は、常に明示的な値を渡してください。
国コードはISOコードであり、国名ではありません。
"DE"を渡し、"Germany"は渡さないでください。APIはコードを完全一致で照合するため、国名を渡すと米国プロキシ経由でレンダリングされ、そのことを通知しません。そのため、サーバーはレンダリングを消費する前に完全な国名を拒否します。strict_country(デフォルトでオン)も同様に、利用できない国を黙って米国スクリーンショットにする代わりにエラーにします — フォールバックを希望する場合はfalseを渡してください。対応国 →
capture_full_page
capture_screenshot と同じですが、全ページキャプチャが有効です。
エージェント自身のブラウザではなく、このサーバーを呼ぶ理由
エージェントがブラウザを操作する場合、ページ自体をスクリーンショットできます — ログインが必要なページやフローをステップ実行する必要があるページには、それが適切なツールです。公開URLの場合、キャプチャをこのサーバーに委任する方が通常は優れたエンジニアリングです。各キャプチャは同じパイプラインで実行され(実行間の再計画なし)、一致するロケールとタイムゾーンを持つ特定の国から取得でき(country + strict_country)、返される前に画像分類器とその背後にあるエスカレーション式リトライラダーによってスコアリングされ、ブラウザセッションと毎回のビジョントークンではなく、1セント未満のコストで済みます。両方向を正直に論じた完全な比較: AIエージェント vs スクリーンショットAPI — 誰がページをキャプチャすべきか。
設定
環境変数 | 必須 | 説明 |
| はい | あなたのSite-Shot APIキー( |
このサーバーは既存のSite-Shot HTTP API(https://api.site-shot.com/)の薄いラッパーです — 別のバックエンドはありません。
ローカル開発
npm install
npm run check # syntax check
npm run smoke # offline tests (stubbed fetch, no API key needed)
SITESHOT_API_KEY=yourkey npm start # run the server on stdio要件
Node.js ≥ 18(組み込みの fetch を使用)。
ライセンス
MIT
Available Tools
2 toolscapture_full_pageCapture full-page website screenshotA
Take a full-page (entire scrollable height) screenshot of a web page with Site-Shot and return it as an image. Convenience wrapper around capture_screenshot with full-page capture enabled.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | The URL of the web page to capture. A bare domain like example.com is accepted (https:// is assumed). | |
| width | No | Viewport width in pixels (default 1280). | |
| height | No | Viewport height in pixels (default 1024). | |
| format | No | Image format. Default: png. | |
| block_ads | No | Remove ads for a cleaner screenshot. Default: true. | |
| block_cookie_banners | No | Remove cookie-consent banners/popups. Default: true. | |
| country | No | Render through a proxy in this country, e.g. "Germany" (auto-sets IP, language, time zone, geolocation). | |
| language | No | Override browser language, e.g. "de". | |
| time_zone | No | Override time zone, e.g. "Europe/Berlin". | |
| geolocation | No | Override geolocation as "lat,lng". | |
| wait_ms | No | Milliseconds to wait after load before capturing (for SPAs/animations). | |
| max_height | No | Cap the captured height in pixels (max 20000). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It describes the action as a wrapper but does not disclose side effects, output format details, or limitations beyond what the schema parameters cover.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with no wasted words; the core purpose is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has 12 parameters (all well-described in schema) and no output schema, the description is adequate as a summary but lacks details on return format and additional behavioral context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% with detailed parameter descriptions, so the description adds little extra meaning. Baseline score of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool takes a full-page screenshot and mentions it's a convenience wrapper around capture_screenshot with full-page capture enabled, distinguishing it from the sibling tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for full-page screenshots and references the sibling tool, but does not explicitly state when not to use it or provide alternative scenarios.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
capture_screenshotCapture website screenshotB
Take a screenshot of a web page with Site-Shot and return it as an image. Renders in a real Chromium browser. Supports viewport/device sizing, full-page capture, country proxies, and automatic ad & cookie-banner removal (cleaner image, fewer vision tokens).
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | The URL of the web page to capture. A bare domain like example.com is accepted (https:// is assumed). | |
| width | No | Viewport width in pixels (default 1280). | |
| height | No | Viewport height in pixels (default 1024). | |
| format | No | Image format. Default: png. | |
| block_ads | No | Remove ads for a cleaner screenshot. Default: true. | |
| block_cookie_banners | No | Remove cookie-consent banners/popups. Default: true. | |
| country | No | Render through a proxy in this country, e.g. "Germany" (auto-sets IP, language, time zone, geolocation). | |
| language | No | Override browser language, e.g. "de". | |
| time_zone | No | Override time zone, e.g. "Europe/Berlin". | |
| geolocation | No | Override geolocation as "lat,lng". | |
| wait_ms | No | Milliseconds to wait after load before capturing (for SPAs/animations). | |
| max_height | No | Cap the captured height in pixels (max 20000). | |
| full_page | No | Capture the entire scrollable page instead of just the viewport. Default: false. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the burden of behavioral disclosure. It states the tool renders in a real Chromium browser and automatically removes ads and cookie banners, which is helpful. However, it does not mention potential side effects, rate limits, execution time, or authentication requirements. It also does not clarify whether the screenshot is destructive or what happens to the browser instance after capture.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences long, making it relatively concise. The first sentence states the core action, and the second lists major features. It avoids extraneous details but could be slightly more compact by combining the two sentences or trimming the feature list slightly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the complexity of the tool (13 parameters, no output schema), the description provides a high-level overview of capabilities but lacks detail on return format (e.g., image type, resolution), error handling, and how features like 'full_page' work in practice. It is adequate for an experienced user but incomplete for a novice.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema covers all 13 parameters with descriptions, achieving 100% coverage. The tool description reiterates some schema concepts (viewport sizing, full-page capture, country proxies) but does not add significant new meaning beyond what the schema already provides. For example, 'country' parameter is explained in the schema; the description only mentions 'country proxies' generically. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states it takes a screenshot of a web page using Site-Shot and returns an image. It mentions features like viewport sizing, full-page capture, and ad removal. However, it does not explicitly distinguish itself from the sibling tool 'capture_full_page', which may cause confusion.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description lists features but provides no guidance on when to use this tool versus alternatives like 'capture_full_page'. It does not mention any prerequisites or conditions for use, nor does it explain when to use the 'full_page' parameter or when to prefer a different tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
2 tool updates
v1.0.1- Changed
capture_full_page3 fields changed- changed
Input schema / properties / url / descriptionPrevious value: -"The URL of the web page to capture."New value: +"The URL of the web page to capture. A bare domain like example.com is accepted (https:// is assumed)." - removed
Input schema / properties / url / formatRemoved value: -"uri" - added
Input schema / properties / url / minLengthAdded value: +1
- Changed
capture_screenshot3 fields changed- changed
Input schema / properties / url / descriptionPrevious value: -"The URL of the web page to capture."New value: +"The URL of the web page to capture. A bare domain like example.com is accepted (https:// is assumed)." - removed
Input schema / properties / url / formatRemoved value: -"uri" - added
Input schema / properties / url / minLengthAdded value: +1
2 tool updates
v0.1.1- First observed
capture_full_page - First observed
capture_screenshot
TDQS
Scored across 2 tools
The two tools are nearly identical; capture_full_page is explicitly a wrapper for capture_screenshot with full-page enabled. An agent would likely misuse them, as the difference is only a parameter.
Both use verb_noun pattern ('capture_screenshot', 'capture_full_page'), but 'full_page' is a qualifier while 'screenshot' is the resource; inconsistent because one tool name specifies a parameter in the name itself.
Two tools is minimal but arguably sufficient for a simple screenshot service. However, the duplication suggests one tool could have been omitted, making the surface slightly too heavy for the scope.
The set covers basic screenshot needs with features like viewport sizing, proxies, and ad removal. However, it lacks tools for specific device emulation or batch processing, which are common in screenshot services.
Maintenance
Related MCP Connectors
Screenshot any public web page from an AI agent. Free without signup, or with an API key.
Generate images, GIFs, videos, and PDFs from HTML, URLs, or templates — from your AI agent.
Desktop and mobile website screenshots plus page context for AI agents and automation workflows.
Screenshot any URL/HTML as PNG/JPEG/WebP, or read it as clean Markdown/text for LLMs.
Related MCP Servers
- AlicenseBqualityBmaintenanceAn official MCP server implementation that allows AI assistants to capture website screenshots through the ScreenshotOne API, enabling visual context from web pages during conversations.127 npm38MIT
- AlicenseAqualityDmaintenanceCaptures screenshots of web pages using Puppeteer, allowing AI agents to visually verify web applications and see their progress when generating web apps.557MIT
- AlicenseAqualityDmaintenanceEnables AI assistants to capture screenshots of web pages using automated browser sessions. Supports full-page and element-specific screenshots, device simulation, and JavaScript execution for comprehensive web testing and monitoring.618 npmMIT
- AlicenseCqualityDmaintenanceEnables taking screenshots of web pages with support for multiple devices (desktop, mobile, tablet), custom dimensions, full-page capture, and various image formats. Built with Playwright for reliable web page rendering and screenshot generation.1MIT