Webpage Screenshot MCP Server
ウェブページのスクリーンショット MCP サーバー
Puppeteerを使用してWebページのスクリーンショットをキャプチャするMCP(Model Context Protocol)サーバー。このサーバーにより、AIエージェントはWebアプリケーションを視覚的に検証し、Webアプリケーション生成の進行状況を確認できます。
特徴
フルページスクリーンショット: Webページ全体またはビューポートのみをキャプチャします
要素のスクリーンショット: CSSセレクターを使用して特定の要素をターゲットにする
複数のフォーマット: PNG、JPEG、WebP フォーマットをサポート
カスタマイズ可能なオプション: ビューポートのサイズ、画質、待機条件、遅延を設定します
Base64 エンコード: スクリーンショットを Base64 エンコードされた画像として返して簡単に統合できるようにします。
認証サポート: 手動ログインとCookieの永続化
デフォルトのブラウザ統合: システムのデフォルトのブラウザを使用して、より自然なエクスペリエンスを実現します
セッションの永続性: 複数ステップのワークフローのためにブラウザセッションを開いたままにします
Related MCP server: MCP Browser Screenshot Server
インストール
# Install globally
npm install -g screenshot-webpage-mcp
# Or use locally in a project
npm install screenshot-webpage-mcp使用法
ツール
この MCP サーバーはいくつかのツールを提供します。
1. ログインして待機する
手動ログイン用に表示されているブラウザ ウィンドウで Web ページを開き、ユーザーがログインを完了するまで待機してから、Cookie を保存します。
{
"url": "https://example.com/login",
"waitMinutes": 5,
"successIndicator": ".dashboard-welcome",
"useDefaultBrowser": true
}url(必須): ログインページのURLwaitMinutes(オプション):ログインを待つ最大時間(デフォルト:5)successIndicator(オプション): ログイン成功を示す CSS セレクターまたは URL パターンuseDefaultBrowser(オプション): システムのデフォルトのブラウザを使用するかどうか (デフォルト: true)
2. スクリーンショットページ
指定された URL のスクリーンショットをキャプチャし、base64 でエンコードされた画像として返します。
{
"url": "https://example.com/dashboard",
"fullPage": true,
"width": 1920,
"height": 1080,
"format": "png",
"quality": 80,
"waitFor": "networkidle2",
"delay": 500,
"useSavedAuth": true,
"reuseAuthPage": true,
"useDefaultBrowser": true,
"visibleBrowser": true
}url(必須): スクリーンショットを撮るウェブページのURLfullPage(オプション):ページ全体をキャプチャするか、ビューポートのみをキャプチャするか(デフォルト:true)width(オプション):ピクセル単位のビューポートの幅(デフォルト:1920)height(オプション):ピクセル単位のビューポートの高さ(デフォルト:1080)format(オプション):画像形式 - 「png」、「jpeg」、または「webp」(デフォルト:「png」)quality(オプション):画像の品質(0〜100)。jpegとwebpにのみ適用されます。waitFor(オプション): ページが読み込まれたとみなすタイミング - 「load」、「domcontentloaded」、「networkidle0」、または「networkidle2」(デフォルト: 「networkidle2」)delay(オプション):ページ読み込み後の追加遅延(ミリ秒)(デフォルト:0)useSavedAuth(オプション): 前回のログインから保存されたCookieを使用するかどうか (デフォルト: true)reuseAuthPage(オプション): 既存の認証済みページを使用するかどうか (デフォルト: false)useDefaultBrowser(オプション): システムのデフォルトのブラウザを使用するかどうか (デフォルト: false)visibleBrowser(オプション): ブラウザウィンドウを表示するかどうか (デフォルト: false)
3. スクリーンショット要素
CSS セレクターを使用して、Web ページ上の特定の要素のスクリーンショットをキャプチャします。
{
"url": "https://example.com/dashboard",
"selector": ".user-profile",
"waitForSelector": true,
"format": "png",
"quality": 80,
"padding": 10,
"useSavedAuth": true,
"useDefaultBrowser": true,
"visibleBrowser": true
}url(必須): ウェブページのURLselector(必須): スクリーンショットを撮る要素の CSS セレクターwaitForSelector(オプション): セレクターが表示されるまで待機するかどうか (デフォルト: true)format(オプション):画像形式 - 「png」、「jpeg」、または「webp」(デフォルト:「png」)quality(オプション):画像の品質(0〜100)。jpegとwebpにのみ適用されます。padding(オプション):要素の周囲のパディング(ピクセル単位)(デフォルト:0)useSavedAuth(オプション): 前回のログインから保存されたCookieを使用するかどうか (デフォルト: true)useDefaultBrowser(オプション): システムのデフォルトのブラウザを使用するかどうか (デフォルト: false)visibleBrowser(オプション): ブラウザウィンドウを表示するかどうか (デフォルト: false)
4. 認証クッキーをクリアする
特定のドメインまたはすべてのドメインの保存された認証 Cookie をクリアします。
{
"url": "https://example.com"
}url(オプション): Cookieを消去するドメインのURL。指定しない場合は、すべてのCookieが消去されます。
デフォルトのブラウザモード
デフォルトのブラウザモードでは、PuppeteerにバンドルされているChromiumの代わりに、システムの通常のブラウザ(Chrome、Edgeなど)を使用できます。これは以下の場合に便利です。
既存のブラウザセッションと拡張機能を使用する
保存した資格情報を使用してウェブサイトに手動でログインする
複数ステップのワークフローでより自然なブラウジング体験を実現
ユーザーと同じブラウザ環境でテストする
デフォルトのブラウザ モードを有効にするには、ツール パラメータでuseDefaultBrowser: trueとvisibleBrowser: trueを設定します。
デフォルトのブラウザモードの仕組み
デフォルトのブラウザ モードを有効にすると、次のようになります。
このツールは、システムのデフォルトのブラウザ (Chrome、Edge など) を見つけようとします。
ランダムなポートでリモートデバッグを有効にしてブラウザを起動します
Puppeteerは独自のブラウザインスタンスを起動する代わりに、このブラウザインスタンスに接続します。
既存のプロファイル、拡張機能、Cookieはセッション中に利用できます
ブラウザウィンドウは表示されたままなので、手動で操作できます。
このモードは、認証や複雑なユーザー操作を必要とするワークフローに特に役立ちます。
ブラウザの永続性
MCP サーバーは、複数のツール呼び出しにわたって永続的なブラウザ セッションを維持できます。
login-and-waitを使用すると、ブラウザセッションは開いたままになります以降の
screenshot-pageまたはscreenshot-elementへのreuseAuthPage: true呼び出しでは同じページが使用されます。これにより、再認証せずに複数ステップのワークフローが可能になります。
クッキー管理
アクセスしたドメインごとに Cookie が自動的に保存されます。
login-and-waitを使用した後、クッキーはホームフォルダの.mcp-screenshot-cookiesディレクトリに保存されます。これらのCookieは
useSavedAuth: trueで同じドメインに再度アクセスしたときに自動的に読み込まれます。clear-auth-cookiesツールを使用してCookieを消去できます。
ワークフローの例: 保護されたページのスクリーンショット
認証が必要なページのスクリーンショットを撮るワークフローの例を次に示します。
手動ログインフェーズ
{
"name": "login-and-wait",
"parameters": {
"url": "https://example.com/login",
"waitMinutes": 3,
"successIndicator": ".dashboard-welcome",
"useDefaultBrowser": true
}
}デフォルトのブラウザが開き、ログインページが表示されます。手動でログインすることもできます。ログインが完了すると(成功インジケーターが表示されるか、ログインページから移動した後)、セッションCookieが保存されます。
保存したセッションを使用してスクリーンショットを撮る
{
"name": "screenshot-page",
"parameters": {
"url": "https://example.com/account",
"fullPage": true,
"useSavedAuth": true,
"reuseAuthPage": true,
"useDefaultBrowser": true,
"visibleBrowser": true
}
}これにより、同じブラウザ ウィンドウに保存されている認証 Cookie を使用して、アカウント ページのスクリーンショットが撮影されます。
特定の要素のスクリーンショットを撮る
{
"name": "screenshot-element",
"parameters": {
"url": "https://example.com/dashboard",
"selector": ".user-profile-section",
"useSavedAuth": true,
"useDefaultBrowser": true,
"visibleBrowser": true
}
}完了したらCookieを消去する
{
"name": "clear-auth-cookies",
"parameters": {
"url": "https://example.com"
}
}このワークフローにより、通常のユーザーと同じように保護されたページを操作でき、デフォルトのブラウザで完全な認証フローを完了できます。
ヘッドレスモードと可視モード
ヘッドレス モード(
visibleBrowser: false): より高速で、ユーザーの操作が不要な自動ワークフローに適しています。表示モード(
visibleBrowser: true):ブラウザウィンドウを表示し、ユーザーによる操作と手動検証を可能にします。useDefaultBrowseruseDefaultBrowser: trueの場合に必須です。
プラットフォームサポート
デフォルトのブラウザ検出は次のブラウザで機能します:
macOS : Chrome、Edge、Safariを検出します
Windows : レジストリまたは一般的なインストールパスを介してChromeとEdgeを検出します
Linux : システムコマンド経由でChromeとChromiumを検出します
トラブルシューティング
よくある問題
デフォルトのブラウザが見つかりません: システムがデフォルトのブラウザを見つけられない場合は、Puppeteer にバンドルされている Chromium にフォールバックします。
接続の問題: ブラウザのデバッグ ポートへの接続に問題がある場合は、別のインスタンスがすでにそのポートを使用しているかどうかを確認してください。
Cookie の問題: 認証が機能しない場合は、
clear-auth-cookiesツールを使用して Cookie をクリアしてみてください。
デバッグ
MCPサーバーは、問題が発生すると、コンソールに役立つエラーメッセージを記録します。これらのメッセージでトラブルシューティング情報を確認してください。
Available Tools
5 toolsclear-auth-cookiesA
Clears saved authentication cookies for a specific domain or all domains
| Name | Required | Description | Default |
|---|---|---|---|
| url | No | URL of the domain to clear cookies for. If not provided, clears all cookies. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It states the action ('clears saved authentication cookies') but does not disclose behavioral traits such as whether this requires specific permissions, if it's reversible, potential side effects (e.g., logging out users), or rate limits. The description is minimal and lacks critical context for a mutation tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence with zero waste, front-loading the core action and scope. It is appropriately sized for a simple tool with one optional parameter.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (simple mutation with one parameter) and lack of annotations or output schema, the description is adequate but has clear gaps. It covers the basic purpose and parameter semantics via the schema, but fails to provide behavioral context needed for safe usage, such as permissions or side effects.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, with the parameter 'url' documented as 'URL of the domain to clear cookies for. If not provided, clears all cookies.' The description adds no additional meaning beyond this, as it only restates the same information. Baseline 3 is appropriate when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose with a specific verb ('clears') and resource ('saved authentication cookies'), and distinguishes its scope ('for a specific domain or all domains'). It directly answers what the tool does without being vague or tautological.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage context by specifying 'for a specific domain or all domains,' but does not explicitly state when to use this tool versus alternatives or provide exclusions. Given the sibling tools (e.g., 'login-and-wait'), it lacks guidance on when to clear cookies relative to login/logout workflows.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
login-and-waitA
Opens a webpage in a visible browser window for manual login, waits for user to complete login, then saves cookies
| Name | Required | Description | Default |
|---|---|---|---|
| successIndicator | No | Optional CSS selector or URL pattern that indicates successful login | |
| url | Yes | The URL of the login page | |
| useDefaultBrowser | No | Whether to use the system's default browser instead of Puppeteer's bundled Chromium | |
| waitMinutes | No | Maximum minutes to wait for login (default: 3) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It describes the tool's behavior well: opening a visible browser, waiting for manual login, and saving cookies. However, it misses details like error handling, what happens after timeout, or how cookies are saved/stored. It does not contradict annotations, as none exist.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence that front-loads the core purpose and steps. Every word earns its place, with no redundancy or unnecessary details, making it highly concise and well-structured.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no annotations and no output schema, the description adequately covers the tool's purpose and high-level behavior. However, for a tool with 4 parameters and no output schema, it lacks details on return values, error cases, or integration with sibling tools like 'signal-login-complete'. It's complete enough for basic understanding but has gaps for full contextual use.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema fully documents all parameters. The description does not add any parameter-specific information beyond what the schema provides (e.g., it doesn't explain 'successIndicator' usage or 'waitMinutes' implications). Baseline 3 is appropriate as the schema handles the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the specific action sequence: 'Opens a webpage in a visible browser window for manual login, waits for user to complete login, then saves cookies.' It uses precise verbs (opens, waits, saves) and identifies the resource (webpage, cookies), distinguishing it from sibling tools like screenshot tools or cookie-clearing tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for manual login scenarios where user interaction is required, but it does not explicitly state when to use this tool versus alternatives like automated login tools or other authentication methods. It provides clear context (manual login in a browser) but lacks explicit exclusions or named alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
screenshot-elementB
Captures a screenshot of a specific element on a webpage using a CSS selector
| Name | Required | Description | Default |
|---|---|---|---|
| format | No | Image format for the screenshot | png |
| padding | No | Padding around the element in pixels | |
| quality | No | Quality of the image (0-100), only applicable for jpeg and webp | |
| selector | Yes | CSS selector for the element to screenshot | |
| url | Yes | The URL of the webpage | |
| useDefaultBrowser | No | Whether to use the system's default browser instead of Puppeteer's bundled Chromium | |
| useSavedAuth | No | Whether to use saved cookies from previous login | |
| visibleBrowser | No | Whether to show the browser window (non-headless mode) | |
| waitForSelector | No | Whether to wait for the selector to appear |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries full burden but only states the basic function. It doesn't disclose important behavioral traits like: whether this navigates to new URLs, requires page loading, handles authentication, has rate limits, or what happens with invalid selectors. The description is minimal beyond the core action.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Single sentence, zero waste words, front-loaded with the core action. Every word earns its place by specifying element-level capture with CSS selector mechanism.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 9-parameter tool with no annotations and no output schema, the description is inadequate. It doesn't explain what the tool returns (image data? file path? error formats?), doesn't mention authentication dependencies despite sibling login tools, and provides minimal behavioral context for a complex screenshot operation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema fully documents all 9 parameters. The description adds no parameter-specific information beyond implying 'selector' and 'url' are involved. Baseline 3 is appropriate when schema does all parameter documentation work.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the specific action ('Captures a screenshot') and target resource ('specific element on a webpage'), using precise terminology ('CSS selector'). It distinguishes from sibling 'screenshot-page' by specifying element-level rather than page-level capture.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No explicit guidance on when to use this tool versus alternatives like 'screenshot-page' or other siblings. The description implies usage for element-specific screenshots but doesn't provide context about prerequisites (e.g., needing authentication via login tools) or exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
screenshot-pageA
Captures a screenshot of a given URL and returns it as base64 encoded image. Can use saved cookies from login-and-wait.
| Name | Required | Description | Default |
|---|---|---|---|
| delay | No | Additional delay in milliseconds to wait after page load | |
| format | No | Image format for the screenshot | png |
| fullPage | No | Whether to capture the full page or just the viewport | |
| height | No | Viewport height in pixels | |
| quality | No | Quality of the image (0-100), only applicable for jpeg and webp | |
| reuseAuthPage | No | Whether to use the existing authenticated page instead of creating a new one | |
| url | Yes | The URL of the webpage to screenshot | |
| useDefaultBrowser | No | Whether to use the system's default browser instead of Puppeteer's bundled Chromium | |
| useSavedAuth | No | Whether to use saved cookies from previous login | |
| visibleBrowser | No | Whether to show the browser window (non-headless mode) | |
| waitFor | No | When to consider the page loaded | networkidle2 |
| width | No | Viewport width in pixels |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It mentions the ability to use saved cookies, which hints at authentication behavior, but doesn't cover other important traits like performance implications (e.g., page load delays), potential failures (e.g., invalid URLs), or side effects (e.g., browser resource usage). The description adds some value but leaves significant gaps for a tool with 12 parameters.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, well-structured sentence that efficiently conveys the core functionality and a key feature (cookie reuse). Every word earns its place with no redundancy or fluff, making it appropriately sized and front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (12 parameters, no annotations, no output schema), the description is incomplete. It covers the basic purpose and authentication context but lacks details on behavioral traits, error handling, or output specifics (beyond base64 encoding). For a screenshot tool with many configuration options, more guidance on usage scenarios or limitations would be helpful.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all 12 parameters thoroughly. The description adds minimal semantic context by mentioning 'saved cookies from login-and-wait,' which loosely relates to the 'useSavedAuth' parameter, but doesn't provide additional meaning beyond what the schema specifies for most parameters. Baseline 3 is appropriate when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the specific action ('captures a screenshot') and resource ('of a given URL'), and distinguishes from sibling tools by mentioning the ability to use saved cookies from 'login-and-wait' (differentiating from 'screenshot-element' which targets specific elements). It also specifies the output format ('returns it as base64 encoded image').
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context by mentioning saved cookies from 'login-and-wait', which implies when to use this tool (for authenticated pages). However, it doesn't explicitly state when NOT to use it or name alternatives like 'screenshot-element' for element-specific captures, leaving some guidance implicit rather than explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
signal-login-completeA
Signals that manual login is complete and the login-and-wait tool should continue
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It describes the behavioral trait of signaling completion to another tool, which is useful context. However, it doesn't disclose other aspects like whether it requires specific permissions, has side effects, or how it interacts with authentication states, leaving some gaps.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence that front-loads the key information: signaling login completion. There is zero waste, and it earns its place by clearly stating the tool's role in the workflow.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's simplicity (0 parameters, no output schema, no annotations), the description is complete enough. It explains the purpose and usage in context with sibling tools. However, it could be slightly more complete by mentioning any prerequisites or effects, but for a signaling tool, this is adequate.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 0 parameters with 100% coverage, so no parameter documentation is needed. The description doesn't add param info, but this is acceptable given the lack of parameters. Baseline is 4 for 0 params, as it doesn't need to compensate for any gaps.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: to signal completion of manual login so another tool (login-and-wait) can continue. It specifies the verb 'signals' and the context 'manual login is complete,' but doesn't explicitly differentiate from all sibling tools like clear-auth-cookies or screenshot tools, which serve different purposes.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly states when to use this tool: 'when manual login is complete' and that it should be used to allow 'login-and-wait tool should continue.' It names the specific alternative tool (login-and-wait) and implies usage in a sequence, providing clear context without exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
5 tool updates
v1.0.0- First observed
clear-auth-cookies - First observed
login-and-wait - First observed
screenshot-element - First observed
screenshot-page - First observed
signal-login-complete
TDQS
Scored across 5 tools
Most tools have distinct purposes: screenshot-element and screenshot-page target different screenshot scopes, while login-and-wait and clear-auth-cookies handle authentication. However, signal-login-complete is tightly coupled with login-and-wait, which could cause confusion about whether to use it separately or as part of the login flow.
The naming is mixed: screenshot-element and screenshot-page follow a verb-noun pattern, but clear-auth-cookies and login-and-wait use hyphens and compound phrases, while signal-login-complete is a full sentence. This inconsistency makes the set less predictable, though the names remain readable.
With 5 tools, the count is well-scoped for a webpage screenshot server. Each tool serves a clear role in the workflow (authentication, screenshot capture, and cleanup), and there are no extraneous tools, making it efficient for agents to navigate.
The toolset covers core screenshot and authentication workflows effectively, including login, cookie management, and element/page capture. A minor gap is the lack of tools for advanced screenshot options (e.g., full-page capture or viewport adjustments), but agents can still accomplish the main tasks without significant workarounds.
Maintenance
Related MCP Connectors
Screenshot any public web page from an AI agent. Free without signup, or with an API key.
Generate images, GIFs, videos, and PDFs from HTML, URLs, or templates — from your AI agent.
- mcpOAuthcom.screenshotink
Screenshot, diff, audit and sitemap-capture any web page — 5 MCP tools for AI agents.
Captured web interfaces, screenshots, flows and structured evidence for AI coding agents.
1
Related MCP Servers
- AlicenseCqualityDmaintenanceEnables LLMs to perform web browsing tasks, take screenshots, and execute JavaScript using Puppeteer for browser automation.435,614 npm1MIT
- AlicenseAqualityDmaintenanceEnables AI assistants to capture screenshots of web pages using automated browser sessions. Supports full-page and element-specific screenshots, device simulation, and JavaScript execution for comprehensive web testing and monitoring.618 npmMIT
- AlicenseCqualityDmaintenanceEnables taking screenshots of web pages with support for multiple devices (desktop, mobile, tablet), custom dimensions, full-page capture, and various image formats. Built with Playwright for reliable web page rendering and screenshot generation.1MIT

site-shot-mcpofficial
AlicenseBqualityBmaintenanceEnables AI agents to capture full-page or viewport screenshots of any web page with options for ad removal, cookie banner blocking, and proxy country selection.273 npm3MIT