Webpage Screenshot MCP Server
Screenshot der Webseite MCP-Server
Ein MCP-Server (Model Context Protocol), der mit Puppeteer Screenshots von Webseiten erstellt. Dieser Server ermöglicht es KI-Agenten, Webanwendungen visuell zu überprüfen und ihren Fortschritt bei der Generierung von Web-Apps zu verfolgen.
Merkmale
Screenshots ganzer Seiten : Erfassen Sie ganze Webseiten oder nur den Ansichtsbereich
Element-Screenshots : Zielen Sie mit CSS-Selektoren auf bestimmte Elemente
Mehrere Formate : Unterstützung für die Formate PNG, JPEG und WebP
Anpassbare Optionen : Festlegen der Ansichtsfenstergröße, Bildqualität, Wartebedingungen und Verzögerungen
Base64-Kodierung : Gibt Screenshots als Base64-kodierte Bilder zurück, für eine einfache Integration
Authentifizierungsunterstützung : Manuelle Anmeldung und Cookie-Persistenz
Standardbrowserintegration : Verwenden Sie den Standardbrowser Ihres Systems für ein natürlicheres Erlebnis
Sitzungspersistenz : Halten Sie Browsersitzungen für mehrstufige Workflows geöffnet
Related MCP server: MCP Browser Screenshot Server
Installation
# Install globally
npm install -g screenshot-webpage-mcp
# Or use locally in a project
npm install screenshot-webpage-mcpVerwendung
Werkzeuge
Dieser MCP-Server bietet mehrere Tools:
1. Anmelden und warten
Öffnet eine Webseite in einem sichtbaren Browserfenster zur manuellen Anmeldung, wartet, bis der Benutzer die Anmeldung abgeschlossen hat, und speichert dann Cookies.
{
"url": "https://example.com/login",
"waitMinutes": 5,
"successIndicator": ".dashboard-welcome",
"useDefaultBrowser": true
}url(erforderlich): Die URL der AnmeldeseitewaitMinutes(optional): Maximale Wartezeit in Minuten für die Anmeldung (Standard: 5)successIndicator(optional): CSS-Selektor oder URL-Muster, das eine erfolgreiche Anmeldung anzeigtuseDefaultBrowser(optional): Ob der Standardbrowser des Systems verwendet werden soll (Standard: true)
2. Screenshot-Seite
Erstellt einen Screenshot einer bestimmten URL und gibt ihn als Base64-codiertes Bild zurück.
{
"url": "https://example.com/dashboard",
"fullPage": true,
"width": 1920,
"height": 1080,
"format": "png",
"quality": 80,
"waitFor": "networkidle2",
"delay": 500,
"useSavedAuth": true,
"reuseAuthPage": true,
"useDefaultBrowser": true,
"visibleBrowser": true
}url(erforderlich): Die URL der Webseite, von der ein Screenshot erstellt werden sollfullPage(optional): Ob die ganze Seite oder nur der Ansichtsbereich erfasst werden soll (Standard: true)width(optional): Ansichtsfensterbreite in Pixeln (Standard: 1920)height(optional): Höhe des Ansichtsfensters in Pixeln (Standard: 1080)format(optional): Bildformat – „png“, „jpeg“ oder „webp“ (Standard: „png“)quality(optional): Qualität des Bildes (0-100), gilt nur für JPEG und WebPwaitFor(optional): Wann soll die Seite als geladen betrachtet werden – „load“, „domcontentloaded“, „networkidle0“ oder „networkidle2“ (Standard: „networkidle2“)delay(optional): Zusätzliche Verzögerung in Millisekunden nach dem Laden der Seite (Standard: 0)useSavedAuth(optional): Ob gespeicherte Cookies vom vorherigen Login verwendet werden sollen (Standard: true)reuseAuthPage(optional): Ob die vorhandene authentifizierte Seite verwendet werden soll (Standard: false)useDefaultBrowser(optional): Ob der Standardbrowser des Systems verwendet werden soll (Standard: false)visibleBrowser(optional): Ob das Browserfenster angezeigt werden soll (Standard: false)
3. Screenshot-Element
Erstellt mithilfe eines CSS-Selektors einen Screenshot eines bestimmten Elements auf einer Webseite.
{
"url": "https://example.com/dashboard",
"selector": ".user-profile",
"waitForSelector": true,
"format": "png",
"quality": 80,
"padding": 10,
"useSavedAuth": true,
"useDefaultBrowser": true,
"visibleBrowser": true
}url(erforderlich): Die URL der Webseiteselector(erforderlich): CSS-Selektor für das Element, für das ein Screenshot erstellt werden sollwaitForSelector(optional): Ob auf das Erscheinen des Selektors gewartet werden soll (Standard: true)format(optional): Bildformat – „png“, „jpeg“ oder „webp“ (Standard: „png“)quality(optional): Qualität des Bildes (0-100), gilt nur für JPEG und WebPpadding(optional): Abstand um das Element in Pixeln (Standard: 0)useSavedAuth(optional): Ob gespeicherte Cookies vom vorherigen Login verwendet werden sollen (Standard: true)useDefaultBrowser(optional): Ob der Standardbrowser des Systems verwendet werden soll (Standard: false)visibleBrowser(optional): Ob das Browserfenster angezeigt werden soll (Standard: false)
4. Authentifizierungscookies löschen
Löscht gespeicherte Authentifizierungscookies für eine bestimmte Domäne oder alle Domänen.
{
"url": "https://example.com"
}url(optional): URL der Domäne, deren Cookies gelöscht werden sollen. Falls nicht angegeben, werden alle Cookies gelöscht.
Standardbrowsermodus
Im Standardbrowsermodus können Sie den regulären Browser Ihres Systems (Chrome, Edge usw.) anstelle des mitgelieferten Chromium von Puppeteer verwenden. Dies ist nützlich für:
Verwenden Ihrer vorhandenen Browsersitzungen und Erweiterungen
Manuelles Anmelden bei Websites mit Ihren gespeicherten Anmeldeinformationen
Ein natürlicheres Browsing-Erlebnis für mehrstufige Workflows
Testen Sie mit derselben Browserumgebung wie Ihre Benutzer
Um den Standardbrowsermodus zu aktivieren, legen Sie in Ihren Toolparametern useDefaultBrowser: true und visibleBrowser: true fest.
So funktioniert der Standardbrowsermodus
Wenn Sie den Standardbrowsermodus aktivieren:
Das Tool versucht, den Standardbrowser Ihres Systems (Chrome, Edge usw.) zu finden.
Es startet Ihren Browser mit aktiviertem Remote-Debugging auf einem zufälligen Port
Puppeteer verbindet sich mit dieser Browserinstanz, anstatt eine eigene zu starten
Ihre bestehenden Profile, Erweiterungen und Cookies sind während der Sitzung verfügbar
Das Browserfenster bleibt sichtbar, sodass Sie manuell damit interagieren können
Dieser Modus ist besonders nützlich für Workflows, die eine Authentifizierung oder komplexe Benutzerinteraktionen erfordern.
Browserpersistenz
Der MCP-Server kann eine dauerhafte Browsersitzung über mehrere Tool-Aufrufe hinweg aufrechterhalten:
Wenn Sie
login-and-waitverwenden, bleibt die Browsersitzung geöffnetNachfolgende Aufrufe von
screenshot-pageoderscreenshot-elementmitreuseAuthPage: trueverwenden dieselbe SeiteDies ermöglicht mehrstufige Workflows ohne erneute Authentifizierung
Cookie-Verwaltung
Für jede von Ihnen besuchte Domain werden automatisch Cookies gespeichert:
Nach der Verwendung
login-and-waitwerden Cookies im Verzeichnis.mcp-screenshot-cookiesin Ihrem Home-Ordner gespeichertDiese Cookies werden beim erneuten Besuch derselben Domain mit
useSavedAuth: trueautomatisch geladen.Sie können Cookies mit dem Tool
clear-auth-cookieslöschen.
Beispiel-Workflow: Screenshots geschützter Seiten
Hier ist ein Beispiel-Workflow zum Erstellen von Screenshots von Seiten, die eine Authentifizierung erfordern:
Manuelle Anmeldephase
{
"name": "login-and-wait",
"parameters": {
"url": "https://example.com/login",
"waitMinutes": 3,
"successIndicator": ".dashboard-welcome",
"useDefaultBrowser": true
}
}Dadurch wird Ihr Standardbrowser mit der Anmeldeseite geöffnet. Sie können sich manuell anmelden. Sobald die Anmeldung abgeschlossen ist (entweder durch Erkennen der Erfolgsanzeige oder nach Verlassen der Anmeldeseite), werden die Sitzungscookies gespeichert.
Screenshots mit gespeicherter Sitzung erstellen
{
"name": "screenshot-page",
"parameters": {
"url": "https://example.com/account",
"fullPage": true,
"useSavedAuth": true,
"reuseAuthPage": true,
"useDefaultBrowser": true,
"visibleBrowser": true
}
}Dadurch wird unter Verwendung Ihrer gespeicherten Authentifizierungs-Cookies im selben Browserfenster ein Screenshot der Kontoseite erstellt.
Machen Sie Screenshots von bestimmten Elementen
{
"name": "screenshot-element",
"parameters": {
"url": "https://example.com/dashboard",
"selector": ".user-profile-section",
"useSavedAuth": true,
"useDefaultBrowser": true,
"visibleBrowser": true
}
}Cookies nach Abschluss löschen
{
"name": "clear-auth-cookies",
"parameters": {
"url": "https://example.com"
}
}Dieser Workflow ermöglicht Ihnen die Interaktion mit geschützten Seiten, als wären Sie ein normaler Benutzer, und führt den vollständigen Authentifizierungsablauf in Ihrem Standardbrowser durch.
Headless-Modus vs. sichtbarer Modus
Headless-Modus (
visibleBrowser: false): Schneller und besser geeignet für automatisierte Arbeitsabläufe, bei denen keine Benutzerinteraktion erforderlich ist.Sichtbarer Modus (
visibleBrowser: true): Zeigt das Browserfenster an und ermöglicht Benutzerinteraktion und manuelle Überprüfung. Erforderlich füruseDefaultBrowser: true.
Plattformunterstützung
Die Standardbrowsererkennung funktioniert auf:
macOS : Erkennt Chrome, Edge und Safari
Windows : Erkennt Chrome und Edge über die Registrierung oder allgemeine Installationspfade
Linux : Erkennt Chrome und Chromium über Systembefehle
Fehlerbehebung
Häufige Probleme
Standardbrowser nicht gefunden : Wenn das System Ihren Standardbrowser nicht finden kann, wird auf das mitgelieferte Chromium von Puppeteer zurückgegriffen.
Verbindungsprobleme : Wenn beim Herstellen einer Verbindung zum Debug-Port des Browsers Probleme auftreten, prüfen Sie, ob dieser Port bereits von einer anderen Instanz verwendet wird.
Cookie-Probleme : Wenn die Authentifizierung nicht funktioniert, versuchen Sie, Cookies mit dem Tool
clear-auth-cookieszu löschen.
Debuggen
Der MCP-Server protokolliert hilfreiche Fehlermeldungen in der Konsole, wenn Probleme auftreten. Überprüfen Sie diese Meldungen auf Informationen zur Fehlerbehebung.
Available Tools
5 toolsclear-auth-cookiesA
Clears saved authentication cookies for a specific domain or all domains
| Name | Required | Description | Default |
|---|---|---|---|
| url | No | URL of the domain to clear cookies for. If not provided, clears all cookies. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It states the action ('clears saved authentication cookies') but does not disclose behavioral traits such as whether this requires specific permissions, if it's reversible, potential side effects (e.g., logging out users), or rate limits. The description is minimal and lacks critical context for a mutation tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence with zero waste, front-loading the core action and scope. It is appropriately sized for a simple tool with one optional parameter.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (simple mutation with one parameter) and lack of annotations or output schema, the description is adequate but has clear gaps. It covers the basic purpose and parameter semantics via the schema, but fails to provide behavioral context needed for safe usage, such as permissions or side effects.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, with the parameter 'url' documented as 'URL of the domain to clear cookies for. If not provided, clears all cookies.' The description adds no additional meaning beyond this, as it only restates the same information. Baseline 3 is appropriate when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose with a specific verb ('clears') and resource ('saved authentication cookies'), and distinguishes its scope ('for a specific domain or all domains'). It directly answers what the tool does without being vague or tautological.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage context by specifying 'for a specific domain or all domains,' but does not explicitly state when to use this tool versus alternatives or provide exclusions. Given the sibling tools (e.g., 'login-and-wait'), it lacks guidance on when to clear cookies relative to login/logout workflows.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
login-and-waitA
Opens a webpage in a visible browser window for manual login, waits for user to complete login, then saves cookies
| Name | Required | Description | Default |
|---|---|---|---|
| successIndicator | No | Optional CSS selector or URL pattern that indicates successful login | |
| url | Yes | The URL of the login page | |
| useDefaultBrowser | No | Whether to use the system's default browser instead of Puppeteer's bundled Chromium | |
| waitMinutes | No | Maximum minutes to wait for login (default: 3) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It describes the tool's behavior well: opening a visible browser, waiting for manual login, and saving cookies. However, it misses details like error handling, what happens after timeout, or how cookies are saved/stored. It does not contradict annotations, as none exist.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence that front-loads the core purpose and steps. Every word earns its place, with no redundancy or unnecessary details, making it highly concise and well-structured.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no annotations and no output schema, the description adequately covers the tool's purpose and high-level behavior. However, for a tool with 4 parameters and no output schema, it lacks details on return values, error cases, or integration with sibling tools like 'signal-login-complete'. It's complete enough for basic understanding but has gaps for full contextual use.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema fully documents all parameters. The description does not add any parameter-specific information beyond what the schema provides (e.g., it doesn't explain 'successIndicator' usage or 'waitMinutes' implications). Baseline 3 is appropriate as the schema handles the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the specific action sequence: 'Opens a webpage in a visible browser window for manual login, waits for user to complete login, then saves cookies.' It uses precise verbs (opens, waits, saves) and identifies the resource (webpage, cookies), distinguishing it from sibling tools like screenshot tools or cookie-clearing tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for manual login scenarios where user interaction is required, but it does not explicitly state when to use this tool versus alternatives like automated login tools or other authentication methods. It provides clear context (manual login in a browser) but lacks explicit exclusions or named alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
screenshot-elementB
Captures a screenshot of a specific element on a webpage using a CSS selector
| Name | Required | Description | Default |
|---|---|---|---|
| format | No | Image format for the screenshot | png |
| padding | No | Padding around the element in pixels | |
| quality | No | Quality of the image (0-100), only applicable for jpeg and webp | |
| selector | Yes | CSS selector for the element to screenshot | |
| url | Yes | The URL of the webpage | |
| useDefaultBrowser | No | Whether to use the system's default browser instead of Puppeteer's bundled Chromium | |
| useSavedAuth | No | Whether to use saved cookies from previous login | |
| visibleBrowser | No | Whether to show the browser window (non-headless mode) | |
| waitForSelector | No | Whether to wait for the selector to appear |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries full burden but only states the basic function. It doesn't disclose important behavioral traits like: whether this navigates to new URLs, requires page loading, handles authentication, has rate limits, or what happens with invalid selectors. The description is minimal beyond the core action.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Single sentence, zero waste words, front-loaded with the core action. Every word earns its place by specifying element-level capture with CSS selector mechanism.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 9-parameter tool with no annotations and no output schema, the description is inadequate. It doesn't explain what the tool returns (image data? file path? error formats?), doesn't mention authentication dependencies despite sibling login tools, and provides minimal behavioral context for a complex screenshot operation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema fully documents all 9 parameters. The description adds no parameter-specific information beyond implying 'selector' and 'url' are involved. Baseline 3 is appropriate when schema does all parameter documentation work.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the specific action ('Captures a screenshot') and target resource ('specific element on a webpage'), using precise terminology ('CSS selector'). It distinguishes from sibling 'screenshot-page' by specifying element-level rather than page-level capture.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No explicit guidance on when to use this tool versus alternatives like 'screenshot-page' or other siblings. The description implies usage for element-specific screenshots but doesn't provide context about prerequisites (e.g., needing authentication via login tools) or exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
screenshot-pageA
Captures a screenshot of a given URL and returns it as base64 encoded image. Can use saved cookies from login-and-wait.
| Name | Required | Description | Default |
|---|---|---|---|
| delay | No | Additional delay in milliseconds to wait after page load | |
| format | No | Image format for the screenshot | png |
| fullPage | No | Whether to capture the full page or just the viewport | |
| height | No | Viewport height in pixels | |
| quality | No | Quality of the image (0-100), only applicable for jpeg and webp | |
| reuseAuthPage | No | Whether to use the existing authenticated page instead of creating a new one | |
| url | Yes | The URL of the webpage to screenshot | |
| useDefaultBrowser | No | Whether to use the system's default browser instead of Puppeteer's bundled Chromium | |
| useSavedAuth | No | Whether to use saved cookies from previous login | |
| visibleBrowser | No | Whether to show the browser window (non-headless mode) | |
| waitFor | No | When to consider the page loaded | networkidle2 |
| width | No | Viewport width in pixels |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It mentions the ability to use saved cookies, which hints at authentication behavior, but doesn't cover other important traits like performance implications (e.g., page load delays), potential failures (e.g., invalid URLs), or side effects (e.g., browser resource usage). The description adds some value but leaves significant gaps for a tool with 12 parameters.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, well-structured sentence that efficiently conveys the core functionality and a key feature (cookie reuse). Every word earns its place with no redundancy or fluff, making it appropriately sized and front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (12 parameters, no annotations, no output schema), the description is incomplete. It covers the basic purpose and authentication context but lacks details on behavioral traits, error handling, or output specifics (beyond base64 encoding). For a screenshot tool with many configuration options, more guidance on usage scenarios or limitations would be helpful.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all 12 parameters thoroughly. The description adds minimal semantic context by mentioning 'saved cookies from login-and-wait,' which loosely relates to the 'useSavedAuth' parameter, but doesn't provide additional meaning beyond what the schema specifies for most parameters. Baseline 3 is appropriate when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the specific action ('captures a screenshot') and resource ('of a given URL'), and distinguishes from sibling tools by mentioning the ability to use saved cookies from 'login-and-wait' (differentiating from 'screenshot-element' which targets specific elements). It also specifies the output format ('returns it as base64 encoded image').
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context by mentioning saved cookies from 'login-and-wait', which implies when to use this tool (for authenticated pages). However, it doesn't explicitly state when NOT to use it or name alternatives like 'screenshot-element' for element-specific captures, leaving some guidance implicit rather than explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
signal-login-completeA
Signals that manual login is complete and the login-and-wait tool should continue
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It describes the behavioral trait of signaling completion to another tool, which is useful context. However, it doesn't disclose other aspects like whether it requires specific permissions, has side effects, or how it interacts with authentication states, leaving some gaps.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence that front-loads the key information: signaling login completion. There is zero waste, and it earns its place by clearly stating the tool's role in the workflow.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's simplicity (0 parameters, no output schema, no annotations), the description is complete enough. It explains the purpose and usage in context with sibling tools. However, it could be slightly more complete by mentioning any prerequisites or effects, but for a signaling tool, this is adequate.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 0 parameters with 100% coverage, so no parameter documentation is needed. The description doesn't add param info, but this is acceptable given the lack of parameters. Baseline is 4 for 0 params, as it doesn't need to compensate for any gaps.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: to signal completion of manual login so another tool (login-and-wait) can continue. It specifies the verb 'signals' and the context 'manual login is complete,' but doesn't explicitly differentiate from all sibling tools like clear-auth-cookies or screenshot tools, which serve different purposes.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly states when to use this tool: 'when manual login is complete' and that it should be used to allow 'login-and-wait tool should continue.' It names the specific alternative tool (login-and-wait) and implies usage in a sequence, providing clear context without exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
5 tool updates
v1.0.0- First observed
clear-auth-cookies - First observed
login-and-wait - First observed
screenshot-element - First observed
screenshot-page - First observed
signal-login-complete
TDQS
Scored across 5 tools
Most tools have distinct purposes: screenshot-element and screenshot-page target different screenshot scopes, while login-and-wait and clear-auth-cookies handle authentication. However, signal-login-complete is tightly coupled with login-and-wait, which could cause confusion about whether to use it separately or as part of the login flow.
The naming is mixed: screenshot-element and screenshot-page follow a verb-noun pattern, but clear-auth-cookies and login-and-wait use hyphens and compound phrases, while signal-login-complete is a full sentence. This inconsistency makes the set less predictable, though the names remain readable.
With 5 tools, the count is well-scoped for a webpage screenshot server. Each tool serves a clear role in the workflow (authentication, screenshot capture, and cleanup), and there are no extraneous tools, making it efficient for agents to navigate.
The toolset covers core screenshot and authentication workflows effectively, including login, cookie management, and element/page capture. A minor gap is the lack of tools for advanced screenshot options (e.g., full-page capture or viewport adjustments), but agents can still accomplish the main tasks without significant workarounds.
Maintenance
Related MCP Connectors
Screenshot any public web page from an AI agent. Free without signup, or with an API key.
Generate images, GIFs, videos, and PDFs from HTML, URLs, or templates — from your AI agent.
- mcpOAuthcom.screenshotink
Screenshot, diff, audit and sitemap-capture any web page — 5 MCP tools for AI agents.
Captured web interfaces, screenshots, flows and structured evidence for AI coding agents.
1
Related MCP Servers
- AlicenseCqualityDmaintenanceEnables LLMs to perform web browsing tasks, take screenshots, and execute JavaScript using Puppeteer for browser automation.435,614 npm1MIT
- AlicenseAqualityDmaintenanceEnables AI assistants to capture screenshots of web pages using automated browser sessions. Supports full-page and element-specific screenshots, device simulation, and JavaScript execution for comprehensive web testing and monitoring.618 npmMIT
- AlicenseCqualityDmaintenanceEnables taking screenshots of web pages with support for multiple devices (desktop, mobile, tablet), custom dimensions, full-page capture, and various image formats. Built with Playwright for reliable web page rendering and screenshot generation.1MIT

site-shot-mcpofficial
AlicenseBqualityBmaintenanceEnables AI agents to capture full-page or viewport screenshots of any web page with options for ad removal, cookie banner blocking, and proxy country selection.273 npm3MIT