mcp-page-capture
Server Configuration
Describes the environment variables required to run the server.
| Name | Required | Description | Default |
|---|---|---|---|
No arguments | |||
Instructions
Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.
This server publishes no instructions, or was last inspected before Glama recorded them.
Capabilities
Features and capabilities supported by this server
Protocol revision2025-11-25
| Capability | Details |
|---|---|
| tools | {
"listChanged": true
} |
Tools
Functions exposed to the LLM to take actions
| Name | Description |
|---|---|
| captureScreenshotA | Capture webpage screenshot with optional pre-capture interactions. PARAMS: • url (required): Page URL • steps (optional): Array of 6 step types (order auto-fixed) • headers (optional): HTTP auth headers • validate (optional): Dry-run step validation 6 STEPS (all optional, order doesn't matter): • viewport: { device: "mobile" } - Set device/screen • wait: { for: ".loaded" } or { duration: 2000 } - Wait for element or time • fill: { target: "#email", value: "a@b.com" } - Fill form field • click: { target: "#btn", waitFor: ".result" } - Click element • scroll: { to: "#footer" } or { y: 500 } - Scroll page • screenshot: { fullPage: true } - Capture (auto-added if omitted) COMMON ERRORS & FIXES: • ELEMENT_NOT_FOUND → Add wait step before the failing step • ELEMENT_NOT_VISIBLE → Add scroll step before the failing step EXAMPLE: { "url": "...", "steps": [{ "type": "fill", "target": "#email", "value": "a@b.com" }, { "type": "click", "target": "#submit", "waitFor": ".dashboard" }] } |
| extractDomA | Extract HTML, text, and DOM structure from a webpage. USE WHEN: Need text content for analysis, DOM structure for selector discovery, or pre-capture page validation. USE captureScreenshot WHEN: Need visual verification or rendered UI. • url (required): Page URL • selector (optional): CSS selector to scope extraction (e.g., 'main', 'article') EXAMPLE: { "url": "https://example.com", "selector": "article" } |
Prompts
Interactive templates invoked by user choice
| Name | Description |
|---|---|
No prompts | |
Resources
Contextual data attached and managed by the client
| Name | Description |
|---|---|
No resources | |
TDQS
Scored across 2 tools
The two tools have clearly distinct purposes: one captures a rendered screenshot with optional interactions, the other extracts raw HTML/text/DOM. There is no overlap in functionality, making it easy for an agent to choose the correct tool.
Both tool names follow the same camelCase verb+noun pattern ('captureScreenshot', 'extractDom'), which is consistent and predictable. The naming clearly conveys the action and target.
With only two tools, the server is very focused and minimal. This is appropriate for a simple 'page capture' utility that handles both visual and textual extraction. It's slightly on the low side, but the scope is narrow enough that two tools feel sufficient.
The tool surface covers the primary needs for capturing webpage content: visual (screenshot) and textual (DOM extraction). A potential minor gap might be a dedicated tool for downloading assets or getting page metadata, but for most use cases, the two tools provide complete coverage.