spa-reader-mcp
# spa-reader-mcp
[](https://github.com/XXO47OXX/spa-reader-mcp/actions/workflows/ci.yml)
[](https://www.npmjs.com/package/spa-reader-mcp)
MCP server that renders JavaScript SPA pages and extracts Markdown via headless Chromium.
Traditional scrapers fail on SPAs because content is rendered client-side. This tool launches Playwright, waits for JS to finish, then extracts clean Markdown using Readability + Turndown.
## Install
```bash
npx playwright install chromium
```
### Claude Desktop
```json
{
"mcpServers": {
"spa-reader": {
"command": "npx",
"args": ["-y", "spa-reader-mcp"]
}
}
}
```
### Claude Code
```bash
claude mcp add spa-reader -- npx -y spa-reader-mcp
```
## Tools
### `spa_read`
Render a page and extract content as Markdown.
| Parameter | Type | Default | Description |
|-----------|------|---------|-------------|
| `url` | string | — | URL to read (required) |
| `waitForSelector` | string | — | CSS selector to wait for |
| `waitTimeout` | number | 30000 | Timeout in ms |
| `includeMetadata` | boolean | true | Add YAML frontmatter |
| `cookies` | array | — | Cookies for auth |
| `headers` | object | — | Custom HTTP headers |
### `spa_screenshot`
Capture a PNG screenshot after JS rendering.
| Parameter | Type | Default | Description |
|-----------|------|---------|-------------|
| `url` | string | — | URL to capture (required) |
| `waitForSelector` | string | — | CSS selector to wait for |
| `waitTimeout` | number | 30000 | Timeout in ms |
| `width` | number | 1280 | Viewport width |
| `height` | number | 720 | Viewport height |
| `fullPage` | boolean | false | Full page capture |
| `cookies` | array | — | Cookies for auth |
| `headers` | object | — | Custom HTTP headers |
## Security
- SSRF protection: blocks private/loopback IPs
- Only `http:` and `https:` schemes allowed
- Selector injection prevention
- Content capped at 100KB
## Dev
```bash
pnpm install && pnpm build
pnpm test
```
## License
MIT
TDQS
Scored across 2 tools
The two tools have clearly distinct purposes: spa_read extracts textual content as Markdown for LLM processing, while spa_screenshot captures visual output as PNG for screenshots. There is no overlap in functionality or ambiguity about which tool to use for a given task.
Both tools follow a consistent 'spa_' prefix pattern with descriptive suffixes (read, screenshot), indicating they belong to the same domain and operate on SPA pages. The naming is uniform, predictable, and clearly communicates each tool's function.
With only 2 tools, the server feels thin for a general-purpose SPA reader domain, as it lacks operations like navigation, interaction simulation, or performance monitoring. However, it covers the core tasks of content extraction and screenshot capture adequately for basic use.
The tools provide essential read-only capabilities for SPAs (extracting content and screenshots), but there are notable gaps: no ability to interact with pages (e.g., click buttons, fill forms), navigate beyond initial URLs, or handle dynamic content beyond rendering. This limits advanced agent workflows.