Skip to main content
Glama
0pen1
by 0pen1
README.md
# scrapingdog-mcp

A [Model Context Protocol](https://modelcontextprotocol.io) server that exposes the
[Scrapingdog](https://www.scrapingdog.com/documentation/) scraping APIs as tools for
MCP-compatible clients (Claude Desktop, Claude Code, and others).

Communicates over **stdio**. Built with TypeScript and the official MCP SDK.

## What it does

Nine tools, each wrapping one Scrapingdog endpoint:

| Tool | Endpoint | What it returns |
|------|----------|-----------------|
| `scrape` | `/scrape` | Raw HTML of any URL (JS rendering, premium proxies, geotargeting, sessions, stealth, wait) |
| `screenshot` | `/screenshot` | Page screenshot metadata (png/jpg/webp, full-page, viewport, quality) |
| `google_search` | `/google` | Google organic results, ads, knowledge graph, SERP features (JSON or HTML) |
| `bing_search` | `/bing/search` | Bing search results (market/geo/pagination/safe-search) |
| `duckduckgo_search` | `/duckduckgo/search` | DuckDuckGo results (region, date filter, pagination token) |
| `baidu_search` | `/baidu/search` | Baidu results (Chinese-language restriction, pagination) |
| `x_profile` | `/x/profile` | X (Twitter) profile data (name, handle, follower counts, bio) |
| `x_post` | `/x/post` | X (Twitter) post (tweet) data |
| `datacenter_proxy` | *(forward proxy)* | Connection details + ready-to-paste curl/Python for `proxy.scrapingdog.com:8081` |

Every tool accepts an optional `api_key` argument to override the configured key for that call.

## Prerequisites

- Node.js **≥ 18.17** (tested on 20.x).
- A Scrapingdog API key — get one from your [Scrapingdog dashboard](https://www.scrapingdog.com/) after signing up.

## Install & build

```bash
npm install
npm run build      # outputs to dist/
```

## Provide your API key

The key is resolved in this order (first wins):

1. The `api_key` argument on an individual tool call.
2. The `SCRAPINGDOG_API_KEY` environment variable.
3. A `.env` file containing `SCRAPINGDOG_API_KEY=...` in the working directory,
   `~/.scrapingdog.env`, or `~/.env`.

Option **2** (env var) is recommended — the key never touches disk beyond your MCP
client config, and nothing ends up in tool-call logs.

## Configure your MCP client

### Claude Desktop / Claude Code (`claude_desktop_config.json`)

```jsonc
{
  "mcpServers": {
    "scrapingdog": {
      "command": "node",
      "args": ["/absolute/path/to/scrcpy/dist/index.js"],
      "env": {
        "SCRAPINGDOG_API_KEY": "your-key-here"
      }
    }
  }
}
```

### Install from npm (no clone needed)

Published as [`@0pen1/scrapingdog-mcp`](https://www.npmjs.com/package/@0pen1/scrapingdog-mcp). Run directly with `npx` (no install):

```bash
npx @0pen1/scrapingdog-mcp
```

The server reads `SCRAPINGDOG_API_KEY` from the environment. Pick your CLI below for a one-line setup.

### Claude Code

```bash
claude mcp add scrapingdog --scope user \
  --env SCRAPINGDOG_API_KEY=your-key-here \
  -- npx -y @0pen1/scrapingdog-mcp
```

`--scope user` makes it available in all your projects (use `local` or `project` to scope it tighter).

### OpenAI Codex CLI

```bash
codex mcp add scrapingdog \
  --env SCRAPINGDOG_API_KEY=your-key-here \
  -- npx -y @0pen1/scrapingdog-mcp
```

This writes to `~/.codex/config.toml`. The equivalent TOML block is:

```toml
[mcp_servers.scrapingdog]
command = "npx"
args = ["-y", "@0pen1/scrapingdog-mcp"]

[mcp_servers.scrapingdog.env]
SCRAPINGDOG_API_KEY = "your-key-here"
```

### Gemini CLI

```bash
gemini mcp add scrapingdog \
  --env SCRAPINGDOG_API_KEY=your-key-here \
  -- npx -y @0pen1/scrapingdog-mcp
```

### Cursor / Windsurf / Claude Desktop (JSON config)

For clients that use a JSON config file, add this block. Typical locations:

- **Claude Desktop** — `~/Library/Application Support/Claude/claude_desktop_config.json` (macOS)
- **Cursor** — `~/.cursor/mcp.json` (or `.cursor/mcp.json` per-project)
- **Windsurf** — `~/.codeium/windsurf/mcp_config.json`

```jsonc
{
  "mcpServers": {
    "scrapingdog": {
      "command": "npx",
      "args": ["-y", "@0pen1/scrapingdog-mcp"],
      "env": { "SCRAPINGDOG_API_KEY": "your-key-here" }
    }
  }
}
```

> Replace `your-key-here` with your Scrapingdog API key from the
> [dashboard](https://www.scrapingdog.com/). After editing a JSON config, restart the client so it picks up the server.

## Run standalone (for testing)

```bash
SCRAPINGDOG_API_KEY=your-key node dist/index.js
```

It speaks JSON-RPC over stdio, so you can pipe messages to it directly.

## Example tool calls

```
scrape          { "url": "https://example.com", "dynamic": true, "country": "us" }
google_search   { "query": "best espresso machines 2026", "results": "10", "country": "us" }
bing_search     { "query": "site:github.com mcp server", "count": "20" }
duckduckgo_search { "query": "rust async runtime", "df": "m" }
baidu_search    { "query": "人工智能", "ct": 2 }
x_profile       { "profileId": "elonmusk" }
x_post          { "tweetId": "1655608985058267139" }
screenshot      { "url": "https://example.com", "fullPage": true, "format": "png" }
datacenter_proxy { "target_url": "https://httpbin.org/ip" }
```

## How responses are handled

- **JSON endpoints** (search, social, etc.): the parsed JSON is returned as text so the
  host can reason about it directly.
- **`/scrape`**: raw HTML is returned as text.
- **`/screenshot`**: the endpoint returns binary image bytes, which can't be carried as
  text cleanly. The tool reports the content type and byte length, and notes how to fetch
  the image directly. (If you need true image delivery, extend the handler to base64-encode
  into an `image` content block.)
- **Non-2xx responses**: returned as MCP error results (`isError: true`) with the upstream
  status, content type, and body. Note Scrapingdog reports invalid keys as HTTP 400 with a
  JSON "Unauthorized request" message — the body makes this clear.

## Credit costs (per request)

Scrape: 1 (rotating proxy) → 25 (JS render + premium). Google/Bing/DDG/Baidu search: ~5.
Screenshot: 5. Profile/Post scrapers: 5–10. See the [pricing/credit table](https://www.scrapingdog.com/documentation/)
for the full breakdown. Requests time out after 60s upstream.

## Project layout

```
src/
  index.ts          # entry point — creates McpServer, connects stdio transport
  scrapingdog.ts    # API key resolution + shared HTTP request helper
  result.ts         # response → MCP tool-result formatting (ok / error / fromResponse)
  tools.ts          # the 9 tool definitions (schema + handler) and registration
```

## License

MIT

TDQS

A3.9/5.0

Scored across 9 tools

Disambiguation5/5

Each tool targets a distinct resource or action: raw HTML scraping, screenshots, four search engines, X profile/post, and proxy details. The search tools are differentiated by the named engine, and the X tools by profile vs post, leaving no ambiguous overlaps.

Naming Consistency3/5

Tool names use a mix of verb forms and noun compounds, such as scrape, screenshot, google_search, x_profile, and datacenter_proxy. While the snake_case style is consistent, there is no uniform verb_noun pattern, making the naming somewhat mixed.

Tool Count5/5

With 9 tools covering scraping, screenshots, search, social media, and proxy access, the count is well within the ideal 3-15 range and each tool earns its place.

Completeness4/5

The set covers core Scrapingdog features (scrape, screenshot, search engines, X data, proxy) with only minor gaps such as missing X search or timeline and other less-used Scrapingdog endpoints. Agents can work around these gaps by using the scrape tool.

Maintenance

ActivityStale
ResponsivenessNo issues