Skip to main content
Glama
codewithkate

Robots Compliance Scrape Workflow MCP Server

by codewithkate
README.md
# Robots Compliance Scrape Workflow — MCP Server

An MCP server that brackets AI-agent web scraping with documented robots.txt
compliance, an explicit user-confirmed plan, and a single provenance-wrapped
data object ready for database loading.

## ⚠️ What this does and does not certify

`check_robots_compliance` produces **verifiable technical evidence** that a
site's robots.txt permits a given user-agent on given paths: the verbatim
applicable directive lines, the file's SHA-256 hash, HTTP status, and a UTC
fetch timestamp. It does **not** certify legality. robots.txt is a voluntary
convention (RFC 9309), not a contract or a license. Terms of Service,
copyright, and data-protection law (GDPR/CCPA) apply independently — every
tool response repeats this disclaimer so it travels with the data.

## Architecture

The server does not drive the browser. The AI agent (e.g., Claude with the
Claude-in-Chrome MCP tools) performs the dynamic navigation *between* the
server's gates:

```
User request
   │
   ▼
[1] check_robots_compliance ──► verdict + evidence (check_id)
   │            └── denied? STOP.
   ▼
Agent confirms with the user in chat:
   navigation steps · scrape scope · exclusions ·
   max pages · rate limit · output format · output schema
   │
   ▼
[2] create_scrape_plan(check_id, ..., user_confirmed=True) ──► plan_id
   │   (rejects unconfirmed plans, denied verdicts, unknown check_ids;
   │    enforces robots.txt Crawl-delay as the rate-limit floor)
   ▼
Agent executes the plan with browser tools (Claude in Chrome):
   navigate → wait for JS render → extract fields → paginate
   │
   ▼
[3] package_results(plan_id, records) ──► ONE data object:
       { data, schema, provenance{robots evidence, plan, pages},
         validation_issues, warnings }
```

## Install & run

```bash
pip install mcp httpx
python server.py            # stdio transport
```

### Claude Desktop / Claude Code config

```json
{
  "mcpServers": {
    "robots-compliance-scraper": {
      "command": "python",
      "args": ["/absolute/path/to/server.py"]
    }
  }
}
```

## Tools

### `check_robots_compliance(site_url, user_agent="ClaudeBot", target_paths=["/"])`
Fetches `/robots.txt`, evaluates each target path per RFC 9309, and returns:
`verdict` (`allowed` | `partially_allowed` | `denied` | `allowed_no_robots` |
`indeterminate`), per-path results, the verbatim applicable user-agent group,
SHA-256 of the file, any `Crawl-delay`, and the legal disclaimer.

### `create_scrape_plan(check_id, navigation_steps, scope_description, output_format, output_schema, max_pages, min_delay_seconds, user_confirmed, exclusions)`
Gated: requires a passing `check_id` and `user_confirmed=True`. The agent must
present the full plan to the user in chat and get explicit approval *before*
setting that flag. `output_format` is chosen per run: `json_object`, `jsonl`,
or `sql_rows`. The effective delay is `max(min_delay_seconds, Crawl-delay)`.

### `package_results(plan_id, records, pages_visited, warnings)`
Validates records against the plan's schema (reports missing/unexpected
fields per record), renders the chosen format, and returns a single
`scrape_result_v1` object embedding the compliance evidence and plan as
provenance — load it into your database and transform downstream.

## Example data object (abridged)

```json
{
  "object_type": "scrape_result_v1",
  "record_count": 2,
  "output_format": "sql_rows",
  "data": { "records": [...], "columns": ["price", "title"], "rows": [[1.5, "A"], ...] },
  "provenance": {
    "site_url": "https://example.com",
    "pages_visited": ["https://example.com/js/"],
    "scrape_plan": { "plan_id": "plan_ab12...", "navigation_steps": [...], ... },
    "robots_compliance": {
      "verdict": "allowed",
      "evidence": {
        "applicable_group_verbatim": ["User-agent: *", "Allow: /"],
        "robots_txt_sha256": "05a5baaa...",
        "crawl_delay_seconds": 5.0
      }
    }
  },
  "validation_issues": [],
  "legal_disclaimer": "robots.txt compliance is a technical/etiquette signal, not legal consent. ..."
}
```

## Agent etiquette baked into every plan

- Honor `Crawl-delay`; default floor 2s between page loads
- Hard page cap per run (`max_pages`, default 25)
- No logins, paywall bypass, or CAPTCHA solving
- Scope + exclusions confirmed by the user before any navigation