Skip to main content
Glama
tymrtn

mcp-firecrawl-licensed

by tymrtn
README.md
# Firecrawl Licensed MCP (Copyright.sh)

Open-source MCP server that wraps Firecrawl and adds Copyright.sh licensing, usage logging, and optional x402 licensed fetch. Tool names match Firecrawl’s OSS MCP server (`firecrawl_*`) so clients can drop this in without prompt changes.

## Features

- License discovery via `ai-license` meta tags
- Usage logging for compensation and audit trails
- Optional x402 licensed fetch on `402 Payment Required`
- Token estimation with `tiktoken`
- Graceful degradation when the ledger is unavailable

## Quick Start

### Install from source

```bash
git clone https://github.com/tymrtn/mcp-firecrawl-licensed.git
cd mcp-firecrawl-licensed
npm install
npm run build
```

### Run via npx (if published)

```bash
npx @copyrightsh/firecrawl-licensed-mcp@latest
```

## Configuration

Copy `env.example` to `.env` and set:

- `FIRECRAWL_API_KEY` (required unless using `FIRECRAWL_API_URL`)
- `FIRECRAWL_API_URL` (optional for self-hosted Firecrawl)
- `COPYRIGHTSH_LEDGER_API_KEY` (recommended for license acquire + usage logging)

### MCP Config Example (npx)

```json
{
  "mcpServers": {
    "firecrawl-licensed": {
      "command": "npx",
      "args": ["-y", "@copyrightsh/firecrawl-licensed-mcp@latest"],
      "env": {
        "FIRECRAWL_API_KEY": "fc-your-key-here",
        "COPYRIGHTSH_LEDGER_API": "https://ledger.copyright.sh",
        "COPYRIGHTSH_LEDGER_API_KEY": "cs-ledger-your-key-here"
      }
    }
  }
}
```

### Environment Variables

| Variable | Required | Default | Description |
| --- | --- | --- | --- |
| `FIRECRAWL_API_KEY` | Yes* | - | Firecrawl API key |
| `FIRECRAWL_API_URL` | No | - | Self-hosted Firecrawl base URL |
| `COPYRIGHTSH_LEDGER_API` | No | `https://ledger.copyright.sh` | Ledger base URL |
| `COPYRIGHTSH_LEDGER_API_KEY` | Recommended | - | API key for license discovery + usage logging |
| `ENABLE_LICENSE_TRACKING` | No | `true` | Enable/disable licensing |
| `ENABLE_LICENSE_CACHE` | No | `false` | Cache license results |
| `LICENSE_CACHE_TTL_SECONDS` | No | `300` | Cache TTL for license lookups |
| `LICENSE_CHECK_TIMEOUT_MS` | No | `5000` | License discovery timeout (ms) |
| `LICENSE_ACQUIRE_TIMEOUT_MS` | No | `8000` | License acquisition timeout (ms) |
| `USAGE_LOG_TIMEOUT_MS` | No | `3000` | Usage logging timeout (ms) |
| `FETCH_TIMEOUT_MS` | No | `12000` | Direct fetch timeout (ms) |

*`FIRECRAWL_API_KEY` is required unless `FIRECRAWL_API_URL` is set.

## Tools

- `firecrawl_scrape`
- `firecrawl_map`
- `firecrawl_search`
- `firecrawl_crawl`
- `firecrawl_check_crawl_status`
- `firecrawl_extract`
- `firecrawl_agent`
- `firecrawl_agent_status`

### x402 Licensed Fetch

For tools that accept URLs (`scrape`, `search`, `check_crawl_status`), set:

- `fetch: true`
- `stage: infer|embed|tune|train`
- `distribution: private|public`
- `estimated_tokens: number`
- `payment_method: account_balance|x402`

When enabled, the server performs a direct fetch that handles `402 Payment Required` + `payment-required: x402` responses, acquires a license token from the ledger, and retries with the returned `licensed_url`.

### Example: Licensed Scrape

```json
{
  "name": "firecrawl_scrape",
  "arguments": {
    "url": "https://example.com/article",
    "options": { "formats": ["markdown"], "onlyMainContent": true },
    "fetch": true,
    "stage": "infer",
    "distribution": "private",
    "estimated_tokens": 1500,
    "payment_method": "account_balance"
  }
}
```

## CLI Helpers

```bash
node build/index.js --list-tools
node build/index.js --doctor
```

## Notes on Paywalled Sources

This server only unlocks sources that implement the Copyright.sh x402 flow (`402` + `payment-required: x402`). It does not bypass login/subscription paywalls.

## Contributing

See [CONTRIBUTING.md](CONTRIBUTING.md).

## License

MIT

TDQS

C2.9/5.0

Scored across 8 tools

Disambiguation5/5

Each tool has a clearly distinct purpose: agent management, crawl management, scraping, searching, mapping, extraction. No overlapping functionality.

Naming Consistency5/5

All tools follow the consistent pattern firecrawl_<action> (or firecrawl_<action>_<sub>), using snake_case throughout.

Tool Count5/5

8 tools is well-scoped for a web scraping/crawling API, covering agent, crawl, scrape, search, map, and extract operations without excess.

Completeness4/5

Covers all core workflows (agent start/status, crawl start/status, scrape, search, map, extract), though missing operations like stop/delete agent.

Maintenance

ActivityInactive
ResponsivenessNo issues