freecrawl
by RistrixOP
README.md
# π₯ Freecrawl
> Open-source web scraping API with anti-bot bypass. Drop-in Firecrawl alternative.
> **AI-first**: MCP server, OpenAI tool definitions, REST API β any AI model can use it.
> No API keys, no credits, no limits.
## Quick Start
```bash
pip install -r requirements.txt
patchright install chromium
python -m freecrawl serve # REST API on :9100
```
## AI Integration
### MCP Server (Claude, Cursor, any MCP client)
```json
{"mcpServers":{"freecrawl":{"command":"python","args":["-m","freecrawl.mcp_server"]}}}
```
### OpenAI Tool Definitions
```python
from freecrawl.ai_tools import get_tools
tools = get_tools() # 9 tool definitions, pass to any LLM
```
### REST API
```bash
curl -X POST http://localhost:9100/v1/scrape \
-H "Content-Type: application/json" \
-d '{"url":"https://example.com","formats":["markdown"]}'
```
## Features
| Feature | Status |
|---------|:------:|
| Scrape (URLβMarkdown/HTML/text) | β
|
| Anti-bot bypass (Cloudflare, DataDome) | β
|
| Auto-escalation engine chain | β
|
| Crawl (BFS + sitemap) | β
|
| Search (DuckDuckGo/SearXNG) | β
|
| Map (fast URL discovery) | β
|
| Batch scrape (parallel) | β
|
| LLM extraction (structured JSON) | β
|
| Change tracking (diff) | β
|
| PDF parsing | β
|
| CSS extraction | β
|
| Screenshots | β
|
| JS actions (click/scroll/type) | β
|
| Product/Menu/Branding extractors | β
|
| Summary (LLM) | β
|
| Proxy rotation | β
|
| MCP server | β
|
| OpenAI tool definitions | β
|
| Docker deployment | β
|
| CLI tool | β
|
## Architecture
```
Engine Chain (auto-escalation):
httpx (fast) β flaresolverr (Cloudflare) β patchright (DataDome/Akamai)
```
## License
MIT
This server cannot be deployed
Maintenance
ActivityStale
ResponsivenessNo issues