web-crawler-mcp
README.md
# web-crawler-mcp
A minimal [crawl4ai](https://github.com/unclecode/crawl4ai) crawler exposed as an MCP server, deployable to Cloud Run.
## What it does
Exposes a single MCP tool, `crawl(url, max_length)`, which fetches a page with crawl4ai's headless-browser crawler and returns its content as markdown.
## Local development
```bash
pip install -r requirements.txt
playwright install --with-deps chromium
crawl4ai-setup
python src/server.py
```
The server listens on `PORT` (default `8080`) using the MCP Streamable HTTP transport, reachable at `http://localhost:8080/mcp`.
## Deployment
`.github/workflows/deploy.yml` builds the Docker image with Cloud Build and deploys it to Cloud Run on every push to `main` that touches `src/`, `Dockerfile`, or `requirements.txt`.
Required GitHub Actions secrets:
- `GCP_SA_KEY` — service account key JSON used to authenticate to Google Cloud.
- `GCP_PROJECT_ID` — the target GCP project ID.
The Cloud Run service (`web-crawler-mcp`, region `us-central1`) is deployed with `--allow-unauthenticated`, matching the reference deployment pattern. Restrict access with IAM invoker bindings or a load balancer in front if the endpoint needs to be private.
## Connecting an MCP client
Point an MCP client (Streamable HTTP transport) at:
```
https://<cloud-run-service-url>/mcp
```
This server cannot be deployed
Maintenance
ActivityMaintained
ResponsivenessSyncing