Skip to main content
Glama
jlehman
by jlehman
README.md
# scribe-me

Scrape web pages to clean Markdown via headless Chromium. Available as a standalone CLI tool and as a Claude Code MCP tool + skill.

## Setup

```bash
npm install
npx playwright install chromium
npm link
```

## CLI Usage

### Single URL

```
scribe-me -p <project> -u <url> [-c <class>]
```

### Batch (file of URLs)

```
scribe-me -p <project> -f <file> [-c <class>]
```

The file is plain text with one URL per line. Lines starting with `#` are ignored. The scraper runs 3 pages concurrently against a single shared Chromium instance.

| Flag | Description |
|------|-------------|
| `-p, --project` | Project name — output is organized under `scribe-me/<project>/` |
| `-u, --url` | Single URL to scrape |
| `-f, --file` | Path to a file with one URL per line (use instead of `-u`) |
| `-c, --class` | CSS class of the content container to extract |

Output files are written to `scribe-me/<project>/` with timestamped filenames:

```
scribe-me/freewheel/2026-02-27 03:53:23-Getting-Started-with-the-Buzz-API.md
```

### Container Detection

If `-c` is provided, the scraper targets that class. Otherwise it walks a fallback list of common content selectors (`main`, `article`, `[role="main"]`, `#content`, etc.) and picks the first with meaningful text content, falling back to `<body>`.

## MCP Server

The same scraping logic is exposed as MCP tools for use inside Claude Code:

- **`scrape-to-markdown`** — scrape a single URL
- **`scrape-batch-to-markdown`** — scrape an array of URLs (3 concurrent, shared browser)

Add to `~/.claude/settings.json`:

```json
{
  "mcpServers": {
    "scribe-me": {
      "command": "node",
      "args": ["/absolute/path/to/src/mcp-server.js"]
    }
  }
}
```

## Claude Code Skill

The `/scrape` skill chains the MCP tool with an AI cleanup pass. It scrapes the page, then has Claude clean up the resulting markdown — removing UI artifacts ("Suggest Edits" links, "Did this page help you?" prompts, empty heading anchors, ToC blocks), fixing malformed headings, and improving code block formatting.

```
/scrape joshlehman https://www.joshlehman.com -c content
```

If any arguments are missing, the skill prompts for them interactively.

### Skill Installation

Copy the skill prompt to your Claude Code commands directory:

```bash
cp commands/scrape.md ~/.claude/commands/scrape.md
```

## Architecture

```
bin/
  scribe-me.js          # CLI entry point (Commander)
src/
  scraper.js            # Core: Playwright + Turndown
  container-finder.js   # Heuristic content container detection
  file-writer.js        # Directory creation + timestamped file writing
  sanitize.js           # Filename sanitization
  mcp-server.js         # MCP server wrapping scraper
commands/
  scrape.md             # Claude Code /scrape skill prompt
```