Skip to main content
Glama
namanxajmera

Reddit Scraper

by namanxajmera
README.md
# mcp-reddit

MCP server for scraping Reddit - **no API keys required**.

Scrapes posts and comments from subreddits and user profiles using Reddit's public RSS/Atom feeds.

> [!IMPORTANT]
> **What changed in v0.4.0:** Reddit shut down unauthenticated `.json` access in May 2026,
> which broke this project (and every other keyless Reddit scraper). v0.4.0 pivots to
> Reddit's still-public `.rss` feeds. Still zero API keys - but RSS exposes less data:
> **no scores, upvote ratios, comment counts, NSFW flags, or comment threading.**
> Those fields were removed from the data model. Comments are flat, up to 100 per post.

## Features

- **No API keys** - Uses public RSS feeds, no Reddit API credentials needed
- **Posts & comments** - Titles, authors, timestamps, selftext, flat comment bodies
- **Image previews** - Downloads preview images (videos unavailable via RSS)
- **Local persistence** - Query scraped data offline
- **Filtering** - By post type and keywords

## Installation

```bash
pip install mcp-reddit
```

Or with uvx:

```bash
uvx mcp-reddit
```

## Usage Modes

### Local (stdio) - Default

For local MCP clients like Claude Desktop and Claude Code:

```bash
uvx mcp-reddit
```

### Remote (HTTP/SSE)

For remote MCP clients that connect via URL:

```bash
uvx mcp-reddit --http --port 8000
```

Options:
- `--http` - Run in HTTP/SSE mode instead of stdio
- `--host` - Host to bind to (default: 0.0.0.0)
- `--port` - Port to listen on (default: 8000, or `PORT` env var)

The server exposes:
- `GET /sse` - SSE endpoint for MCP connection
- `POST /messages/` - Message endpoint
- `GET /health` - Health check

## Configuration

Add to your Claude Desktop or Claude Code settings:

### Claude Desktop (`~/Library/Application Support/Claude/claude_desktop_config.json`)

Claude Desktop doesn't inherit your shell PATH, so you need the full path to `uvx`:

```bash
# Find your uvx path
which uvx
```

Then use the full path in your config:

```json
{
  "mcpServers": {
    "reddit": {
      "command": "/Users/YOUR_USERNAME/.local/bin/uvx",
      "args": ["mcp-reddit"]
    }
  }
}
```

Replace `/Users/YOUR_USERNAME/.local/bin/uvx` with the output from `which uvx`.

### Claude Code

```bash
claude mcp add reddit -- uvx mcp-reddit
```

Or manually in `~/.claude.json`:

```json
{
  "mcpServers": {
    "reddit": {
      "command": "uvx",
      "args": ["mcp-reddit"]
    }
  }
}
```

## Available Tools

| Tool                   | Description                                                  |
| ---------------------- | ------------------------------------------------------------ |
| `scrape_subreddit`     | Scrape newest posts from a subreddit                         |
| `scrape_user`          | Scrape posts from a user's profile                           |
| `scrape_post`          | Fetch a specific post by URL (supports image download)       |
| `get_posts`            | Query stored posts with filters                              |
| `get_comments`         | Query stored comments                                        |
| `search_reddit`        | Search across all scraped data                               |
| `list_scraped_sources` | List all scraped subreddits/users                            |

> `get_top_posts` was removed in v0.4.0 - RSS feeds expose no scores to rank by.
> Calling it returns a JSON error pointing to `get_posts` (newest first) or `search_reddit`.

## Example Usage

```
"Scrape the latest 50 posts from r/LocalLLaMA"

"Fetch this post and download the image: https://reddit.com/r/ClaudeAI/comments/abc123/title"

"Search my scraped data for posts about 'fine-tuning'"
```

## Data Storage

Data is stored in `~/.mcp-reddit/data/` by default.

Data scraped with pre-v0.4.0 versions is migrated to the new (narrower) schema
automatically on the next scrape of that subreddit/user - existing posts and
comments stay queryable, and the original file is preserved as `*.pre-rss.bak`.

Set `MCP_REDDIT_DATA_DIR` environment variable to customize:

```json
{
  "mcpServers": {
    "reddit": {
      "command": "/Users/YOUR_USERNAME/.local/bin/uvx",
      "args": ["mcp-reddit"],
      "env": {
        "MCP_REDDIT_DATA_DIR": "/path/to/your/data"
      }
    }
  }
}
```

## Credits

Built on top of [reddit-universal-scraper](https://github.com/ksanjeev284/reddit-universal-scraper)
by [@ksanjeev284](https://github.com/ksanjeev284) - a full-featured Reddit scraper with
analytics dashboard, REST API, and plugin system.

## License

MIT

TDQS

A3.7/5.0

Scored across 8 tools

Disambiguation5/5

Each tool targets a distinct action (scrape vs. retrieve) and resource (subreddit, user, post, comments, sources). The scrape_* tools clearly handle live fetching while get_* and search_* handle local data.

Naming Consistency5/5

Tool names follow a consistent verb_noun pattern: scrape_subreddit, scrape_user, scrape_post, get_posts, get_comments, get_top_posts, list_scraped_sources, and search_reddit. The naming scheme is predictable and uniform.

Tool Count5/5

8 tools is well-scoped for a Reddit scraper server, covering both ingestion (scraping) and retrieval (querying local data) without unnecessary bloat.

Completeness4/5

The tool set covers core scraping from subreddits, users, and individual posts, plus retrieval/search of stored data. Minor gaps like deleting scraped data or listing comments per post directly are absent, but the main workflows are complete.

Maintenance

ActivityInactive
ResponsivenessNo issues