substack_scraper
Extract per-post metadata from any Substack publication's public archive API: title, author, date, paywall status, word count, likes, comments, with optional full article text.
Instructions
Scrapes any Substack publication's own public JSON archive API — title, author, publish date, paywall status, word count, likes and comments per post, one row per post, with an RSS fallback and optional full article text. Billed to your own Apify account: ~$0.0005 per Post (Apify free-plan price, lower on paid plans).
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| freeOnly | No | Free posts only — Keep this on to skip posts marked paywalled/subscriber-only (isPaywalled) and only return posts anyone can read. | |
| maxPosts | No | Max posts per publication — Enter how many of the most recent posts to fetch per publication, e.g. 50. The Actor pages through the publication's own archive API in batches of 50 until this many posts are collected. | |
| includeBody | No | Include full article text — Keep this off for metadata only (fast). Turn it on to also fetch each post's own page and extract its full article text into bodyText — one extra request per post. | |
| publications | Yes | Publications — Enter one Substack publication per row: a subdomain handle (bariweiss), a full subdomain (bariweiss.substack.com), or any post/publication URL on a custom domain (https://www.thefp.com or https://www.thefp.com/p/some-post). Example: ["bariweiss"]. |