Apify Public Data & Leads
mcp-name: io.github.jlucasmcrell/apify-scrapers
Apify Public Data Scrapers & Extractors
A curated collection of reliable, production-ready scrapers and public-data extractors hosted on the Apify Store.
Each actor is built with strict schema validation, deterministic field mapping, self-healing DOM selectors, and pay-per-event pricing starting at $0.0002 / start.
Quick Navigation
Available Extractors & Store Listings
Tool | Store Link | Key Output Fields | Best For |
Google Maps Business Leads | Name, phone, website, rating, reviews, address, coordinates, hours | B2B lead generation, local agency prospecting | |
Glassdoor Jobs & Salaries | Title, company, salary estimate, rating, location, job URL, posting date | Hiring intelligence, compensation benchmarking | |
Airbnb Vacation Rentals | Title, room type, nightly price, rating, reviews count, listing URL | Real estate research, market rate tracking | |
SEC EDGAR Corporate Filings | Ticker, CIK, form (10-K, 10-Q, 8-K), filing date, primary document URL | Financial diligence, equity research, compliance | |
USAspending Federal Awards | Recipient vendor, award amount, awarding agency, description, dates | Government contracting, procurement intel | |
LinkedIn Public Jobs | Job title, employer, location, direct apply URL, posting age | Recruitment, tech talent monitoring | |
Google Play App Reviews | Review text, star score, thumbs up, date, reviewer name | App store sentiment, competitor feedback | |
YouTube Video Search | Title, video URL, channel, views count, duration, publish date | Content tracking, creator outreach | |
Twitch Live Streams | Streamer username, title, viewer count, language, category | Esports analytics, live stream monitoring | |
US Contractor Licenses | Contractor name, license number, classification, status, state | Trades verification, subcontractor diligence | |
US Business Entity Registries | Legal entity name, filing number, jurisdiction, status | Legal due diligence, corporate registration checks |
Python Quickstart
1. Install dependencies
pip install apify-client pandas python-dotenv2. Export 50 Google Maps Leads to CSV
import os
from apify_client import ApifyClient
import pandas as pd
# Get your API token from https://console.apify.com/account/integrations
client = ApifyClient(os.getenv("APIFY_TOKEN"))
# Run the actor
run = client.actor("captainhandsome/google-maps-business-search").call(run_input={
"search_query": "commercial electricians",
"location": "Dallas, Texas",
"max_items": 50,
"include_details": True,
})
# Fetch dataset items and export to CSV
items = list(client.dataset(run["defaultDatasetId"]).iterate_items())
df = pd.DataFrame(items)
df.to_csv("dallas_electricians.csv", index=False)
print(f"Exported {len(df)} leads to dallas_electricians.csv")See examples/google_maps_leads_to_csv.py for the full script.
Node.js Quickstart
1. Install dependencies
npm install apify-client2. Query SEC EDGAR Filings
import { ApifyClient } from 'apify-client';
const client = new ApifyClient({ token: process.env.APIFY_TOKEN });
const run = await client.actor('captainhandsome/sec-edgar-filings-search').call({
companies: ['AAPL', 'NVDA', 'MSFT'],
forms: ['10-K'],
max_items: 15,
});
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach(filing => {
console.log(`[${filing.ticker}] ${filing.form} (${filing.filing_date}): ${filing.primary_document_url}`);
});See examples/sec_filings.js for the full script.
No-Code & Automation Workflows
If you automate via n8n, Make, Zapier, or Google Sheets, ready-to-import blueprints are included in workflows/:
Google Maps Leads to Google Sheets (n8n): Daily automated cron scrape piping HVAC/trade leads directly into Google Sheets with deduplication.
SEC EDGAR 10-K & 8-K Alerts to Slack (n8n): Hourly monitor alerting Slack or Discord when watchlisted public companies drop new filings.
Pre-Built Example Tasks (Zero Code)
If you prefer runnable web UI tasks without writing any code, each actor includes pre-configured tasks published on Apify Store:
Google Maps Leads
Glassdoor Jobs
Airbnb Rentals
YouTube & Google Play
Free Sample Datasets
Looking for clean data to benchmark, analyze, or train models? Verified sample bundles with metadata schemas are available in datasets/ and hosted publicly on Hugging Face Datasets:
Phoenix HVAC Contractor Leads:
datasets/phoenix_hvac_leads/| Hugging Face Hub (20 verified HVAC contractor profiles with ratings, addresses, and phone numbers).California Licensed Contractors:
datasets/california_solar_contractors/| Hugging Face Hub (Active C-46 and B licensed solar installers with state verification numbers).Austin Software Engineer Postings:
datasets/austin_software_jobs/| Hugging Face Hub (Normalized job listings with estimated posting dates and salary ranges).
AI Agent & MCP Integration
All actors in this repository conform to OpenAPI and JSON Schema standards, making them directly callable by AI agents via the Model Context Protocol (MCP):
Option A: Hosted Apify MCP Server (Claude Desktop / Cursor)
Add this to your claude_desktop_config.json or Cursor MCP settings:
{
\"mcpServers\": {
\"apify\": {
\"command\": \"npx\",
\"args\": [\"-y\", \"@apify/mcp-server\"],
\"env\": {
\"APIFY_TOKEN\": \"YOUR_APIFY_API_TOKEN\"
}
}
}
}Option B: Local Lightweight Python MCP Server
For local agent workflows without Node.js dependencies, a direct Python MCP server is included:
export APIFY_TOKEN=\"your_token_here\"
python mcp_server.pyInspect tools and capabilities via mcp.json.
Agent Prompts That Work Out-of-the-Box:
"Search Google Maps for 50 commercial roofers in Atlanta with phone numbers and websites."
"Retrieve Apple and Microsoft Form 10-K filings from SEC EDGAR for the last 2 years."
"Find the 30 newest reviews for Duolingo on Google Play and analyze negative feedback."
In-Depth Engineering Guides
Technical case studies and problem-solution writeups are located in articles/:
Bypassing Playwright Headless Pagination Hurdles on Airbnb: How to solve sticky overlay modal interruptions and viewport boundary clipping in large headless browser crawls.
Extracting & Normalizing Clean Job Posting Dates from Glassdoor: Overcoming relative timestamp drift ("24h", "3d", "30d+") with deterministic parsing and ISO-8601 boundary tracking.
Repository Structure
apify-scrapers/
README.md # Documentation and quickstart
LICENSE # MIT License
requirements.txt # Python client dependencies
package.json # Node.js dependencies
mcp.json # MCP tool registry specification
mcp_server.py # Native Python stdio MCP server
articles/ # In-depth engineering case studies
airbnb_playwright_pagination_guide.md
glassdoor_posting_dates_guide.md
reddit_community_responses.md # Reference technical answers for forums
datasets/ # Sample benchmark datasets
phoenix_hvac_leads/
california_solar_contractors/
austin_software_jobs/
workflows/ # No-code automation templates
n8n_google_maps_to_sheets.json
n8n_sec_edgar_to_slack.json
README.md
examples/ # Standalone developer scripts
google_maps_leads_to_csv.py
sec_edgar_filings_downloader.py
glassdoor_jobs_tracker.py
airbnb_market_scraper.py
usaspending_defense_awards.py
twitch_live_stream_monitor.py
google_maps_leads.js
sec_filings.jsAuthor & Support
Maintained by Joseph McRell.
Apify Store: https://apify.com/captainhandsome
GitHub: @jlucasmcrell
Hugging Face: @joeygambino
Issues & Requests: Please open an issue on this repository or submit a ticket on the respective Apify Actor Store page.
License
This project is licensed under the MIT License - see the LICENSE file for details.