Skip to main content
Glama
paulet4a-commits

WebDataTools Developer, app & research data MCP server

hacker_news_scraper

Search Hacker News stories, comments, Show HN and Ask HN by keyword or fetch the current front page, returning points, authors and comment counts without an API key.

Instructions

Search Hacker News stories, comments, Show HN and Ask HN posts by keyword, or pull the current front page — via the free Algolia HN API, no API key. One row per story or comment with points, author, comment count and text. Billed to your own Apify account: ~$0.0005 per Story (Apify free-plan price, lower on paid plans).

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
authorNoAuthor username — Enter a Hacker News username, e.g. pg, to only return items posted by that user. Leave empty to include every author.
sortByNoSort by — relevance uses Algolia's default ranking (Hacker News' own /search), date returns the newest items first (/search_by_date), and points sorts the matched items by score, highest first. Ignored for front_page queries. Options: relevance = Relevance; date = Newest first; points = Highest points first.relevance
queriesYesSearch queries — Enter one search term per row, e.g. apify or rust. Leave a row empty, or type front_page, to fetch the current Hacker News front page instead of running a search. Example: ["front_page"].
minPointsNoMinimum points — Enter the minimum score an item must have to be included, e.g. 50. Set to 0 to disable this filter.
sinceDaysNoOnly items from the last N days — Enter how many days back to search, e.g. 7 for the last week. Set to 0 to disable this filter and search all time.
maxResultsNoMax results per query — Enter how many rows to return for each entry in queries, e.g. 50. Algolia serves at most 1000 hits per query.
contentTypeNoContent type — Choose which kind of Hacker News item to search: story (links and text posts), comment, show_hn, ask_hn, poll, or all to search every type. Ignored for front_page queries, which always return front-page stories. Options: story = Stories; comment = Comments; show_hn = Show HN; ask_hn = Ask HN; poll = Polls; all = All types.story

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.1.0

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does solid work: it discloses the data source (Algolia HN API), that no API key is needed, the billing model (~$0.0005 per Story on Apify's free plan), and the returned row shape. Gaps remain around rate limits, pagination behavior, and error handling.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core capability, then output shape, then cost. Dense but every clause earns its place; no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 7-parameter scraper with full schema coverage and no output schema, the description supplies the missing return-value context (row fields) plus sourcing and pricing. Only minor operational details (pagination, cost at scale) are absent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and every parameter (author, sortBy, queries, minPoints, sinceDays, maxResults, contentType) is documented in the schema itself. The description repeats the high-level search semantics without adding syntax, defaults, or edge cases beyond what the schema already provides, so baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Specific verb+resource+source: 'Search Hacker News stories, comments, Show HN and Ask HN posts by keyword, or pull the current front page — via the free Algolia HN API.' It also states the output shape (one row per story/comment with points, author, comment count, text), so an agent can distinguish it from siblings like stackexchange_scraper or github_trending_scraper at a glance.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Clearly frames the two operating modes (keyword search vs. front-page pull) and notes the front_page sentinel. It does not name alternatives or exclusions, but sibling tools target entirely different platforms, so the routing need is minimal.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.