Skip to main content
Glama

google_news_scraper

Fetch Google News headlines by keyword, site: or when: query or topic, and resolve each article's real publisher URL—one row per article for RAG feeds.

Instructions

Google News Scraper returns headlines from Google News RSS search and topic feeds for any keyword, site: or when: query, and resolves each article's real publisher URL — one row per article. Billed to your own Apify account: ~$0.0005 per article (Apify free-plan price, lower on paid plans).

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
topicsNoSection topics — Enter Google News section topics to fetch instead of (or in addition to) a search, e.g. TECHNOLOGY or BUSINESS. Accepted values: WORLD, NATION, BUSINESS, TECHNOLOGY, ENTERTAINMENT, SCIENCE, SPORTS, HEALTH. Unknown values are ignored.
countryNoCountry — Enter the 2-letter country/edition code, e.g. US for the United States or GB for the United Kingdom. Combined with Language to build the hl/gl/ceid feed parameters.US
queriesYesSearch queries — Enter Google News search terms, one row is returned per matching article. Supports Google's search operators, e.g. web scraping, "exact phrase", site:reuters.com AI, or -unwanted. Leave empty and use Topics below instead if you only want section feeds. Example: ["web scraping"].
languageNoLanguage — Enter the 2-letter interface language code Google News should use, e.g. en for English or fr for French. Combined with Country to build the hl/gl/ceid feed parameters.en
sinceDaysNoOnly articles from the last N days — Enter how many days back to search, e.g. 7 for the last week. Appends Google's when:Nd search operator to every query (topic feeds ignore this — they are always "latest"). Enter 0 to disable and return whatever Google's default ranking gives.
maxItemsPerQueryNoMax items per query/topic — Enter the maximum number of articles to keep per query or topic, e.g. 50. Google's own feed rarely returns more than ~100 items for a single request no matter how high this is set.
resolvePublisherUrlsNoResolve publisher URLs — Keep this on to follow each article's news.google.com redirect link and fetch the publisher's real URL (2 extra requests per article, using an undocumented Google endpoint — best-effort, see README). Turn it off for a much faster run that only returns the Google News link.

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.1.0

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does disclose real behavioral traits: billing to the user's own Apify account at ~$0.0005/article, that resolvePublisherUrls costs 2 extra requests per article via an undocumented best-effort endpoint, and that topic feeds ignore sinceDays. It does not cover rate limits or failure behavior beyond the 'best-effort' note, keeping it short of a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two dense sentences: purpose first, then cost and row granularity. No filler, and pricing is placed where a cost-sensitive agent will see it. Slightly overstuffed with parenthetical pricing detail, but every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 7-parameter scraper with no output schema and no annotations, the description covers what matters: row-per-article output shape, cost model, and the performance/accuracy tradeoff of URL resolution. It leaves some return-field detail unstated, but the essential selection and invocation context is present.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all 7 parameters in detail (operator support, country/language feed construction, bounds). The description adds no per-parameter syntax beyond reinforcing the query types, so the baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a specific verb+resource ('returns headlines from Google News RSS search and topic feeds') and adds a concrete scope detail ('one row per article'). It implicitly separates itself from generic siblings like ai_web_search by narrowing to Google News RSS, but it never names an alternative tool, so it stops short of full sibling differentiation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It indicates supported query forms (keyword, site:, when:) and notes that topic feeds are an alternative to searching, which implies usage. However, it never states when to pick this tool over ai_web_search, article_extractor, or press_release_monitor, and there are no explicit exclusion conditions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.