Skip to main content
Glama
outscraper

Outscraper MCP

Official
by outscraper

AI Scraper

ai_scraper

Extract structured JSON from any web page by providing a URL, a natural-language prompt, and an optional schema. Turn company, people, product, or document metadata into clean, queryable data.

Instructions

Extract structured information from a web page with Outscraper AI Scraper.

Best for:

  • scraping one page and turning it into structured JSON

  • extracting company, people, product, or document metadata from a site

  • guiding extraction with both a prompt and a JSON-schema-like shape

This tool is best for extracting structured data from a single page.

How schema works:

  • schema describes the shape of the output you want back

  • use type="object" with properties for named fields

  • use type="array" with items when a field should be a list

  • add required when some fields must be present

Example schema: { "type": "object", "required": [], "properties": { "company_name": { "type": "string" }, "company_description": { "type": "string" }, "people": { "type": "array", "items": { "type": "string" } } } }

Execution notes:

  • execution_mode="sync" requests a direct response

  • execution_mode="async" returns a request id for polling with requests_get

  • if both prompt and schema are provided, prompt guides the extraction and schema shapes the output

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
asyncNoDeprecated compatibility flag. Prefer execution_mode.
queryYesOne URL to scrape, for example https://outscraper.com.
promptNoNatural-language extraction instructions, for example what to summarize or pull from the page.
schemaNoExtraction schema describing the shape of the result. This is typically a JSON-schema-like object with type, properties, and optional required fields.
execution_modeNoExecution strategy. Use auto to let the MCP server choose between sync and async.

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault
dataYes
metaYes
asyncNo

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed2 schema fields changedv0.2.5
    • changedInput schema / $schema
      Previous value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema"
    • changedOutput schema / $schema
      Previous value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema"
  2. First observedv0.2.3

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and largely delivers. It explains sync vs async behavior, that async returns a request id for polling via requests_get, and how prompt and schema jointly affect extraction. It stops short of disclosing costs, rate limits, or failure behavior, but what it does disclose is substantial and non-obvious.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is longer than typical, but every section earns its place given the nested schema and mode selection. The use of 'Best for', 'How schema works', and 'Execution notes' improves scannability. A slight redundancy exists because 'best for structured data from a single page' is repeated after the earlier best-for list.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with a nested schema, no annotations, and multiple execution modes, the description covers the key operational details: what the tool does, when to use it, how to write the schema, and how to handle sync vs async. It does not mention authentication, account balance, or failure/error handling, but the presence of an output schema lessens the need to describe return shapes.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds significant meaning beyond the schema. The example schema clarifies how to shape output, and the execution notes explain the semantic difference between sync and async plus the combined effect of prompt and schema. This directly helps an agent construct correct parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Extract structured information from a web page with Outscraper AI Scraper.' It further clarifies this is for turning a single page into structured JSON, which clearly distinguishes it from the many search/review/listing sibling tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The 'Best for' section gives concrete use cases and the description states it is best for structured data from a single page. It also explains how async execution connects to requests_get for polling. It does not explicitly name sibling tools to avoid, but the single-page scoping makes the intended usage clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.