Skip to main content
Glama

structured_data_extractor

Extract Schema.org JSON-LD and Open Graph tags from any URL, returning one clean row per page for RAG pipelines and SEO audits.

Instructions

Structured Data & JSON-LD Extractor reads every Schema.org JSON-LD block and Open Graph tag on a page and returns one clean row per URL — product, price, rating, article, job posting, event, recipe, FAQ and breadcrumb data, ready for RAG pipelines and SEO rich-result audits. Billed to your own Apify account: ~$0.002 per result (Apify free-plan price, lower on paid plans).

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
urlsYesURLs — Enter the page URLs to extract structured data from, one row is returned per URL, e.g. https://www.allbirds.com/products/mens-tree-dashers. Works on product, article, job, event, recipe and FAQ pages — anything that publishes Schema.org JSON-LD or Open Graph tags. Example: ["https://www.apify.com"].
followRedirectsNoFollow redirects — Keep this on to follow HTTP redirects to the final page, e.g. a shortened or tracking URL. Turn it off to get an error row instead when a URL 301/302s, useful for auditing which URLs redirect.
includeOpenGraphNoInclude Open Graph / Twitter Card data — Keep this on to return the openGraph field (og:title, og:description, og:image, og:type, og:site_name, twitter:card and related tags). Most pages publish these even without JSON-LD.
includeRawJsonLdNoInclude raw JSON-LD — Keep this on to also return the page's raw parsed JSON-LD documents in the jsonLd field (capped at 400 KB per row), useful when you need a schema type the mapped fields do not cover. Turn it off for a smaller, cheaper-to-store dataset.

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.1.0

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the behavioral burden and delivers non-obvious operational facts: billing is charged to the caller's own Apify account at roughly $0.002 per result. It omits run-time behavior such as auth setup, rate limits, and pagination, but the cost/billing disclosure is genuine value beyond structured fields.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with the core action before the cost note. The first sentence is dense but every clause lists concrete extracted types rather than filler; the pricing sentence is short and earns its place as actionable cost context.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, so the description must convey the return shape — and it does ('one clean row per URL' with the enumerated field categories). For a read-only scraping tool with fully documented parameters, this is nearly complete; only the absence of return-format details like the row field names holds it below a 5.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all four parameters (urls, followRedirects, includeOpenGraph, includeRawJsonLd) are already documented in detail. The description adds only the 'one row per URL' output-shape hint and no parameter-level syntax or format detail, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('reads every Schema.org JSON-LD block and Open Graph tag on a page') and enumerates the payload it returns (product, price, rating, article, job posting, event, recipe, FAQ, breadcrumb). An agent can distinguish it from siblings like article_extractor or website_to_markdown without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Clear context is supplied — 'ready for RAG pipelines and SEO rich-result audits' tells the agent the intended scenarios. However, it never names an alternative tool or states a when-not-to-use condition (e.g., when to pick article_extractor instead), so it falls short of explicit routing guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.