Skip to main content
Glama

Extract structured data (AI)

extract_data
Read-only

Extract specified fields from a web page, PDF, or Office file as JSON, using AI. Provide 'fields' plus ONE source: a 'url', a 'pdf' (file_id or base64), or a 'file' (base64 .docx/.xlsx/.csv). Returns a JSON object mapping each field to its value, or null when absent.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
pdfNoA file_id from a prior tool result, or a base64-encoded PDF.
urlNoA public http/https URL to read.
fileNoA base64-encoded .docx, .xlsx or .csv file (converted to text first).
fieldsYesField names to extract, e.g. ['invoice total','due date'].

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed1 schema field changed
    • addedInput schema / properties / file
      Added value: +{
      +  "description": "A base64-encoded .docx, .xlsx or .csv file (converted to text first).",
      +  "type": "string"
      +}
  2. Added

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description discloses that the tool returns a JSON object mapping each requested field to its value, with null for absent fields. This adds useful behavioral detail beyond the readOnlyHint and openWorldHint annotations, though it does not mention potential model limitations, latency, or cost.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is brief, well-organized, and free of redundant detail. It conveys the core behavior, required inputs, and output format in two clear sentences.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the moderate complexity and lack of an output schema, the description provides enough context for an agent to call the tool correctly: which sources are supported, how to specify fields, and what the response will look like. Error cases like multiple sources are not mentioned, but the 'ONE source' instruction mitigates that gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema already covers 100% of parameters, and the description reinforces the meanings by specifying that 'pdf' can be a file_id or base64, 'file' is base64 for Office formats, and 'fields' are the names to extract. It adds valuable one-of guidance not fully captured by the schema's structure.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states a specific action: extract specified fields from a web page, PDF, or Office file and return them as JSON. It distinguishes this tool from sibling tools focused on scraping, conversion, or full-text extraction by emphasizing custom field-level extraction using AI.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explains that users must provide 'fields' plus exactly one source, and it enumerates valid source types. It does not explicitly name alternative tools or state when not to use this tool, but the 'using AI' and custom-field wording makes the intended usage clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.

Resources