Skip to main content
Glama

Extract structured fields from a document

forcedream_extract_data

Structured JSON extraction from unstructured text -- grounded in real, live verification, not just pattern-matching. Pulls requested fields, nulls anything missing, never guesses. Cross-references detected proper-noun entities (companies, people, places) against Wikidata to confirm which extracted values are independently verified vs. unconfirmed. SPENDS your balance -- requires authentication (OAuth). Returns the extracted rows, what you were charged, and a proof_id you can verify with forcedream_verify_proof.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
fieldsYesThe field names to extract, e.g. ["company_name", "ceo_name"].
documentYesThe unstructured document text to extract from.
budget_penceNoOptional max spend in pence for this extraction.

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault
outputNo
statusYes'completed' or 'error'.
verifyNo
task_idNo
proof_idNo
balance_penceNo
charged_penceNo

Schema Changelog

Changes observed during successful MCP inspections.

  1. Added

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description provides rich behavioral context beyond annotations: it spends balance, requires authentication, cross-references Wikidata for verification, and returns a proof_id for verification. Annotations already indicate openWorldHint=true and readOnlyHint=false, but the description adds critical details about cost, authentication needs, and verification workflow.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single paragraph that efficiently conveys key points. It is slightly marketing-heavy ('real, live verification') but every sentence adds value. It could be slightly more concise, but the structure is logical and front-loaded with the core purpose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given that an output schema exists (not shown but mentioned), the description succinctly summarizes return values (extracted rows, charge, proof_id). It also covers authentication, cost, and verification procedure. For a tool with 3 parameters and no enums, this is complete and well-contextualized.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, with each parameter described. The description adds meaningful semantics beyond schema: it states that nulls are returned for missing fields and that extraction is 'grounded in real, live verification', which explains 'never guesses' behavior. This adds value, though the schema already provides good descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: structured JSON extraction from unstructured text with grounding in verification. It specifies the verb (extract), resource (structured fields from a document), and scope (grounded in verification, nulls missing fields). It distinguishes itself from sibling forcedream_verify_proof by mentioning the proof_id usage.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description notes that it spends the user's balance and requires OAuth authentication, giving clear context on when it's appropriate to use. It also mentions that extracted values are cross-referenced against Wikidata and that the tool never guesses. However, it does not explicitly state when not to use this tool or alternatives beyond forcedream_verify_proof.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.