Skip to main content
Glama

Extract structured fields from a document

forcedream_extract_data

Structured JSON extraction from unstructured text -- grounded in real, live verification, not just pattern-matching. Pulls requested fields, nulls anything missing, never guesses. Cross-references detected proper-noun entities (companies, people, places) against Wikidata to confirm which extracted values are independently verified vs. unconfirmed. SPENDS your balance -- requires authentication (OAuth). Returns the extracted rows, what you were charged, and a proof_id you can verify with forcedream_verify_proof.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
fieldsYesThe field names to extract, e.g. ["company_name", "ceo_name"].
documentYesThe unstructured document text to extract from.
budget_penceNoOptional max spend in pence for this extraction.

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault
outputNo
statusYes'completed' or 'error'.
verifyNo
task_idNo
proof_idNo
balance_penceNo
charged_penceNo

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description provides rich behavioral context beyond annotations: it spends balance, requires authentication, cross-references Wikidata for verification, and returns a proof_id for verification. Annotations already indicate openWorldHint=true and readOnlyHint=false, but the description adds critical details about cost, authentication needs, and verification workflow.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single paragraph that efficiently conveys key points. It is slightly marketing-heavy ('real, live verification') but every sentence adds value. It could be slightly more concise, but the structure is logical and front-loaded with the core purpose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given that an output schema exists (not shown but mentioned), the description succinctly summarizes return values (extracted rows, charge, proof_id). It also covers authentication, cost, and verification procedure. For a tool with 3 parameters and no enums, this is complete and well-contextualized.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, with each parameter described. The description adds meaningful semantics beyond schema: it states that nulls are returned for missing fields and that extraction is 'grounded in real, live verification', which explains 'never guesses' behavior. This adds value, though the schema already provides good descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: structured JSON extraction from unstructured text with grounding in verification. It specifies the verb (extract), resource (structured fields from a document), and scope (grounded in verification, nulls missing fields). It distinguishes itself from sibling forcedream_verify_proof by mentioning the proof_id usage.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description notes that it spends the user's balance and requires OAuth authentication, giving clear context on when it's appropriate to use. It also mentions that extracted values are cross-referenced against Wikidata and that the tool never guesses. However, it does not explicitly state when not to use this tool or alternatives beyond forcedream_verify_proof.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.

TDQS

A4/5.0
Disambiguation3/5

Most tools have clearly distinct purposes (fraud vs extract vs generate vs sentiment vs lead scoring vs quote vs proof verification). However, there is notable overlap among the search_* discovery tools: forcedream_search_agents, forcedream_search_reliability, and forcedream_search_costs all surface overlapping agent metadata (success_rate appears in both search_agents and search_reliability), which could cause misselection. Additionally, forcedream_extract_data vs forcedream_extract_entities vs forcedream_extract_action_items overlap somewhat in the extraction domain despite distinct outputs (JSON fields vs raw entities vs action items).

Naming Consistency4/5

The forcedream_ prefix is used consistently throughout, and most tools follow a forcedream_<verb>_<object> pattern (extract_data, generate_code, score_lead, security_scan). However, there is inconsistency in verb style: check vs extract vs generate vs invoke vs search vs verify vs summarize are all different verb types, and the objects don't follow a uniform noun convention (some are actions like invole_agent, others resources like market_quote). The naming is readable and discoverable but not perfectly uniform.

Tool Count4/5

At 17 tools, this is slightly above the ideal range but justifiable given the broad multi-service scope (fraud, extraction, generation, discovery, verification). Each tool maps to a reasonably distinct service capability, and none feel like padding. The count borders on heavy but earns its place given the diverse domain coverage.

Completeness4/5

The tool surface is comprehensive for a multi-purpose AI/ML service platform, covering fraud detection, data extraction, code generation, sentiment analysis, embeddings, lead scoring, security scanning, summarization, market quotes, agent discovery, and proof verification. Missing are update/delete operations, but this appears to be a stateless service rather than a CRUD resource store. The discovery tools (search_* variants) and meta capabilities (verify_proof) round out the lifecycle well, though there's no clear cleanup or batch-processing tool.