Skip to main content
Glama
mysleekdesigns

CrawlForge MCP Server

extract_with_llm

Read-only

Extract data from a URL or text using a natural-language prompt; default local Ollama needs no API key, with OpenAI/Anthropic optional.

Instructions

Extract data from a URL or text using a natural-language prompt. Defaults to a local Ollama model (http://localhost:11434, no API key required) and picks an installed model itself; pass model to choose one, or provider: "openai" or "anthropic" with the matching API key to use a cloud model instead. Not for known selectors (scrape_structured) or a schema-shaped result (extract_structured). Call list_ollama_models only when a model name is rejected. Cost: 3 credits plus your LLM provider's own charge.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
urlNoURL to fetch and extract from (one of url/content required)
modelNoOverride the model. For ollama, pass a name returned by list_ollama_models (e.g. 'llama3.2', 'qwen2.5:7b'). Defaults: openai='gpt-4o-mini', anthropic='claude-haiku-4-5-20251001', ollama=$OLLAMA_DEFAULT_MODEL, else the best installed model for extraction (gemma3:12b, then gemma3:4b, gpt-oss:20b, mistral:7b, llama3.2, qwen2.5:3b).
promptYesNatural-language extraction instruction
schemaNoOptional JSON-schema for output shape (used as Ollama structured-outputs format when provider is 'ollama')
contentNoPre-fetched text to extract from (one of url/content required)
providerNoLLM provider. Defaults to 'ollama' (local, no key, http://localhost:11434). Use 'openai' or 'anthropic' for cloud models (requires the matching API key).auto
maxTokensNoMaximum output tokens
user_agentNoOverride the outbound User-Agent. CrawlForge identifies itself honestly by default; use this only for targets you have your own agreement with.
respect_robotsNoRespect the target site's robots.txt (default: true). Setting this to false is honoured, returns a warning in the response, and is recorded against your API key — it is your decision, not a silent default.
verify_numbersNoNumeric provenance guard (default: true): every price or numeric value the LLM returns must appear literally in the page source, else it is returned as null with a reason in `provenance.unverified`. Set false to get the model's raw numbers back, including ones it derived (a count, a sum, a total) rather than read off the page.

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed1 schema field changedv6.19.2
    • changedInput schema / properties / model / description
      Previous value: -"Override the model. For ollama, pass a name returned by list_ollama_models (e.g. 'llama3.2', 'qwen2.5:7b'). Defaults: openai='gpt-4o-mini', anthropic='claude-haiku-4-5-20251001', ollama='llama3.2' or $OLLAMA_DEFAULT_MODEL."New value: +"Override the model. For ollama, pass a name returned by list_ollama_models (e.g. 'llama3.2', 'qwen2.5:7b'). Defaults: openai='gpt-4o-mini', anthropic='claude-haiku-4-5-20251001', ollama=$OLLAMA_DEFAULT_MODEL, else the best installed model for extraction (gemma3:12b, then gemma3:4b, gpt-oss:20b, mistral:7b, llama3.2, qwen2.5:3b)."
  2. Changed2 schema fields changedv6.0.0
    • removedInput schema / additionalProperties
      Removed value: -false
    • addedInput schema / properties / schema / propertyNames
      Added value: +{
      +  "type": "string"
      +}
  3. Changed3 schema fields changedv5.4.0
    • addedInput schema / properties / respect_robots
      Added value: +{
      +  "description": "Respect the target site's robots.txt (default: true). Setting this to false is honoured, returns a warning in the response, and is recorded against your API key — it is your decision, not a silent default.",
      +  "type": "boolean"
      +}
    • addedInput schema / properties / user_agent
      Added value: +{
      +  "description": "Override the outbound User-Agent. CrawlForge identifies itself honestly by default; use this only for targets you have your own agreement with.",
      +  "type": "string"
      +}
    • addedInput schema / properties / verify_numbers
      Added value: +{
      +  "default": true,
      +  "description": "Numeric provenance guard (default: true): every price or numeric value the LLM returns must appear literally in the page source, else it is returned as null with a reason in `provenance.unverified`. Set false to get the model's raw numbers back, including ones it derived (a count, a sum, a total) rather than read off the page.",
      +  "type": "boolean"
      +}
  4. Changed1 schema field changedv5.0.4
    • changedInput schema / $schema
      Previous value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema"
  5. First observedv4.10.0

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, openWorldHint, idempotentHint=false, destructiveHint=false, so the safety profile is covered. The description adds genuinely non-structured context: default local Ollama endpoint with no API key, self-selection of an installed model, the need for a matching cloud API key on provider switch, and a credit cost (3 credits plus provider charge). It does not describe return shape or rate limits, but cost and provider prerequisites are the salient behavioral facts.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, front-loaded with purpose, then provider mechanics, then exclusions, then cost. No filler and no repetition of schema fields beyond the minimal cross-reference needed for routing.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 10-parameter, nested-object tool with no output schema, the description covers purpose, defaults, provider/auth prerequisites, sibling routing, and cost. The main remaining gap is the response shape, but read-only annotations and the absence of an output schema make that a minor omission rather than a blocking one.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema baseline is 3. The description goes slightly beyond it by linking the `model` and `provider` parameters into one decision (local Ollama by default, or 'provider: "openai" or "anthropic" with the matching API key'), which the schema documents per-field but does not connect.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Extract data from a URL or text') plus the mechanism (natural-language prompt), and explicitly distinguishes itself from the nearest siblings by naming scrape_structured and extract_structured as the wrong choices. An agent can route correctly without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit when-not conditions ('Not for known selectors (scrape_structured) or a schema-shaped result (extract_structured)') and a narrow conditional for list_ollama_models ('only when a model name is rejected'). Alternatives and their selection criteria are stated, not implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.