Skip to main content
Glama
Crawlora-org

Crawlora MCP

Official

extract

Scrape a public URL into clean Markdown and return JSON that strictly conforms to a supplied bounded JSON Schema; use render mode and instructions as needed.

Instructions

Extract schema-conforming JSON from a URL. Scrapes a public URL into clean Markdown, then returns data that strictly conforms to the supplied bounded JSON Schema.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
extractOptionYesExtraction options

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed2 schema fields changedv1.17.5
    • addedInput schema / properties / extractOption / properties
      Added value: +{
      +  "instructions": {
      +    "type": "string"
      +  },
      +  "render": {
      +    "enum": [
      +      "auto",
      +      "http",
      +      "browser"
      +    ],
      +    "type": "string"
      +  },
      +  "schema": {
      +    "additionalProperties": {},
      +    "type": "object"
      +  },
      +  "url": {
      +    "type": "string"
      +  }
      +}
    • addedInput schema / properties / extractOption / required
      Added value: +[
      +  "schema",
      +  "url"
      +]
  2. Addedv1.5.0

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It does reveal the internal steps (scraping to Markdown before extraction) and emphasizes strict conformance to the schema, which is useful. However, it does not disclose potential side effects, error behavior, or any resource limitations. While scraping public URLs implies read-only behavior, this is never explicitly stated.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise and well-structured, with the core functionality presented in the first sentence and the method elaborated in the second. Every sentence adds value, and there is no redundant or irrelevant content. It is appropriately front-loaded for quick understanding.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a generic extraction tool, the description covers the main inputs and outputs (URL, schema, resulting JSON). However, it lacks details on error cases, pagination, rate limits, or how to handle complex schemas. Given the extensive sibling list, it would benefit from explicitly stating that it is a universal fallback, but it does not. The absence of an output schema means the description is relied upon more heavily, yet it leaves some gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The description adds little meaning to the parameters beyond what the schema already provides. It mentions the 'schema' parameter implicitly but does not explain the purpose of 'url', 'render', or 'instructions'. The render enum values (auto, http, browser) are not described at all. Even though schema coverage is reported as 100%, the inner properties lack individual descriptions, so the description should have compensated but does not.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action (Extract), the resource (schema-conforming JSON from a URL), and the process (scrapes URL to Markdown, then returns data conforming to a provided JSON Schema). It distinguishes itself from the numerous domain-specific sibling tools by emphasizing that it is a generic, schema-driven extractor for any public URL.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus the many specific scrapers available. It does not state that this should be used as a fallback for arbitrary URLs without a dedicated tool, nor does it mention any exclusions or prerequisites. The only implicit hint is the generic nature, but it is not explicitly communicated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Deploy Server

Other Tools