CrawlForge MCP Server
CrawlForge MCP Server is an MCP-native web scraping, crawling, research and browser-automation toolkit that gives AI clients 31 credit-metered tools for turning any site into clean Markdown or structured JSON.
Single-page reading:
scrapereturns markdown/html/rawHtml/text/links/metadata/branding/screenshot plus schema-driven JSON or query-basedhighlights/questionin one fetch;fetch_urlgrabs raw API/JSON bodies;extract_text,extract_links,extract_metadata,extract_contenthandle targeted reads.Structured extraction:
scrape_structured(CSS selectors, row-aligned records),extract_structured(LLM/schema-driven),extract_with_llm(natural language, local Ollama by default or OpenAI/Anthropic),scrape_template(10+ prebuilt site/API templates),extract_embedded_state(Next.js/Nuxt/SvelteKit/Redux/ytInitialData payloads).Multi-page & batch:
crawl_deepfollows links across a site with depth/page caps and PageRank analysis;batch_scrapedoes 2–50 URLs with sync or async+webhook;map_sitediscovers URLs from sitemaps/links;get_batch_resultspages through finished jobs.Search & research:
search_web(up to 10 queries per call),deep_research(multi-source verified reports),serp_rank(real Google organic positions via DataForSEO),reddit_search(Arctic Shift archive posts/comments/threads).Browser interaction:
scrape_with_actionsruns click/type/scroll/form actions;browser_sessionkeeps a page, cookies and login alive across calls with stable @e refs;stealth_moderenders with anti-detect fingerprints;agentautonomously plans and answers from a prompt with no URLs.Content & documents:
analyze_content(sentiment, topics, entities, readability),summarize_content,process_document(PDF/DOCX),track_changes(baselines, diffs, scheduled monitors, alerts).Utilities:
localization(country locale settings),generate_llms_txt,list_ollama_models, andread_resultto slice/search results too large to return inline.Safety & plumbing: SSRF validation on every URL, backend host allow-list, browser action allow-list, per-tool credit gating, optional PII redaction, robots.txt respect, and MCP structured output, self-correctable errors and experimental async tasks on long-running tools.
Provides tools to scrape structured data from Amazon product pages, including details, reviews, and more.
Provides tools to scrape structured data from GitHub repositories, user profiles, and activity.
Provides web search capabilities using Google Search API, allowing retrieval of search results.
Provides tools to scrape structured data from npm package pages and registry.
Integrates with local Ollama models for LLM-powered content extraction without external API calls.
Offers integration with OpenAI models as an optional provider for LLM-powered extraction and analysis.
Provides tools to scrape structured data from Reddit posts, comments, and subreddits.
Provides tools to scrape structured data from YouTube videos, channels, and playlists.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@CrawlForge MCP Servercrawl docs.python.org and return the table of contents as markdown"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Table of Contents
Related MCP server: Crawl4AI MCP
🎯 Why CrawlForge?
31 MCP-native tools — scraping, crawling, search, real Google SERP rank tracking, deep research, an autonomous
agent, a unified multi-formatscrape, document processing, stealth browsing, stateful browser sessions, and more, callable directly from your AI assistant.Generous free tier — 1,000 credits to start instantly, no credit card. The grant is one-time rather than monthly, and the credits never expire.
Local-LLM by default —
extract_with_llmruns against a local Ollama model out of the box: no LLM API key, no per-token cost, and your data never leaves your machine. Cloud (OpenAI/Anthropic) is opt-in.LLM-ready output — clean Markdown, structured JSON (schema-driven), screenshots, links, and metadata from a single fetch.
Autonomous
agent— describe what you need in natural language; it plans, gathers, and shapes an answer under orchestrator-enforced hard stops (max steps/URLs/wall-clock) — no URLs required.Security-hardened — SSRF protection on every request, a fail-closed backend allow-list, a vetted action allowlist for browser automation, and per-tool credit gating.
Works everywhere MCP does — Claude Desktop, Claude Code, Cursor, and any other MCP-enabled client, configured in one command.
📊 CrawlForge vs. alternatives
CrawlForge MCP | Firecrawl | Raw scraping API | |
Native MCP server | ✅ 31 tools | ✅ | ❌ |
Free tier | ✅ 1,000 credits, one-time, never expire | Limited | Varies |
Self-hosted / local LLM extraction (Ollama) | ✅ default, $0/token | ❌ | ❌ |
Autonomous agent (no URLs needed) | ✅ | ✅ | ❌ |
Deep research with source verification | ✅ | Partial | ❌ |
Browser automation / actions | ✅ | ✅ | Varies |
Stealth / anti-detection engines | ✅ Chromium + Camoufox | ✅ | Add-on |
Pre-built site templates | ✅ 10 sites | ❌ | ❌ |
License | MIT | AGPL-3.0 | Proprietary |
Comparison reflects publicly documented capabilities at time of writing. CrawlForge is MIT-licensed and MCP-first — built to plug straight into AI coding assistants.
🚀 Quick Start (2 Minutes)
1. Install from NPM
npm install -g crawlforge-mcp-server2. Setup Your API Key (required)
Every tool requires a CrawlForge API key — new accounts get 1,000 free trial credits to start. The recommended path signs you in through the browser, so the key is never pasted into a terminal (a coding agent can run this for you and relay the URL):
crawlforge loginIt prints an approval URL; open it, approve, and the key is stored in ~/.crawlforge/config.json. Then run crawlforge init to register the MCP server with your client. Or use the interactive wizard, which also configures your clients:
npx crawlforge-setupThis will:
Guide you through getting your free API key
Configure your credentials securely
Auto-configure Claude Code and Cursor (if installed)
Verify your setup is working
Don't have an API key? Get one free at https://www.crawlforge.dev/signup
One-step setup (v4.6.0+):
crawlforge initdetects your API key, installs the agent skill, and idempotently merges the MCP config stanza into Claude Code, Claude Desktop, and Cursor. Usecrawlforge init --all --yesto configure every detected client non-interactively.
3. Configure Your IDE (if not auto-configured)
Add to claude_desktop_config.json:
{
"mcpServers": {
"crawlforge": {
"command": "npx",
"args": ["-y", "crawlforge-mcp-server"]
}
}
}Location:
macOS:
~/Library/Application Support/Claude/claude_desktop_config.jsonWindows:
%APPDATA%/Claude/claude_desktop_config.jsonLinux:
~/.config/Claude/claude_desktop_config.json
Restart Claude Desktop to activate.
The setup wizard automatically configures Claude Code by adding to ~/.claude.json:
{
"mcpServers": {
"crawlforge": {
"type": "stdio",
"command": "crawlforge-mcp"
}
}
}After setup, restart Claude Code to activate.
The setup wizard automatically configures Cursor by adding to ~/.cursor/mcp.json:
{
"mcpServers": {
"crawlforge": {
"type": "stdio",
"command": "crawlforge-mcp"
}
}
}Restart Cursor to activate.
n8n's built-in MCP Client Tool node connects over Streamable HTTP (works on n8n Cloud and self-hosted). Run the server in HTTP mode:
export CRAWLFORGE_API_KEY=your_api_key
npm run start:http # Streamable HTTP endpoint at http://localhost:10000/mcpThen point the MCP Client Tool node at http://<host>:10000/mcp with transport HTTP Streamable and a Bearer credential set to the same API key. On self-hosted n8n you can instead use the community n8n-nodes-mcp node over STDIO (npx -y crawlforge-mcp-server).
Full guide: docs/n8n-integration.md
Which launch command?
npx -y crawlforge-mcp-serverneeds no global install and always runs the published version (recommended for Claude Desktop). For a global install (npm i -g crawlforge-mcp-server), use the dedicatedcrawlforge-mcpbin — it resolves on yourPATH, so it survives Node/nvm version switches. The barecrawlforgecommand still launches the server when an MCP client spawns it over stdio (backward compatibility for configs created before v4.2.5); interactively it's the CLI — runcrawlforge mcpto start the server by hand.
📊 Available Tools
CrawlForge requires a CrawlForge API key — every tool is metered and consumes credits. New accounts get 1,000 free trial credits to start. Get a key at crawlforge.dev/signup.
All Tools (API key required)
Tool | Credits | What it does |
| 1 | Fetch content from any URL |
| 1 (6 projected with | Extract clean text from web pages |
| 1 (6 projected with | Get all links from a page |
| 1 | Extract page metadata (title, OG tags, schema.org) |
| 1 | Structured data from well-known sites (Amazon, GitHub, LinkedIn, YouTube, Reddit, Hacker News, npm, and more) without writing selectors |
| 1 | List the Ollama models installed locally (helps you pick a |
| 1 | Retrieve paginated results for a |
| 1 | Search, slice, read lines or a JSON path from a result a tool returned with |
| 2 | Unified single-fetch, multi-format extraction. Pass a |
| 2 | Extract structured data with CSS selectors |
| 2 (7 projected with | Read a page's embedded JavaScript state — |
| 2 | Enhanced content extraction |
| 2 | Discover and map website structure (optional |
| 2 | Multi-format document processing |
| 2 | Look up a country's locale settings (Accept-Language, timezone, currency) to pass to other tools; returns values, applies none |
| 3 | Monitor content changes over time |
| 3 | Comprehensive content analysis |
| 3 | LLM-powered schema-driven extraction (your own LLM key or local Ollama) |
| 3 | Natural-language extraction. Defaults to a local Ollama model; pass |
| 3 | A browser page that stays open across calls, keeping its cookies and its login in between. |
| 4 | Generate intelligent summaries |
| 4 | Deep crawl entire websites |
| 5 | Search the web using Google Search API |
| 5 | Search Reddit posts/comments or read a full thread — reddit.com blocks direct scraping, so this reads the Arctic Shift community archive (free, no Reddit credentials). A Reddit-wide search spends a web search to discover posts, so it is priced with |
| 5 | Check where a domain ranks in Google's real organic SERP for a keyword (the position |
| 5 | Process multiple URLs simultaneously |
| 5 | Browser automation chains |
| 5 | Generate AI interaction guidelines |
| 5 | Anti-detection browser management |
| 8 (+5 per stealth retry that gets the page, max 2) | Autonomous research/extraction from a natural-language prompt — no URLs required. Plans, gathers, and shapes an answer under hard safety stops (max steps/URLs/wall-clock enforced by the orchestrator, never the LLM). A page that walls the plain fetch is retried in the stealth browser automatically — URLs you name first — and that evidence is marked |
| 10 | Multi-stage research with source verification |
Ten tools (scrape, fetch_url, extract_content, crawl_deep, batch_scrape, stealth_mode, scrape_with_actions, process_document, deep_research, extract_embedded_state) accept max_inline_chars (default 40,000; env CRAWLFORGE_MAX_INLINE_CHARS): a result over it comes back as a preview plus a result_handle for read_result, with the full result kept for 1 hour under ~/.crawlforge/results/ on your own machine — nothing is uploaded.
For the full canonical capabilities reference (all tools, CLI commands, stealth engines, research workflow), see SKILL.md.
💳 Pricing
Every tool is metered and requires an API key. New accounts get 1,000 free trial credits — no credit card required to start.
Plan | Credits | Best For |
Free | 1,000 one-time | Testing & personal projects |
Hobby ($19) | 5,000 / month | Small projects & development |
Professional ($99) | 50,000 / month | Professional use & production |
Business ($399) | 250,000 / month | Large scale operations |
All plans include:
Access to all 31 tools
Credits never expire; paid-plan credits roll over month to month
API access and webhook notifications
🔧 Advanced Configuration
Environment Variables
# Optional: Set API key via environment
export CRAWLFORGE_API_KEY="cf_live_your_api_key_here"
# Optional: Custom API endpoint (for enterprise)
export CRAWLFORGE_API_URL="https://api.crawlforge.dev"
# As of v3.0.18, this variable is validated against an allow-list of CrawlForge backend hosts.
# Optional: Local LLM (Ollama) overrides — extract_with_llm, extract_structured
# and deep_research all use Ollama when no cloud key is set
export OLLAMA_BASE_URL="http://localhost:11434" # default; set https://ollama.com for Ollama Cloud
export OLLAMA_DEFAULT_MODEL="gemma3:4b" # optional; unset = pick the best installed model automatically
# deep_research judges claims with gemma3:12b when it is installed (ollama pull gemma3:12b);
# conflict detection is on only with that model, or a cloud provider
export OLLAMA_EMBEDDING_MODEL="nomic-embed-text" # default: OLLAMA_DEFAULT_MODEL; used for semantic ranking in deep_research
export OLLAMA_API_KEY="..." # only for authenticated endpoints (required by Ollama Cloud; a local instance needs none)
export DISABLE_OLLAMA="true" # skip Ollama entirely and use CSS/keyword fallbacks
# Optional: Cloud LLM keys — only needed when you pass provider: "openai" or "anthropic"
export OPENAI_API_KEY="sk-..."
export ANTHROPIC_API_KEY="sk-ant-..."
# Optional: limit which tools this client sees — by name, by group, or both (comma-separated)
export CRAWLFORGE_TOOLS="scrape,search_web,extract_content"
export CRAWLFORGE_TOOL_GROUPS="basic,search,scrape" # unset = all tools; unknown names/groups are ignored with a warning
# Optional: deep_research stealth extraction fallback (v4.6.6) — see below
export RESEARCH_STEALTH_ENGINE="auto" # auto (default) | camoufox | chromium
export RESEARCH_STEALTH_FALLBACK="true" # set to "false" to disable entirely
export RESEARCH_MAX_STEALTH_RETRIES="8" # cap on stealth retries per research run
# Optional: your own proxies for the stealth paths that have no caller to ask —
# the scrape escalation stage, the deep_research retry, the agent, browser_session.
# Comma-separated; a proxy passed on a call always wins. CrawlForge supplies none.
export CRAWLFORGE_STEALTH_PROXIES="http://user:pass@gw.provider.net:8080"
# Unrelated to the PROXY_ROTATION_* variables, which belong to the localization tool.
# Optional: turn off the clearance jar (on by default) — see "Stealth engines and proxies"
export CRAWLFORGE_CLEARANCE_JAR="off"MCP Spec Features
CrawlForge tracks the current MCP spec (2025-06-18) plus select experimental extensions:
Structured output —
scrape,map_site,serp_rank,reddit_search,search_web,extract_structured, andcrawl_deepreturn machine-parseablestructuredContentalongside the usual text, validated against a publishedoutputSchema; legacy clients keep working off the text.Self-correctable errors — invalid tool input now comes back as an
isError: trueresult the calling model can read and retry from, instead of a raw JSON-RPC protocol error.JSON Schema 2020-12 tool schemas, deterministic
tools/listordering (client prompt-cache friendly), and cacheable-result hints on read-only tools.Icons on the server, its tools, and its prompts.
Async tasks (experimental) on the four long-running tools —
crawl_deep,batch_scrape,deep_research,agent— for clients that support polling; synchronous results are still returned for clients that don't.
See docs/mcp-spec-adoption.md for wire-level examples and client-compatibility notes.
Local-LLM quickstart (extract_with_llm with Ollama)
extract_with_llm defaults to a local Ollama model — no LLM-provider key, no per-token LLM costs, and no data leaving your machine (the CrawlForge credit cost still applies).
# 1. Install Ollama: https://ollama.com
# 2. Pull any model from https://ollama.com/library
ollama pull llama3.2
# 3. Discover what's installed (from your MCP client)
# list_ollama_models()
# 4. Extract — defaults to Ollama with the model from step 2
# extract_with_llm({ url: "https://example.com", prompt: "…", model: "llama3.2" })Stealth engines and proxies
Every stealth entry point — scrape with escalate: true, stealth_mode, browser_session, scrape_with_actions and the deep_research retry — now defaults to auto: use Camoufox (Firefox anti-detect) when its binary is installed, fall back to Chromium when it is not, with the reason in the result's warnings[]. Every result names the engine that actually ran. Naming an engine explicitly behaves as it always has — camoufox fails rather than falling back, and chromium and playwright are the same engine (on scrape's escalate_engine, spell it playwright).
On browser_session and scrape_with_actions the engine applies only with stealth: true — their ordinary path is the standard Chromium pool.
auto is not free. Measured once each on an Apple Silicon Mac on 2026-09-21 (launch plus a context and a blank page): Chromium 148 ms and 253 MB, Camoufox 946 ms and 667 MB — roughly +0.8 s and +400 MB per stealth call. Pin engine: "playwright" for high-volume work on sites that do not block.
CrawlForge supplies no proxies, and a datacenter proxy does not fix a block: Cloudflare scores the IP's ASN and the TLS/HTTP2 handshake before any JavaScript runs, so no browser-side patch compensates for a datacenter address. Bring your own residential exit with CRAWLFORGE_STEALTH_PROXIES (comma-separated URLs), which the escalation stage, stealth_mode, browser_session, scrape_with_actions, the deep_research retry and the agent tool's automatic stealth retry use when the caller passes none; a proxy passed on the call always wins. With a proxy, Camoufox derives its timezone, locale and geolocation from the exit IP.
A challenge solved once is not solved again. When a stealth render gets past a Cloudflare or DataDome wall, the vendor's clearance cookies (cf_clearance, __cf_bm, datadome — nothing else) are kept in ~/.crawlforge/stealth-clearance.json (mode 0600) and replayed to the next stealth context with the same engine, User-Agent and proxy, for at most the cookie's own lifetime and never more than 24 hours. A render that still meets the wall drops them. Set CRAWLFORGE_CLEARANCE_JAR=off to disable.
Full detail, measurements and sources: docs/stealth-engines.md.
Stealth extraction for deep_research (Camoufox)
deep_research automatically retries sources that block the normal fetch path (Reddit, Quora, forums, and Cloudflare/DataDome-protected pages return HTTP 403) through a real fingerprinted browser, then re-extracts from the rendered HTML. It's bounded (RESEARCH_MAX_STEALTH_RETRIES, default 8, plus a per-page timeout) and lazy — the browser stack only loads when a source is actually blocked.
Engine selection (RESEARCH_STEALTH_ENGINE):
auto(default) — prefer Camoufox (Firefox anti-detect), fall back to Chromium stealth, then plain fetch.camoufox— force Camoufox.chromium— force the Chromium stealth engine.
Camoufox gets through walls that stop headless Chromium: in a benchmark run on 2026-09-21 from a residential IP it was the only engine to clear Cloudflare Turnstile (indeed.com) and Akamai (harrods.com), both of which blocked Chromium stealth. Neither engine cleared DataDome or an interactive Turnstile, so this is a better fetch path, not a bypass. To enable it, install the optional dependency and run its one-time binary fetch:
# Camoufox is declared as an optional dependency, so a normal install already pulls it.
# If you installed with --no-optional, add it explicitly:
npm install camoufox
# One-time download of the Camoufox Firefox binary (~130 MB):
npx camoufox fetchWithout the Camoufox binary, deep_research silently falls back to Chromium stealth and then to plain fetch — no errors, just lower recovery on heavily-protected sites. Disable the whole fallback with RESEARCH_STEALTH_FALLBACK=false.
Note: Hard IP-reputation blocks (e.g. Reddit's edge
403) resist headless stealth from any IP and require residential/mobile proxies, which CrawlForge does not provide — point the server at your own withCRAWLFORGE_STEALTH_PROXIES. See docs/stealth-engines.md for details.
Manual Configuration
Your configuration is stored at ~/.crawlforge/config.json:
{
"apiKey": "cf_live_...",
"userId": "user_...",
"email": "you@example.com"
}📖 Usage Examples
Once configured, use these tools in your AI assistant:
"Search for the latest AI news"
"Extract all links from example.com"
"Crawl the documentation site and summarize it"
"Monitor this page for changes"
"Extract product prices from this e-commerce site"🔒 Security & Privacy
Secure Authentication: API keys required for all metered tools
Local Storage: API keys stored securely at
~/.crawlforge/config.jsonHTTPS Only: All connections use encrypted HTTPS
No Data Retention: We don't store scraped data, only usage logs
Rate Limiting: Built-in protection against abuse
Compliance: Respects robots.txt and GDPR requirements
Security & Approvals
SSRF enforcement: Every scraped URL is validated before the request is sent — http/https only; blocks loopback, RFC1918, IPv6 private/link-local ranges, cloud metadata endpoints (GCP, Azure), and dangerous ports (SSH, SMTP, DNS, MySQL, Postgres, Redis, MongoDB, etc.). Redirects are re-validated each hop, capped at 5.
Backend endpoint guard (v3.0.18): The server's own calls to CrawlForge.dev use a separate fail-closed allow-list (
{crawlforge.dev, www.crawlforge.dev, api.crawlforge.dev}, HTTPS required). SettingCRAWLFORGE_API_URLto an arbitrary host is blocked at parse time.Action allowlist:
scrape_with_actionsaccepts only 11 action types (snapshot,wait,click,type,press,scroll,screenshot,executeJavaScript,select,hover,navigate). No download, file-write, or arbitrary cross-page navigation primitives exist —navigategoes through the same SSRF and robots.txt gate as the initial URL.JavaScript gate: The
executeJavaScriptaction throws by default. SetALLOW_JAVASCRIPT_EXECUTION=trueat deploy time to enable (not recommended in production).MCP Elicitation (v3.6.0): Four tools request user confirmation before executing expensive operations —
deep_research(>50 URLs),batch_scrape(sync mode, >25 URLs),crawl_deep(projected >500 pages),extract_structured(schema has >3 required fields with no LLM configured). Credit-low situations also elicit. Confirmation is best-effort: if the MCP client does not support elicitation the tool proceeds (fail-open).Per-tool credit gating: Every tool is wrapped with
withAuth()and is metered — credits are checked and deducted before execution, and a valid API key is required for every tool (fail-closed since v3.0.18).
See docs/sandboxing-and-approvals.md for the full reference.
Security Updates
v3.0.3 (2025-10-01): Removed authentication bypass vulnerability. All users must authenticate with valid API keys.
For the full security policy and how to report a vulnerability, see SECURITY.md.
🆘 Support
Documentation: https://www.crawlforge.dev/docs
Issues: GitHub Issues
Email: support@crawlforge.dev
Discord: Join our community
📄 License
MIT License - see LICENSE file for details.
🤝 Contributing
Contributions are welcome! Please read our Contributing Guide first.
Built with ❤️ by the CrawlForge team
Available Tools
31 toolsagentARead-only
Use this when you need an autonomous agent to research, navigate, and synthesise an answer from the web - no URLs required. The agent plans search queries, fetches and filters relevant pages, and returns a prose or structured answer. model:"pro" uses deep multi-source research. Hard limits: maxSteps<=10, maxUrls<=20, 120s wall-clock. Confirms before pro runs. Degraded-but-useful output if no LLM keys/Ollama. Not for a URL you already have (scrape) or a question one search answers (search_web). Pages that block a plain fetch are retried in the stealth browser automatically (at most 2 a run, URLs you name first; evidence marked via:"stealth"). Cost: 18 credits at most - 8, plus 5 per stealth retry that gets the page; a retry that is blocked again is free. Example: agent({prompt:"What are the top 5 MCP servers in 2025?", maxUrls:10})
| Name | Required | Description | Default |
|---|---|---|---|
| urls | No | Optional seed URLs to include (max 20) | |
| model | No | "default" = SamplingClient loop (no keys needed); "pro" = full ResearchOrchestrator | default |
| prompt | Yes | Natural-language task or question | |
| schema | No | Optional JSON schema for structured output | |
| maxUrls | No | Max URLs to fetch (hard cap: 20) | |
| maxSteps | No | Max fetch iterations (hard cap: 10) | |
| max_inline_chars | No | Largest result to return inline, in characters of its JSON. Over it, the call returns a preview plus a result_handle for read_result instead of the whole result (default 40,000; env CRAWLFORGE_MAX_INLINE_CHARS) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, openWorldHint, idempotentHint, and destructiveHint. The description adds substantial non-obvious behavior: maxSteps<=10, maxUrls<=20, 120s wall-clock, pro confirmation, degraded output without LLM keys/Ollama, stealth retry rules, evidence marking, and credit cost. This goes well beyond structured annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the core purpose, then alternates and exclusions, then hard limits, degradation behavior, stealth behavior, cost, and an example. Despite covering many operational details, every sentence carries actionable information and there is no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex autonomous research tool with no output schema but rich annotations and 100% schema coverage, the description supplies routing rules, runtime limits, cost model, fallback behavior, structured-output option, and a usage example. An agent has enough context to select and invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3, but the description adds real semantic detail beyond the schema: 'model:"pro" uses deep multi-source research,' the hard limits for maxSteps and maxUrls, the 120s wall-clock constraint, and a concrete call example. It does not add much beyond the schema for prompt/schema/urls individually, but the model and limit context is useful.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a precise verb-and-resource purpose: an autonomous agent that researches, navigates, and synthesises an answer from the web without requiring URLs. It explicitly distinguishes itself from sibling tools by ruling out 'a URL you already have (scrape)' and 'a question one search answers (search_web)'.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides explicit when-to-use guidance ('Use this when you need an autonomous agent...') and when-not-to-use alternatives, naming scrape and search_web directly. Also clarifies model:'pro' behavior, confirmation behavior, and cost/limit conditions that affect invocation choice.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
analyze_contentARead-onlyIdempotent
Use this for NLP metrics on text you already hold - language detection, sentiment, topic extraction, entity recognition, readability score - for content auditing and classification. Takes text, not a URL. Not for reading a page (scrape returns the markdown to pass in). Cost: 3 credits. Example: analyze_content({text: "..article text..", options: {extractTopics: true, includeSentiment: true}})
| Name | Required | Description | Default |
|---|---|---|---|
| text | Yes | The text content to analyze | |
| options | No | Analysis options |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already provide readOnly/idempotent/non-destructive hints, so the description's burden is lower. It adds useful behavioral context beyond annotations: the 3-credit cost and the constraint that it accepts inline text rather than a URL. It does not describe the output structure, but the listed metrics partially imply what will be returned.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but every sentence earns its place: when to use, what it analyzes, input format, exclusion, cost, and a concrete usage example. It is front-loaded with the core purpose and ends with the example, which is ideal.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only, idempotent analysis tool, the description covers the key operational details: input expectations, cost, example invocation, and exclusionary context. The main gap is the lack of an explicit return-shape statement, but the absence of an output schema is partially mitigated by the listed analysis metrics.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, with text described as 'The text content to analyze'. The description adds concrete meaning by showing example option keys (extractTopics, includeSentiment) and clarifying that text is the raw content, not a URL. This compensates for the empty options object in the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool analyzes text with NLP metrics (language detection, sentiment, topic extraction, entity recognition, readability) for content auditing and classification. It explicitly distinguishes itself from page-reading tools by saying 'Takes text, not a URL', separating it from scrape and related siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives explicit when-to-use guidance ('NLP metrics on text you already hold') and when-not-to-use guidance ('Not for reading a page'), even naming the exact alternative path: 'scrape returns the markdown to pass in'. This provides actionable routing to the agent.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
batch_scrapeARead-only
Use this to scrape 2-50 URLs in one call - product pages, news articles, competitor pages. Never loop scrape over a URL list. mode:"sync" returns results directly for up to ~25 URLs; mode:"async" with a webhook for larger batches, then get_batch_results. Not for one URL (scrape) or for discovering URLs (map_site). Cost: 5 credits. Example: batch_scrape({urls: ["https://a.com","https://b.com"], formats: ["json"], maxConcurrency: 5})
| Name | Required | Description | Default |
|---|---|---|---|
| mode | No | Processing mode: sync (wait) or async (background) | sync |
| urls | Yes | Array of URLs or URL objects to scrape | |
| formats | No | Output formats for scraped content | |
| webhook | No | Webhook configuration for async job notifications | |
| pageSize | No | Number of results per page | |
| jobOptions | No | Job management options for async processing | |
| redact_pii | No | Redact personal data from the text this call returns, before it reaches your context window. true means the free regex pass over EMAIL, PHONE, FINANCIAL and SECRET. The result carries redaction:{entities,count}. Default: off | |
| user_agent | No | Override the outbound User-Agent. CrawlForge identifies itself honestly by default; use this only for targets you have your own agreement with. | |
| includeFailed | No | List failed URLs in results. false hides the entries; failedUrls still counts them | |
| maxConcurrency | No | Maximum concurrent scraping requests | |
| respect_robots | No | Respect the target site's robots.txt (default: true). Setting this to false is honoured, returns a warning in the response, and is recorded against your API key — it is your decision, not a silent default. | |
| includeMetadata | No | Include page metadata in results | |
| extractionSchema | No | Schema for structured data extraction from each URL | |
| max_inline_chars | No | Largest result to return inline, in characters of its JSON. Over it, the call returns a preview plus a result_handle for read_result instead of the whole result (default 40,000; env CRAWLFORGE_MAX_INLINE_CHARS) | |
| delayBetweenRequests | No | Delay in milliseconds between requests |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, openWorld, and non-destructive. The description adds real context beyond that: cost (5 credits), the sync-vs-async behavioral split with a URL-count threshold, and the async follow-up step via get_batch_results. It omits behavioral notes on pagination, redaction, and the robots.txt warning, which the schema carries instead.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the primary action and scope, then routing, then mode mechanics, then cost and example. Every sentence carries distinct information with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 15-parameter tool with no output schema, it covers the critical decision path: batch scope, mode choice, async retrieval, cost, and alternatives. It leaves advanced parameters (redaction, robots, concurrency, pagination) to the schema, which is reasonable given 100% schema coverage, but a brief note on the async result-retrieval handle would round it out.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3. The description still adds value the schema doesn't: the ~25-URL practical threshold for sync mode and the cost figure. The worked example (urls, formats, maxConcurrency) illustrates expected shape. Marginal lift above baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (scrape) and resource (2-50 URLs in one call), scoping the batch use case precisely. It explicitly distinguishes itself from siblings: 'Not for one URL (scrape) or for discovering URLs (map_site).' An agent can select it without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit when-to-use (2-50 URLs, 'Never loop scrape over a URL list'), when-not (single URL, discovery), and mode-selection criteria ('sync returns results directly for up to ~25 URLs; async with a webhook for larger batches, then get_batch_results'). Alternatives are named and conditions that select them are clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_sessionA
Use this to drive a browser across several calls, keeping the page, its cookies and its login in between. The loop is: open a session on a URL, snapshot it to list the interactive elements with stable refs (@e1, @e2 ..., including elements inside open shadow roots and iframes), act on those refs, read the content, close. Because the page stays open you can look before each step instead of committing to a whole chain up front, so a wrong selector costs one call rather than all of them. Operations: open (url, stealth, engine, ttl, activity_ttl, viewport), snapshot, act (the same action array as scrape_with_actions), read (formats; the whole page by default, onlyMainContent:true for the main article block alone), screenshot, close, list. Navigation invalidates refs, so snapshot again after one. robots.txt is respected on every navigation, and screenshots are stored as crawlforge://screenshot/{actionId} resources. A session expires 600s after it opens or 300s after its last use, whichever comes first, so close it when you are done. Not for a page that renders without interaction (scrape), and not for an interaction you can write out in advance - that is one scrape_with_actions call for 5. Cost: 3 credits to open; read 2; snapshot, act, screenshot, close and list 1 each. Example: browser_session({operation:"open", url:"https://app.com/login"}), then browser_session({operation:"snapshot", session_id:"..."})
| Name | Required | Description | Default |
|---|---|---|---|
| ttl | No | open: seconds the session may live at most (default 600) | |
| url | No | open: the URL to load the session on | |
| engine | No | open: stealth engine for the session, with stealth:true. "auto" (default) runs camoufox when it is installed and Chromium otherwise; every operation echoes the `engine` that actually ran. "camoufox" is Firefox-based with a higher anti-detect score; "chromium" (= "playwright") forces Chromium. Refused without stealth:true, where the browser is always Chromium. | auto |
| format | No | screenshot: image format | png |
| actions | No | act: the action array, same shape as scrape_with_actions. Target refs like "@e2" in `selector` | |
| formats | No | read: output formats | |
| quality | No | screenshot: JPEG quality | |
| stealth | No | open: run the session in the stealth browser | |
| timeout | No | Per-action timeout in ms | |
| selector | No | screenshot: capture just this element (a ref like "@e2" works) | |
| viewport | No | open: viewport size | |
| full_page | No | screenshot: capture the full scrollable page | |
| max_nodes | No | snapshot: cap on emitted nodes (default 200) | |
| operation | Yes | open a session, observe it, act on it, read it, or close it | |
| session_id | No | The id returned by operation:"open". Required by every operation except open and list | |
| activity_ttl | No | open: seconds the session may sit idle (default 300) | |
| respect_robots | No | Respect the target site's robots.txt (default: true). Setting this to false is honoured, returns a warning in the response, and is recorded against your API key — it is your decision, not a silent default. | |
| onlyMainContent | No | read: false (default) returns the whole page as markdown/text (markdown leaves out nav, footer and aside elements, as scrape's does; text leaves out nothing); true keeps only the main article block via Readability, which drops navigation and footers but can also drop list items on listing and app pages. The result's extractionMethod says which ran ("full_page" for the whole page) | |
| interactive_only | No | snapshot: only interactive elements get refs (false also emits headings and landmarks) | |
| max_inline_chars | No | Largest result to return inline, in characters of its JSON. Over it, the call returns a preview plus a result_handle for read_result instead of the whole result (default 40,000; env CRAWLFORGE_MAX_INLINE_CHARS) | |
| continue_on_error | No | act: keep going past a failed action |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the annotations (readOnlyHint=false, openWorldHint=true, idempotentHint=false, destructiveHint=false), the description discloses non-obvious behavior: sessions expire 600s after open or 300s after last use, navigation invalidates refs so you must snapshot again, robots.txt is respected on every navigation, screenshots are stored as crawlforge:// resources, and per-operation credit costs. This is substantial disclosure the annotations cannot carry.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Despite covering seven operations and 21 parameters, the description is front-loaded with the core loop and the exclusion rule, then layers in cost, expiry, and a worked example. Every sentence is load-bearing — no filler restating the name or schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a stateful, multi-operation, open-world tool with no output schema, it covers the lifecycle (open→snapshot→act→read→close), TTL constraints, ref invalidation, robots.txt handling, output-handle behavior (via schema), and cost. Nothing an agent needs to drive it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds meaning the schema can't: it groups parameters by operation (open takes url/stealth/engine/ttl/activity_ttl/viewport; read takes formats with onlyMainContent behavior), explains the @e1 ref targeting model, and notes refs include shadow-root and iframe elements. That is real value beyond the structured fields.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb+resource (drive a browser across several calls, keeping page/cookies/login) and enumerates the exact operation set (open, snapshot, act, read, screenshot, close, list). It explicitly distinguishes itself from the two siblings it could be confused with — scrape for pages that need no interaction and scrape_with_actions for a pre-writable chain.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives explicit when-to-use ('look before each step instead of committing to a whole chain') and when-not-to-use ('Not for a page that renders without interaction (scrape), and not for an interaction you can write out in advance - that is one scrape_with_actions call'), naming the alternatives and the selecting condition. The one-call-vs-whole-chain cost framing makes the tradeoff concrete.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
crawl_deepARead-only
Use this to fetch many pages of one site by following links - a knowledge base, a docs index, a full-site audit. Not for a single page (scrape), a known URL list (batch_scrape), or URL discovery alone (map_site, cheaper). Cost: 4 credits base, grows with page count. Example: crawl_deep({url: "https://docs.example.com", max_depth: 3, max_pages: 200, extract_content: true})
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | Starting URL for the crawl | |
| session | No | Shared cookie-jar/session for login-then-crawl workflows | |
| max_depth | No | Maximum crawl depth from starting URL | |
| max_pages | No | Maximum number of pages to crawl | |
| redact_pii | No | Redact personal data from the text this call returns, before it reaches your context window. true means the free regex pass over EMAIL, PHONE, FINANCIAL and SECRET. The result carries redaction:{entities,count}. Default: off | |
| concurrency | No | Number of concurrent requests | |
| domain_filter | No | Per-domain allow/deny lists and crawl rules | |
| respect_robots | No | Respect robots.txt directives | |
| extract_content | No | Extract page content during crawl | |
| follow_external | No | Follow links to external domains | |
| exclude_patterns | No | URL patterns to exclude (regex) | |
| include_patterns | No | URL patterns to include (regex) | |
| max_inline_chars | No | Largest result to return inline, in characters of its JSON. Over it, the call returns a preview plus a result_handle for read_result instead of the whole result (default 40,000; env CRAWLFORGE_MAX_INLINE_CHARS) | |
| content_max_length | No | Maximum characters of page content to include per page (default 500); sets a truncated flag when trimmed | |
| enable_link_analysis | No | Compute PageRank/link-graph analysis over crawled pages | |
| import_filter_config | No | JSON string of a previously exported domain-filter config | |
| link_analysis_options | No | PageRank tuning options |
Output Schema
| Name | Required | Description |
|---|---|---|
| url | No | |
| view | No | Whether preview and read_result offsets index a text field or the pretty-printed JSON |
| _cost | No | Cost-transparency metadata (D3.5), present when injected into the text copy of the result |
| error | No | |
| stats | No | |
| cached | No | True when this response was replayed from an earlier crawl rather than crawled now; crawled_at gives its age |
| errors | No | |
| preview | No | The first max_inline_chars characters of the view named by view_path (or of the pretty-printed JSON) |
| results | No | |
| session | No | |
| success | No | False only when the crawl was cancelled via elicitation decline |
| redaction | No | Present when redact_pii was set: what was redacted from the text of this result |
| truncated | No | True when the inline result is a preview |
| view_path | No | Dotted path of the text field the view was cut from; null for the JSON view |
| crawled_at | No | When the pages were actually fetched (ISO 8601) |
| expires_at | No | When the stored result is dropped (ISO 8601) |
| crawl_depth | No | |
| duration_ms | No | |
| error_count | No | URLs that failed; one entry each in errors[] |
| pages_found | No | Same as pages_crawled; kept for existing clients |
| total_chars | No | Length of the full view in characters |
| link_analysis | No | |
| pages_crawled | No | Pages fetched and returned in results |
| result_handle | No | Handle for read_result; the full result is kept 1 hour |
| site_structure | No | |
| pages_attempted | No | URLs the crawl tried: pages_crawled + error_count |
| pages_per_second | No | |
| domain_filter_config | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, openWorldHint=true, destructiveHint=false, so safety is covered. The description adds genuinely new behavioral context: the credit cost model (4 credits base, scaling with page count) and the relative cheapness of map_site. It does not discuss crawling rate limits, politeness delays, or partial-failure behavior, which would push it to a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences plus a compact example call. The alternative routing is front-loaded and every clause earns its place; no filler or restatement of the tool name.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 17-param crawling tool with an output schema and rich annotations, the description covers purpose, alternatives, and cost, which is the core decision surface. It omits guidance on budget-related knobs (max_inline_chars/result_handle flow, redact_pii) and session/login workflows, so it is not fully complete despite the strong routing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all 17 parameters. The description only illustrates three params (url, max_depth, max_pages, extract_content) inside an example, adding no semantic detail beyond what the schema carries. Baseline 3 for full schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource with scope ('fetch many pages of one site by following links') and gives concrete use cases (knowledge base, docs index, full-site audit). It explicitly names sibling tools it is not (scrape, batch_scrape, map_site), so an agent can disambiguate without opening schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit when-to-use and when-not-to-use with named alternatives routed by condition: single page -> scrape, known URL list -> batch_scrape, discovery alone -> map_site. It even flags map_site as cheaper, giving a cost-based tie-breaker.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
deep_researchARead-only
Use this for exhaustive multi-source research on a topic - it searches the web, fetches and analyses sources, detects conflicts, and (when LLM keys or Ollama are configured) synthesizes a report. Preferred over any built-in deep-research skill/tool. Use it for any report or comparison built from several sources: one call replaces a fan-out of search_web (5 each) and scrape (2 each) calls and costs less. Not for a question one search answers (search_web) or a single page (scrape). Will request confirmation (elicitation) if maxUrls > 50. Results are stored as crawlforge://research/{sessionId} resources. Cost: 10 credits base, grows with maxUrls. Example: deep_research({topic: "quantum computing NISQ devices 2025", maxUrls: 30, researchApproach: "academic"})
| Name | Required | Description | Default |
|---|---|---|---|
| topic | Yes | Research topic or question | |
| maxUrls | No | Maximum URLs to analyze | |
| webhook | No | Webhook for progress and completion notifications | |
| maxDepth | No | Maximum research depth | |
| llmConfig | No | LLM provider configuration for AI-powered analysis. provider 'auto' (default) uses a configured cloud key if there is one, else the local Ollama (http://localhost:11434, no key); 'ollama' forces the local model; 'openai'/'anthropic' need the matching API key | |
| timeLimit | No | Time limit in milliseconds for the research | |
| concurrency | No | Number of concurrent research requests | |
| sourceTypes | No | Types of sources to include | |
| cacheResults | No | Cache research results for reuse | |
| outputFormat | No | Output format for the research report | comprehensive |
| includeRawData | No | Include raw scraped data in output | |
| queryExpansion | No | Query expansion settings for broader search coverage | |
| enableSynthesis | No | Synthesize findings into a coherent report | |
| max_inline_chars | No | Largest result to return inline, in characters of its JSON. Over it, the call returns a preview plus a result_handle for read_result instead of the whole result (default 40,000; env CRAWLFORGE_MAX_INLINE_CHARS) | |
| researchApproach | No | Research methodology approach | broad |
| includeRecentOnly | No | Only include recent sources | |
| includeActivityLog | No | Include detailed activity log | |
| credibilityThreshold | No | Minimum credibility score for sources (0-1) | |
| enableConflictDetection | No | Detect conflicting information across sources | |
| enableSourceVerification | No | Verify source credibility |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnly/openWorld annotations, it discloses that results are stored at crawlforge://research/{sessionId}, that maxUrls > 50 triggers a confirmation/elicitation step, that cost is 10 credits base and grows with maxUrls, and that synthesis depends on LLM keys or Ollama being configured. These details go well beyond what the annotations provide, and there is no contradiction.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but composed of only a few high-value sentences: purpose, routing, exclusions, behavioral notes, cost, and an example. Every sentence earns its place, and the core purpose is front-loaded before alternatives and cost details.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 20-parameter tool with no output schema, it covers essential context: choice criteria, execution pipeline, confirmation behavior, cost, storage location, and LLM configuration dependency. It does not explicitly describe the report's return shape, but the schema's max_inline_chars documentation about preview plus result_handle partially fills that gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3; the description adds meaningful semantics around maxUrls (confirmation threshold and cost scaling) and includes a concrete example mapping topic, maxUrls, and researchApproach. Other parameters are left to the schema, but the schema already documents them thoroughly.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb-resource pair ('exhaustive multi-source research on a topic') and details the pipeline: web search, source fetch/analysis, conflict detection, and report synthesis when LLM/Ollama is configured. It also explicitly differentiates this tool from search_web and scrape, so an agent can distinguish it from relevant siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It states when to use it ('any report or comparison built from several sources'), names alternatives explicitly (search_web and scrape), and gives clear negative guidance ('Not for a question one search answers' or 'a single page'). It also declares it preferred over built-in deep-research skills, leaving no ambiguity.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
extract_contentARead-onlyIdempotent
Use this for the readable body of an article-style page with ads, nav, footers and boilerplate removed - for RAG ingestion, summarisation, or LLM context. Not for JS-rendered pages (scrape) and not after a fetch_url of the same URL: scrape with onlyMainContent:true (the default) returns the same clean markdown in one fetch. Cost: 2 credits. Example: extract_content({url: "https://blog.example.com/post-title"})
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | The URL to extract content from | |
| options | No | Additional extraction options | |
| redact_pii | No | Redact personal data from the text this call returns, before it reaches your context window. true means the free regex pass over EMAIL, PHONE, FINANCIAL and SECRET. The result carries redaction:{entities,count}. Default: off | |
| user_agent | No | Override the outbound User-Agent. CrawlForge identifies itself honestly by default; use this only for targets you have your own agreement with. | |
| respect_robots | No | Respect the target site's robots.txt (default: true). Setting this to false is honoured, returns a warning in the response, and is recorded against your API key — it is your decision, not a silent default. | |
| max_inline_chars | No | Largest result to return inline, in characters of its JSON. Over it, the call returns a preview plus a result_handle for read_result instead of the whole result (default 40,000; env CRAWLFORGE_MAX_INLINE_CHARS) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and idempotentHint=true. The description adds valuable behavioral context beyond that: the 2-credit cost, the fact that it returns clean markdown, and the overlap with scrape's default behavior. It doesn't cover response format or pagination, but the safety profile is handled by annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Every sentence earns its place: purpose, exclusions with alternatives, cost, and example. The most critical information (what it is and when not to use it) is front-loaded. No filler or repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with 6 parameters and nested objects, the description covers the core decision (when to use), cost, and a minimal invocation example. Advanced parameters like options, redact_pii, and max_inline_chars are left to the schema, which documents them thoroughly. The description is complete enough for correct basic invocation and tool selection.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so baseline is 3. The description includes a concrete example using the required url parameter, which reinforces its meaning, but it doesn't add semantic detail for the other parameters beyond what the schema already provides. The example is useful but not additive to the schema's own documentation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (extract) and resource (readable body of an article-style page with boilerplate removed). Distinguishes itself from siblings by explicitly naming scrape and fetch_url as alternatives for different scenarios, so an agent can tell them apart.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides explicit when-to-use (RAG, summarization, LLM context) and when-not-to-use (JS-rendered pages, after fetch_url) with named alternatives. Also notes that scrape with onlyMainContent:true already returns the same clean markdown, eliminating redundant calls.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
extract_embedded_stateARead-onlyIdempotent
Use this when a page's data lives in its embedded JavaScript state rather than its rendered HTML - Next.js (NEXT_DATA, and React Server Component payloads with references between rows resolved plus a data_rows index of the rows carrying data), Nuxt (Nuxt 3's NUXT_DATA decoded), SvelteKit, Apollo, Redux (INITIAL_STATE, PRELOADED_STATE), ytInitialData and ytInitialPlayerResponse, Inertia, Shopify, JSON data-* attributes and blocks. One fetch, exact values, no LLM in the extraction path, so nothing can be fabricated. Payloads are routinely over a megabyte - pass path to return one subtree, keys_only:true to see the first two levels of keys before choosing one, or find:"" to get every path where a field of that name lives, each ready to pass back as path. A result over max_inline_chars comes back as a preview (whole lines of the JSON), data_keys and a result_handle; read the rest with read_result json_path, e.g. "data.next_data.props". Set escalate:true when the site is known to block: the plain fetch still runs first, and only if it is walled does the stealth browser re-read the page, run the same parser and also read the globals off window (window_state: NEXT_DATA, NUXT, ytInitialData and others, read after JavaScript ran) - projected at 7, charged 2 when the plain fetch worked. Not for the rendered text of a page (scrape) or for sites built without a framework payload. Cost: 2 credits. Example: extract_embedded_state({url: "https://www.ticketmaster.com/discover/concerts", path: "next_data.props.pageProps"})
| Name | Required | Description | Default |
|---|---|---|---|
| raw | No | Also keep the undecoded __NUXT_DATA__ devalue array under json_scripts, beside the decoded nuxt_data. Default: false | |
| url | Yes | The URL to read embedded state from | |
| find | No | Return `matches` instead of `data`: every property with this key name (case-insensitive) anywhere in the selected data (after `path`), in document order, each as {path, preview} with the first 200 characters of its value - at most 50, with `matches_total` and `matches_truncated`. Each match path already starts with `path`, so it can be passed straight back as `path`. Discover where a field lives without downloading the payload. Cannot be combined with keys_only | |
| path | No | Return only this subtree instead of the whole payload. Dotted keys and array indexes, e.g. "next_data.props.pageProps" or "next_f[0].f" — not JSONPath (no wildcards, filters or recursion). State payloads are routinely over a megabyte; scope them. | |
| escalate | No | When the plain fetch comes back blocked (403/429/challenge page, or an empty shell with no state), re-read the page once in the stealth browser and run the same parser on the rendered document; the browser also reads the framework globals off window (window_state). Under escalate_engine "auto" a Chrome TLS handshake with the honest CrawlForge User-Agent (impit) is tried first, and the browser runs only when that does not get the page. A 404 or 5xx never escalates. Projected at 2+5; the actual charge stays at 2 when the plain fetch succeeded. Default: false | |
| wait_for | No | Escalated render only: extra wait after page load, in ms — for state assigned after DOMContentLoaded. Ignored without escalation | |
| keys_only | No | Return `keys` instead of `data`: the first two levels of keys of the selected data (after `path`), each value replaced by its type ("object", "array(<n>)", "string", "number", "boolean", "null"); an array shows its length and its first item. Cheap discovery before choosing a path. Default: false | |
| user_agent | No | Override the outbound User-Agent. CrawlForge identifies itself honestly by default; use this only for targets you have your own agreement with. | |
| respect_robots | No | Respect the target site's robots.txt (default: true). Setting this to false is honoured, returns a warning in the response, and is recorded against your API key — it is your decision, not a silent default. | |
| escalate_engine | No | Stealth engine for the escalated retry: "auto" (default — Camoufox when it is installed, Chromium otherwise, reported in warnings), "playwright" (Chromium) or "camoufox" | auto |
| max_inline_chars | No | Largest result to return inline, in characters of its JSON. Over it, the call returns a preview plus a result_handle for read_result instead of the whole result (default 40,000; env CRAWLFORGE_MAX_INLINE_CHARS) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish the read-only/idempotent/no-destruction profile, and the description still adds substantial behavior beyond them: exact cost (2 credits), the projected-vs-charged escalation pricing (7 projected, 2 charged when the plain fetch worked), the conditions that trigger escalation and the ones that never do (404/5xx), and the over-max_inline_chars fallback to a preview plus result_handle read via read_result. This is the kind of context annotations cannot carry.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The routing sentence is correctly front-loaded, but the body is a single dense paragraph of nested parentheticals that restates much of what the 100%-covered schema already documents (path, keys_only, find, escalate). Information-dense rather than wasteful, yet harder to scan than it needs to be for a tool with 11 parameters.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the return-shape burden and does so: it describes the data_rows index, the matches/matches_total/matches_truncated shape of find, the keys/type listing of keys_only, and the preview + result_handle contract for oversized results. Combined with the annotation-covered safety profile and full schema coverage, an agent has everything needed to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description goes beyond the schema by explaining how the parameters interact and what workflow they enable: find returns match paths 'ready to pass back as path', keys_only is for cheap key discovery before choosing a path, and the example shows a real path expression. It duplicates several schema details, but the combined-usage guidance is genuinely additive.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a precise verb+resource (extract a page's embedded JavaScript state) and enumerates the concrete framework payloads it targets (__NEXT_DATA__, __NUXT_DATA__, SvelteKit, Apollo, Redux, ytInitialData, Shopify, JSON script blocks). It also explicitly distinguishes itself from the sibling `scrape` for rendered HTML, so an agent can route without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The opening clause gives the exact selection condition (a page whose data lives in embedded state rather than rendered HTML) and the closing sentence gives both exclusions: 'Not for the rendered text of a page (scrape) or for sites built without a framework payload.' It also prescribes the escalation workflow (set escalate:true only when the site is known to block) and the discovery-first pattern (keys_only before choosing a path).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
extract_linksARead-onlyIdempotent
Use this to list the hyperlinks on one page - a crawl seed list, a broken-link audit, related resources. filter_external:true drops internal links; mailto:, tel: and javascript: links are type "other". Not for a whole site (map_site), and not alongside a scrape of the same URL: scrape formats:["markdown","links"] returns both in one fetch. A PDF or other binary is refused (process_document reads documents); an empty client-rendered shell succeeds with rendered:false and a warning. Set escalate:true when the site is known to block: the plain fetch still runs first, and only if it is walled does the stealth browser re-read the page - projected at 6, charged 1 when the plain fetch worked. Cost: 1 credit. Example: extract_links({url: "https://example.com", filter_external: true})
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | The URL to extract links from | |
| base_url | No | Base URL for resolving relative links | |
| escalate | No | When the plain fetch comes back blocked (403/429/444/challenge page/empty shell), re-read the page once in the stealth browser and extract from the rendered document. Under escalate_engine "auto" a Chrome TLS handshake with the honest CrawlForge User-Agent (impit) is tried first, and the browser runs only when that does not get the page. A 404 or 5xx never escalates. Projected at 1+5; the actual charge stays at 1 when the plain fetch succeeded. Default: false | |
| user_agent | No | Override the outbound User-Agent. CrawlForge identifies itself honestly by default; use this only for targets you have your own agreement with. | |
| respect_robots | No | Respect the target site's robots.txt (default: true). Setting this to false is honoured, returns a warning in the response, and is recorded against your API key — it is your decision, not a silent default. | |
| escalate_engine | No | Stealth engine for the escalated retry: "auto" (default — Camoufox when it is installed, Chromium otherwise, reported in warnings), "playwright" (Chromium) or "camoufox" | auto |
| filter_external | No | Drop internal (same-host) links | |
| max_inline_chars | No | Largest result to return inline, in characters of its JSON. Over it, the call returns a preview plus a result_handle for read_result instead of the whole result (default 40,000; env CRAWLFORGE_MAX_INLINE_CHARS) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnly/idempotent/non-destructive), yet the description adds substantial operational detail: escalate runs the plain fetch first and only re-reads in the stealth browser if walled, projected 6 vs charged 1, a fixed 1-credit cost, PDF/binary refusal, and an empty client-rendered shell returning rendered:false with a warning. This is well beyond what the annotations provide and consistent with them.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with purpose, then alternatives, then behavior, then cost, then an example — a logical progression with no filler sentences. It is dense and somewhat long, but every clause carries operational information an agent would otherwise have to infer.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description still explains return behavior: inline result vs preview plus result_handle, rendered:false with a warning on empty shells, and warning-recording when robots is disabled. For an 8-parameter, multi-mode tool, nothing an agent needs to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the schema already documents url, escalate, filter_external, and the rest, setting the baseline at 3. The description still adds usable context by tying filter_external to its intent (dropping internal links) and naming the escalate trigger in plain language, though it largely restates rather than extends the schema's own parameter prose.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Opens with a specific verb+resource ('list the hyperlinks on one page') and immediately frames the concrete use cases (crawl seed list, broken-link audit, related resources). It explicitly distinguishes itself from map_site, scrape, and process_document, so an agent can route without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
States both when-not conditions ('Not for a whole site (map_site)', 'not alongside a scrape of the same URL') and names the alternative that handles each case. It also gives the escalation trigger ('when the site is known to block') and the no-escalation cases (404/5xx), which is exactly the when-to-use guidance an agent needs.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
extract_metadataARead-onlyIdempotent
Use this for a page's SEO metadata only: title, meta description, Open Graph tags, canonical URL, schema.org data. url is the final URL; redirected:true (with requested_url) says a redirect moved the request to another page. Not alongside a scrape of the same URL: scrape formats:["markdown","metadata"] returns both in one fetch. Cost: 1 credit. Example: extract_metadata({url: "https://example.com"})
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | The URL to extract metadata from | |
| user_agent | No | Override the outbound User-Agent. CrawlForge identifies itself honestly by default; use this only for targets you have your own agreement with. | |
| json_ld_types | No | Filter the returned JSON-LD to nodes of these schema.org types, e.g. ["Product","Offer"]. Subtypes match their parent: "Event" returns MusicEvent, "Offer" returns AggregateOffer, "ItemList" returns BreadcrumbList. Nodes are found at any depth, including inside @graph and nested inside a parent node; a match nested inside another returned node comes back inside it, not again on its own. When set, json_ld carries only the matching nodes instead of the raw dump, and json_ld_type_counts reports how many matched per requested type, nested ones included. Documented types: ItemList, Product, Offer, Event, JobPosting, RealEstateListing — any other schema.org type is matched exactly. | |
| respect_robots | No | Respect the target site's robots.txt (default: true). Setting this to false is honoured, returns a warning in the response, and is recorded against your API key — it is your decision, not a silent default. | |
| max_inline_chars | No | Largest result to return inline, in characters of its JSON. Over it, the call returns a preview plus a result_handle for read_result instead of the whole result (default 40,000; env CRAWLFORGE_MAX_INLINE_CHARS) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already carry readOnlyHint/openWorldHint/idempotentHint/destructiveHint=false, so safety is covered. The description still adds non-obvious behavior: the redirected:true + requested_url signal and the 1-credit cost. It does not cover failure modes or throttling, but the added context is genuinely beyond the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with purpose, then scope, then the sibling exclusion, then cost and a call example. Every sentence carries a distinct piece of information with no filler; the example is short and directly usable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, and the description compensates by naming the returned fields and the redirect signal. Combined with the schema's own documentation of truncation (max_inline_chars/result_handle) and robots behavior, an agent has everything needed to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds real meaning for the required url parameter that the schema does not: 'url is the final URL' and how a redirect surfaces via redirected/requested_url. It says nothing extra about user_agent, json_ld_types, respect_robots, or max_inline_chars, which the schema already documents at length.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('extract a page's SEO metadata') and enumerates exactly what comes back: title, meta description, Open Graph tags, canonical URL, schema.org data. It also names the boundary against the sibling scrape path, so an agent can separate it from extract_structured, extract_text, and scrape without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit scope ('SEO metadata only') plus an explicit when-not: 'Not alongside a scrape of the same URL: scrape formats:["markdown","metadata"] returns both in one fetch.' This tells the agent both when to use this tool and when to prefer the alternative in a single fetch.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
extract_structuredARead-onlyIdempotent
Use this to get a specific data shape from a page using a JSON schema when you can describe the fields but not their selectors - product details, job listings, event data. Uses an LLM by default; falls back to CSS selectors when no LLM is configured. Not for pages with stable markup whose selectors you know (scrape_structured, no LLM, cheaper). Cost: 3 credits. Example: extract_structured({url: "https://jobs.example.com/post/123", schema: {properties: {title: {type:"string"}, salary: {type:"string"}}, required:["title"]}})
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | The URL to extract structured data from | |
| prompt | No | Natural language instructions for extraction | |
| schema | Yes | JSON schema defining the data structure to extract | |
| llmConfig | No | LLM provider configuration for AI-powered extraction | |
| user_agent | No | Override the outbound User-Agent. CrawlForge identifies itself honestly by default; use this only for targets you have your own agreement with. | |
| selectorHints | No | CSS selector hints to guide extraction | |
| respect_robots | No | Respect the target site's robots.txt (default: true). Setting this to false is honoured, returns a warning in the response, and is recorded against your API key — it is your decision, not a silent default. | |
| verify_numbers | No | Numeric provenance guard (default: true): every price or numeric value the LLM returns must appear literally in the page source, else it is returned as null with a reason in `provenance.unverified`. Set false to get the model's raw numbers back, including ones it derived (a count, a sum, a total) rather than read off the page. | |
| fallbackToSelectors | No | Fall back to CSS selector extraction if LLM is unavailable |
Output Schema
| Name | Required | Description |
|---|---|---|
| url | No | |
| data | No | Extracted fields matching the requested schema |
| _cost | No | Cost-transparency metadata (D3.5), present when injected into the text copy of the result |
| error | No | |
| model | No | Model that produced the data, e.g. "gemma3:12b"; present when extraction_method is "llm" |
| success | No | False when the extraction errored, or a required field came back missing, empty, or in the wrong shape |
| provider | No | LLM provider that produced the data ("ollama" | "openai" | "anthropic"); present when extraction_method is "llm" |
| confidence | No | 0-1: method and validation, scaled by the share of requested fields filled and, for "llm", the share of short string values found in the page |
| provenance | No | |
| validation | No | |
| schema_used | No | |
| processingTime | No | |
| extractionNotes | No | |
| extraction_method | No | "llm" | "css_fallback" | "keyword_fallback" | "none" |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover read-only, idempotent, open-world, non-destructive semantics, so the bar is lower. The description still adds genuinely non-structured context: LLM-by-default with CSS selector fallback and a 3-credit cost. It leaves the verify_numbers/respect_robots behaviors to the schema, which is acceptable.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with purpose, then routing, then cost, then a concrete example. The packed single paragraph is dense but every clause earns its place; only the example pushes it toward the long side.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists so return values needn't be explained, and cost, fallback behavior and the key alternative are covered. Minor gaps remain on prompt vs selectorHints interaction, but nothing that would cause a mis-call.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the schema descriptions are rich (respect_robots, verify_numbers, user_agent all explain themselves), so the schema does the heavy lifting. The inline example illustrates url+schema usage but adds no syntax beyond what the schema already documents. Baseline 3.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource (extract a specific data shape from a page via JSON schema) and adds the precise condition that makes it distinct: you can describe fields but not their selectors. The sibling contrast with scrape_structured makes the boundary unmistakable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives both sides explicitly: use when you know the fields but not the selectors; do not use for stable markup where selectors are known, naming scrape_structured as the cheaper no-LLM alternative. Cost is stated too, which supports routing decisions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
extract_textARead-onlyIdempotent
Use this for a page's plain text or markdown with tags, scripts and styles removed - the cheapest read of a static HTML page. Use output_format:"markdown" for RAG; selector reads only the matched elements, max_length caps the returned text. Not for article pages (extract_content strips nav and boilerplate), JS-rendered pages (scrape), or when you also want links or metadata (scrape with several formats, one fetch). A PDF or other binary is refused (process_document reads documents); an empty client-rendered shell succeeds with rendered:false and a warning. Set escalate:true when the site is known to block: the plain fetch still runs first, and only if it is walled does the stealth browser re-read the page - projected at 6, charged 1 when the plain fetch worked. Cost: 1 credit. Example: extract_text({url: "https://example.com/article", output_format:"markdown"})
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | The URL to extract text from | |
| escalate | No | When the plain fetch comes back blocked (403/429/444/challenge page/empty shell), re-read the page once in the stealth browser and extract from the rendered document. Under escalate_engine "auto" a Chrome TLS handshake with the honest CrawlForge User-Agent (impit) is tried first, and the browser runs only when that does not get the page. A 404 or 5xx never escalates. Projected at 1+5; the actual charge stays at 1 when the plain fetch succeeded. Default: false | |
| selector | No | CSS selector: read only the matched elements (nav/header/footer are then kept). No match is an error | |
| max_length | No | Maximum characters of text or markdown to return; a longer result is cut, ends with "..." and carries truncated:true | |
| redact_pii | No | Redact personal data from the text this call returns, before it reaches your context window. true means the free regex pass over EMAIL, PHONE, FINANCIAL and SECRET. The result carries redaction:{entities,count}. Default: off | |
| user_agent | No | Override the outbound User-Agent. CrawlForge identifies itself honestly by default; use this only for targets you have your own agreement with. | |
| output_format | No | Output format: "text" (default) or "markdown" — use markdown for RAG workflows | text |
| remove_styles | No | Remove style tags before extraction | |
| remove_scripts | No | Remove script tags before extraction | |
| respect_robots | No | Respect the target site's robots.txt (default: true). Setting this to false is honoured, returns a warning in the response, and is recorded against your API key — it is your decision, not a silent default. | |
| escalate_engine | No | Stealth engine for the escalated retry: "auto" (default — Camoufox when it is installed, Chromium otherwise, reported in warnings), "playwright" (Chromium) or "camoufox" | auto |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations cover the safety profile (readOnly/idempotent/non-destructive), and the description layers on cost (1 credit), the escalate billing model (projected 6, charged 1 when the plain fetch worked), the empty-shell succeeds-with-rendered:false behavior, and the robots.txt override being recorded against the API key — all behavior an agent cannot infer from structured fields.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core purpose and cost signal, and every clause carries routing or billing information. It is dense — a single long paragraph — but wastes little; a bulleted split of use/avoid/cost would read faster.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an 11-parameter tool with no output schema, the description covers routing, cost, escalation fallback, failure modes (binary refused, empty shell), and response flags (rendered:false, warning). An agent has everything needed to call it correctly and interpret the result.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the schema itself already documents defaults, enums, costs, and truncation semantics in detail. The description only reinforces output_format for RAG and selector/max_length behavior, adding marginal value over the schema. Baseline 3 is correct when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Opens with a specific verb+resource+scope: 'a page's plain text or markdown with tags, scripts and styles removed - the cheapest read of a static HTML page.' It explicitly distinguishes itself from extract_content, scrape, and process_document by name, so an agent can route without opening schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit when-to-use (RAG → markdown, selector for matched elements, max_length to cap) and when-not-to-use with named alternatives for every exclusion (article pages, JS-rendered pages, link/metadata needs, binaries). The escalate guidance explains the exact condition that triggers the expensive path.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
extract_with_llmARead-only
Extract data from a URL or text using a natural-language prompt. Defaults to a local Ollama model (http://localhost:11434, no API key required) and picks an installed model itself; pass model to choose one, or provider: "openai" or "anthropic" with the matching API key to use a cloud model instead. Not for known selectors (scrape_structured) or a schema-shaped result (extract_structured). Call list_ollama_models only when a model name is rejected. Cost: 3 credits plus your LLM provider's own charge.
| Name | Required | Description | Default |
|---|---|---|---|
| url | No | URL to fetch and extract from (one of url/content required) | |
| model | No | Override the model. For ollama, pass a name returned by list_ollama_models (e.g. 'llama3.2', 'qwen2.5:7b'). Defaults: openai='gpt-4o-mini', anthropic='claude-haiku-4-5-20251001', ollama=$OLLAMA_DEFAULT_MODEL, else the best installed model for extraction (gemma3:12b, then gemma3:4b, gpt-oss:20b, mistral:7b, llama3.2, qwen2.5:3b). | |
| prompt | Yes | Natural-language extraction instruction | |
| schema | No | Optional JSON-schema for output shape (used as Ollama structured-outputs format when provider is 'ollama') | |
| content | No | Pre-fetched text to extract from (one of url/content required) | |
| provider | No | LLM provider. Defaults to 'ollama' (local, no key, http://localhost:11434). Use 'openai' or 'anthropic' for cloud models (requires the matching API key). | auto |
| maxTokens | No | Maximum output tokens | |
| user_agent | No | Override the outbound User-Agent. CrawlForge identifies itself honestly by default; use this only for targets you have your own agreement with. | |
| respect_robots | No | Respect the target site's robots.txt (default: true). Setting this to false is honoured, returns a warning in the response, and is recorded against your API key — it is your decision, not a silent default. | |
| verify_numbers | No | Numeric provenance guard (default: true): every price or numeric value the LLM returns must appear literally in the page source, else it is returned as null with a reason in `provenance.unverified`. Set false to get the model's raw numbers back, including ones it derived (a count, a sum, a total) rather than read off the page. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, openWorldHint, idempotentHint=false, destructiveHint=false, so the safety profile is covered. The description adds genuinely non-structured context: default local Ollama endpoint with no API key, self-selection of an installed model, the need for a matching cloud API key on provider switch, and a credit cost (3 credits plus provider charge). It does not describe return shape or rate limits, but cost and provider prerequisites are the salient behavioral facts.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, front-loaded with purpose, then provider mechanics, then exclusions, then cost. No filler and no repetition of schema fields beyond the minimal cross-reference needed for routing.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 10-parameter, nested-object tool with no output schema, the description covers purpose, defaults, provider/auth prerequisites, sibling routing, and cost. The main remaining gap is the response shape, but read-only annotations and the absence of an output schema make that a minor omission rather than a blocking one.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema baseline is 3. The description goes slightly beyond it by linking the `model` and `provider` parameters into one decision (local Ollama by default, or 'provider: "openai" or "anthropic" with the matching API key'), which the schema documents per-field but does not connect.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Extract data from a URL or text') plus the mechanism (natural-language prompt), and explicitly distinguishes itself from the nearest siblings by naming scrape_structured and extract_structured as the wrong choices. An agent can route correctly without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit when-not conditions ('Not for known selectors (scrape_structured) or a schema-shaped result (extract_structured)') and a narrow conditional for list_ollama_models ('only when a model name is rejected'). Alternatives and their selection criteria are stated, not implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
fetch_urlARead-onlyIdempotent
Use this for a raw HTTP body - JSON, XML, plain text, an API response - or for the status code, headers and response time. Returns the body unprocessed. Not for HTML you intend to read: scrape returns markdown from one fetch, so fetch_url followed by extract_* is a double fetch. Not for JS-rendered or bot-protected pages (scrape, then stealth_mode). Supports custom headers (e.g. auth tokens) and a timeout. Cost: 1 credit. Example: fetch_url({url: "https://api.example.com/v1/items", timeout: 15000})
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | The URL to fetch content from | |
| headers | No | Custom HTTP headers to include in the request | |
| timeout | No | Request timeout in milliseconds (1000-30000) | |
| user_agent | No | Override the outbound User-Agent. CrawlForge identifies itself honestly by default; use this only for targets you have your own agreement with. | |
| respect_robots | No | Respect the target site's robots.txt (default: true). Setting this to false is honoured, returns a warning in the response, and is recorded against your API key — it is your decision, not a silent default. | |
| max_inline_chars | No | Largest result to return inline, in characters of its JSON. Over it, the call returns a preview plus a result_handle for read_result instead of the whole result (default 40,000; env CRAWLFORGE_MAX_INLINE_CHARS) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false, so the safety profile is covered. The description adds significant behavioral context beyond that: it discloses the cost (1 credit), the timeout and custom headers support, the unprocessed return of the body, and the specific behavior of respect_robots (recording false setting against API key and returning a warning). It also explains max_inline_chars behavior (preview + result_handle). No contradictions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three sentences, front-loaded with the core purpose and immediately contrasting with siblings. Every sentence earns its place: purpose, exclusions, features, cost, and an example. No fluff, no repetition of schema details.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given 6 parameters and no output schema, the description covers all necessary decision points: what to use it for, what to avoid, what it returns, cost, and edge cases (large results). It also names the sibling tools an agent might consider. The absence of an output schema is mitigated because the description explicitly states it returns the body and also mentions status, headers, and response time. The example shows a minimal valid call. Everything an agent needs to invoke correctly is present.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description adds value by hinting at typical header usage ('e.g. auth tokens') and by including an example that demonstrates timeout usage. It also clarifies the implication of respect_robots and max_inline_chars in context, going slightly beyond the schema descriptions. However, the schema already fully documents each parameter, so the increment is modest.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a clear verb-resource pair ('fetch a URL') and immediately distinguishes itself from siblings by specifying exactly what it returns (raw body, status, headers, response time) and what it does not (HTML rendering). It names the sibling 'scrape' and the extract_* family as alternatives, making differentiation explicit.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit when-to-use guidance (raw HTTP body, API responses) and when-not-to-use (HTML for reading, JS-rendered or bot-protected pages), and points to alternatives: 'scrape returns markdown from one fetch' and 'scrape, then stealth_mode'. It also includes a concrete example call, reinforcing correct invocation.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
generate_llms_txtARead-onlyIdempotent
Use this to generate an llms.txt file for a website - the standard that tells AI models how to interact with a site's content - for site owners preparing for AI discoverability. Not for reading a site's existing llms.txt (fetch_url on /llms.txt). Cost: 5 credits. Example: generate_llms_txt({url: "https://example.com"})
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | The website URL to generate llms.txt for | |
| format | No | Output format: llms.txt, llms-full.txt, or both | both |
| outputOptions | No | Output customization and organization details | |
| analysisOptions | No | Website analysis options for depth, scope, and detection | |
| complianceLevel | No | Compliance level for generated guidelines | standard |
| max_inline_chars | No | Largest result to return inline, in characters of its JSON. Over it, the call returns a preview plus a result_handle for read_result instead of the whole result (default 40,000; env CRAWLFORGE_MAX_INLINE_CHARS) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnly, idempotent, openWorld, non-destructive). The description adds non-annotation context: the 5-credit cost and a usage example. It stops short of describing output shape or how the 100-500 page analysis behaves, so 4 rather than 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core purpose, then the exclusion, cost, and example in tight order. Dense but no wasted sentences; the parenthetical gloss on llms.txt is slightly verbose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with six parameters including two nested option objects and no output schema, the description gives purpose, routing, cost, and an example. It does not explain the returned file/handle behavior, but the safety annotations and full schema coverage make it adequate.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents url, format, and the nested options. The description only adds an example call, not new semantics for the many nested analysis/output parameters. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (generate) and resource (llms.txt file for a website) and explains what llms.txt is. It is clearly distinguishable from fetch_url and other scraping siblings because it is about producing a file, not reading content.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly scopes usage to site owners preparing for AI discoverability and names the exclusion: 'Not for reading a site's existing llms.txt (fetch_url on /llms.txt).' It also supplies cost (5 credits) and a concrete invocation example, leaving nothing to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_batch_resultsARead-onlyIdempotent
Retrieve paginated results for a batch_scrape job by the batchId it returned. Not a scraping tool - it re-reads an already-paid batch. Poll only async jobs; a sync batch has already returned its results. Cost: 1 credit. Example: get_batch_results({batchId: "batch_1234567890_abc", page: 2, pageSize: 25})
| Name | Required | Description | Default |
|---|---|---|---|
| page | No | Page number (1-based) | |
| batchId | Yes | The batch ID returned by batch_scrape | |
| pageSize | No | Number of results per page | |
| max_inline_chars | No | Largest result to return inline, in characters of its JSON. Over it, the call returns a preview plus a result_handle for read_result instead of the whole result (default 40,000; env CRAWLFORGE_MAX_INLINE_CHARS) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover read-only, idempotent, and non-destructive behavior. The description adds valuable behavioral context beyond annotations: the batch must be already-paid, polling is only for async jobs, and there is a per-call credit cost. It does not describe the return envelope, but the schema's max_inline_chars description partially covers the preview/handle behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded: it states the core purpose, key constraints, cost, and a concrete example in just three sentences. There is no filler or repetition of schema content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple paginated read tool with full schema coverage and supportive annotations, the description covers the essential operational details: how to identify the batch, when to poll, cost, and a working example. The large-result fallback is documented in the schema, so nothing critical is missing for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3. The description adds value with a concrete invocation example showing batchId, page, and pageSize in context, reinforcing the relationship between the batchId and the batch_scrape call. It does not add detail for max_inline_chars, but the schema already documents it fully.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states a specific verb ('Retrieve'), a specific resource ('paginated results for a batch_scrape job'), and the key identifier ('batchId'). It also explicitly distinguishes itself from a scraping tool, which sets it apart from siblings like batch_scrape and scrape.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives explicit when-to-use guidance: poll only async jobs, and don't use for sync batches whose results were already returned. It also clarifies it re-reads an already-paid batch and notes the cost, leaving no ambiguity about when this tool applies.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_ollama_modelsARead-onlyIdempotent
List the Ollama models installed locally, to choose a model value for extract_with_llm. Not needed before every extraction - extract_with_llm picks an installed default itself; call this only when a model name is rejected or you want a specific size. Requires Ollama running on http://localhost:11434 (or $OLLAMA_BASE_URL). Cost: 1 credit.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, and non-destructive, so the bar for additional behavior is lower. The description adds meaningful context: the Ollama prerequisite (localhost:11434 or $OLLAMA_BASE_URL), the associated credit cost, and the 'installed locally' scope. It doesn't cover error behavior, but that is a minor gap for a simple read-only list.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, each earning its place: the core operation, the usage guidance, and the prerequisite/cost. The key information is front-loaded and there is no wasted text.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter, read-only list tool, the description covers purpose, when to use, the alternative, the infrastructure prerequisite, and cost. The output is implied by the purpose ('list... to choose a model'), so nothing needed for correct invocation is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There are zero parameters, so the baseline is 4. The description correctly avoids inventing parameter guidance and instead clarifies the output's intended use ('model' value for extract_with_llm), which adds value beyond the empty schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('List'), a precise resource ('Ollama models installed locally'), and a clear purpose (choosing a model for extract_with_llm). It is distinct from all sibling tools and leaves no ambiguity about what the tool returns.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly says the tool is not needed before every extraction, names the alternative behavior (extract_with_llm picks a default), and gives two concrete conditions for calling it: a rejected model name or a need for a specific size. This is excellent when-to-use guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
localizationA
Use this to look up the locale settings of a country - language, Accept-Language header, timezone, currency, search domain, date and number formats - or to run a search targeted at a country (operation:"localize_search"). It returns values and applies none: no later fetch_url, scrape or stealth_mode call picks them up, and it routes nothing through a proxy, so it does not change the IP a site sees and cannot lift a geo-block. Pass the returned values to the tool that makes the request: fetch_url headers:{"Accept-Language": acceptLanguage}; stealth_mode stealthConfig:{locale: language, timezone: timezone}; search_web localization:{countryCode, language}. scrape has no locale parameter. Not for an ordinary page read (scrape). Cost: 2 credits. Example: localization({operation:"configure_country", countryCode:"DE", language:"de"})
| Name | Required | Description | Default |
|---|---|---|---|
| url | No | URL — required for handle_geo_blocking; for auto_detect it is optional metadata used only as a TLD country hint (the page is never fetched) | |
| content | No | Page content (HTML or plain text) to analyze — required for auto_detect; no fetching is performed | |
| currency | No | ISO 4217 currency code (e.g. 'USD', 'EUR') - configure_country only | |
| language | No | Language code (e.g. 'en', 'fr', 'de-CH') - configure_country only; other values are refused | |
| response | No | HTTP response for geo-blocking analysis | |
| timezone | No | IANA timezone identifier (e.g. 'America/New_York') - configure_country and generate_timezone_spoof | |
| operation | No | Localization operation to perform. Every operation returns data; none changes how another tool behaves. configure_country: the country's settings (language, acceptLanguage, timezone, currency, searchDomain, browserLocale). localize_search: runs search_web for searchParams.query in the country and its language. localize_browser: Playwright-style context options for the country (locale, timezoneId, geolocation, extraHTTPHeaders) - no browser is launched. generate_timezone_spoof: a JavaScript snippet that overrides Date and Intl for the timezone - nothing is injected. handle_geo_blocking: classifies a response you supply (url + response) and lists suggestions - nothing is fetched or bypassed. auto_detect: language and country detected in content you supply. get_stats: counters for this server process. get_supported_countries: the accepted country codes. | configure_country |
| userAgent | No | Not used by any operation. To carry a user agent through localize_browser, set browserOptions.userAgent | |
| countryCode | No | ISO 3166-1 alpha-2 country code, upper or lower case; operation:"get_supported_countries" lists the accepted ones. When omitted, the other operations use the country of the last configure_country call in this server process (US at start) - pass it on every call | |
| geoLocation | No | Coordinates echoed back in the configure_country result; nothing is emulated | |
| searchParams | No | Search parameters for localized search queries | |
| customHeaders | No | Headers echoed back in the configure_country result; nothing is sent | |
| proxySettings | No | Echoed back in the configure_country result. No request is routed through it: CrawlForge supplies no proxies, and your own go on stealth_mode stealthConfig.proxyRotation | |
| acceptLanguage | No | Accept-Language value to return from configure_country in place of the one built from the language and country | |
| browserOptions | No | Browser context options for localize_browser to build on: returned with the country's locale, timezoneId, geolocation and Accept-Language set, and your userAgent and other extraHTTPHeaders kept |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Discloses the single most surprising trait — the tool applies nothing: no later fetch_url/scrape/stealth_mode call picks up the values, nothing routes through a proxy, the IP is unchanged and geo-blocks are not lifted. It also states the cost (2 credits). These are exactly the side-effect expectations an agent would otherwise wrongly assume from a 'configure' operation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the purpose and the critical no-side-effect constraint, then the value-routing recipe, the exclusion, and cost. Dense but every clause earns its place for a 15-parameter, 8-operation tool; nothing is redundant.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given 15 parameters, nested objects, 8 enum operations, no output schema and an open-world hint, the description covers return semantics, side-effect profile, cost, and how values flow to named siblings. An agent has everything needed to select and invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the schema carries the per-parameter detail and the baseline is 3. The description goes beyond it with the credential-relevant advice to 'pass countryCode on every call' (since omitted calls silently reuse the last configure_country state), plus a worked example and the per-operation return descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource — 'look up the locale settings of a country' and 'run a search targeted at a country' — and immediately differentiates from siblings by naming the two operations and their outputs (language, Accept-Language, timezone, currency, etc.). The closing example call pins the shape down concretely.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit about when to use it (setting up locale for a request), when not to ('Not for an ordinary page read (scrape)'), and which alternatives to route results to (fetch_url headers, stealth_mode stealthConfig, search_web localization). It even notes scrape has no locale parameter, ruling out a misuse path.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
map_siteARead-onlyIdempotent
Use this to list a site's URLs without fetching page bodies - reads sitemap.xml when available, otherwise follows links. Not for page content (scrape, or crawl_deep for many pages) and not for the links on one page (extract_links). Cost: 2 credits. Example: map_site({url: "https://example.com", include_sitemap: true, max_urls: 500})
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | The website URL to map | |
| search | No | When set, rank discovered URLs by relevance to this string and emit ranked_urls:[{url,score}] | |
| max_urls | No | Maximum number of URLs to discover | |
| user_agent | No | Override the outbound User-Agent. CrawlForge identifies itself honestly by default; use this only for targets you have your own agreement with. | |
| domain_filter | No | Per-domain allow/deny lists and URL include/exclude patterns | |
| group_by_path | No | Group URLs by path segments | |
| respect_robots | No | Respect the target site's robots.txt (default: true). Setting this to false is honoured, returns a warning in the response, and is recorded against your API key — it is your decision, not a silent default. | |
| include_sitemap | No | Include sitemap.xml data in results | |
| include_metadata | No | Include page metadata for each URL | |
| max_inline_chars | No | Largest result to return inline, in characters of its JSON. Over it, the call returns a preview plus a result_handle for read_result instead of the whole result (default 40,000; env CRAWLFORGE_MAX_INLINE_CHARS) | |
| import_filter_config | No | JSON string of a previously exported domain-filter config |
Output Schema
| Name | Required | Description |
|---|---|---|
| urls | No | Flat array of URLs, or grouped-by-path object when group_by_path=true (default) |
| view | No | Whether preview and read_result offsets index a text field or the pretty-printed JSON |
| _cost | No | Cost-transparency metadata (D3.5), present when injected into the text copy of the result |
| preview | No | The first max_inline_chars characters of the view named by view_path (or of the pretty-printed JSON) |
| base_url | No | |
| metadata | No | Per-URL metadata when include_metadata=true |
| site_map | No | The shape of the site as counts; `urls` is the one list of URLs |
| warnings | No | Present when the map is partial, e.g. the start page could not be read and the URLs come from the sitemap only, or when the result is over max_inline_chars |
| truncated | No | True when the inline result is a preview |
| view_path | No | Dotted path of the text field the view was cut from; null for the JSON view |
| expires_at | No | When the stored result is dropped (ISO 8601) |
| statistics | No | |
| total_urls | No | |
| ranked_urls | No | Present only when the `search` param was set |
| total_chars | No | Length of the full view in characters |
| filter_stats | No | |
| result_handle | No | Handle for read_result; the full result is kept 1 hour |
| domain_filter_config | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover read-only, idempotency, non-destructiveness, and open-world access. The description adds meaningful behavior beyond that: it reads sitemap.xml when available and otherwise follows links, does not fetch page bodies, and costs 2 credits. This is operational context the agent can use when choosing among scraping tools.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded: purpose first, then exclusions, then cost, then an example. Every sentence adds useful information, and no space is wasted restating the name or schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex 11-parameter tool with an output schema and full annotation coverage, the description supplies the routing context an agent needs: what it does, what it does not do, how it discovers URLs, and its cost. Return-value details are appropriately left to the output schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all 11 parameters are already documented in the schema. The description mentions include_sitemap and max_urls only in an example and does not add syntax, defaults, or constraints beyond what the schema provides. This meets the baseline of 3 when the schema carries parameter semantics.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: list a site's URLs without fetching page bodies. It also distinguishes the tool from siblings by naming scrape, crawl_deep, and extract_links as wrong alternatives. An agent can identify the correct tool without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives explicit when-to-use guidance ('list a site's URLs without fetching page bodies') and when-not-to-use guidance ('Not for page content', 'not for the links on one page'), naming the alternatives for each excluded case. The sitemap-first fallback behavior also tells the agent how the tool will go about discovery.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
process_documentARead-onlyIdempotent
Use this to extract text from a PDF or DOCX URL or file - research papers, contracts, reports. The body decides how it is read: a PDF or Word document served under sourceType "url" still reaches its parser, and a body this tool cannot read (an image, an archive) is refused by name. Returns structured sections, metadata, and word count; for a PDF, pagesRead names the pages read (maxPages defaults to 100). Not for ordinary web pages (scrape), though an HTML URL is accepted. Cost: 2 credits. Example: process_document({source: "https://example.com/report.pdf", sourceType: "pdf_url"})
| Name | Required | Description | Default |
|---|---|---|---|
| source | Yes | Document source - URL or file path | |
| options | No | Additional processing options (maxPages, pageRange:{start,end}, extractText, extractMetadata, outputFormat, ...) | |
| redact_pii | No | Redact personal data from the text this call returns, before it reaches your context window. true means the free regex pass over EMAIL, PHONE, FINANCIAL and SECRET. The result carries redaction:{entities,count}. Default: off | |
| sourceType | No | Type of document source | |
| user_agent | No | Override the outbound User-Agent. CrawlForge identifies itself honestly by default; use this only for targets you have your own agreement with. | |
| respect_robots | No | Respect the target site's robots.txt (default: true). Setting this to false is honoured, returns a warning in the response, and is recorded against your API key — it is your decision, not a silent default. | |
| max_inline_chars | No | Largest result to return inline, in characters of its JSON. Over it, the call returns a preview plus a result_handle for read_result instead of the whole result (default 40,000; env CRAWLFORGE_MAX_INLINE_CHARS) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only cover safety profile (readOnly, idempotent, non-destructive, openWorld); the description adds substantive behavioral context beyond them: 2-credit cost, parser routing, refusal behavior, page-read reporting with maxPages defaulting to 100, and return shape (sections, metadata, word count). This is exactly the extra context annotations cannot convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with purpose and cost, followed by routing rules, return shape, and an example. Dense but every sentence carries signal; the 'body decides how it is read' phrasing is slightly convoluted but still earns its place by clarifying parser routing.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description correctly compensates by describing returns (structured sections, metadata, word count, pagesRead). Combined with full parameter coverage and annotation safety hints, an agent has everything needed to call this correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds genuine meaning the schema does not: sourceType 'url' still routes a PDF to its parser, maxPages defaults to 100, and a concrete call example with source/sourceType. It does not explain options/redact_pii/user_agent beyond the schema, hence not a 5.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource (extract text from a PDF or DOCX) and enumerates the document types (research papers, contracts, reports). It explicitly distinguishes itself from the sibling 'scrape' by declaring what it is not for, so an agent can route correctly without opening schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit when-to-use (PDF/DOCX URLs or files), what is refused (images, archives, refused by name), an edge case (an HTML URL is accepted but ordinary web pages belong to 'scrape'), and names the alternative tool. Nothing is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
read_resultARead-onlyIdempotent
Use this to read a result that came back with truncated: true and a result_handle - the tool kept the whole result for 1 hour and returned a preview. operation:"search" finds a literal query with offsets and context, "slice" returns characters from an offset, "lines" pages by line, "json_path" reads one subtree of a JSON result (crawl_deep pages, batch results, a fetch_url JSON body). Not a fetching tool: never call the original tool again while the handle is valid, and not for a result that arrived whole. Cost: 1 credit. Example: read_result({handle: "res_…", operation: "search", query: "pricing"})
| Name | Required | Description | Default |
|---|---|---|---|
| path | No | json_path: dotted keys and array indexes, e.g. "results[3].content" — not JSONPath | |
| query | No | search: the text to find, matched literally, case-insensitive | |
| handle | Yes | The result_handle a truncated result returned (res_… or a batch id) | |
| length | No | slice: characters to return (default 10,000); lines: lines to return (default 200, max 5,000) | |
| offset | No | slice: first character (default 0); lines: first line index (default 0) | |
| operation | Yes | slice: characters from offset; search: case-insensitive literal query with context and offsets; lines: a page of lines; json_path: one subtree of a JSON result | |
| max_matches | No | search: matches to return (default 20) | |
| max_inline_chars | No | Largest result to return inline, in characters of its JSON. Over it, the call returns a preview plus a result_handle for read_result instead of the whole result (default 40,000; env CRAWLFORGE_MAX_INLINE_CHARS) |
Output Schema
| Name | Required | Description |
|---|---|---|
| path | No | |
| text | No | slice: verbatim view.slice(offset, offset + length) |
| tool | No | The tool that produced the stored result |
| view | No | |
| _cost | No | Cost-transparency metadata (D3.5), present when injected into the text copy of the result |
| lines | No | |
| query | No | |
| value | No | json_path: the subtree; null with a preview when it is over max_inline_chars |
| handle | No | |
| length | No | slice: characters returned |
| offset | No | slice: first character returned |
| matches | No | search: matches with 200 chars of context each side |
| preview | No | |
| has_more | No | slice/lines: more follows the returned range |
| warnings | No | |
| operation | No | |
| truncated | No | search: more matches than returned; json_path: value replaced by a preview |
| view_path | No | |
| expires_at | No | |
| first_line | No | |
| line_count | No | |
| char_offset | No | lines: view offset of the first returned line |
| total_chars | No | Length of the full view |
| total_lines | No | |
| value_chars | No | |
| total_matches | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint, so the safety profile is covered. The description adds valuable behavioral context: the 1-hour retention window, the preview behavior, the credit cost, and the semantics of each operation (search, slice, lines, json_path). This goes beyond the annotations without contradicting them.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single dense paragraph that front-loads the primary use case and ends with an example. It is somewhat lengthy but every sentence earns its place by covering scope, operations, exclusions, cost, and an example. A more structured layout could improve skimmability, but it remains efficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given an output schema exists (not shown but present), the description need not detail return formats. It covers the main trigger, all operations, exclusions, cost, and handle validity. It does not explicitly mention error cases (e.g., expired handle), but the output schema likely handles those. Overall, it's sufficiently complete for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so every parameter is documented. The description adds operation-specific meaning (e.g., 'search finds a literal query with offsets and context', 'json_path reads one subtree') and a usage example that clarifies how parameters combine. This enriches understanding beyond the schema's basic field descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool reads truncated results, identifies the trigger condition (truncated: true with a result_handle), and explicitly distinguishes itself from fetching tools by instructing not to call the original tool again. It also names sibling tools indirectly and lists distinct operations, making its scope unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides explicit when-to-use (truncated results) and when-not-to-use (whole results, not a fetching tool) guidance. It also gives a concrete example call and notes the 1-hour handle validity, which helps the agent decide when to invoke this tool versus alternatives like fetch_url or crawl_deep.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
reddit_searchARead-onlyIdempotent
Use this to search Reddit posts or comments, or read a full comment thread - reddit.com blocks direct scraping, so this reads the Arctic Shift community archive instead (free, no Reddit credentials). Modes: 'posts' (default) and 'comments' search; 'thread' returns a post plus its nested comment tree by link_id. A subreddit/author-scoped search queries the archive directly. A keyword search across ALL of Reddit finds posts with a site-restricted web search and then reads those posts from the archive, because Arctic Shift can only keyword-search within a scope; results come back as real archive rows, ordered by search relevance. An unscoped COMMENT search discovers posts the same way and then searches each post's comments for the keywords. A scoped comment search Arctic Shift times out on is retried over narrower windows (7d, 3d, 1d) and reports window_applied. Arctic Shift is tried first and the PullPush archive second for posts/comments searches (fallback_used says so; PullPush has refused automated clients since August 2026). Not for reddit.com URLs via scrape or fetch_url (blocked) - use mode:'thread' with the post's link_id. Cost: 5 credits. Example: reddit_search({query: "best mechanical keyboard", subreddit: "MechanicalKeyboards", limit: 10})
| Name | Required | Description | Default |
|---|---|---|---|
| mode | No | What to search: posts (default), comments, or thread (full comment tree — requires link_id) | |
| sort | No | Sort by post date (default desc = newest first) | |
| after | No | Only content posted after this date — ISO 8601, epoch seconds, or an offset like '7d' | |
| limit | No | Max results (default 25; thread mode: max comments returned) | |
| query | No | Keyword search. Posts: matches title+selftext; comments: matches body. Supports "quoted phrases", OR, -exclusion | |
| author | No | Limit to one author (with or without the u/ prefix) | |
| before | No | Only content posted before this date — same formats as after | |
| source | No | Backend: auto routes + falls back (default). web_discovery serves only unscoped keyword searches (web search finds the posts, the archive supplies the rows). reddit_api uses the official Reddit Data API — only when REDDIT_CLIENT_ID/REDDIT_CLIENT_SECRET are set; serves posts/thread, not comment search | |
| link_id | No | Post ID (e.g. '1twm1zh' or 't3_1twm1zh') — required for thread mode, optional filter for comments mode | |
| subreddit | No | Limit to one subreddit (with or without the r/ prefix) |
Output Schema
| Name | Required | Description |
|---|---|---|
| mode | No | |
| post | No | thread mode: the post itself |
| _cost | No | Cost-transparency metadata (D3.5), present when injected into the text copy of the result |
| count | No | |
| notes | No | Data-provenance caveats (archive freshness, coverage gaps) |
| query | No | |
| author | No | |
| source | No | Which backend served this result — an archive, or web discovery (site-restricted web search hydrated from the archive) for Reddit-wide keyword search |
| link_id | No | Present in thread mode |
| results | No | posts/comments modes |
| comments | No | thread mode: nested comment tree ({...comment, replies:[...]}); a collapsed branch appears as a {more_count, more_ids} stub, where more_count is how many comments it hides, replies included, and more_ids lists only the hidden direct replies, so more_ids can be shorter than more_count |
| checkedAt | No | |
| subreddit | No | |
| discovered | No | web_discovery: how many post ids the site-restricted web search surfaced before archive hydration |
| comment_count | No | thread mode: comments in `comments` at every depth, stubs excluded; at most limit |
| fallback_used | No | Present when the primary archive failed and the fallback served the result |
| posts_searched | No | web_discovery comments mode: how many discovered posts had their comments searched before limit was reached |
| window_applied | No | arctic_shift comments mode: the after-window ("7d"/"3d"/"1d") the search was narrowed to after the full-history search timed out; absent when the caller set after or no narrowing was needed |
| comments_collapsed | No | thread mode: the sum of every stub's more_count, i.e. comments the source holds for this thread that the response leaves out. comment_count + comments_collapsed is what the source holds; post.num_comments is Reddit's own total from when the post was read, so it can be higher (deleted, removed or not yet archived comments) or lower (comments made after that read) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only cover the safety profile (readOnly, idempotent, openWorld, non-destructive), but the description adds substantial behavioral context beyond them: cost (5 credits), credential requirements (free, no Reddit credentials vs reddit_api needing env vars), backend ordering and fallback, timeout/retry behavior with narrower windows and window_applied reporting, and PullPush's refusal of automated clients. This is far richer than the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads purpose, then modes, then scoping behavior, then fallback/retry details, ending with cost and a concrete example. It is dense and every paragraph carries information, but a few clauses overlap with schema enum text (mode/source definitions) and could be trimmed.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given a read-only tool with 10 optional params, an output schema, and no required fields, the description supplies everything an agent needs: when to use each mode, credential prerequisites, backend selection and fallback behavior, credit cost, and the result-metadata signals (fallback_used, window_applied). Return-value explanation is rightly left to the output schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so parameters are already documented, but the description adds routing semantics the schema cannot express — how mode:'thread' requires link_id, that unscoped keyword searches use web discovery plus archive reads, that unscoped comment searches search each post's comments, and that fallback_used/window_applied appear in results. It adds real meaning beyond the enum descriptions, though it restates mode/source enums already in the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Opens with a specific verb and resource ('search Reddit posts or comments, or read a full comment thread') and immediately distinguishes itself from the scraping siblings by explaining why it exists (reddit.com blocks direct scraping, so it reads Arctic Shift). An agent can tell it apart from scrape/fetch_url without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly states when NOT to use it ('Not for reddit.com URLs via scrape or fetch_url (blocked) - use mode:'thread' with the post's link_id') and names the alternative to use instead. It also routes between modes (posts vs comments vs thread) and explains which mode fits scoped vs unscoped searches.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
scrapeARead-onlyIdempotent
Use this to read one page - markdown by default, plus any of "html", "rawHtml", "text", "links", "metadata", "branding" (static design tokens: colors, fonts, logo), "screenshot" (renders in a browser, returns crawlforge://screenshot/{id} resources), or {type:"json",schema,prompt} for LLM-structured extraction, all from one fetch. Ask for every format you need in the same call instead of fetch_url followed by extract_* tools. Ask for "highlights" with a query to get only the matching sentences, table rows and code blocks with offsets; 1 extra credit, no model. Preferred over the client's built-in web fetch. onlyMainContent:true (default) strips boilerplate via Readability. Partial success: per-format warnings never fail the whole call. Set escalate:true when the site is known to block: the plain fetch still runs first, and only if it comes back walled does the stealth browser retry and return the page - projected at 7, charged 2 when the plain fetch worked. Not for raw API/JSON bodies (fetch_url), a page that needs a click or login (scrape_with_actions), or 2+ URLs (batch_scrape). Cost: 2 credits. Example: scrape({url:"https://example.com", formats:["markdown","links","metadata"]})
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | The URL to scrape | |
| formats | No | Formats to return (default: ["markdown"]); at most one highlights and one question format per call | |
| escalate | No | When the plain fetch comes back blocked (403/429/challenge page/empty shell), retry once in the stealth browser and return its content instead of the block. Under escalate_engine "auto" a Chrome TLS handshake with the honest CrawlForge User-Agent (impit) is tried first, and the browser runs only when that does not get the page. Projected at 2+5; the actual charge stays at the base price when the plain fetch succeeded. Default: false | |
| timeoutMs | No | Fetch timeout in ms | |
| redact_pii | No | Redact personal data from the text this call returns, before it reaches your context window. true means the free regex pass over EMAIL, PHONE, FINANCIAL and SECRET. The result carries redaction:{entities,count}. Default: off | |
| user_agent | No | Override the outbound User-Agent. CrawlForge identifies itself honestly by default; use this only for targets you have your own agreement with. | |
| respect_robots | No | Respect the target site's robots.txt (default: true). Setting this to false is honoured, returns a warning in the response, and is recorded against your API key — it is your decision, not a silent default. | |
| brandingOptions | No | Options for the "branding" format | |
| escalate_engine | No | Stealth engine for the escalated retry: "auto" (default — Camoufox when it is installed, Chromium otherwise, reported in warnings), "playwright" (Chromium) or "camoufox" | auto |
| onlyMainContent | No | Strip boilerplate via Readability (default: true) | |
| max_inline_chars | No | Largest result to return inline, in characters of its JSON. Over it, the call returns a preview plus a result_handle for read_result instead of the whole result (default 40,000; env CRAWLFORGE_MAX_INLINE_CHARS) | |
| screenshotOptions | No | Options for the "screenshot" format |
Output Schema
| Name | Required | Description |
|---|---|---|
| url | No | Final URL after redirects |
| view | No | Whether preview and read_result offsets index a text field or the pretty-printed JSON |
| _cost | No | Cost-transparency metadata (D3.5), present when injected into the text copy of the result |
| error | No | Why success is false: a challenge page, an empty shell or an error placeholder was served instead of the content |
| title | No | Document title; present when success is false |
| status | No | HTTP status of the fetch; present when success is false |
| blocked | No | Present when a bot-defence vendor served a challenge page; the fallback hint names the tool to try next |
| content | No | One key per requested format |
| preview | No | The first max_inline_chars characters of the view named by view_path (or of the pretty-printed JSON) |
| stealth | No | Present when escalated is true: the stealth engine that ran ("impit" when the Chrome TLS handshake got the page without a browser), and the bot-defence vendor the plain fetch hit (null when the block named none) |
| success | No | Whether the scrape completed |
| warnings | No | Per-format warnings; partial success never fails the whole call |
| escalated | No | Present only when escalate:true was passed: whether the blocked plain fetch was retried in the stealth browser. False means the plain fetch sufficed and the call is charged at the base price |
| redaction | No | Present when redact_pii was set: what was redacted from the text of this result |
| truncated | No | True when the inline result is a preview |
| view_path | No | Dotted path of the text field the view was cut from; null for the JSON view |
| expires_at | No | When the stored result is dropped (ISO 8601) |
| total_chars | No | Length of the full view in characters |
| result_handle | No | Handle for read_result; the full result is kept 1 hour |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly/idempotent/non-destructive, so the bar is lower, yet the description adds genuinely new behavior: per-format warnings never fail the whole call, escalate runs the plain fetch first and only retries on a wall, the projected-vs-actual charge behavior, and the 2-credit base cost. That is context an agent cannot derive from the hints.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core purpose and default, then progressively adds format details, alternatives, escalation, and cost. It is a single dense paragraph of mostly load-bearing sentences, though the length and lack of breaks make it slightly harder to scan than it needs to be.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 12-parameter, nested-format tool with a rich schema, output schema, and annotations, the description covers the decision points an agent needs: default format, combining formats, cost, escalation semantics, redaction, robots behavior, and which sibling to pick. Nothing material is left to inference.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3; the description still adds meaning beyond the schema by explaining highlights (+1 credit, no model, returns verbatim units with offsets), branding as static design tokens, and that one extra fetch is avoided by combining formats. It does not add much on escalate/timeout/redact beyond what the schema already states.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Opens with a specific verb+resource ("read one page") and enumerates exactly what it can return (markdown default plus html, rawHtml, text, links, metadata, branding, screenshot, structured json, highlights). It explicitly separates itself from fetch_url, extract_* tools, scrape_with_actions, and batch_scrape, so an agent can route without opening another schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit when-to-use guidance ("Ask for every format you need in the same call instead of fetch_url followed by extract_* tools") and an explicit exclusion list: not for raw API/JSON bodies (fetch_url), click/login pages (scrape_with_actions), or 2+ URLs (batch_scrape). It even states a preference over the client's built-in web fetch.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
scrape_structuredARead-onlyIdempotent
Use this when you know the exact CSS selectors for the data you want - e.g. a pricing table or product list with consistent markup. More reliable than LLM extraction for well-structured pages. By default each selector is matched independently across the whole page, so the returned arrays are NOT row-aligned: data.price[0] need not belong to the same row as data.name[0]. Pass row_selector to get aligned records instead - one object per row, null for a field the row lacks. Not for pages whose markup varies or where you cannot name the selectors (extract_structured, LLM-driven). Cost: 2 credits. Example: scrape_structured({url: "https://shop.com/products", row_selector: ".product-card", selectors: {price: ".price", name: ".product-title"}})
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | The URL to scrape | |
| selectors | Yes | CSS selectors mapping field names to selectors. Append @attr to extract an attribute instead of text (e.g. "a.link@href", "img@src") | |
| user_agent | No | Override the outbound User-Agent. CrawlForge identifies itself honestly by default; use this only for targets you have your own agreement with. | |
| max_results | No | Maximum number of matches to return per field when a selector matches multiple elements, or the maximum number of rows when row_selector is set | |
| row_selector | No | CSS selector for the repeating row/container element. When set, each field in selectors is matched inside each row and data is an array of row-aligned records ({field: value|null}) instead of parallel arrays | |
| respect_robots | No | Respect the target site's robots.txt (default: true). Setting this to false is honoured, returns a warning in the response, and is recorded against your API key — it is your decision, not a silent default. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnly/idempotent annotations, it discloses the critical row-alignment gotcha (data.price[0] need not belong to the same row as data.name[0]), the cost of 2 credits, and the meaning of row_selector. This is exactly the kind of non-obvious behavior an agent needs before invoking.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description leads with the decision rule, then covers exclusions, the key behavioral warning, cost, and an example with no filler. Every sentence contributes necessary operational information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite having no output schema, the description explains the shape of returned data (parallel arrays vs row-aligned objects) and gives enough context for an agent to invoke correctly. Together with the rich input schema, this is complete for a read-only extraction tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema already covers all 6 parameters in detail, so the baseline is 3. The description adds value with a concrete usage example and clarifies how selectors and row_selector work together, but it does not substantially redefine the parameter meanings beyond the schema descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb and resource (scrape structured data with exact CSS selectors) and differentiates it from LLM-driven extraction. It is immediately clear this tool is for well-structured pages with known markup.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives an explicit when-to-use condition (you know exact selectors, consistent markup) and states what it is not for (varying markup or unnamed selectors), pointing to extract_structured / LLM-driven extraction as the alternative. This lets an agent select correctly among many siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
scrape_templateARead-onlyIdempotent
Use this when you want structured data from a well-known site or platform API without writing custom selectors. Three modes: a template id with a url (scrape_template({template:"github-repo", url:"https://github.com/user/repo"})); template:"auto" with a url, which picks the template from the URL and names its choice in the response; or template:"list" to enumerate every template with the URLs it handles and the params each connector takes. Page templates return one record - e-commerce, social, developer and news sites (shopify-product, amazon-product, github-repo, youtube-video, reddit-thread, hacker-news-front-page, producthunt-launch, stackoverflow-question, npm-package; reddit-thread reads the post from the Arctic Shift archive and reddit_search reads the comment tree). linkedin-profile and tweet are retired - those sites' robots.txt disallow every keyless path - and naming one returns the reason. List connectors return N records from one call and are driven by params instead of a url: job boards (Greenhouse, Lever, Ashby, Workable, Recruitee, Teamtailor) return a company's whole careers board, US government APIs (NHTSA VIN decode, NPI provider registry) answer keyless lookups, and shopify-collection returns a whole collection. Not for a site without a template (scrape) - template:"list" shows what exists. Cost: 1 credit. Example: scrape_template({template:"greenhouse-jobs", params:{company:"stripe"}})
| Name | Required | Description | Default |
|---|---|---|---|
| url | No | URL to scrape — required unless template is list, or params drive a list connector | |
| params | No | Parameters for a list connector, e.g. {company:"stripe"} for greenhouse-jobs or {store:"www.allbirds.com", collection:"mens"} for shopify-collection. Use template:"list" to see which templates take params | |
| timeout | No | Request timeout in milliseconds | |
| template | Yes | Template ID (e.g. github-repo), "auto" to detect one from the url, or "list" to enumerate available templates | |
| user_agent | No | Override the outbound User-Agent. CrawlForge identifies itself honestly by default; use this only for targets you have your own agreement with. | |
| respect_robots | No | Respect the target site's robots.txt (default: true). Setting this to false is honoured, returns a warning in the response, and is recorded against your API key — it is your decision, not a silent default. | |
| max_inline_chars | No | Largest result to return inline, in characters of its JSON. Over it, the call returns a preview plus a result_handle for read_result instead of the whole result (default 40,000; env CRAWLFORGE_MAX_INLINE_CHARS) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly/openWorld/idempotent/non-destructive, so the bar is lower. The description still adds real context: 1 credit cost, page templates returning one record vs list connectors returning N, the Arctic Shift archive behavior for reddit-thread, and the retired templates (linkedin-profile, tweet) rejected by robots.txt. Minor gap: inline-preview/result_handle behavior lives in the schema, not here.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the usage trigger and mode enumeration; each sentence carries weight. The long parenthetical inventory of template names and job-board vendors is dense but earns its place as discoverability aid. Slightly overlong, but no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 7-param tool with a nested `params` object and no output schema, the description covers modes, cost, required-vs-optional param logic, example calls, and failure/exception cases. An agent has everything needed to select and invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3. The description goes beyond the schema by grouping the parameters into three modes and showing worked invocations for template+url, auto, and params-driven list connectors, which clarifies how `url` vs `params` are selected — information the schema documents only per-field.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource (fetch structured data via pre-built site/API templates) and immediately contrasts itself with the sibling `scrape` for sites without templates. The three operating modes (template id, auto, list) make the tool's scope unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit when-to-use (well-known site or platform API without writing selectors), when-not (no template → use `scrape`, and `template:"list"` to check), and it names concrete alternatives/enumeration paths. Includes the retired-template behavior so the agent knows a failed name returns a reason rather than an error.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
scrape_with_actionsARead-only
Use this when you must interact with a page before scraping - login, click buttons, fill forms, scroll, or wait for dynamic content to load - for SPAs, login-gated content, or multi-step flows. Actions: snapshot, wait, click, type, press, scroll, screenshot, executeJavaScript (disabled unless the server runs with ALLOW_JAVASCRIPT_EXECUTION=true; refused on the hosted API), select (dropdowns), hover, navigate. Start a chain with {type:"snapshot"} to list the page's interactive elements, including those in open shadow roots and iframes, with stable refs (@e1, @e2 ... in document order), then target those refs in later actions instead of guessing CSS selectors; navigation invalidates refs, so snapshot again after one. Set browserOptions.consent:"reject" (or "accept") to answer a cookie/consent banner before the first action; it is off by default. Set browserOptions.stealth:true to run the chain in the stealth browser, and browserOptions.engine to pick its engine ("auto" by default - camoufox when it is installed, Chromium otherwise, and the result says which ran). robots.txt is respected on every navigation, and each navigation is checked: a bot wall is reported as blocked with the vendor named, per navigation in navigations and for the final page at the top level. A click or form submit that loads another page is a navigation too: it gets its own navigations entry (trigger names the action type), and finalUrl is where the chain ended. browserOptions.proxyRotation routes a stealth chain through your own proxies. Screenshots are taken by {type:"screenshot"} actions (and on error) and stored as crawlforge://screenshot/{actionId} resources. Not for pages that render without interaction (scrape) and not as the first attempt on a blocked site (stealth_mode operation:"scrape"). Cost: 5 credits. Example: scrape_with_actions({url: "https://app.com/dashboard", actions: [{type:"snapshot"},{type:"type",selector:"@e2",text:"user@a.com"},{type:"click",selector:"@e4"}]})
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | The URL to scrape | |
| actions | Yes | Browser actions to perform before scraping | |
| formats | No | Output formats for scraped content. For an image, add a {type:"screenshot"} action where you want it taken. | |
| maxRetries | No | Whole-chain retries on failure (0-3). A retry re-navigates to the starting URL and replays every action; each attempt is reported under attempts[]. | |
| redact_pii | No | Redact personal data from the text this call returns, before it reaches your context window. true means the free regex pass over EMAIL, PHONE, FINANCIAL and SECRET. The result carries redaction:{entities,count}. Default: off | |
| formAutoFill | No | Form auto-fill configuration | |
| browserOptions | No | Browser configuration options | |
| respect_robots | No | Respect the target site's robots.txt (default: true). Setting this to false is honoured, returns a warning in the response, and is recorded against your API key — it is your decision, not a silent default. | |
| max_inline_chars | No | Largest result to return inline, in characters of its JSON. Over it, the call returns a preview plus a result_handle for read_result instead of the whole result (default 40,000; env CRAWLFORGE_MAX_INLINE_CHARS) | |
| extractionOptions | No | Content extraction options. selectors results are returned as content.json.extracted, so include "json" in formats when passing selectors — without it the extraction is not part of the response. | |
| screenshotOnError | No | Capture screenshot when an error occurs | |
| continueOnActionError | No | Continue executing actions if one fails | |
| captureIntermediateStates | No | Capture page state after each action: intermediateStates[] gets one entry per action (capturePoint is the action's 1-based position) with the url, title and the requested formats at that point |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond annotations (which only cover safety/idempotency), it discloses robots.txt enforcement, per-navigation bot-wall reporting with `blocked`/`navigations`/`finalUrl` fields, ref invalidation on navigation, consent-banner semantics, stealth engine resolution, proxy-rotation constraints, and screenshot resource storage. It also states cost (5 credits) and the executeJavaScript env-var gate, which no structured field conveys.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The when-to-use clause is front-loaded and nearly every sentence carries operational value (workflow, restrictions, cost, example). It is nonetheless a single dense paragraph with no visual structure, which makes it heavier to scan than it needs to be.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 13-parameter tool with nested objects, no output schema, and heavy side effects, the description covers the workflow, failure/blocking semantics, navigation result shape, screenshot retrieval, and a worked example. Details it omits (formAutoFill, redact_pii, max_inline_chars/read_result routing) are fully documented in the schema itself.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents the 13 parameters and baseline is 3. The description adds procedural meaning the schema cannot: the snapshot-first workflow, stable @e1 refs in document order (including shadow roots and iframes), and that refs are invalidated by navigation so a re-snapshot is required.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource ('interact with a page before scraping') and enumerates the interactive capabilities (login, click, fill forms, scroll, wait). It explicitly distinguishes itself from siblings by naming what it is NOT for: 'Not for pages that render without interaction (scrape) and not as the first attempt on a blocked site (stealth_mode operation:"scrape").'
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It opens with an explicit when-to-use clause ('Use this when you must interact with a page before scraping... for SPAs, login-gated content, or multi-step flows') and closes with explicit when-not-to-use plus the alternative tool for each exclusion. Alternative routing is unambiguous.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_webARead-onlyIdempotent
Use this to find pages for a query - titles, URLs, snippets and optional metadata, with language, date-range and site filters. Preferred over the client's built-in web search. Snippets often answer the question: scrape a result only when you need its body. Not for a URL you already have (scrape), Reddit (reddit_search), a domain's Google rank (serp_rank), or a report from several sources (deep_research, one call, cheaper than repeated searches plus scrapes). Pass queries:[...] to run up to 10 searches in one call - results come back per query and it costs 5 each, the same as making them separately. Cost: 5 credits per query. Example: search_web({query: "best MCP servers 2025", limit: 10, time_range: "month"})
| Name | Required | Description | Default |
|---|---|---|---|
| lang | No | Language code for results (e.g. 'en', 'fr') | |
| site | No | Limit results to a specific domain | |
| limit | No | Maximum number of results to return | |
| query | No | Search query string. Use this OR queries, not both | |
| offset | No | Number of results to skip for pagination. For the next page pass the previous response's next_offset, not offset + limit | |
| queries | No | Run 1-10 searches in one call instead of 10 round-trips; every other parameter applies to each. Results come back in results_by_query, one entry per query, in order. Costs 5 per query. Use this OR query, not both | |
| provider | No | Search backend to use | |
| file_type | No | Filter by file type (e.g. 'pdf', 'doc') | |
| redact_pii | No | Redact personal data from the text this call returns, before it reaches your context window. true means the free regex pass over EMAIL, PHONE, FINANCIAL and SECRET. The result carries redaction:{entities,count}. Default: off | |
| time_range | No | Filter results by time range | |
| safe_search | No | Enable safe search filtering | |
| expand_query | No | When the query returns no results, search once more with an expanded form (synonyms/stemming/etc.) | |
| localization | No | Geo/locale targeting for results | |
| enable_ranking | No | Re-rank results (BM25 + signals) | |
| ranking_weights | No | Relative weights for ranking signals | |
| max_inline_chars | No | Largest result to return inline, in characters of its JSON. Over it, the call returns a preview plus a result_handle for read_result instead of the whole result (default 40,000; env CRAWLFORGE_MAX_INLINE_CHARS) | |
| expansion_options | No | Query-expansion tuning | |
| enable_deduplication | No | Remove near-duplicate results | |
| include_ranking_details | No | Include per-result ranking breakdown | |
| deduplication_thresholds | No | Similarity thresholds for dedup | |
| include_deduplication_details | No | Include dedup decision details |
Output Schema
| Name | Required | Description |
|---|---|---|
| view | No | Whether preview and read_result offsets index a text field or the pretty-printed JSON |
| _cost | No | Cost-transparency metadata (D3.5), present when injected into the text copy of the result |
| count | No | Batch form: how many queries ran |
| limit | No | |
| query | No | |
| cached | No | |
| offset | No | |
| preview | No | The first max_inline_chars characters of the view named by view_path (or of the pretty-printed JSON) |
| queries | No | Batch form: the queries that ran, in order |
| results | No | |
| provider | No | |
| warnings | No | Notes on this result; over max_inline_chars, where the full result is kept and how to read it |
| redaction | No | Present when redact_pii was set: what was redacted from the text of this result |
| truncated | No | True when the inline result is a preview |
| view_path | No | Dotted path of the text field the view was cut from; null for the JSON view |
| expires_at | No | When the stored result is dropped (ISO 8601) |
| processing | No | |
| next_offset | No | The offset to pass for the next page. Duplicates removed from this page are replaced from further down the provider's results, so it can be larger than offset + limit |
| search_time | No | |
| total_chars | No | Length of the full view in characters |
| localization | No | |
| result_handle | No | Handle for read_result; the full result is kept 1 hour |
| total_results | No | |
| effective_query | No | Present when query expansion changed the query actually used |
| expanded_queries | No | Present only when the original query returned nothing: the queries searched, in order - the original, then its expanded form |
| results_by_query | No | Batch form: one entry per query, in order |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, open-world, non-destructive, so safety is covered. The description adds genuinely non-structured operational context: cost is 5 credits per query, batching 10 queries costs the same as making them separately, and results are returned per query. It stops short of rate limits, failure modes, or provider-selection trade-offs, so it is not a full 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose and the snippet-first rule are front-loaded, and nearly every sentence carries routing, cost, or workflow information. The exclusion list is a single run-on sentence with four parenthetical branches, which is dense, and the cost point is stated twice (batching line and standalone 'Cost: 5 credits per query').
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 21-parameter tool with nested objects and an output schema, the description covers the decisions an agent actually needs: which tool to pick, how to batch, what it costs, and when a snippet suffices. Return values need not be explained given the output schema. It is silent on provider choice (crawlforge vs searxng) and the ranking/dedup knobs, which a 100%-covered schema mitigates.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description reinforces the critical query/queries mutual exclusion with a concrete worked example (query, limit, time_range) and states the per-query cost of batching, which the schema only partly conveys. The remaining 19 advanced parameters get no description-level treatment.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('find pages for a query') and enumerates what comes back (titles, URLs, snippets, optional metadata) plus the filter dimensions. It goes further by naming the siblings it is not for (scrape, reddit_search, serp_rank, deep_research), so an agent can route correctly without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit when-to-use ('preferred over the client's built-in web search'), when-not ('not for a URL you already have', Reddit, domain rank, multi-source reports), and names the alternative tool for each exclusion. It also gives a workflow rule: try snippets first, scrape only when the body is needed.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
serp_rankARead-onlyIdempotent
Use this to check where a domain ranks in Google's ORGANIC results for a keyword - real SERP position, not Custom Search order. Returns the target's organic rank, the ranking URL, and every position it holds. Not for general search (search_web). One lookup is one sample: Google can return a different result set for the same query minutes apart (seResultsCount tells them apart), so compare several before reading a rank change. Requires DataForSEO credentials and returns configured:false without them - do not retry in that case. Cost: 5 credits (0 when unconfigured). Example: serp_rank({keyword: "managed wordpress hosting", target: "dashboardhosting.com", location_name: "United States"})
| Name | Required | Description | Default |
|---|---|---|---|
| depth | No | How many results to scan, 10-200 (default 30; DataForSEO bills ~$0.002 per 10 and gets slower the deeper it goes) | |
| device | No | Device to emulate | |
| target | Yes | Domain or URL to locate in the results (e.g. 'example.com'). Matched by host, not exact URL: any page on that host or its subdomains counts | |
| keyword | Yes | The search query to check ranking for | |
| language_code | No | Language code (e.g. 'en') | |
| location_code | No | Numeric DataForSEO location code (overrides location_name) | |
| location_name | No | Location, e.g. 'United States' or 'London,England,United Kingdom' |
Output Schema
| Name | Required | Description |
|---|---|---|
| url | No | URL of the target's best-ranking result |
| cost | No | USD charged by DataForSEO for this lookup (separate from CrawlForge credits) |
| note | No | Present when configured=false, explains how to enable |
| _cost | No | Cost-transparency metadata (D3.5), present when injected into the text copy of the result |
| found | No | Whether the target appeared anywhere in the scanned SERP |
| title | No | |
| device | No | |
| target | No | Host the SERP was matched against: a URL target is reduced to its host, and any page on it or its subdomains counts |
| keyword | No | |
| results | No | Top organic competitors as Google actually ranks them (capped) |
| checkUrl | No | Link to view the real SERP on DataForSEO |
| location | No | |
| position | No | Best (lowest) organic rank; null = not within top `depth` |
| checkedAt | No | |
| configured | No | False when DATAFORSEO_LOGIN/PASSWORD are unset — no rank was fabricated |
| allPositions | No | Every position the target holds on this SERP |
| depthScanned | No | |
| rankAbsolute | No | |
| organicResults | No | |
| seResultsCount | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnly, openWorld, idempotent, non-destructive), and the description adds rich context beyond them: the sampling variability of Google results, the seResultsCount field to distinguish samples, credential dependency, the configured:false outcome, and exact credit cost. This is unusually complete behavioral disclosure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but front-loaded with the core purpose and exclusion, followed by sampling, credential, and cost notes, ending with a concrete example. Every sentence earns its place and no redundant phrasing is present.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given an output schema exists, the description needn't explain return values, and it covers the remaining gaps an agent needs: credential prerequisites, cost, sampling caveats, sibling exclusion, and a usage example. Nothing material is missing for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all seven parameters in detail. The description provides a concrete usage example but adds no parameter-specific meaning beyond what the schema already states (e.g., host matching, depth billing, location codes). Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: check where a domain ranks in Google's organic results for a keyword, and explicitly distinguishes this from Custom Search order. It also names the sibling it is not (search_web), so an agent can route correctly without inspecting schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says when not to use it ('Not for general search (search_web)'), gives operational guidance about comparing multiple samples before reading a rank change, and states credential requirements plus the exact behavior when unconfigured (returns configured:false, do not retry). Alternatives and exclusions are unambiguous.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
stealth_modeA
Use this when a site blocks normal scraping - Cloudflare, Datadome, or other bot-detection systems. Renders in a real browser with randomized fingerprints, human behavior simulation, WebRTC/canvas spoofing - Camoufox (Firefox) when it is installed, Chromium otherwise, and the result names the one that ran. operation:"scrape" is the one-shot path: it creates a context, navigates, returns the requested formats and tears down. The create_context -> create_page -> cleanup operations remain for multi-step work. operation:"configure" only validates a stealthConfig and returns it with the defaults filled in - it stores nothing, so pass stealthConfig on each scrape or create_context call that should use it. robots.txt is respected on every navigation. Not a first choice: try scrape first and switch here after a 403/429/CAPTCHA/challenge page or an empty shell. Cost: 5 credits per browser operation; configure, get_stats and cleanup cost 1. Example: stealth_mode({operation:"scrape", url:"https://example.com", formats:["markdown","links"]})
| Name | Required | Description | Default |
|---|---|---|---|
| url | No | URL to scrape — required for operation:"scrape" | |
| engine | No | Browser engine: "auto" (default — camoufox when it is installed, Chromium otherwise, and the result says which), "camoufox" (Firefox-based, higher anti-detect score; fails if not installed), or "chromium" ("playwright" is the same engine under its old name) | auto |
| formats | No | Formats to return from operation:"scrape" (default: ["markdown"]). "screenshot" returns a crawlforge://screenshot/{id} resource URI. | |
| verbose | No | Return the full generated fingerprint from create_context instead of a summary | |
| wait_for | No | Extra wait after page load, in ms — for content that renders after DOMContentLoaded | |
| contextId | No | Browser context ID for page operations | |
| operation | No | Stealth operation to perform. scrape: render one URL and return the formats. configure: validate stealthConfig and return it with defaults filled in; nothing is stored. create_context / create_page: a context to open pages in. get_stats: this server process's open contexts, browser and proxy state. cleanup: close the browser and every context. | configure |
| urlToTest | No | URL to navigate to when creating a page | |
| redact_pii | No | Redact personal data from the text this call returns, before it reaches your context window. true means the free regex pass over EMAIL, PHONE, FINANCIAL and SECRET. The result carries redaction:{entities,count}. Default: off | |
| stealthConfig | No | Stealth browser configuration with anti-detection settings. Applies to the call it is passed on and is not remembered. Chromium applies customUserAgent, customViewport, locale and timezone per call. On camoufox customUserAgent and customViewport are not applied, locale is fixed when the browser launches (run operation:"cleanup" first to change it), and behind a proxy locale and timezone come from the exit IP; each of these is reported in `warnings`. | |
| respect_robots | No | Respect the target site's robots.txt (default: true). Setting this to false is honoured, returns a warning in the response, and is recorded against your API key — it is your decision, not a silent default. | |
| max_inline_chars | No | Largest result to return inline, in characters of its JSON. Over it, the call returns a preview plus a result_handle for read_result instead of the whole result (default 40,000; env CRAWLFORGE_MAX_INLINE_CHARS) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Goes well beyond the annotations: discloses cost (5 credits per browser op; 1 for configure/get_stats/cleanup), that robots.txt is respected on every navigation, that configure stores nothing so stealthConfig must be re-passed, and which engine ran. Annotations only cover the safety/idempotency profile, so this context is genuinely additive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the trigger condition and the anti-detection value before the operational details, and every sentence carries information. It is dense and long, but not padded; the cost and example are the only borderline lines.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 12-param, nested-object tool with no output schema, it covers trigger, alternatives, cost, engine selection, config persistence, and robots behavior. It partially signals return content (engine name, redaction counts) but does not fully characterize output shape, which is acceptable given the depth provided.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3. The description adds the operation lifecycle framing (scrape tears down; create_context/create_page/cleanup for multi-step; configure validates and stores nothing), which clarifies how parameters interact across calls beyond the enum's own text.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource+scope: renders in a real browser with randomized fingerprints when a site blocks normal scraping, naming the anti-detect stack (Camoufox/Chromium). It explicitly distinguishes itself from the sibling 'scrape' tool, so an agent can route without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit when-to-use (site blocks: Cloudflare, Datadome, bot-detection), when-not ('Not a first choice: try scrape first'), and the trigger to switch (403/429/CAPTCHA/challenge page or empty shell). It also maps operations to workflows: scrape as the one-shot path vs create_context/create_page/cleanup for multi-step work.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
summarize_contentARead-onlyIdempotent
Use this to condense text you already hold into a briefing, comparison, or shorter LLM context - extractive (sentence selection) or abstractive (rewrite via Ollama/sampling). Takes text, not a URL: pass the markdown from a scrape result. Not needed for text short enough to summarise in context yourself. Cost: 4 credits. Example: summarize_content({text: "..long article..", options: {summaryLength: "short", summaryType: "abstractive"}})
| Name | Required | Description | Default |
|---|---|---|---|
| text | Yes | The text content to summarize | |
| options | No | Summarization options |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare the safety profile (readOnlyHint, idempotentHint, destructiveHint false). The description adds meaningful behavioral context: it accepts text not URLs, performs extractive or abstractive summarization, leverages Ollama/sampling for abstractive rewrites, and costs 4 credits. This goes beyond what annotations provide, though it doesn't disclose return format or edge-case behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but organized: it starts with the core purpose, then input constraint, a usage heuristic, cost, and a concrete example. No sentence is wasted; the structure front-loads the main action and defers cost and example details.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter tool with high schema coverage but an empty options schema, the description covers how to invoke it (text plus optional options), what input type to pass, and an example. It lacks an explicit statement of return value shape, which matters because there is no output schema; however, the core invocation requirements are adequately specified.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% for both properties, but the 'options' property is an empty object in the schema, providing no usable semantics. The description compensates with an example showing summaryLength and summaryType, and clarifies that 'text' should be markdown from a scrape result. This adds real meaning beyond the schema, especially where the schema is empty.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('condense') and names the resource ('text you already hold'), and distinguishes itself from URL-input tools by stating 'Takes text, not a URL.' It also lists output forms (briefing, comparison, shorter LLM context) and methods (extractive/abstractive), which clarifies exactly what it does.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides explicit when-to-use context: summarizing text already held, and an explicit when-not-to-use: 'Not needed for text short enough to summarise in context yourself.' It also directs users to pass markdown from a scrape result, implying the preceding step. It doesn't name a specific sibling tool as an alternative, but it gives enough routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
track_changesA
Use this to monitor a URL for content changes over time - competitor pricing, regulation updates, product availability. Start with operation:"create_baseline", then periodically use operation:"compare" to diff; repeated compare calls on the same URL are expected. Supports webhooks and scheduled monitoring, and scheduledMonitorOptions.hosted:true runs the monitor on CrawlForge's servers with email and signed webhooks. Not for a one-off read (scrape). Cost: 3 credits. Example: track_changes({url: "https://example.com/pricing", operation: "create_baseline"})
| Name | Required | Description | Default |
|---|---|---|---|
| url | No | The URL to track changes for (optional for list_scheduled_monitors) | |
| html | No | HTML content to compare against baseline | |
| content | No | Content to compare against baseline | |
| operation | No | Tracking operation to perform | compare |
| user_agent | No | Override the outbound User-Agent. CrawlForge identifies itself honestly by default; use this only for targets you have your own agreement with. | |
| queryOptions | No | Query options for history and stats retrieval | |
| exportOptions | No | Export options for change history data | |
| respect_robots | No | Respect the target site's robots.txt (default: true). Setting this to false is honoured, returns a warning in the response, and is recorded against your API key — it is your decision, not a silent default. | |
| storageOptions | No | Storage and history retention settings | |
| trackingOptions | No | Options for how changes are tracked and compared | |
| alertRuleOptions | No | Alert rule configuration for change notifications | |
| dashboardOptions | No | Dashboard display options | |
| max_inline_chars | No | Largest result to return inline, in characters of its JSON. Over it, the call returns a preview plus a result_handle for read_result instead of the whole result (default 40,000; env CRAWLFORGE_MAX_INLINE_CHARS) | |
| monitoringOptions | No | Monitoring schedule and notification settings | |
| notificationOptions | No | Notification configuration for webhooks, Slack and email (email is sent by hosted monitors only) | |
| scheduledMonitorOptions | No | Scheduled monitoring: recurring compare + notify, optional plain-English goal |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotations declare readOnlyHint=false, idempotentHint=false, and destructiveHint=false, so the description doesn't need to re-state safety. It does add valuable behavioral context beyond annotations: webhook/scheduled monitoring support, hosted execution on CrawlForge's servers with email and signed webhooks, a cost of 3 credits, and that repeated compares are expected. It does not discuss rate limits, auth prerequisites, or result-size behavior in depth, so it falls just short of a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the primary purpose, then workflow, then differentiation, then cost, then example. No sentence is wasted; it covers a complex tool efficiently in a compact block of text.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given 16 parameters, deep nested objects, and no output schema, the description does a good job covering the operational workflow and the hosted-versus-local distinction. It doesn't explain return shapes (acceptable since no output schema exists, though some hint of result format would help) and omits coverage of less-common operations like get_history, export_history, or create_alert_rule that appear in the schema union. Adequate for most calls but not exhaustive.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all 16 parameters in detail (including nested objects like monitoringOptions, scheduledMonitorOptions, trackingOptions). The description adds a small amount of meaning by naming two specific operations (create_baseline, compare) and the hosted option semantics, but largely repeats what the schema already provides. Baseline 3 is appropriate when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb+resource: monitor a URL for content changes over time, with concrete use cases (competitor pricing, regulation updates, product availability). It explicitly distinguishes itself from the sibling 'scrape' by stating 'Not for a one-off read (scrape)', which is the key differentiator an agent needs.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides explicit workflow guidance: start with operation:"create_baseline", then periodically use operation:"compare", and notes that repeated compare calls are expected. It also names the alternative for one-off reads ('scrape') and gives exclusion guidance, which is exactly the when-to-use/when-not-to-use clarity that helps an agent select correctly.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
14 tool updates
v6.19.2- Changed
agent1 field changed- added
Input schema / properties / max_inline_charsAdded value: +{ + "description": "Largest result to return inline, in characters of its JSON. Over it, the call returns a preview plus a result_handle for read_result instead of the whole result (default 40,000; env CRAWLFORGE_MAX_INLINE_CHARS)", + "maximum": 10000000, + "minimum": 1000, + "type": "integer" +}
- Changed
extract_links1 field changed- added
Input schema / properties / max_inline_charsAdded value: +{ + "description": "Largest result to return inline, in characters of its JSON. Over it, the call returns a preview plus a result_handle for read_result instead of the whole result (default 40,000; env CRAWLFORGE_MAX_INLINE_CHARS)", + "maximum": 10000000, + "minimum": 1000, + "type": "integer" +}
- Changed
extract_metadata2 fields changed- changed
Input schema / properties / json_ld_types / descriptionPrevious value: -"Filter the returned JSON-LD to nodes of these schema.org types, e.g. [\"Product\",\"Offer\"]. Subtypes match their parent: \"Event\" returns MusicEvent, \"Offer\" returns AggregateOffer, \"ItemList\" returns BreadcrumbList. Nodes are found at any depth, including inside @graph and nested inside a parent node. When set, json_ld carries only the matching nodes instead of the raw dump, and json_ld_type_counts reports how many matched per requested type. Documented types: ItemList, Product, Offer, Event, JobPosting, RealEstateListing — any other schema.org type is matched exactly."New value: +"Filter the returned JSON-LD to nodes of these schema.org types, e.g. [\"Product\",\"Offer\"]. Subtypes match their parent: \"Event\" returns MusicEvent, \"Offer\" returns AggregateOffer, \"ItemList\" returns BreadcrumbList. Nodes are found at any depth, including inside @graph and nested inside a parent node; a match nested inside another returned node comes back inside it, not again on its own. When set, json_ld carries only the matching nodes instead of the raw dump, and json_ld_type_counts reports how many matched per requested type, nested ones included. Documented types: ItemList, Product, Offer, Event, JobPosting, RealEstateListing — any other schema.org type is matched exactly." - added
Input schema / properties / max_inline_charsAdded value: +{ + "description": "Largest result to return inline, in characters of its JSON. Over it, the call returns a preview plus a result_handle for read_result instead of the whole result (default 40,000; env CRAWLFORGE_MAX_INLINE_CHARS)", + "maximum": 10000000, + "minimum": 1000, + "type": "integer" +}
- Changed
extract_structured5 fields changed- added
Output schema / properties / confidence / descriptionAdded value: +"0-1: method and validation, scaled by the share of requested fields filled and, for \"llm\", the share of short string values found in the page" - added
Output schema / properties / modelAdded value: +{ + "description": "Model that produced the data, e.g. \"gemma3:12b\"; present when extraction_method is \"llm\"", + "type": "string" +} - changed
Output schema / properties / provenance / properties / nulled / descriptionPrevious value: -"Numeric values replaced with null because the source does not contain them"New value: +"Values replaced with null: numbers the source does not contain, and string fields whose value cannot be what the field names" - changed
Output schema / properties / provenance / properties / unverified / items / properties / reason / descriptionPrevious value: -"\"not_found_in_source\""New value: +"\"not_found_in_source\" | \"not_a_version\" (a version field with no number, e.g. \"latest\") | \"not_an_identifier\" (a sku/isbn/gtin/upc/ean/mpn field holding a phrase)" - added
Output schema / properties / providerAdded value: +{ + "description": "LLM provider that produced the data (\"ollama\" | \"openai\" | \"anthropic\"); present when extraction_method is \"llm\"", + "type": "string" +}
- Changed
extract_text1 field changed- changed
Input schema / properties / max_length / descriptionPrevious value: -"Maximum characters of text or markdown to return; a longer result is cut and ends with \"...\""New value: +"Maximum characters of text or markdown to return; a longer result is cut, ends with \"...\" and carries truncated:true"
- Changed
extract_with_llm1 field changed- changed
Input schema / properties / model / descriptionPrevious value: -"Override the model. For ollama, pass a name returned by list_ollama_models (e.g. 'llama3.2', 'qwen2.5:7b'). Defaults: openai='gpt-4o-mini', anthropic='claude-haiku-4-5-20251001', ollama='llama3.2' or $OLLAMA_DEFAULT_MODEL."New value: +"Override the model. For ollama, pass a name returned by list_ollama_models (e.g. 'llama3.2', 'qwen2.5:7b'). Defaults: openai='gpt-4o-mini', anthropic='claude-haiku-4-5-20251001', ollama=$OLLAMA_DEFAULT_MODEL, else the best installed model for extraction (gemma3:12b, then gemma3:4b, gpt-oss:20b, mistral:7b, llama3.2, qwen2.5:3b)."
- Changed
generate_llms_txt1 field changed- added
Input schema / properties / max_inline_charsAdded value: +{ + "description": "Largest result to return inline, in characters of its JSON. Over it, the call returns a preview plus a result_handle for read_result instead of the whole result (default 40,000; env CRAWLFORGE_MAX_INLINE_CHARS)", + "maximum": 10000000, + "minimum": 1000, + "type": "integer" +}
- Changed
map_site19 fields changed- added
Input schema / properties / max_inline_charsAdded value: +{ + "description": "Largest result to return inline, in characters of its JSON. Over it, the call returns a preview plus a result_handle for read_result instead of the whole result (default 40,000; env CRAWLFORGE_MAX_INLINE_CHARS)", + "maximum": 10000000, + "minimum": 1000, + "type": "integer" +} - added
Output schema / properties / expires_atAdded value: +{ + "description": "When the stored result is dropped (ISO 8601)", + "type": "string" +} - added
Output schema / properties / previewAdded value: +{ + "description": "The first max_inline_chars characters of the view named by view_path (or of the pretty-printed JSON)", + "type": "string" +} - added
Output schema / properties / result_handleAdded value: +{ + "description": "Handle for read_result; the full result is kept 1 hour", + "type": "string" +} - added
Output schema / properties / site_map / descriptionAdded value: +"The shape of the site as counts; `urls` is the one list of URLs" - added
Output schema / properties / site_map / properties / depth_levels / additionalProperties / typeAdded value: +"number" - added
Output schema / properties / site_map / properties / depth_levels / descriptionAdded value: +"URLs per path depth" - added
Output schema / properties / site_map / properties / root / descriptionAdded value: +"How many URLs are the site root itself" - removed
Output schema / properties / site_map / properties / root / itemsRemoved value: -{ - "type": "string" -} - changed
Output schema / properties / site_map / properties / root / typePrevious value: -"array"New value: +"number" - added
Output schema / properties / site_map / properties / sections / additionalProperties / additionalPropertiesAdded value: +{} - added
Output schema / properties / site_map / properties / sections / additionalProperties / propertiesAdded value: +{ + "count": { + "description": "URLs under this first path segment", + "type": "number" + }, + "subsections": { + "additionalProperties": { + "type": "number" + }, + "description": "URLs per second path segment", + "propertyNames": { + "type": "string" + }, + "type": "object" + } +} - added
Output schema / properties / site_map / properties / sections / additionalProperties / typeAdded value: +"object" - added
Output schema / properties / site_map / properties / sections / descriptionAdded value: +"Counts by first path segment; the URLs themselves are in `urls`" - added
Output schema / properties / total_charsAdded value: +{ + "description": "Length of the full view in characters", + "type": "number" +} - added
Output schema / properties / truncatedAdded value: +{ + "description": "True when the inline result is a preview", + "type": "boolean" +} - added
Output schema / properties / viewAdded value: +{ + "description": "Whether preview and read_result offsets index a text field or the pretty-printed JSON", + "enum": [ + "text", + "json" + ], + "type": "string" +} - added
Output schema / properties / view_pathAdded value: +{ + "description": "Dotted path of the text field the view was cut from; null for the JSON view", + "type": [ + "string", + "null" + ] +} - changed
Output schema / properties / warnings / descriptionPrevious value: -"Present when the map is partial, e.g. the start page could not be read and the URLs come from the sitemap only"New value: +"Present when the map is partial, e.g. the start page could not be read and the URLs come from the sitemap only, or when the result is over max_inline_chars"
- Changed
reddit_search3 fields changed- added
Output schema / properties / comment_count / descriptionAdded value: +"thread mode: comments in `comments` at every depth, stubs excluded; at most limit" - changed
Output schema / properties / comments / descriptionPrevious value: -"thread mode: nested comment tree ({...comment, replies:[...]}); collapsed branches appear as {more_count, more_ids}"New value: +"thread mode: nested comment tree ({...comment, replies:[...]}); a collapsed branch appears as a {more_count, more_ids} stub, where more_count is how many comments it hides, replies included, and more_ids lists only the hidden direct replies, so more_ids can be shorter than more_count" - added
Output schema / properties / comments_collapsedAdded value: +{ + "description": "thread mode: the sum of every stub's more_count, i.e. comments the source holds for this thread that the response leaves out. comment_count + comments_collapsed is what the source holds; post.num_comments is Reddit's own total from when the post was read, so it can be higher (deleted, removed or not yet archived comments) or lower (comments made after that read)", + "type": "number" +}
- Changed
scrape7 fields changed- changed
Input schema / properties / formats / descriptionPrevious value: -"Formats to return (default: [\"markdown\"])"New value: +"Formats to return (default: [\"markdown\"]); at most one highlights and one question format per call" - added
Output schema / properties / content / properties / answer / properties / evidence / items / properties / kind / descriptionAdded value: +"\"heading\" occurs only in question evidence" - changed
Output schema / properties / content / properties / answer / properties / evidence / items / properties / kind / enumPrevious value: -[ - "sentence", - "table_row", - "code_block" -]New value: +[ + "sentence", + "table_row", + "code_block", + "heading" +] - changed
Output schema / properties / content / properties / answer / properties / grounded / descriptionPrevious value: -"True when every number and proper noun in text appears in the evidence or the question; always true in extractive mode"New value: +"False when no unit matched (text is empty) or, in model mode, when the model found no answer in the evidence (text is empty) or used a number or proper noun that appears in neither the evidence nor the question" - added
Output schema / properties / content / properties / highlights / items / properties / kind / descriptionAdded value: +"\"heading\" occurs only in question evidence" - changed
Output schema / properties / content / properties / highlights / items / properties / kind / enumPrevious value: -[ - "sentence", - "table_row", - "code_block" -]New value: +[ + "sentence", + "table_row", + "code_block", + "heading" +] - added
Output schema / properties / content / properties / metadata / properties / title / descriptionAdded value: +"The document <title>; og:title, then the first H1, only when the page has none. og:title itself is og_tags.title"
- Changed
scrape_template1 field changed- added
Input schema / properties / max_inline_charsAdded value: +{ + "description": "Largest result to return inline, in characters of its JSON. Over it, the call returns a preview plus a result_handle for read_result instead of the whole result (default 40,000; env CRAWLFORGE_MAX_INLINE_CHARS)", + "maximum": 10000000, + "minimum": 1000, + "type": "integer" +}
- Changed
search_web15 fields changed- changed
Input schema / properties / expand_query / descriptionPrevious value: -"Expand the query with synonyms/stemming/etc."New value: +"When the query returns no results, search once more with an expanded form (synonyms/stemming/etc.)" - added
Input schema / properties / max_inline_charsAdded value: +{ + "description": "Largest result to return inline, in characters of its JSON. Over it, the call returns a preview plus a result_handle for read_result instead of the whole result (default 40,000; env CRAWLFORGE_MAX_INLINE_CHARS)", + "maximum": 10000000, + "minimum": 1000, + "type": "integer" +} - changed
Input schema / properties / offset / descriptionPrevious value: -"Number of results to skip for pagination"New value: +"Number of results to skip for pagination. For the next page pass the previous response's next_offset, not offset + limit" - added
Output schema / properties / expanded_queries / descriptionAdded value: +"Present only when the original query returned nothing: the queries searched, in order - the original, then its expanded form" - added
Output schema / properties / expires_atAdded value: +{ + "description": "When the stored result is dropped (ISO 8601)", + "type": "string" +} - added
Output schema / properties / next_offsetAdded value: +{ + "description": "The offset to pass for the next page. Duplicates removed from this page are replaced from further down the provider's results, so it can be larger than offset + limit", + "type": "number" +} - added
Output schema / properties / previewAdded value: +{ + "description": "The first max_inline_chars characters of the view named by view_path (or of the pretty-printed JSON)", + "type": "string" +} - added
Output schema / properties / processing / properties / query_expansion / descriptionAdded value: +"{original_query, used_query, search_attempts} when the original query returned nothing and its expanded form was searched too; null otherwise" - removed
Output schema / properties / provider / properties / capabilitiesRemoved value: -{ - "additionalProperties": {}, - "propertyNames": { - "type": "string" - }, - "type": "object" -} - added
Output schema / properties / result_handleAdded value: +{ + "description": "Handle for read_result; the full result is kept 1 hour", + "type": "string" +} - added
Output schema / properties / total_charsAdded value: +{ + "description": "Length of the full view in characters", + "type": "number" +} - added
Output schema / properties / truncatedAdded value: +{ + "description": "True when the inline result is a preview", + "type": "boolean" +} - added
Output schema / properties / viewAdded value: +{ + "description": "Whether preview and read_result offsets index a text field or the pretty-printed JSON", + "enum": [ + "text", + "json" + ], + "type": "string" +} - added
Output schema / properties / view_pathAdded value: +{ + "description": "Dotted path of the text field the view was cut from; null for the JSON view", + "type": [ + "string", + "null" + ] +} - added
Output schema / properties / warningsAdded value: +{ + "description": "Notes on this result; over max_inline_chars, where the full result is kept and how to read it", + "items": { + "type": "string" + }, + "type": "array" +}
- Changed
serp_rank3 fields changed- changed
Input schema / properties / depth / descriptionPrevious value: -"How many results to scan, 10-200 (default 20; DataForSEO bills ~$0.002 per 10 and gets slower the deeper it goes)"New value: +"How many results to scan, 10-200 (default 30; DataForSEO bills ~$0.002 per 10 and gets slower the deeper it goes)" - changed
Input schema / properties / target / descriptionPrevious value: -"Domain or URL to locate in the results (e.g. 'example.com')"New value: +"Domain or URL to locate in the results (e.g. 'example.com'). Matched by host, not exact URL: any page on that host or its subdomains counts" - changed
Output schema / properties / target / descriptionPrevious value: -"Bare target domain, normalized"New value: +"Host the SERP was matched against: a URL target is reduced to its host, and any page on it or its subdomains counts"
- Changed
track_changes1 field changed- added
Input schema / properties / max_inline_charsAdded value: +{ + "description": "Largest result to return inline, in characters of its JSON. Over it, the call returns a preview plus a result_handle for read_result instead of the whole result (default 40,000; env CRAWLFORGE_MAX_INLINE_CHARS)", + "maximum": 10000000, + "minimum": 1000, + "type": "integer" +}
11 tool updates
v6.18.1- Changed
batch_scrape1 field changed- changed
Input schema / properties / includeFailed / descriptionPrevious value: -"Include failed URLs in results"New value: +"List failed URLs in results. false hides the entries; failedUrls still counts them"
- Changed
browser_session1 field changed- added
Input schema / properties / onlyMainContentAdded value: +{ + "default": false, + "description": "read: false (default) returns the whole page as markdown/text (markdown leaves out nav, footer and aside elements, as scrape's does; text leaves out nothing); true keeps only the main article block via Readability, which drops navigation and footers but can also drop list items on listing and app pages. The result's extractionMethod says which ran (\"full_page\" for the whole page)", + "type": "boolean" +}
- Changed
crawl_deep5 fields changed- added
Output schema / properties / error_count / descriptionAdded value: +"URLs that failed; one entry each in errors[]" - added
Output schema / properties / pages_attemptedAdded value: +{ + "description": "URLs the crawl tried: pages_crawled + error_count", + "type": "number" +} - added
Output schema / properties / pages_crawled / descriptionAdded value: +"Pages fetched and returned in results" - added
Output schema / properties / pages_found / descriptionAdded value: +"Same as pages_crawled; kept for existing clients" - added
Output schema / properties / site_structure / properties / total_pages / descriptionAdded value: +"Equals pages_crawled; failed URLs are not counted"
- Changed
extract_embedded_state2 fields changed- added
Input schema / properties / findAdded value: +{ + "description": "Return `matches` instead of `data`: every property with this key name (case-insensitive) anywhere in the selected data (after `path`), in document order, each as {path, preview} with the first 200 characters of its value - at most 50, with `matches_total` and `matches_truncated`. Each match path already starts with `path`, so it can be passed straight back as `path`. Discover where a field lives without downloading the payload. Cannot be combined with keys_only", + "maxLength": 100, + "minLength": 1, + "type": "string" +} - added
Input schema / properties / rawAdded value: +{ + "default": false, + "description": "Also keep the undecoded __NUXT_DATA__ devalue array under json_scripts, beside the decoded nuxt_data. Default: false", + "type": "boolean" +}
- Changed
extract_links3 fields changed- added
Input schema / properties / escalateAdded value: +{ + "default": false, + "description": "When the plain fetch comes back blocked (403/429/444/challenge page/empty shell), re-read the page once in the stealth browser and extract from the rendered document. Under escalate_engine \"auto\" a Chrome TLS handshake with the honest CrawlForge User-Agent (impit) is tried first, and the browser runs only when that does not get the page. A 404 or 5xx never escalates. Projected at 1+5; the actual charge stays at 1 when the plain fetch succeeded. Default: false", + "type": "boolean" +} - added
Input schema / properties / escalate_engineAdded value: +{ + "default": "auto", + "description": "Stealth engine for the escalated retry: \"auto\" (default — Camoufox when it is installed, Chromium otherwise, reported in warnings), \"playwright\" (Chromium) or \"camoufox\"", + "enum": [ + "auto", + "playwright", + "camoufox" + ], + "type": "string" +} - changed
Input schema / properties / filter_external / descriptionPrevious value: -"Only return external links"New value: +"Drop internal (same-host) links"
- Changed
extract_text4 fields changed- added
Input schema / properties / escalateAdded value: +{ + "default": false, + "description": "When the plain fetch comes back blocked (403/429/444/challenge page/empty shell), re-read the page once in the stealth browser and extract from the rendered document. Under escalate_engine \"auto\" a Chrome TLS handshake with the honest CrawlForge User-Agent (impit) is tried first, and the browser runs only when that does not get the page. A 404 or 5xx never escalates. Projected at 1+5; the actual charge stays at 1 when the plain fetch succeeded. Default: false", + "type": "boolean" +} - added
Input schema / properties / escalate_engineAdded value: +{ + "default": "auto", + "description": "Stealth engine for the escalated retry: \"auto\" (default — Camoufox when it is installed, Chromium otherwise, reported in warnings), \"playwright\" (Chromium) or \"camoufox\"", + "enum": [ + "auto", + "playwright", + "camoufox" + ], + "type": "string" +} - added
Input schema / properties / max_lengthAdded value: +{ + "description": "Maximum characters of text or markdown to return; a longer result is cut and ends with \"...\"", + "maximum": 1000000, + "minimum": 1, + "type": "integer" +} - added
Input schema / properties / selectorAdded value: +{ + "description": "CSS selector: read only the matched elements (nav/header/footer are then kept). No match is an error", + "type": "string" +}
- Changed
localization11 fields changed- changed
Input schema / properties / acceptLanguage / descriptionPrevious value: -"Accept-Language header value"New value: +"Accept-Language value to return from configure_country in place of the one built from the language and country" - changed
Input schema / properties / browserOptions / descriptionPrevious value: -"Browser context options for locale emulation"New value: +"Browser context options for localize_browser to build on: returned with the country's locale, timezoneId, geolocation and Accept-Language set, and your userAgent and other extraHTTPHeaders kept" - changed
Input schema / properties / countryCode / descriptionPrevious value: -"ISO 3166-1 alpha-2 country code"New value: +"ISO 3166-1 alpha-2 country code, upper or lower case; operation:\"get_supported_countries\" lists the accepted ones. When omitted, the other operations use the country of the last configure_country call in this server process (US at start) - pass it on every call" - changed
Input schema / properties / currency / descriptionPrevious value: -"ISO 4217 currency code (e.g. 'USD', 'EUR')"New value: +"ISO 4217 currency code (e.g. 'USD', 'EUR') - configure_country only" - changed
Input schema / properties / customHeaders / descriptionPrevious value: -"Custom HTTP headers for localized requests"New value: +"Headers echoed back in the configure_country result; nothing is sent" - changed
Input schema / properties / geoLocation / descriptionPrevious value: -"GPS coordinates for geolocation emulation"New value: +"Coordinates echoed back in the configure_country result; nothing is emulated" - changed
Input schema / properties / language / descriptionPrevious value: -"Language code (e.g. 'en', 'fr', 'de')"New value: +"Language code (e.g. 'en', 'fr', 'de-CH') - configure_country only; other values are refused" - changed
Input schema / properties / operation / descriptionPrevious value: -"Localization operation to perform"New value: +"Localization operation to perform. Every operation returns data; none changes how another tool behaves. configure_country: the country's settings (language, acceptLanguage, timezone, currency, searchDomain, browserLocale). localize_search: runs search_web for searchParams.query in the country and its language. localize_browser: Playwright-style context options for the country (locale, timezoneId, geolocation, extraHTTPHeaders) - no browser is launched. generate_timezone_spoof: a JavaScript snippet that overrides Date and Intl for the timezone - nothing is injected. handle_geo_blocking: classifies a response you supply (url + response) and lists suggestions - nothing is fetched or bypassed. auto_detect: language and country detected in content you supply. get_stats: counters for this server process. get_supported_countries: the accepted country codes." - changed
Input schema / properties / proxySettings / descriptionPrevious value: -"Proxy configuration for geo-targeted requests"New value: +"Echoed back in the configure_country result. No request is routed through it: CrawlForge supplies no proxies, and your own go on stealth_mode stealthConfig.proxyRotation" - changed
Input schema / properties / timezone / descriptionPrevious value: -"IANA timezone identifier (e.g. 'America/New_York')"New value: +"IANA timezone identifier (e.g. 'America/New_York') - configure_country and generate_timezone_spoof" - changed
Input schema / properties / userAgent / descriptionPrevious value: -"Custom user agent string"New value: +"Not used by any operation. To carry a user agent through localize_browser, set browserOptions.userAgent"
- Changed
map_site1 field changed- added
Output schema / properties / warningsAdded value: +{ + "description": "Present when the map is partial, e.g. the start page could not be read and the URLs come from the sitemap only", + "items": { + "type": "string" + }, + "type": "array" +}
- Changed
scrape_with_actions6 fields changed- added
Input schema / properties / browserOptions / properties / proxyRotationAdded value: +{ + "description": "Route a stealth chain through your own proxies; requires stealth:true and the Chromium engine, and is refused without stealth or on camoufox (including \"auto\" when it resolves to camoufox), which shares one launch-time proxy across calls. Each entry is a proxy URL — \"http://user:pass@host:port\" (percent-encode a password containing @ : or /), or a bare \"host:port\" for an unauthenticated HTTP proxy; http, https, socks4 and socks5 are accepted. rotationInterval is the minimum ms on one proxy before the list advances. A proxy on a loopback, link-local or cloud-metadata address is refused, and credentials are removed from the result. Cloudflare scores the IP before it serves a challenge, so a residential proxy is what gets past a block that no fingerprint fixes. CrawlForge supplies no proxies.", + "properties": { + "enabled": { + "default": false, + "type": "boolean" + }, + "proxies": { + "items": { + "type": "string" + }, + "type": "array" + }, + "rotationInterval": { + "default": 300000, + "type": "number" + } + }, + "type": "object" +} - changed
Input schema / properties / captureIntermediateStates / descriptionPrevious value: -"Capture page state after each action"New value: +"Capture page state after each action: intermediateStates[] gets one entry per action (capturePoint is the action's 1-based position) with the url, title and the requested formats at that point" - removed
Input schema / properties / captureScreenshotsRemoved value: -{ - "default": true, - "description": "Take screenshots during action execution", - "type": "boolean" -} - added
Input schema / properties / formAutoFill / properties / fields / items / properties / type / descriptionAdded value: +"How the field is filled: text types the value, select picks the option by value or label, checkbox/radio check the input (a selector matching a group checks the member whose value attribute equals value; a checkbox with value \"false\" is unchecked). file is not supported and fails that action" - changed
Input schema / properties / formats / descriptionPrevious value: -"Output formats for scraped content"New value: +"Output formats for scraped content. For an image, add a {type:\"screenshot\"} action where you want it taken." - changed
Input schema / properties / formats / items / enumPrevious value: -[ - "markdown", - "html", - "json", - "text", - "screenshots" -]New value: +[ + "markdown", + "html", + "json", + "text" +]
- Changed
stealth_mode4 fields changed- changed
Input schema / properties / operation / descriptionPrevious value: -"Stealth operation to perform"New value: +"Stealth operation to perform. scrape: render one URL and return the formats. configure: validate stealthConfig and return it with defaults filled in; nothing is stored. create_context / create_page: a context to open pages in. get_stats: this server process's open contexts, browser and proxy state. cleanup: close the browser and every context." - changed
Input schema / properties / operation / enumPrevious value: -[ - "scrape", - "configure", - "enable", - "disable", - "create_context", - "create_page", - "get_stats", - "cleanup" -]New value: +[ + "scrape", + "configure", + "create_context", + "create_page", + "get_stats", + "cleanup" +] - changed
Input schema / properties / stealthConfig / descriptionPrevious value: -"Stealth browser configuration with anti-detection settings"New value: +"Stealth browser configuration with anti-detection settings. Applies to the call it is passed on and is not remembered. Chromium applies customUserAgent, customViewport, locale and timezone per call. On camoufox customUserAgent and customViewport are not applied, locale is fixed when the browser launches (run operation:\"cleanup\" first to change it), and behind a proxy locale and timezone come from the exit IP; each of these is reported in `warnings`." - changed
Input schema / properties / stealthConfig / properties / proxyRotation / descriptionPrevious value: -"Route the browser through your own proxies. Each entry is a proxy URL — \"http://user:pass@host:port\" (percent-encode a password containing @ : or /), or a bare \"host:port\" for an unauthenticated HTTP proxy; http, https, socks4 and socks5 are accepted. rotationInterval is the minimum ms on one proxy before the list advances. Cloudflare scores the IP before it serves a challenge, so a residential proxy is what gets past a block that no fingerprint fixes. CrawlForge supplies no proxies."New value: +"Route the browser through your own proxies. Each entry is a proxy URL — \"http://user:pass@host:port\" (percent-encode a password containing @ : or /), or a bare \"host:port\" for an unauthenticated HTTP proxy; http, https, socks4 and socks5 are accepted. rotationInterval is the minimum ms on one proxy before the list advances. Cloudflare scores the IP before it serves a challenge, so a residential proxy is what gets past a block that no fingerprint fixes. Requires the Chromium engine: refused on camoufox (including \"auto\" when it resolves to camoufox), which shares one launch-time proxy across calls. CrawlForge supplies no proxies."
- Changed
track_changes17 fields changed- added
Input schema / properties / alertRuleOptions / properties / condition / descriptionAdded value: +"significance <operator> <level>, e.g. significance >= \"moderate\" (default: significance === \"major\"). Operators: ===, ==, !==, !=, >=, <=, >, <. Levels: none, minor, moderate, major, critical; quotes optional. Anything else is rejected" - changed
Input schema / properties / monitoringOptions / properties / enabled / defaultPrevious value: -falseNew value: +true - added
Input schema / properties / monitoringOptions / properties / enabled / descriptionAdded value: +"operation:\"monitor\" starts polling the URL; false stops the URL's polling monitor and starts nothing" - changed
Input schema / properties / scheduledMonitorOptions / properties / monitorId / descriptionPrevious value: -"Monitor id for stop_scheduled_monitor"New value: +"Monitor id for stop_scheduled_monitor, as list_scheduled_monitors shows it (poll:<url> for a polling monitor)" - added
Input schema / properties / scheduledMonitorOptions / properties / templateId / descriptionAdded value: +"Preset from get_monitoring_templates; it sets interval, goal, notificationThreshold and trackingOptions, and any of those passed explicitly wins" - removed
Input schema / properties / trackingOptions / properties / granularity / defaultRemoved value: -"section" - added
Input schema / properties / trackingOptions / properties / granularity / descriptionAdded value: +"Default: section" - removed
Input schema / properties / trackingOptions / properties / ignoreCase / defaultRemoved value: -false - added
Input schema / properties / trackingOptions / properties / ignoreCase / descriptionAdded value: +"Default: false" - removed
Input schema / properties / trackingOptions / properties / ignoreWhitespace / defaultRemoved value: -true - added
Input schema / properties / trackingOptions / properties / ignoreWhitespace / descriptionAdded value: +"Default: true" - removed
Input schema / properties / trackingOptions / properties / trackLinks / defaultRemoved value: -true - added
Input schema / properties / trackingOptions / properties / trackLinks / descriptionAdded value: +"Default: true" - removed
Input schema / properties / trackingOptions / properties / trackStructure / defaultRemoved value: -true - added
Input schema / properties / trackingOptions / properties / trackStructure / descriptionAdded value: +"Default: true" - removed
Input schema / properties / trackingOptions / properties / trackText / defaultRemoved value: -true - added
Input schema / properties / trackingOptions / properties / trackText / descriptionAdded value: +"Default: true"
3 tool updates
v6.14.0- Changed
extract_embedded_state4 fields changed- added
Input schema / properties / escalateAdded value: +{ + "default": false, + "description": "When the plain fetch comes back blocked (403/429/challenge page, or an empty shell with no state), re-read the page once in the stealth browser and run the same parser on the rendered document; the browser also reads the framework globals off window (window_state). Under escalate_engine \"auto\" a Chrome TLS handshake with the honest CrawlForge User-Agent (impit) is tried first, and the browser runs only when that does not get the page. A 404 or 5xx never escalates. Projected at 2+5; the actual charge stays at 2 when the plain fetch succeeded. Default: false", + "type": "boolean" +} - added
Input schema / properties / escalate_engineAdded value: +{ + "default": "auto", + "description": "Stealth engine for the escalated retry: \"auto\" (default — Camoufox when it is installed, Chromium otherwise, reported in warnings), \"playwright\" (Chromium) or \"camoufox\"", + "enum": [ + "auto", + "playwright", + "camoufox" + ], + "type": "string" +} - added
Input schema / properties / keys_onlyAdded value: +{ + "default": false, + "description": "Return `keys` instead of `data`: the first two levels of keys of the selected data (after `path`), each value replaced by its type (\"object\", \"array(<n>)\", \"string\", \"number\", \"boolean\", \"null\"); an array shows its length and its first item. Cheap discovery before choosing a path. Default: false", + "type": "boolean" +} - added
Input schema / properties / wait_forAdded value: +{ + "description": "Escalated render only: extra wait after page load, in ms — for state assigned after DOMContentLoaded. Ignored without escalation", + "maximum": 30000, + "minimum": 0, + "type": "number" +}
- Changed
scrape2 fields changed- changed
Input schema / properties / escalate / descriptionPrevious value: -"When the plain fetch comes back blocked (403/429/challenge page/empty shell), retry once in the stealth browser and return its content instead of the block. Projected at 2+5; the actual charge stays at the base price when the plain fetch succeeded. Default: false"New value: +"When the plain fetch comes back blocked (403/429/challenge page/empty shell), retry once in the stealth browser and return its content instead of the block. Under escalate_engine \"auto\" a Chrome TLS handshake with the honest CrawlForge User-Agent (impit) is tried first, and the browser runs only when that does not get the page. Projected at 2+5; the actual charge stays at the base price when the plain fetch succeeded. Default: false" - changed
Output schema / properties / stealth / descriptionPrevious value: -"Present when escalated is true: the stealth engine that ran, and the bot-defence vendor the plain fetch hit (null when the block named none)"New value: +"Present when escalated is true: the stealth engine that ran (\"impit\" when the Chrome TLS handshake got the page without a browser), and the bot-defence vendor the plain fetch hit (null when the block named none)"
- Changed
scrape_with_actions5 fields changed- changed
Input schema / properties / actions / items / properties / selector / descriptionPrevious value: -"A CSS selector, or a @e1 ref from an earlier snapshot action in this chain"New value: +"A CSS selector, or a ref from an earlier snapshot action in this chain (@e1, including elements inside shadow roots and iframes)" - added
Input schema / properties / actions / items / properties / type / descriptionAdded value: +"executeJavaScript is disabled unless the server runs with ALLOW_JAVASCRIPT_EXECUTION=true; on the hosted API it is refused." - added
Input schema / properties / browserOptions / properties / consentAdded value: +{ + "default": "off", + "description": "Cookie/consent banner handling, using DuckDuckGo autoconsent's rules for known consent platforms (OneTrust, Sourcepoint, Didomi ...). \"reject\" declines non-essential cookies and \"accept\" accepts them, once after the initial page load and again after each navigate action, for at most 2s each; the result's `consent` ({cmp, action, ms}) and each navigate result say what was found and done. A banner no rule matches is left in place (action \"none\") and never fails the chain. Default \"off\": the page is left as it loads, so a chain that clicks the banner itself works unchanged.", + "enum": [ + "off", + "reject", + "accept" + ], + "type": "string" +} - changed
Input schema / properties / maxRetries / defaultPrevious value: -1New value: +0 - changed
Input schema / properties / maxRetries / descriptionPrevious value: -"Maximum retry attempts on failure"New value: +"Whole-chain retries on failure (0-3). A retry re-navigates to the starting URL and replays every action; each attempt is reported under attempts[]."
4 tool updates
v6.8.0- Changed
browser_session1 field changed- added
Input schema / properties / engineAdded value: +{ + "default": "auto", + "description": "open: stealth engine for the session, with stealth:true. \"auto\" (default) runs camoufox when it is installed and Chromium otherwise; every operation echoes the `engine` that actually ran. \"camoufox\" is Firefox-based with a higher anti-detect score; \"chromium\" (= \"playwright\") forces Chromium. Refused without stealth:true, where the browser is always Chromium.", + "enum": [ + "auto", + "chromium", + "camoufox", + "playwright" + ], + "type": "string" +}
- Changed
scrape3 fields changed- changed
Input schema / properties / escalate_engine / defaultPrevious value: -"playwright"New value: +"auto" - changed
Input schema / properties / escalate_engine / descriptionPrevious value: -"Stealth engine for the escalated retry (default: \"playwright\")"New value: +"Stealth engine for the escalated retry: \"auto\" (default — Camoufox when it is installed, Chromium otherwise, reported in warnings), \"playwright\" (Chromium) or \"camoufox\"" - changed
Input schema / properties / escalate_engine / enumPrevious value: -[ - "playwright", - "camoufox" -]New value: +[ + "auto", + "playwright", + "camoufox" +]
- Changed
scrape_with_actions1 field changed- added
Input schema / properties / browserOptions / properties / engineAdded value: +{ + "default": "auto", + "description": "Stealth engine for the chain, with stealth:true. \"auto\" (default) runs camoufox when it is installed and Chromium otherwise; the result's `engine` says which ran. \"camoufox\" is Firefox-based with a higher anti-detect score; \"chromium\" (= \"playwright\") forces Chromium. Refused without stealth:true, where the browser is always Chromium.", + "enum": [ + "auto", + "chromium", + "camoufox", + "playwright" + ], + "type": "string" +}
- Changed
stealth_mode4 fields changed- changed
Input schema / properties / engine / defaultPrevious value: -"playwright"New value: +"auto" - changed
Input schema / properties / engine / descriptionPrevious value: -"Browser engine: \"playwright\" (Chromium, default) or \"camoufox\" (Firefox-based, higher anti-detect score — install with npm install camoufox)"New value: +"Browser engine: \"auto\" (default — camoufox when it is installed, Chromium otherwise, and the result says which), \"camoufox\" (Firefox-based, higher anti-detect score; fails if not installed), or \"chromium\" (\"playwright\" is the same engine under its old name)" - changed
Input schema / properties / engine / enumPrevious value: -[ - "playwright", - "camoufox" -]New value: +[ + "auto", + "chromium", + "camoufox", + "playwright" +] - added
Input schema / properties / stealthConfig / properties / proxyRotation / descriptionAdded value: +"Route the browser through your own proxies. Each entry is a proxy URL — \"http://user:pass@host:port\" (percent-encode a password containing @ : or /), or a bare \"host:port\" for an unauthenticated HTTP proxy; http, https, socks4 and socks5 are accepted. rotationInterval is the minimum ms on one proxy before the list advances. Cloudflare scores the IP before it serves a challenge, so a residential proxy is what gets past a block that no fingerprint fixes. CrawlForge supplies no proxies."
2 tool updates
v6.6.0- Added
browser_session - Changed
scrape_with_actions4 fields changed- added
Input schema / properties / actions / items / properties / interactiveOnlyAdded value: +{ + "description": "snapshot: only interactive elements (default true)", + "type": "boolean" +} - added
Input schema / properties / actions / items / properties / maxNodesAdded value: +{ + "description": "snapshot: cap on nodes listed (default 200, max 1000); the result says truncated when the cap stopped the walk", + "maximum": 1000, + "minimum": 1, + "type": "number" +} - added
Input schema / properties / actions / items / properties / selector / descriptionAdded value: +"A CSS selector, or a @e1 ref from an earlier snapshot action in this chain" - changed
Input schema / properties / actions / items / properties / type / enumPrevious value: -[ - "wait", - "click", - "type", - "press", - "scroll", - "screenshot", - "executeJavaScript", - "select", - "hover", - "navigate" -]New value: +[ + "snapshot", + "wait", + "click", + "type", + "press", + "scroll", + "screenshot", + "executeJavaScript", + "select", + "hover", + "navigate" +]
4 tool updates
v6.5.0- Changed
extract_structured1 field changed- changed
Output schema / properties / success / descriptionPrevious value: -"False when the extraction errored or a required field came back missing or empty"New value: +"False when the extraction errored, or a required field came back missing, empty, or in the wrong shape"
- Changed
generate_llms_txt1 field changed- changed
Input schema / properties / outputOptions / properties / contactEmail / patternPrevious value: -"^(?!\\.)(?!.*\\.\\.)([A-Za-z0-9_'+\\-\\.]*)[A-Za-z0-9_+-]@([A-Za-z0-9][A-Za-z0-9\\-]*\\.)+[A-Za-z]{2,}$"New value: +"^(?:[A-Za-z0-9_'+\\-]+\\.)*[A-Za-z0-9_'+\\-]*[A-Za-z0-9_+-]@(?:[A-Za-z0-9][A-Za-z0-9\\-]*\\.)+[A-Za-z]{2,}$"
- Changed
get_batch_results1 field changed- added
Input schema / properties / max_inline_charsAdded value: +{ + "description": "Largest result to return inline, in characters of its JSON. Over it, the call returns a preview plus a result_handle for read_result instead of the whole result (default 40,000; env CRAWLFORGE_MAX_INLINE_CHARS)", + "maximum": 10000000, + "minimum": 1000, + "type": "integer" +}
- Changed
track_changes9 fields changed- added
Input schema / properties / monitoringOptions / defaultAdded value: +{} - changed
Input schema / properties / notificationOptions / descriptionPrevious value: -"Notification configuration for webhooks and Slack"New value: +"Notification configuration for webhooks, Slack and email (email is sent by hosted monitors only)" - added
Input schema / properties / notificationOptions / properties / emailAdded value: +{ + "properties": { + "enabled": { + "default": false, + "type": "boolean" + }, + "includeDetails": { + "default": true, + "type": "boolean" + }, + "recipients": { + "items": { + "format": "email", + "pattern": "^(?:[A-Za-z0-9_'+\\-]+\\.)*[A-Za-z0-9_'+\\-]*[A-Za-z0-9_+-]@(?:[A-Za-z0-9][A-Za-z0-9\\-]*\\.)+[A-Za-z]{2,}$", + "type": "string" + }, + "type": "array" + }, + "subject": { + "type": "string" + } + }, + "type": "object" +} - added
Input schema / properties / queryOptions / defaultAdded value: +{} - added
Input schema / properties / scheduledMonitorOptions / properties / hostedAdded value: +{ + "default": false, + "description": "Run the monitor on CrawlForge's servers: it fires from the hosted scheduler whether or not this process is alive and sends email and signed webhooks. Each check bills 3 credits per compared target from the account; blocked and errored targets are free. Default false = local, in-process.", + "type": "boolean" +} - added
Input schema / properties / scheduledMonitorOptions / properties / nameAdded value: +{ + "description": "Display name for a hosted monitor (default: the URL host)", + "maxLength": 80, + "minLength": 1, + "type": "string" +} - added
Input schema / properties / storageOptions / defaultAdded value: +{} - added
Input schema / properties / trackingOptions / defaultAdded value: +{} - added
Input schema / properties / trackingOptions / properties / excludeSelectors / defaultAdded value: +[ + "script", + "style", + "noscript", + ".advertisement", + ".ad", + "#comments" +]
29 tool updates
v6.0.0- Changed
agent2 fields changed- removed
Input schema / additionalPropertiesRemoved value: -false - added
Input schema / properties / schema / propertyNamesAdded value: +{ + "type": "string" +}
- Changed
analyze_content2 fields changed- removed
Input schema / additionalPropertiesRemoved value: -false - changed
Input schema / properties / options / additionalPropertiesPrevious value: -trueNew value: +{}
- Changed
batch_scrape8 fields changed- removed
Input schema / additionalPropertiesRemoved value: -false - added
Input schema / properties / extractionSchema / propertyNamesAdded value: +{ + "type": "string" +} - removed
Input schema / properties / jobOptions / additionalPropertiesRemoved value: -false - added
Input schema / properties / max_inline_charsAdded value: +{ + "description": "Largest result to return inline, in characters of its JSON. Over it, the call returns a preview plus a result_handle for read_result instead of the whole result (default 40,000; env CRAWLFORGE_MAX_INLINE_CHARS)", + "maximum": 10000000, + "minimum": 1000, + "type": "integer" +} - added
Input schema / properties / redact_piiAdded value: +{ + "anyOf": [ + { + "type": "boolean" + }, + { + "properties": { + "entities": { + "description": "Which classes to redact, case-insensitive: EMAIL, PHONE, FINANCIAL, SECRET, plus PERSON and LOCATION when mode is \"model\". Omitted or empty means all four regex classes (and both model classes in \"model\" mode). An unknown name, or a model-only name without mode:\"model\", is rejected", + "items": { + "type": "string" + }, + "type": "array" + }, + "mode": { + "description": "\"fast\" (default) is regex only and free; \"model\" adds an Ollama NER pass for PERSON and LOCATION (+3 credits once per call)", + "enum": [ + "fast", + "model" + ], + "type": "string" + }, + "replace_style": { + "description": "\"tag\" (default) writes <EMAIL>, \"mask\" writes [REDACTED], \"remove\" deletes the value", + "enum": [ + "tag", + "mask", + "remove" + ], + "type": "string" + } + }, + "type": "object" + } + ], + "description": "Redact personal data from the text this call returns, before it reaches your context window. true means the free regex pass over EMAIL, PHONE, FINANCIAL and SECRET. The result carries redaction:{entities,count}. Default: off" +} - changed
Input schema / properties / urls / items / anyOfPrevious value: -[ - { - "format": "uri", - "type": "string" - }, - { - "additionalProperties": false, - "properties": { - "headers": { - "additionalProperties": { - "type": "string" - }, - "type": "object" - }, - "metadata": { - "additionalProperties": {}, - "type": "object" - }, - "selectors": { - "additionalProperties": { - "type": "string" - }, - "type": "object" - }, - "timeout": { - "maximum": 30000, - "minimum": 1000, - "type": "number" - }, - "url": { - "format": "uri", - "type": "string" - } - }, - "required": [ - "url" - ], - "type": "object" - } -]New value: +[ + { + "format": "uri", + "type": "string" + }, + { + "properties": { + "headers": { + "additionalProperties": { + "type": "string" + }, + "propertyNames": { + "type": "string" + }, + "type": "object" + }, + "metadata": { + "additionalProperties": {}, + "propertyNames": { + "type": "string" + }, + "type": "object" + }, + "selectors": { + "additionalProperties": { + "type": "string" + }, + "propertyNames": { + "type": "string" + }, + "type": "object" + }, + "timeout": { + "maximum": 30000, + "minimum": 1000, + "type": "number" + }, + "url": { + "format": "uri", + "type": "string" + } + }, + "required": [ + "url" + ], + "type": "object" + } +] - removed
Input schema / properties / webhook / additionalPropertiesRemoved value: -false - added
Input schema / properties / webhook / properties / headers / propertyNamesAdded value: +{ + "type": "string" +}
- Changed
crawl_deep28 fields changed- removed
Input schema / additionalPropertiesRemoved value: -false - removed
Input schema / properties / domain_filter / additionalPropertiesRemoved value: -false - added
Input schema / properties / domain_filter / properties / blacklist / itemsAdded value: +{} - added
Input schema / properties / domain_filter / properties / domain_rules / propertyNamesAdded value: +{ + "type": "string" +} - added
Input schema / properties / domain_filter / properties / whitelist / itemsAdded value: +{} - removed
Input schema / properties / link_analysis_options / additionalPropertiesRemoved value: -false - added
Input schema / properties / max_inline_charsAdded value: +{ + "description": "Largest result to return inline, in characters of its JSON. Over it, the call returns a preview plus a result_handle for read_result instead of the whole result (default 40,000; env CRAWLFORGE_MAX_INLINE_CHARS)", + "maximum": 10000000, + "minimum": 1000, + "type": "integer" +} - added
Input schema / properties / redact_piiAdded value: +{ + "anyOf": [ + { + "type": "boolean" + }, + { + "properties": { + "entities": { + "description": "Which classes to redact, case-insensitive: EMAIL, PHONE, FINANCIAL, SECRET, plus PERSON and LOCATION when mode is \"model\". Omitted or empty means all four regex classes (and both model classes in \"model\" mode). An unknown name, or a model-only name without mode:\"model\", is rejected", + "items": { + "type": "string" + }, + "type": "array" + }, + "mode": { + "description": "\"fast\" (default) is regex only and free; \"model\" adds an Ollama NER pass for PERSON and LOCATION (+3 credits once per call)", + "enum": [ + "fast", + "model" + ], + "type": "string" + }, + "replace_style": { + "description": "\"tag\" (default) writes <EMAIL>, \"mask\" writes [REDACTED], \"remove\" deletes the value", + "enum": [ + "tag", + "mask", + "remove" + ], + "type": "string" + } + }, + "type": "object" + } + ], + "description": "Redact personal data from the text this call returns, before it reaches your context window. true means the free regex pass over EMAIL, PHONE, FINANCIAL and SECRET. The result carries redaction:{entities,count}. Default: off" +} - removed
Input schema / properties / session / additionalPropertiesRemoved value: -false - added
Input schema / properties / session / properties / headers / propertyNamesAdded value: +{ + "type": "string" +} - removed
Input schema / properties / session / properties / initialRequest / additionalPropertiesRemoved value: -false - added
Input schema / properties / session / properties / initialRequest / properties / headers / propertyNamesAdded value: +{ + "type": "string" +} - changed
Output schema / properties / _cost / additionalPropertiesPrevious value: -trueNew value: +{} - added
Output schema / properties / expires_atAdded value: +{ + "description": "When the stored result is dropped (ISO 8601)", + "type": "string" +} - added
Output schema / properties / previewAdded value: +{ + "description": "The first max_inline_chars characters of the view named by view_path (or of the pretty-printed JSON)", + "type": "string" +} - added
Output schema / properties / redactionAdded value: +{ + "additionalProperties": {}, + "description": "Present when redact_pii was set: what was redacted from the text of this result", + "properties": { + "count": { + "description": "Total spans replaced", + "type": "number" + }, + "entities": { + "additionalProperties": { + "type": "number" + }, + "description": "How many spans were replaced, by entity class; a class with no hits is omitted", + "propertyNames": { + "type": "string" + }, + "type": "object" + }, + "mode": { + "description": "\"fast\" is the free regex pass; \"model\" added an Ollama NER pass for PERSON and LOCATION", + "enum": [ + "fast", + "model" + ], + "type": "string" + }, + "model_ran": { + "description": "mode \"model\" only: whether a model actually answered. False means no LLM route existed and the model surcharge was not charged", + "type": "boolean" + } + }, + "type": "object" +} - added
Output schema / properties / result_handleAdded value: +{ + "description": "Handle for read_result; the full result is kept 1 hour", + "type": "string" +} - changed
Output schema / properties / results / items / additionalPropertiesPrevious value: -trueNew value: +{} - changed
Output schema / properties / session / additionalPropertiesPrevious value: -trueNew value: +{} - changed
Output schema / properties / site_structure / additionalPropertiesPrevious value: -trueNew value: +{} - added
Output schema / properties / site_structure / properties / depth_distribution / propertyNamesAdded value: +{ + "type": "string" +} - added
Output schema / properties / site_structure / properties / file_types / propertyNamesAdded value: +{ + "type": "string" +} - added
Output schema / properties / site_structure / properties / path_depth_distribution / propertyNamesAdded value: +{ + "type": "string" +} - added
Output schema / properties / site_structure / properties / path_patterns / propertyNamesAdded value: +{ + "type": "string" +} - added
Output schema / properties / total_charsAdded value: +{ + "description": "Length of the full view in characters", + "type": "number" +} - added
Output schema / properties / truncatedAdded value: +{ + "description": "True when the inline result is a preview", + "type": "boolean" +} - added
Output schema / properties / viewAdded value: +{ + "description": "Whether preview and read_result offsets index a text field or the pretty-printed JSON", + "enum": [ + "text", + "json" + ], + "type": "string" +} - added
Output schema / properties / view_pathAdded value: +{ + "description": "Dotted path of the text field the view was cut from; null for the JSON view", + "type": [ + "string", + "null" + ] +}
- Changed
deep_research9 fields changed- removed
Input schema / additionalPropertiesRemoved value: -false - removed
Input schema / properties / llmConfig / additionalPropertiesRemoved value: -false - removed
Input schema / properties / llmConfig / properties / anthropic / additionalPropertiesRemoved value: -false - removed
Input schema / properties / llmConfig / properties / ollama / additionalPropertiesRemoved value: -false - removed
Input schema / properties / llmConfig / properties / openai / additionalPropertiesRemoved value: -false - added
Input schema / properties / max_inline_charsAdded value: +{ + "description": "Largest result to return inline, in characters of its JSON. Over it, the call returns a preview plus a result_handle for read_result instead of the whole result (default 40,000; env CRAWLFORGE_MAX_INLINE_CHARS)", + "maximum": 10000000, + "minimum": 1000, + "type": "integer" +} - removed
Input schema / properties / queryExpansion / additionalPropertiesRemoved value: -false - removed
Input schema / properties / webhook / additionalPropertiesRemoved value: -false - added
Input schema / properties / webhook / properties / headers / propertyNamesAdded value: +{ + "type": "string" +}
- Changed
extract_content4 fields changed- removed
Input schema / additionalPropertiesRemoved value: -false - added
Input schema / properties / max_inline_charsAdded value: +{ + "description": "Largest result to return inline, in characters of its JSON. Over it, the call returns a preview plus a result_handle for read_result instead of the whole result (default 40,000; env CRAWLFORGE_MAX_INLINE_CHARS)", + "maximum": 10000000, + "minimum": 1000, + "type": "integer" +} - changed
Input schema / properties / options / additionalPropertiesPrevious value: -trueNew value: +{} - added
Input schema / properties / redact_piiAdded value: +{ + "anyOf": [ + { + "type": "boolean" + }, + { + "properties": { + "entities": { + "description": "Which classes to redact, case-insensitive: EMAIL, PHONE, FINANCIAL, SECRET, plus PERSON and LOCATION when mode is \"model\". Omitted or empty means all four regex classes (and both model classes in \"model\" mode). An unknown name, or a model-only name without mode:\"model\", is rejected", + "items": { + "type": "string" + }, + "type": "array" + }, + "mode": { + "description": "\"fast\" (default) is regex only and free; \"model\" adds an Ollama NER pass for PERSON and LOCATION (+3 credits once per call)", + "enum": [ + "fast", + "model" + ], + "type": "string" + }, + "replace_style": { + "description": "\"tag\" (default) writes <EMAIL>, \"mask\" writes [REDACTED], \"remove\" deletes the value", + "enum": [ + "tag", + "mask", + "remove" + ], + "type": "string" + } + }, + "type": "object" + } + ], + "description": "Redact personal data from the text this call returns, before it reaches your context window. true means the free regex pass over EMAIL, PHONE, FINANCIAL and SECRET. The result carries redaction:{entities,count}. Default: off" +}
- Changed
extract_embedded_state2 fields changed- removed
Input schema / additionalPropertiesRemoved value: -false - added
Input schema / properties / max_inline_charsAdded value: +{ + "description": "Largest result to return inline, in characters of its JSON. Over it, the call returns a preview plus a result_handle for read_result instead of the whole result (default 40,000; env CRAWLFORGE_MAX_INLINE_CHARS)", + "maximum": 10000000, + "minimum": 1000, + "type": "integer" +}
- Changed
extract_links1 field changed- removed
Input schema / additionalPropertiesRemoved value: -false
- Changed
extract_metadata1 field changed- removed
Input schema / additionalPropertiesRemoved value: -false
- Changed
extract_structured11 fields changed- removed
Input schema / additionalPropertiesRemoved value: -false - removed
Input schema / properties / llmConfig / additionalPropertiesRemoved value: -false - removed
Input schema / properties / schema / additionalPropertiesRemoved value: -false - added
Input schema / properties / schema / properties / properties / propertyNamesAdded value: +{ + "type": "string" +} - added
Input schema / properties / selectorHints / propertyNamesAdded value: +{ + "type": "string" +} - changed
Output schema / properties / _cost / additionalPropertiesPrevious value: -trueNew value: +{} - added
Output schema / properties / data / propertyNamesAdded value: +{ + "type": "string" +} - changed
Output schema / properties / provenance / additionalPropertiesPrevious value: -trueNew value: +{} - changed
Output schema / properties / provenance / properties / unverified / items / additionalPropertiesPrevious value: -trueNew value: +{} - added
Output schema / properties / schema_used / propertyNamesAdded value: +{ + "type": "string" +} - changed
Output schema / properties / validation / additionalPropertiesPrevious value: -trueNew value: +{}
- Changed
extract_text2 fields changed- removed
Input schema / additionalPropertiesRemoved value: -false - added
Input schema / properties / redact_piiAdded value: +{ + "anyOf": [ + { + "type": "boolean" + }, + { + "properties": { + "entities": { + "description": "Which classes to redact, case-insensitive: EMAIL, PHONE, FINANCIAL, SECRET, plus PERSON and LOCATION when mode is \"model\". Omitted or empty means all four regex classes (and both model classes in \"model\" mode). An unknown name, or a model-only name without mode:\"model\", is rejected", + "items": { + "type": "string" + }, + "type": "array" + }, + "mode": { + "description": "\"fast\" (default) is regex only and free; \"model\" adds an Ollama NER pass for PERSON and LOCATION (+3 credits once per call)", + "enum": [ + "fast", + "model" + ], + "type": "string" + }, + "replace_style": { + "description": "\"tag\" (default) writes <EMAIL>, \"mask\" writes [REDACTED], \"remove\" deletes the value", + "enum": [ + "tag", + "mask", + "remove" + ], + "type": "string" + } + }, + "type": "object" + } + ], + "description": "Redact personal data from the text this call returns, before it reaches your context window. true means the free regex pass over EMAIL, PHONE, FINANCIAL and SECRET. The result carries redaction:{entities,count}. Default: off" +}
- Changed
extract_with_llm2 fields changed- removed
Input schema / additionalPropertiesRemoved value: -false - added
Input schema / properties / schema / propertyNamesAdded value: +{ + "type": "string" +}
- Changed
fetch_url3 fields changed- removed
Input schema / additionalPropertiesRemoved value: -false - added
Input schema / properties / headers / propertyNamesAdded value: +{ + "type": "string" +} - added
Input schema / properties / max_inline_charsAdded value: +{ + "description": "Largest result to return inline, in characters of its JSON. Over it, the call returns a preview plus a result_handle for read_result instead of the whole result (default 40,000; env CRAWLFORGE_MAX_INLINE_CHARS)", + "maximum": 10000000, + "minimum": 1000, + "type": "integer" +}
- Changed
generate_llms_txt4 fields changed- removed
Input schema / additionalPropertiesRemoved value: -false - removed
Input schema / properties / analysisOptions / additionalPropertiesRemoved value: -false - removed
Input schema / properties / outputOptions / additionalPropertiesRemoved value: -false - added
Input schema / properties / outputOptions / properties / contactEmail / patternAdded value: +"^(?!\\.)(?!.*\\.\\.)([A-Za-z0-9_'+\\-\\.]*)[A-Za-z0-9_+-]@([A-Za-z0-9][A-Za-z0-9\\-]*\\.)+[A-Za-z]{2,}$"
- Changed
get_batch_results1 field changed- removed
Input schema / additionalPropertiesRemoved value: -false
- Changed
localization11 fields changed- removed
Input schema / additionalPropertiesRemoved value: -false - removed
Input schema / properties / browserOptions / additionalPropertiesRemoved value: -false - added
Input schema / properties / browserOptions / properties / extraHTTPHeaders / propertyNamesAdded value: +{ + "type": "string" +} - added
Input schema / properties / customHeaders / propertyNamesAdded value: +{ + "type": "string" +} - removed
Input schema / properties / geoLocation / additionalPropertiesRemoved value: -false - removed
Input schema / properties / proxySettings / additionalPropertiesRemoved value: -false - removed
Input schema / properties / proxySettings / properties / fallback / additionalPropertiesRemoved value: -false - removed
Input schema / properties / proxySettings / properties / rotation / additionalPropertiesRemoved value: -false - removed
Input schema / properties / response / additionalPropertiesRemoved value: -false - removed
Input schema / properties / searchParams / additionalPropertiesRemoved value: -false - added
Input schema / properties / searchParams / properties / headers / propertyNamesAdded value: +{ + "type": "string" +}
- Changed
map_site12 fields changed- removed
Input schema / additionalPropertiesRemoved value: -false - removed
Input schema / properties / domain_filter / additionalPropertiesRemoved value: -false - changed
Output schema / properties / _cost / additionalPropertiesPrevious value: -trueNew value: +{} - added
Output schema / properties / metadata / propertyNamesAdded value: +{ + "type": "string" +} - changed
Output schema / properties / ranked_urls / items / additionalPropertiesPrevious value: -trueNew value: +{} - changed
Output schema / properties / site_map / additionalPropertiesPrevious value: -trueNew value: +{} - added
Output schema / properties / site_map / properties / depth_levels / propertyNamesAdded value: +{ + "type": "string" +} - added
Output schema / properties / site_map / properties / sections / propertyNamesAdded value: +{ + "type": "string" +} - changed
Output schema / properties / statistics / additionalPropertiesPrevious value: -trueNew value: +{} - added
Output schema / properties / statistics / properties / file_extensions / propertyNamesAdded value: +{ + "type": "string" +} - changed
Output schema / properties / statistics / properties / url_lengths / additionalPropertiesPrevious value: -trueNew value: +{} - changed
Output schema / properties / urls / anyOfPrevious value: -[ - { - "items": { - "type": "string" - }, - "type": "array" - }, - { - "additionalProperties": { - "items": { - "type": "string" - }, - "type": "array" - }, - "type": "object" - } -]New value: +[ + { + "items": { + "type": "string" + }, + "type": "array" + }, + { + "additionalProperties": { + "items": { + "type": "string" + }, + "type": "array" + }, + "propertyNames": { + "type": "string" + }, + "type": "object" + } +]
- Changed
process_document4 fields changed- removed
Input schema / additionalPropertiesRemoved value: -false - added
Input schema / properties / max_inline_charsAdded value: +{ + "description": "Largest result to return inline, in characters of its JSON. Over it, the call returns a preview plus a result_handle for read_result instead of the whole result (default 40,000; env CRAWLFORGE_MAX_INLINE_CHARS)", + "maximum": 10000000, + "minimum": 1000, + "type": "integer" +} - changed
Input schema / properties / options / additionalPropertiesPrevious value: -trueNew value: +{} - added
Input schema / properties / redact_piiAdded value: +{ + "anyOf": [ + { + "type": "boolean" + }, + { + "properties": { + "entities": { + "description": "Which classes to redact, case-insensitive: EMAIL, PHONE, FINANCIAL, SECRET, plus PERSON and LOCATION when mode is \"model\". Omitted or empty means all four regex classes (and both model classes in \"model\" mode). An unknown name, or a model-only name without mode:\"model\", is rejected", + "items": { + "type": "string" + }, + "type": "array" + }, + "mode": { + "description": "\"fast\" (default) is regex only and free; \"model\" adds an Ollama NER pass for PERSON and LOCATION (+3 credits once per call)", + "enum": [ + "fast", + "model" + ], + "type": "string" + }, + "replace_style": { + "description": "\"tag\" (default) writes <EMAIL>, \"mask\" writes [REDACTED], \"remove\" deletes the value", + "enum": [ + "tag", + "mask", + "remove" + ], + "type": "string" + } + }, + "type": "object" + } + ], + "description": "Redact personal data from the text this call returns, before it reaches your context window. true means the free regex pass over EMAIL, PHONE, FINANCIAL and SECRET. The result carries redaction:{entities,count}. Default: off" +}
- Added
read_result - Changed
reddit_search4 fields changed- removed
Input schema / additionalPropertiesRemoved value: -false - changed
Output schema / properties / _cost / additionalPropertiesPrevious value: -trueNew value: +{} - changed
Output schema / properties / post / anyOfPrevious value: -[ - { - "$ref": "#/properties/results/items/anyOf/0" - }, - { - "type": "null" - } -]New value: +[ + { + "additionalProperties": {}, + "properties": { + "author": { + "type": [ + "string", + "null" + ] + }, + "created_iso": { + "type": [ + "string", + "null" + ] + }, + "created_utc": { + "type": [ + "number", + "null" + ] + }, + "id": { + "type": [ + "string", + "null" + ] + }, + "num_comments": { + "type": [ + "number", + "null" + ] + }, + "permalink": { + "description": "Full reddit.com URL of the post", + "type": [ + "string", + "null" + ] + }, + "score": { + "type": [ + "number", + "null" + ] + }, + "selftext": { + "type": [ + "string", + "null" + ] + }, + "selftext_truncated": { + "type": "boolean" + }, + "subreddit": { + "type": [ + "string", + "null" + ] + }, + "title": { + "type": [ + "string", + "null" + ] + }, + "url": { + "type": [ + "string", + "null" + ] + } + }, + "type": "object" + }, + { + "type": "null" + } +] - changed
Output schema / properties / results / items / anyOfPrevious value: -[ - { - "additionalProperties": true, - "properties": { - "author": { - "type": [ - "string", - "null" - ] - }, - "created_iso": { - "type": [ - "string", - "null" - ] - }, - "created_utc": { - "type": [ - "number", - "null" - ] - }, - "id": { - "type": [ - "string", - "null" - ] - }, - "num_comments": { - "type": [ - "number", - "null" - ] - }, - "permalink": { - "description": "Full reddit.com URL of the post", - "type": [ - "string", - "null" - ] - }, - "score": { - "type": [ - "number", - "null" - ] - }, - "selftext": { - "type": [ - "string", - "null" - ] - }, - "selftext_truncated": { - "type": "boolean" - }, - "subreddit": { - "type": [ - "string", - "null" - ] - }, - "title": { - "type": [ - "string", - "null" - ] - }, - "url": { - "type": [ - "string", - "null" - ] - } - }, - "type": "object" - }, - { - "additionalProperties": true, - "properties": { - "author": { - "type": [ - "string", - "null" - ] - }, - "body": { - "type": [ - "string", - "null" - ] - }, - "body_truncated": { - "type": "boolean" - }, - "created_iso": { - "type": [ - "string", - "null" - ] - }, - "created_utc": { - "type": [ - "number", - "null" - ] - }, - "id": { - "type": [ - "string", - "null" - ] - }, - "link_id": { - "type": [ - "string", - "null" - ] - }, - "parent_id": { - "type": [ - "string", - "null" - ] - }, - "permalink": { - "type": [ - "string", - "null" - ] - }, - "score": { - "type": [ - "number", - "null" - ] - }, - "subreddit": { - "type": [ - "string", - "null" - ] - } - }, - "type": "object" - } -]New value: +[ + { + "additionalProperties": {}, + "properties": { + "author": { + "type": [ + "string", + "null" + ] + }, + "created_iso": { + "type": [ + "string", + "null" + ] + }, + "created_utc": { + "type": [ + "number", + "null" + ] + }, + "id": { + "type": [ + "string", + "null" + ] + }, + "num_comments": { + "type": [ + "number", + "null" + ] + }, + "permalink": { + "description": "Full reddit.com URL of the post", + "type": [ + "string", + "null" + ] + }, + "score": { + "type": [ + "number", + "null" + ] + }, + "selftext": { + "type": [ + "string", + "null" + ] + }, + "selftext_truncated": { + "type": "boolean" + }, + "subreddit": { + "type": [ + "string", + "null" + ] + }, + "title": { + "type": [ + "string", + "null" + ] + }, + "url": { + "type": [ + "string", + "null" + ] + } + }, + "type": "object" + }, + { + "additionalProperties": {}, + "properties": { + "author": { + "type": [ + "string", + "null" + ] + }, + "body": { + "type": [ + "string", + "null" + ] + }, + "body_truncated": { + "type": "boolean" + }, + "created_iso": { + "type": [ + "string", + "null" + ] + }, + "created_utc": { + "type": [ + "number", + "null" + ] + }, + "id": { + "type": [ + "string", + "null" + ] + }, + "link_id": { + "type": [ + "string", + "null" + ] + }, + "parent_id": { + "type": [ + "string", + "null" + ] + }, + "permalink": { + "type": [ + "string", + "null" + ] + }, + "score": { + "type": [ + "number", + "null" + ] + }, + "subreddit": { + "type": [ + "string", + "null" + ] + } + }, + "type": "object" + } +]
- Changed
scrape33 fields changed- removed
Input schema / additionalPropertiesRemoved value: -false - removed
Input schema / properties / brandingOptions / additionalPropertiesRemoved value: -false - added
Input schema / properties / escalateAdded value: +{ + "default": false, + "description": "When the plain fetch comes back blocked (403/429/challenge page/empty shell), retry once in the stealth browser and return its content instead of the block. Projected at 2+5; the actual charge stays at the base price when the plain fetch succeeded. Default: false", + "type": "boolean" +} - added
Input schema / properties / escalate_engineAdded value: +{ + "default": "playwright", + "description": "Stealth engine for the escalated retry (default: \"playwright\")", + "enum": [ + "playwright", + "camoufox" + ], + "type": "string" +} - changed
Input schema / properties / formats / items / anyOfPrevious value: -[ - { - "enum": [ - "markdown", - "html", - "rawHtml", - "text", - "links", - "metadata", - "screenshot", - "branding" - ], - "type": "string" - }, - { - "additionalProperties": false, - "properties": { - "prompt": { - "description": "Extraction instruction for the LLM", - "type": "string" - }, - "schema": { - "additionalProperties": {}, - "description": "JSON schema for extraction", - "type": "object" - }, - "type": { - "const": "json", - "type": "string" - } - }, - "required": [ - "type" - ], - "type": "object" - } -]New value: +[ + { + "enum": [ + "markdown", + "html", + "rawHtml", + "text", + "links", + "metadata", + "screenshot", + "branding" + ], + "type": "string" + }, + { + "properties": { + "prompt": { + "description": "Extraction instruction for the LLM", + "type": "string" + }, + "schema": { + "additionalProperties": {}, + "description": "JSON schema for extraction", + "propertyNames": { + "type": "string" + }, + "type": "object" + }, + "type": { + "const": "json", + "type": "string" + } + }, + "required": [ + "type" + ], + "type": "object" + }, + { + "properties": { + "max_highlights": { + "default": 10, + "description": "How many units to return (default 10)", + "maximum": 50, + "minimum": 1, + "type": "integer" + }, + "mode": { + "default": "extractive", + "description": "\"extractive\" (default) returns verbatim page text, no model; \"model\" adds an LLM step (+3 credits)", + "enum": [ + "extractive", + "model" + ], + "type": "string" + }, + "query": { + "description": "What to look for; the matching sentences, table rows and code blocks come back verbatim with offsets into the markdown", + "maxLength": 500, + "minLength": 1, + "type": "string" + }, + "type": { + "const": "highlights", + "type": "string" + } + }, + "required": [ + "type", + "query" + ], + "type": "object" + }, + { + "properties": { + "mode": { + "default": "extractive", + "description": "\"extractive\" (default) returns verbatim page text, no model; \"model\" adds an LLM step (+3 credits)", + "enum": [ + "extractive", + "model" + ], + "type": "string" + }, + "question": { + "description": "The question to answer from the page; the evidence units come back verbatim with offsets", + "maxLength": 500, + "minLength": 1, + "type": "string" + }, + "type": { + "const": "question", + "type": "string" + } + }, + "required": [ + "type", + "question" + ], + "type": "object" + } +] - added
Input schema / properties / max_inline_charsAdded value: +{ + "description": "Largest result to return inline, in characters of its JSON. Over it, the call returns a preview plus a result_handle for read_result instead of the whole result (default 40,000; env CRAWLFORGE_MAX_INLINE_CHARS)", + "maximum": 10000000, + "minimum": 1000, + "type": "integer" +} - added
Input schema / properties / redact_piiAdded value: +{ + "anyOf": [ + { + "type": "boolean" + }, + { + "properties": { + "entities": { + "description": "Which classes to redact, case-insensitive: EMAIL, PHONE, FINANCIAL, SECRET, plus PERSON and LOCATION when mode is \"model\". Omitted or empty means all four regex classes (and both model classes in \"model\" mode). An unknown name, or a model-only name without mode:\"model\", is rejected", + "items": { + "type": "string" + }, + "type": "array" + }, + "mode": { + "description": "\"fast\" (default) is regex only and free; \"model\" adds an Ollama NER pass for PERSON and LOCATION (+3 credits once per call)", + "enum": [ + "fast", + "model" + ], + "type": "string" + }, + "replace_style": { + "description": "\"tag\" (default) writes <EMAIL>, \"mask\" writes [REDACTED], \"remove\" deletes the value", + "enum": [ + "tag", + "mask", + "remove" + ], + "type": "string" + } + }, + "type": "object" + } + ], + "description": "Redact personal data from the text this call returns, before it reaches your context window. true means the free regex pass over EMAIL, PHONE, FINANCIAL and SECRET. The result carries redaction:{entities,count}. Default: off" +} - removed
Input schema / properties / screenshotOptions / additionalPropertiesRemoved value: -false - changed
Output schema / properties / _cost / additionalPropertiesPrevious value: -trueNew value: +{} - added
Output schema / properties / blockedAdded value: +{ + "additionalProperties": {}, + "description": "Present when a bot-defence vendor served a challenge page; the fallback hint names the tool to try next", + "properties": { + "evidence": { + "type": "string" + }, + "vendor": { + "type": "string" + } + }, + "type": "object" +} - changed
Output schema / properties / content / additionalPropertiesPrevious value: -trueNew value: +{} - added
Output schema / properties / content / properties / answerAdded value: +{ + "additionalProperties": {}, + "description": "Result of the {type:\"question\"} format", + "properties": { + "evidence": { + "description": "The units the answer rests on, verbatim with offsets", + "items": { + "additionalProperties": {}, + "properties": { + "kind": { + "enum": [ + "sentence", + "table_row", + "code_block" + ], + "type": "string" + }, + "length": { + "type": "number" + }, + "offset": { + "description": "JS string index into the markdown format of this call", + "type": "number" + }, + "score": { + "description": "BM25 relevance to the query, higher is better", + "type": "number" + }, + "text": { + "description": "Verbatim page text: markdown.slice(offset, offset + length) === text", + "type": "string" + } + }, + "type": "object" + }, + "type": "array" + }, + "grounded": { + "description": "True when every number and proper noun in text appears in the evidence or the question; always true in extractive mode", + "type": "boolean" + }, + "text": { + "description": "Extractive mode: the evidence texts joined; model mode: the model's answer", + "type": "string" + } + }, + "type": "object" +} - added
Output schema / properties / content / properties / branding / propertyNamesAdded value: +{ + "type": "string" +} - added
Output schema / properties / content / properties / highlightsAdded value: +{ + "description": "Result of the {type:\"highlights\"} format: the units matching the query, best first, verbatim with offsets", + "items": { + "additionalProperties": {}, + "properties": { + "kind": { + "enum": [ + "sentence", + "table_row", + "code_block" + ], + "type": "string" + }, + "length": { + "type": "number" + }, + "offset": { + "description": "JS string index into the markdown format of this call", + "type": "number" + }, + "score": { + "description": "BM25 relevance to the query, higher is better", + "type": "number" + }, + "text": { + "description": "Verbatim page text: markdown.slice(offset, offset + length) === text", + "type": "string" + } + }, + "type": "object" + }, + "type": "array" +} - changed
Output schema / properties / content / properties / links / additionalPropertiesPrevious value: -trueNew value: +{} - changed
Output schema / properties / content / properties / links / properties / links / items / additionalPropertiesPrevious value: -trueNew value: +{} - changed
Output schema / properties / content / properties / metadata / additionalPropertiesPrevious value: -trueNew value: +{} - added
Output schema / properties / content / properties / metadata / properties / og_tags / propertyNamesAdded value: +{ + "type": "string" +} - added
Output schema / properties / content / properties / metadata / properties / twitter_tags / propertyNamesAdded value: +{ + "type": "string" +} - changed
Output schema / properties / content / properties / screenshots / items / additionalPropertiesPrevious value: -trueNew value: +{} - added
Output schema / properties / errorAdded value: +{ + "description": "Why success is false: a challenge page, an empty shell or an error placeholder was served instead of the content", + "type": "string" +} - added
Output schema / properties / escalatedAdded value: +{ + "description": "Present only when escalate:true was passed: whether the blocked plain fetch was retried in the stealth browser. False means the plain fetch sufficed and the call is charged at the base price", + "type": "boolean" +} - added
Output schema / properties / expires_atAdded value: +{ + "description": "When the stored result is dropped (ISO 8601)", + "type": "string" +} - added
Output schema / properties / previewAdded value: +{ + "description": "The first max_inline_chars characters of the view named by view_path (or of the pretty-printed JSON)", + "type": "string" +} - added
Output schema / properties / redactionAdded value: +{ + "additionalProperties": {}, + "description": "Present when redact_pii was set: what was redacted from the text of this result", + "properties": { + "count": { + "description": "Total spans replaced", + "type": "number" + }, + "entities": { + "additionalProperties": { + "type": "number" + }, + "description": "How many spans were replaced, by entity class; a class with no hits is omitted", + "propertyNames": { + "type": "string" + }, + "type": "object" + }, + "mode": { + "description": "\"fast\" is the free regex pass; \"model\" added an Ollama NER pass for PERSON and LOCATION", + "enum": [ + "fast", + "model" + ], + "type": "string" + }, + "model_ran": { + "description": "mode \"model\" only: whether a model actually answered. False means no LLM route existed and the model surcharge was not charged", + "type": "boolean" + } + }, + "type": "object" +} - added
Output schema / properties / result_handleAdded value: +{ + "description": "Handle for read_result; the full result is kept 1 hour", + "type": "string" +} - added
Output schema / properties / statusAdded value: +{ + "description": "HTTP status of the fetch; present when success is false", + "type": "number" +} - added
Output schema / properties / stealthAdded value: +{ + "additionalProperties": {}, + "description": "Present when escalated is true: the stealth engine that ran, and the bot-defence vendor the plain fetch hit (null when the block named none)", + "properties": { + "engine": { + "type": "string" + }, + "vendor_detected": { + "type": [ + "string", + "null" + ] + } + }, + "type": "object" +} - added
Output schema / properties / titleAdded value: +{ + "description": "Document title; present when success is false", + "type": "string" +} - added
Output schema / properties / total_charsAdded value: +{ + "description": "Length of the full view in characters", + "type": "number" +} - added
Output schema / properties / truncatedAdded value: +{ + "description": "True when the inline result is a preview", + "type": "boolean" +} - added
Output schema / properties / viewAdded value: +{ + "description": "Whether preview and read_result offsets index a text field or the pretty-printed JSON", + "enum": [ + "text", + "json" + ], + "type": "string" +} - added
Output schema / properties / view_pathAdded value: +{ + "description": "Dotted path of the text field the view was cut from; null for the JSON view", + "type": [ + "string", + "null" + ] +}
- Changed
scrape_structured3 fields changed- removed
Input schema / additionalPropertiesRemoved value: -false - added
Input schema / properties / max_results / maximumAdded value: +9007199254740991 - added
Input schema / properties / selectors / propertyNamesAdded value: +{ + "type": "string" +}
- Changed
scrape_template2 fields changed- removed
Input schema / additionalPropertiesRemoved value: -false - added
Input schema / properties / params / propertyNamesAdded value: +{ + "type": "string" +}
- Changed
scrape_with_actions11 fields changed- removed
Input schema / additionalPropertiesRemoved value: -false - removed
Input schema / properties / actions / items / additionalPropertiesRemoved value: -false - added
Input schema / properties / actions / items / properties / args / itemsAdded value: +{} - removed
Input schema / properties / actions / items / properties / position / additionalPropertiesRemoved value: -false - removed
Input schema / properties / browserOptions / additionalPropertiesRemoved value: -false - removed
Input schema / properties / extractionOptions / additionalPropertiesRemoved value: -false - added
Input schema / properties / extractionOptions / properties / selectors / propertyNamesAdded value: +{ + "type": "string" +} - removed
Input schema / properties / formAutoFill / additionalPropertiesRemoved value: -false - removed
Input schema / properties / formAutoFill / properties / fields / items / additionalPropertiesRemoved value: -false - added
Input schema / properties / max_inline_charsAdded value: +{ + "description": "Largest result to return inline, in characters of its JSON. Over it, the call returns a preview plus a result_handle for read_result instead of the whole result (default 40,000; env CRAWLFORGE_MAX_INLINE_CHARS)", + "maximum": 10000000, + "minimum": 1000, + "type": "integer" +} - added
Input schema / properties / redact_piiAdded value: +{ + "anyOf": [ + { + "type": "boolean" + }, + { + "properties": { + "entities": { + "description": "Which classes to redact, case-insensitive: EMAIL, PHONE, FINANCIAL, SECRET, plus PERSON and LOCATION when mode is \"model\". Omitted or empty means all four regex classes (and both model classes in \"model\" mode). An unknown name, or a model-only name without mode:\"model\", is rejected", + "items": { + "type": "string" + }, + "type": "array" + }, + "mode": { + "description": "\"fast\" (default) is regex only and free; \"model\" adds an Ollama NER pass for PERSON and LOCATION (+3 credits once per call)", + "enum": [ + "fast", + "model" + ], + "type": "string" + }, + "replace_style": { + "description": "\"tag\" (default) writes <EMAIL>, \"mask\" writes [REDACTED], \"remove\" deletes the value", + "enum": [ + "tag", + "mask", + "remove" + ], + "type": "string" + } + }, + "type": "object" + } + ], + "description": "Redact personal data from the text this call returns, before it reaches your context window. true means the free regex pass over EMAIL, PHONE, FINANCIAL and SECRET. The result carries redaction:{entities,count}. Default: off" +}
- Changed
search_web25 fields changed- removed
Input schema / additionalPropertiesRemoved value: -false - removed
Input schema / properties / deduplication_thresholds / additionalPropertiesRemoved value: -false - removed
Input schema / properties / expansion_options / additionalPropertiesRemoved value: -false - removed
Input schema / properties / localization / additionalPropertiesRemoved value: -false - removed
Input schema / properties / localization / properties / customLocation / additionalPropertiesRemoved value: -false - added
Input schema / properties / queriesAdded value: +{ + "description": "Run 1-10 searches in one call instead of 10 round-trips; every other parameter applies to each. Results come back in results_by_query, one entry per query, in order. Costs 5 per query. Use this OR query, not both", + "items": { + "minLength": 1, + "type": "string" + }, + "maxItems": 10, + "minItems": 1, + "type": "array" +} - changed
Input schema / properties / query / descriptionPrevious value: -"Search query string"New value: +"Search query string. Use this OR queries, not both" - removed
Input schema / properties / ranking_weights / additionalPropertiesRemoved value: -false - added
Input schema / properties / redact_piiAdded value: +{ + "anyOf": [ + { + "type": "boolean" + }, + { + "properties": { + "entities": { + "description": "Which classes to redact, case-insensitive: EMAIL, PHONE, FINANCIAL, SECRET, plus PERSON and LOCATION when mode is \"model\". Omitted or empty means all four regex classes (and both model classes in \"model\" mode). An unknown name, or a model-only name without mode:\"model\", is rejected", + "items": { + "type": "string" + }, + "type": "array" + }, + "mode": { + "description": "\"fast\" (default) is regex only and free; \"model\" adds an Ollama NER pass for PERSON and LOCATION (+3 credits once per call)", + "enum": [ + "fast", + "model" + ], + "type": "string" + }, + "replace_style": { + "description": "\"tag\" (default) writes <EMAIL>, \"mask\" writes [REDACTED], \"remove\" deletes the value", + "enum": [ + "tag", + "mask", + "remove" + ], + "type": "string" + } + }, + "type": "object" + } + ], + "description": "Redact personal data from the text this call returns, before it reaches your context window. true means the free regex pass over EMAIL, PHONE, FINANCIAL and SECRET. The result carries redaction:{entities,count}. Default: off" +} - removed
Input schema / requiredRemoved value: -[ - "query" -] - changed
Output schema / properties / _cost / additionalPropertiesPrevious value: -trueNew value: +{} - added
Output schema / properties / countAdded value: +{ + "description": "Batch form: how many queries ran", + "type": "number" +} - changed
Output schema / properties / localization / anyOfPrevious value: -[ - { - "additionalProperties": true, - "properties": { - "applied": { - "type": "boolean" - }, - "countryCode": { - "type": "string" - }, - "geoTargeting": { - "type": "boolean" - }, - "language": { - "type": "string" - }, - "searchDomain": { - "type": "string" - } - }, - "type": "object" - }, - { - "type": "null" - } -]New value: +[ + { + "additionalProperties": {}, + "properties": { + "applied": { + "type": "boolean" + }, + "countryCode": { + "type": "string" + }, + "geoTargeting": { + "type": "boolean" + }, + "language": { + "type": "string" + }, + "searchDomain": { + "type": "string" + } + }, + "type": "object" + }, + { + "type": "null" + } +] - changed
Output schema / properties / processing / additionalPropertiesPrevious value: -trueNew value: +{} - changed
Output schema / properties / processing / properties / deduplication / anyOfPrevious value: -[ - { - "additionalProperties": {}, - "type": "object" - }, - { - "type": "null" - } -]New value: +[ + { + "additionalProperties": {}, + "propertyNames": { + "type": "string" + }, + "type": "object" + }, + { + "type": "null" + } +] - changed
Output schema / properties / processing / properties / query_expansion / anyOfPrevious value: -[ - { - "additionalProperties": {}, - "type": "object" - }, - { - "type": "null" - } -]New value: +[ + { + "additionalProperties": {}, + "propertyNames": { + "type": "string" + }, + "type": "object" + }, + { + "type": "null" + } +] - changed
Output schema / properties / processing / properties / ranking / anyOfPrevious value: -[ - { - "additionalProperties": {}, - "type": "object" - }, - { - "type": "null" - } -]New value: +[ + { + "additionalProperties": {}, + "propertyNames": { + "type": "string" + }, + "type": "object" + }, + { + "type": "null" + } +] - changed
Output schema / properties / provider / additionalPropertiesPrevious value: -trueNew value: +{} - added
Output schema / properties / provider / properties / capabilities / propertyNamesAdded value: +{ + "type": "string" +} - added
Output schema / properties / queriesAdded value: +{ + "description": "Batch form: the queries that ran, in order", + "items": { + "type": "string" + }, + "type": "array" +} - added
Output schema / properties / redactionAdded value: +{ + "additionalProperties": {}, + "description": "Present when redact_pii was set: what was redacted from the text of this result", + "properties": { + "count": { + "description": "Total spans replaced", + "type": "number" + }, + "entities": { + "additionalProperties": { + "type": "number" + }, + "description": "How many spans were replaced, by entity class; a class with no hits is omitted", + "propertyNames": { + "type": "string" + }, + "type": "object" + }, + "mode": { + "description": "\"fast\" is the free regex pass; \"model\" added an Ollama NER pass for PERSON and LOCATION", + "enum": [ + "fast", + "model" + ], + "type": "string" + }, + "model_ran": { + "description": "mode \"model\" only: whether a model actually answered. False means no LLM route existed and the model surcharge was not charged", + "type": "boolean" + } + }, + "type": "object" +} - changed
Output schema / properties / results / items / additionalPropertiesPrevious value: -trueNew value: +{} - added
Output schema / properties / results / items / properties / metadata / propertyNamesAdded value: +{ + "type": "string" +} - added
Output schema / properties / results / items / properties / pagemap / propertyNamesAdded value: +{ + "type": "string" +} - added
Output schema / properties / results_by_queryAdded value: +{ + "description": "Batch form: one entry per query, in order", + "items": { + "additionalProperties": {}, + "properties": { + "error": { + "description": "Present when this query failed; the other queries in the batch are unaffected", + "type": "string" + }, + "query": { + "type": "string" + }, + "results": { + "items": { + "additionalProperties": {}, + "properties": { + "displayLink": { + "type": "string" + }, + "formattedUrl": { + "type": "string" + }, + "htmlSnippet": { + "type": "string" + }, + "link": { + "type": "string" + }, + "metadata": { + "additionalProperties": {}, + "propertyNames": { + "type": "string" + }, + "type": "object" + }, + "pagemap": { + "additionalProperties": {}, + "propertyNames": { + "type": "string" + }, + "type": "object" + }, + "snippet": { + "type": "string" + }, + "title": { + "type": "string" + } + }, + "type": "object" + }, + "type": "array" + } + }, + "type": "object" + }, + "type": "array" +}
- Changed
serp_rank7 fields changed- removed
Input schema / additionalPropertiesRemoved value: -false - changed
Output schema / properties / _cost / additionalPropertiesPrevious value: -trueNew value: +{} - changed
Output schema / properties / allPositions / items / additionalPropertiesPrevious value: -trueNew value: +{} - removed
Output schema / properties / results / items / $refRemoved value: -"#/properties/allPositions/items" - added
Output schema / properties / results / items / additionalPropertiesAdded value: +{} - added
Output schema / properties / results / items / propertiesAdded value: +{ + "domain": { + "type": "string" + }, + "position": { + "type": [ + "number", + "null" + ] + }, + "rankAbsolute": { + "type": [ + "number", + "null" + ] + }, + "snippet": { + "type": [ + "string", + "null" + ] + }, + "title": { + "type": [ + "string", + "null" + ] + }, + "url": { + "type": [ + "string", + "null" + ] + } +} - added
Output schema / properties / results / items / typeAdded value: +"object"
- Changed
stealth_mode8 fields changed- removed
Input schema / additionalPropertiesRemoved value: -false - added
Input schema / properties / max_inline_charsAdded value: +{ + "description": "Largest result to return inline, in characters of its JSON. Over it, the call returns a preview plus a result_handle for read_result instead of the whole result (default 40,000; env CRAWLFORGE_MAX_INLINE_CHARS)", + "maximum": 10000000, + "minimum": 1000, + "type": "integer" +} - added
Input schema / properties / redact_piiAdded value: +{ + "anyOf": [ + { + "type": "boolean" + }, + { + "properties": { + "entities": { + "description": "Which classes to redact, case-insensitive: EMAIL, PHONE, FINANCIAL, SECRET, plus PERSON and LOCATION when mode is \"model\". Omitted or empty means all four regex classes (and both model classes in \"model\" mode). An unknown name, or a model-only name without mode:\"model\", is rejected", + "items": { + "type": "string" + }, + "type": "array" + }, + "mode": { + "description": "\"fast\" (default) is regex only and free; \"model\" adds an Ollama NER pass for PERSON and LOCATION (+3 credits once per call)", + "enum": [ + "fast", + "model" + ], + "type": "string" + }, + "replace_style": { + "description": "\"tag\" (default) writes <EMAIL>, \"mask\" writes [REDACTED], \"remove\" deletes the value", + "enum": [ + "tag", + "mask", + "remove" + ], + "type": "string" + } + }, + "type": "object" + } + ], + "description": "Redact personal data from the text this call returns, before it reaches your context window. true means the free regex pass over EMAIL, PHONE, FINANCIAL and SECRET. The result carries redaction:{entities,count}. Default: off" +} - removed
Input schema / properties / stealthConfig / additionalPropertiesRemoved value: -false - removed
Input schema / properties / stealthConfig / properties / antiDetection / additionalPropertiesRemoved value: -false - removed
Input schema / properties / stealthConfig / properties / customViewport / additionalPropertiesRemoved value: -false - removed
Input schema / properties / stealthConfig / properties / fingerprinting / additionalPropertiesRemoved value: -false - removed
Input schema / properties / stealthConfig / properties / proxyRotation / additionalPropertiesRemoved value: -false
- Changed
summarize_content2 fields changed- removed
Input schema / additionalPropertiesRemoved value: -false - changed
Input schema / properties / options / additionalPropertiesPrevious value: -trueNew value: +{}
- Changed
track_changes14 fields changed- removed
Input schema / additionalPropertiesRemoved value: -false - removed
Input schema / properties / alertRuleOptions / additionalPropertiesRemoved value: -false - removed
Input schema / properties / dashboardOptions / additionalPropertiesRemoved value: -false - removed
Input schema / properties / exportOptions / additionalPropertiesRemoved value: -false - removed
Input schema / properties / monitoringOptions / additionalPropertiesRemoved value: -false - removed
Input schema / properties / notificationOptions / additionalPropertiesRemoved value: -false - removed
Input schema / properties / notificationOptions / properties / slack / additionalPropertiesRemoved value: -false - removed
Input schema / properties / notificationOptions / properties / webhook / additionalPropertiesRemoved value: -false - added
Input schema / properties / notificationOptions / properties / webhook / properties / headers / propertyNamesAdded value: +{ + "type": "string" +} - removed
Input schema / properties / queryOptions / additionalPropertiesRemoved value: -false - removed
Input schema / properties / scheduledMonitorOptions / additionalPropertiesRemoved value: -false - removed
Input schema / properties / storageOptions / additionalPropertiesRemoved value: -false - removed
Input schema / properties / trackingOptions / additionalPropertiesRemoved value: -false - removed
Input schema / properties / trackingOptions / properties / significanceThresholds / additionalPropertiesRemoved value: -false
3 tool updates
v5.6.6- Changed
deep_research3 fields changed- changed
Input schema / properties / llmConfig / descriptionPrevious value: -"LLM provider configuration for AI-powered analysis"New value: +"LLM provider configuration for AI-powered analysis. provider 'auto' (default) uses a configured cloud key if there is one, else the local Ollama (http://localhost:11434, no key); 'ollama' forces the local model; 'openai'/'anthropic' need the matching API key" - added
Input schema / properties / llmConfig / properties / ollamaAdded value: +{ + "additionalProperties": false, + "properties": { + "embeddingModel": { + "type": "string" + }, + "model": { + "type": "string" + } + }, + "type": "object" +} - changed
Input schema / properties / llmConfig / properties / provider / enumPrevious value: -[ - "auto", - "openai", - "anthropic" -]New value: +[ + "auto", + "openai", + "anthropic", + "ollama" +]
- Changed
reddit_search5 fields changed- changed
Input schema / properties / source / descriptionPrevious value: -"Backend: auto routes + falls back (default). reddit_api uses the official Reddit Data API — only when REDDIT_CLIENT_ID/REDDIT_CLIENT_SECRET are set; serves posts/thread, not comment search"New value: +"Backend: auto routes + falls back (default). web_discovery serves only unscoped keyword searches (web search finds the posts, the archive supplies the rows). reddit_api uses the official Reddit Data API — only when REDDIT_CLIENT_ID/REDDIT_CLIENT_SECRET are set; serves posts/thread, not comment search" - changed
Input schema / properties / source / enumPrevious value: -[ - "auto", - "arctic_shift", - "pullpush", - "reddit_api" -]New value: +[ + "auto", + "arctic_shift", + "pullpush", + "reddit_api", + "web_discovery" +] - added
Output schema / properties / discoveredAdded value: +{ + "description": "web_discovery: how many post ids the site-restricted web search surfaced before archive hydration", + "type": "number" +} - added
Output schema / properties / posts_searchedAdded value: +{ + "description": "web_discovery comments mode: how many discovered posts had their comments searched before limit was reached", + "type": "number" +} - added
Output schema / properties / window_appliedAdded value: +{ + "description": "arctic_shift comments mode: the after-window (\"7d\"/\"3d\"/\"1d\") the search was narrowed to after the full-history search timed out; absent when the caller set after or no narrowing was needed", + "type": "string" +}
- Changed
scrape_with_actions1 field changed- changed
Input schema / properties / extractionOptions / descriptionPrevious value: -"Content extraction options"New value: +"Content extraction options. selectors results are returned as content.json.extracted, so include \"json\" in formats when passing selectors — without it the extraction is not part of the response."
18 tool updates
v5.4.0- Changed
batch_scrape2 fields changed- added
Input schema / properties / respect_robotsAdded value: +{ + "description": "Respect the target site's robots.txt (default: true). Setting this to false is honoured, returns a warning in the response, and is recorded against your API key — it is your decision, not a silent default.", + "type": "boolean" +} - added
Input schema / properties / user_agentAdded value: +{ + "description": "Override the outbound User-Agent. CrawlForge identifies itself honestly by default; use this only for targets you have your own agreement with.", + "type": "string" +}
- Changed
crawl_deep2 fields changed- added
Output schema / properties / site_structure / properties / depth_distribution / descriptionAdded value: +"Pages per crawl depth (links from the start URL)" - added
Output schema / properties / site_structure / properties / path_depth_distributionAdded value: +{ + "additionalProperties": { + "type": "number" + }, + "description": "Pages per URL path-segment depth", + "type": "object" +}
- Changed
extract_content2 fields changed- added
Input schema / properties / respect_robotsAdded value: +{ + "description": "Respect the target site's robots.txt (default: true). Setting this to false is honoured, returns a warning in the response, and is recorded against your API key — it is your decision, not a silent default.", + "type": "boolean" +} - added
Input schema / properties / user_agentAdded value: +{ + "description": "Override the outbound User-Agent. CrawlForge identifies itself honestly by default; use this only for targets you have your own agreement with.", + "type": "string" +}
- Added
extract_embedded_state - Changed
extract_links2 fields changed- added
Input schema / properties / respect_robotsAdded value: +{ + "description": "Respect the target site's robots.txt (default: true). Setting this to false is honoured, returns a warning in the response, and is recorded against your API key — it is your decision, not a silent default.", + "type": "boolean" +} - added
Input schema / properties / user_agentAdded value: +{ + "description": "Override the outbound User-Agent. CrawlForge identifies itself honestly by default; use this only for targets you have your own agreement with.", + "type": "string" +}
- Changed
extract_metadata3 fields changed- added
Input schema / properties / json_ld_typesAdded value: +{ + "description": "Filter the returned JSON-LD to nodes of these schema.org types, e.g. [\"Product\",\"Offer\"]. Subtypes match their parent: \"Event\" returns MusicEvent, \"Offer\" returns AggregateOffer, \"ItemList\" returns BreadcrumbList. Nodes are found at any depth, including inside @graph and nested inside a parent node. When set, json_ld carries only the matching nodes instead of the raw dump, and json_ld_type_counts reports how many matched per requested type. Documented types: ItemList, Product, Offer, Event, JobPosting, RealEstateListing — any other schema.org type is matched exactly.", + "items": { + "type": "string" + }, + "type": "array" +} - added
Input schema / properties / respect_robotsAdded value: +{ + "description": "Respect the target site's robots.txt (default: true). Setting this to false is honoured, returns a warning in the response, and is recorded against your API key — it is your decision, not a silent default.", + "type": "boolean" +} - added
Input schema / properties / user_agentAdded value: +{ + "description": "Override the outbound User-Agent. CrawlForge identifies itself honestly by default; use this only for targets you have your own agreement with.", + "type": "string" +}
- Changed
extract_structured5 fields changed- added
Input schema / properties / respect_robotsAdded value: +{ + "description": "Respect the target site's robots.txt (default: true). Setting this to false is honoured, returns a warning in the response, and is recorded against your API key — it is your decision, not a silent default.", + "type": "boolean" +} - added
Input schema / properties / user_agentAdded value: +{ + "description": "Override the outbound User-Agent. CrawlForge identifies itself honestly by default; use this only for targets you have your own agreement with.", + "type": "string" +} - added
Input schema / properties / verify_numbersAdded value: +{ + "default": true, + "description": "Numeric provenance guard (default: true): every price or numeric value the LLM returns must appear literally in the page source, else it is returned as null with a reason in `provenance.unverified`. Set false to get the model's raw numbers back, including ones it derived (a count, a sum, a total) rather than read off the page.", + "type": "boolean" +} - added
Output schema / properties / provenanceAdded value: +{ + "additionalProperties": true, + "properties": { + "enabled": { + "description": "Whether the numeric provenance guard ran", + "type": "boolean" + }, + "nulled": { + "description": "Numeric values replaced with null because the source does not contain them", + "type": "number" + }, + "skipped": { + "description": "\"empty_source\" when there was nothing to check against", + "type": "string" + }, + "unverified": { + "items": { + "additionalProperties": true, + "properties": { + "path": { + "description": "Path to the field, e.g. configurations[2].price", + "type": "string" + }, + "reason": { + "description": "\"not_found_in_source\"", + "type": "string" + }, + "value": { + "description": "The value that was removed" + } + }, + "type": "object" + }, + "type": "array" + }, + "verified": { + "description": "Numeric values found literally in the page source", + "type": "number" + } + }, + "type": "object" +} - added
Output schema / properties / successAdded value: +{ + "description": "False when the extraction errored or a required field came back missing or empty", + "type": "boolean" +}
- Changed
extract_text2 fields changed- added
Input schema / properties / respect_robotsAdded value: +{ + "description": "Respect the target site's robots.txt (default: true). Setting this to false is honoured, returns a warning in the response, and is recorded against your API key — it is your decision, not a silent default.", + "type": "boolean" +} - added
Input schema / properties / user_agentAdded value: +{ + "description": "Override the outbound User-Agent. CrawlForge identifies itself honestly by default; use this only for targets you have your own agreement with.", + "type": "string" +}
- Changed
extract_with_llm3 fields changed- added
Input schema / properties / respect_robotsAdded value: +{ + "description": "Respect the target site's robots.txt (default: true). Setting this to false is honoured, returns a warning in the response, and is recorded against your API key — it is your decision, not a silent default.", + "type": "boolean" +} - added
Input schema / properties / user_agentAdded value: +{ + "description": "Override the outbound User-Agent. CrawlForge identifies itself honestly by default; use this only for targets you have your own agreement with.", + "type": "string" +} - added
Input schema / properties / verify_numbersAdded value: +{ + "default": true, + "description": "Numeric provenance guard (default: true): every price or numeric value the LLM returns must appear literally in the page source, else it is returned as null with a reason in `provenance.unverified`. Set false to get the model's raw numbers back, including ones it derived (a count, a sum, a total) rather than read off the page.", + "type": "boolean" +}
- Changed
fetch_url2 fields changed- added
Input schema / properties / respect_robotsAdded value: +{ + "description": "Respect the target site's robots.txt (default: true). Setting this to false is honoured, returns a warning in the response, and is recorded against your API key — it is your decision, not a silent default.", + "type": "boolean" +} - added
Input schema / properties / user_agentAdded value: +{ + "description": "Override the outbound User-Agent. CrawlForge identifies itself honestly by default; use this only for targets you have your own agreement with.", + "type": "string" +}
- Changed
map_site2 fields changed- added
Input schema / properties / respect_robotsAdded value: +{ + "description": "Respect the target site's robots.txt (default: true). Setting this to false is honoured, returns a warning in the response, and is recorded against your API key — it is your decision, not a silent default.", + "type": "boolean" +} - added
Input schema / properties / user_agentAdded value: +{ + "description": "Override the outbound User-Agent. CrawlForge identifies itself honestly by default; use this only for targets you have your own agreement with.", + "type": "string" +}
- Changed
process_document2 fields changed- added
Input schema / properties / respect_robotsAdded value: +{ + "description": "Respect the target site's robots.txt (default: true). Setting this to false is honoured, returns a warning in the response, and is recorded against your API key — it is your decision, not a silent default.", + "type": "boolean" +} - added
Input schema / properties / user_agentAdded value: +{ + "description": "Override the outbound User-Agent. CrawlForge identifies itself honestly by default; use this only for targets you have your own agreement with.", + "type": "string" +}
- Changed
scrape2 fields changed- added
Input schema / properties / respect_robotsAdded value: +{ + "description": "Respect the target site's robots.txt (default: true). Setting this to false is honoured, returns a warning in the response, and is recorded against your API key — it is your decision, not a silent default.", + "type": "boolean" +} - added
Input schema / properties / user_agentAdded value: +{ + "description": "Override the outbound User-Agent. CrawlForge identifies itself honestly by default; use this only for targets you have your own agreement with.", + "type": "string" +}
- Changed
scrape_structured4 fields changed- changed
Input schema / properties / max_results / descriptionPrevious value: -"Maximum number of matches to return per field when a selector matches multiple elements"New value: +"Maximum number of matches to return per field when a selector matches multiple elements, or the maximum number of rows when row_selector is set" - added
Input schema / properties / respect_robotsAdded value: +{ + "description": "Respect the target site's robots.txt (default: true). Setting this to false is honoured, returns a warning in the response, and is recorded against your API key — it is your decision, not a silent default.", + "type": "boolean" +} - added
Input schema / properties / row_selectorAdded value: +{ + "description": "CSS selector for the repeating row/container element. When set, each field in selectors is matched inside each row and data is an array of row-aligned records ({field: value|null}) instead of parallel arrays", + "type": "string" +} - added
Input schema / properties / user_agentAdded value: +{ + "description": "Override the outbound User-Agent. CrawlForge identifies itself honestly by default; use this only for targets you have your own agreement with.", + "type": "string" +}
- Changed
scrape_template5 fields changed- added
Input schema / properties / paramsAdded value: +{ + "additionalProperties": {}, + "description": "Parameters for a list connector, e.g. {company:\"stripe\"} for greenhouse-jobs or {store:\"www.allbirds.com\", collection:\"mens\"} for shopify-collection. Use template:\"list\" to see which templates take params", + "type": "object" +} - added
Input schema / properties / respect_robotsAdded value: +{ + "description": "Respect the target site's robots.txt (default: true). Setting this to false is honoured, returns a warning in the response, and is recorded against your API key — it is your decision, not a silent default.", + "type": "boolean" +} - changed
Input schema / properties / template / descriptionPrevious value: -"Template ID (e.g. github-repo) or list to enumerate available templates"New value: +"Template ID (e.g. github-repo), \"auto\" to detect one from the url, or \"list\" to enumerate available templates" - changed
Input schema / properties / url / descriptionPrevious value: -"URL to scrape — required unless template is list"New value: +"URL to scrape — required unless template is list, or params drive a list connector" - added
Input schema / properties / user_agentAdded value: +{ + "description": "Override the outbound User-Agent. CrawlForge identifies itself honestly by default; use this only for targets you have your own agreement with.", + "type": "string" +}
- Changed
scrape_with_actions7 fields changed- changed
Input schema / properties / actions / items / properties / type / enumPrevious value: -[ - "wait", - "click", - "type", - "press", - "scroll", - "screenshot", - "executeJavaScript" -]New value: +[ + "wait", + "click", + "type", + "press", + "scroll", + "screenshot", + "executeJavaScript", + "select", + "hover", + "navigate" +] - added
Input schema / properties / actions / items / properties / urlAdded value: +{ + "description": "navigate: URL to navigate to — goes through the same SSRF and robots.txt gate as the initial URL", + "format": "uri", + "type": "string" +} - added
Input schema / properties / actions / items / properties / valueAdded value: +{ + "description": "select: option to choose, matched by value or label", + "type": "string" +} - added
Input schema / properties / actions / items / properties / valuesAdded value: +{ + "description": "select: options to choose in a multi-select, matched by value or label", + "items": { + "type": "string" + }, + "type": "array" +} - added
Input schema / properties / actions / items / properties / waitUntilAdded value: +{ + "description": "navigate: when to consider navigation complete", + "enum": [ + "load", + "domcontentloaded", + "networkidle", + "commit" + ], + "type": "string" +} - added
Input schema / properties / browserOptions / properties / stealthAdded value: +{ + "default": false, + "description": "Run the action chain in the stealth browser (randomized fingerprint, WebRTC/canvas spoofing) instead of the standard browser pool. Renders JavaScript; it does not solve challenges.", + "type": "boolean" +} - added
Input schema / properties / respect_robotsAdded value: +{ + "description": "Respect the target site's robots.txt (default: true). Setting this to false is honoured, returns a warning in the response, and is recorded against your API key — it is your decision, not a silent default.", + "type": "boolean" +}
- Changed
stealth_mode6 fields changed- added
Input schema / properties / formatsAdded value: +{ + "default": [ + "markdown" + ], + "description": "Formats to return from operation:\"scrape\" (default: [\"markdown\"]). \"screenshot\" returns a crawlforge://screenshot/{id} resource URI.", + "items": { + "enum": [ + "markdown", + "html", + "text", + "links", + "metadata", + "screenshot" + ], + "type": "string" + }, + "type": "array" +} - changed
Input schema / properties / operation / enumPrevious value: -[ - "configure", - "enable", - "disable", - "create_context", - "create_page", - "get_stats", - "cleanup" -]New value: +[ + "scrape", + "configure", + "enable", + "disable", + "create_context", + "create_page", + "get_stats", + "cleanup" +] - added
Input schema / properties / respect_robotsAdded value: +{ + "description": "Respect the target site's robots.txt (default: true). Setting this to false is honoured, returns a warning in the response, and is recorded against your API key — it is your decision, not a silent default.", + "type": "boolean" +} - added
Input schema / properties / urlAdded value: +{ + "description": "URL to scrape — required for operation:\"scrape\"", + "format": "uri", + "type": "string" +} - added
Input schema / properties / verboseAdded value: +{ + "default": false, + "description": "Return the full generated fingerprint from create_context instead of a summary", + "type": "boolean" +} - added
Input schema / properties / wait_forAdded value: +{ + "description": "Extra wait after page load, in ms — for content that renders after DOMContentLoaded", + "maximum": 30000, + "minimum": 0, + "type": "number" +}
- Changed
track_changes2 fields changed- added
Input schema / properties / respect_robotsAdded value: +{ + "description": "Respect the target site's robots.txt (default: true). Setting this to false is honoured, returns a warning in the response, and is recorded against your API key — it is your decision, not a silent default.", + "type": "boolean" +} - added
Input schema / properties / user_agentAdded value: +{ + "description": "Override the outbound User-Agent. CrawlForge identifies itself honestly by default; use this only for targets you have your own agreement with.", + "type": "string" +}
3 tool updates
v5.1.0- Changed
crawl_deep2 fields changed- added
Output schema / properties / cachedAdded value: +{ + "description": "True when this response was replayed from an earlier crawl rather than crawled now; crawled_at gives its age", + "type": "boolean" +} - added
Output schema / properties / crawled_atAdded value: +{ + "description": "When the pages were actually fetched (ISO 8601)", + "type": "string" +}
- Changed
extract_structured1 field changed- changed
Output schema / properties / extraction_method / descriptionPrevious value: -"\"llm\" | \"css_fallback\" | \"none\""New value: +"\"llm\" | \"css_fallback\" | \"keyword_fallback\" | \"none\""
- Added
reddit_search
1 tool update
v5.0.5- Changed
serp_rank1 field changed- changed
Input schema / properties / depth / descriptionPrevious value: -"How many results to scan, 10-200 (100 = 1 page of cost)"New value: +"How many results to scan, 10-200 (default 20; DataForSEO bills ~$0.002 per 10 and gets slower the deeper it goes)"
27 tool updates
v5.0.4- Changed
agent1 field changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema"
- Changed
analyze_content2 fields changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema" - changed
Input schema / properties / options / additionalPropertiesPrevious value: -falseNew value: +true
- Changed
batch_scrape1 field changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema"
- Changed
crawl_deep2 fields changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema" - changed
Output schema / (root)Previous value: -nullNew value: +{ + "$schema": "https://json-schema.org/draft/2020-12/schema", + "additionalProperties": false, + "properties": { + "_cost": { + "additionalProperties": true, + "description": "Cost-transparency metadata (D3.5), present when injected into the text copy of the result", + "properties": { + "actual": { + "description": "Credits actually charged (0 in creator mode, half-rate on error)", + "type": "number" + }, + "projected": { + "description": "Credits projected for this call before execution", + "type": "number" + }, + "projection_note": { + "description": "Human-readable note about how the cost was projected", + "type": "string" + }, + "remaining_credits": { + "description": "Credits remaining on the account after this call, if known", + "type": [ + "number", + "null" + ] + } + }, + "type": "object" + }, + "crawl_depth": { + "type": "number" + }, + "domain_filter_config": { + "anyOf": [ + {}, + { + "type": "null" + } + ] + }, + "duration_ms": { + "type": "number" + }, + "error": { + "type": "string" + }, + "error_count": { + "type": "number" + }, + "errors": { + "items": {}, + "type": "array" + }, + "link_analysis": { + "anyOf": [ + {}, + { + "type": "null" + } + ] + }, + "pages_crawled": { + "type": "number" + }, + "pages_found": { + "type": "number" + }, + "pages_per_second": { + "type": "number" + }, + "results": { + "items": { + "additionalProperties": true, + "properties": { + "content": { + "type": "string" + }, + "content_length": { + "type": "number" + }, + "depth": { + "type": "number" + }, + "links_count": { + "type": "number" + }, + "metadata": {}, + "timestamp": { + "type": [ + "string", + "number" + ] + }, + "title": { + "type": "string" + }, + "truncated": { + "type": "boolean" + }, + "url": { + "type": "string" + } + }, + "type": "object" + }, + "type": "array" + }, + "session": { + "additionalProperties": true, + "properties": { + "cookies_captured": { + "type": "number" + }, + "enabled": { + "type": "boolean" + } + }, + "type": "object" + }, + "site_structure": { + "additionalProperties": true, + "properties": { + "depth_distribution": { + "additionalProperties": { + "type": "number" + }, + "type": "object" + }, + "file_types": { + "additionalProperties": { + "type": "number" + }, + "type": "object" + }, + "path_patterns": { + "additionalProperties": { + "type": "number" + }, + "type": "object" + }, + "subdomains": { + "items": { + "type": "string" + }, + "type": "array" + }, + "total_pages": { + "type": "number" + } + }, + "type": "object" + }, + "stats": {}, + "success": { + "description": "False only when the crawl was cancelled via elicitation decline", + "type": "boolean" + }, + "url": { + "type": "string" + } + }, + "type": "object" +}
- Changed
deep_research1 field changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema"
- Changed
extract_content2 fields changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema" - changed
Input schema / properties / options / additionalPropertiesPrevious value: -falseNew value: +true
- Changed
extract_links1 field changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema"
- Changed
extract_metadata1 field changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema"
- Changed
extract_structured2 fields changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema" - changed
Output schema / (root)Previous value: -nullNew value: +{ + "$schema": "https://json-schema.org/draft/2020-12/schema", + "additionalProperties": false, + "properties": { + "_cost": { + "additionalProperties": true, + "description": "Cost-transparency metadata (D3.5), present when injected into the text copy of the result", + "properties": { + "actual": { + "description": "Credits actually charged (0 in creator mode, half-rate on error)", + "type": "number" + }, + "projected": { + "description": "Credits projected for this call before execution", + "type": "number" + }, + "projection_note": { + "description": "Human-readable note about how the cost was projected", + "type": "string" + }, + "remaining_credits": { + "description": "Credits remaining on the account after this call, if known", + "type": [ + "number", + "null" + ] + } + }, + "type": "object" + }, + "confidence": { + "type": "number" + }, + "data": { + "additionalProperties": {}, + "description": "Extracted fields matching the requested schema", + "type": "object" + }, + "error": { + "type": "string" + }, + "extractionNotes": { + "items": { + "type": "string" + }, + "type": "array" + }, + "extraction_method": { + "description": "\"llm\" | \"css_fallback\" | \"none\"", + "type": "string" + }, + "processingTime": { + "type": "number" + }, + "schema_used": { + "additionalProperties": {}, + "type": "object" + }, + "url": { + "type": "string" + }, + "validation": { + "additionalProperties": true, + "properties": { + "errors": { + "items": { + "type": "string" + }, + "type": "array" + }, + "valid": { + "type": "boolean" + } + }, + "type": "object" + } + }, + "type": "object" +}
- Changed
extract_text1 field changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema"
- Changed
extract_with_llm1 field changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema"
- Changed
fetch_url1 field changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema"
- Changed
generate_llms_txt4 fields changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema" - changed
Input schema / properties / analysisOptions / properties / checkSecurity / defaultPrevious value: -trueNew value: +false - added
Input schema / properties / analysisOptions / properties / probeRateLimitAdded value: +{ + "default": false, + "type": "boolean" +} - added
Input schema / properties / outputOptions / properties / robotsStyleAdded value: +{ + "default": false, + "type": "boolean" +}
- Changed
get_batch_results1 field changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema"
- Changed
list_ollama_models1 field changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema"
- Changed
localization3 fields changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema" - changed
Input schema / properties / content / descriptionPrevious value: -"Content for auto-detection of language and locale"New value: +"Page content (HTML or plain text) to analyze — required for auto_detect; no fetching is performed" - changed
Input schema / properties / url / descriptionPrevious value: -"URL for geo-blocking detection or auto-detection"New value: +"URL — required for handle_geo_blocking; for auto_detect it is optional metadata used only as a TLD country hint (the page is never fetched)"
- Changed
map_site2 fields changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema" - changed
Output schema / (root)Previous value: -nullNew value: +{ + "$schema": "https://json-schema.org/draft/2020-12/schema", + "additionalProperties": false, + "properties": { + "_cost": { + "additionalProperties": true, + "description": "Cost-transparency metadata (D3.5), present when injected into the text copy of the result", + "properties": { + "actual": { + "description": "Credits actually charged (0 in creator mode, half-rate on error)", + "type": "number" + }, + "projected": { + "description": "Credits projected for this call before execution", + "type": "number" + }, + "projection_note": { + "description": "Human-readable note about how the cost was projected", + "type": "string" + }, + "remaining_credits": { + "description": "Credits remaining on the account after this call, if known", + "type": [ + "number", + "null" + ] + } + }, + "type": "object" + }, + "base_url": { + "type": "string" + }, + "domain_filter_config": { + "anyOf": [ + {}, + { + "type": "null" + } + ] + }, + "filter_stats": { + "anyOf": [ + {}, + { + "type": "null" + } + ] + }, + "metadata": { + "additionalProperties": {}, + "description": "Per-URL metadata when include_metadata=true", + "type": "object" + }, + "ranked_urls": { + "description": "Present only when the `search` param was set", + "items": { + "additionalProperties": true, + "properties": { + "score": { + "type": "number" + }, + "url": { + "type": "string" + } + }, + "type": "object" + }, + "type": "array" + }, + "site_map": { + "additionalProperties": true, + "properties": { + "depth_levels": { + "additionalProperties": {}, + "type": "object" + }, + "root": { + "items": { + "type": "string" + }, + "type": "array" + }, + "sections": { + "additionalProperties": {}, + "type": "object" + } + }, + "type": "object" + }, + "statistics": { + "additionalProperties": true, + "properties": { + "average_depth": { + "type": "number" + }, + "file_extensions": { + "additionalProperties": { + "type": "number" + }, + "type": "object" + }, + "max_depth": { + "type": "number" + }, + "query_parameters": { + "type": "number" + }, + "secure_urls": { + "type": "number" + }, + "total_urls": { + "type": "number" + }, + "unique_paths": { + "type": "number" + }, + "url_lengths": { + "additionalProperties": true, + "properties": { + "average": { + "type": "number" + }, + "max": { + "type": "number" + }, + "min": { + "type": [ + "number", + "null" + ] + } + }, + "type": "object" + } + }, + "type": "object" + }, + "total_urls": { + "type": "number" + }, + "urls": { + "anyOf": [ + { + "items": { + "type": "string" + }, + "type": "array" + }, + { + "additionalProperties": { + "items": { + "type": "string" + }, + "type": "array" + }, + "type": "object" + } + ], + "description": "Flat array of URLs, or grouped-by-path object when group_by_path=true (default)" + } + }, + "type": "object" +}
- Changed
process_document2 fields changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema" - changed
Input schema / properties / options / descriptionPrevious value: -"Additional processing options (maxPages, pageRange:{start,end}, extractText, extractMetadata, password, outputFormat, ...)"New value: +"Additional processing options (maxPages, pageRange:{start,end}, extractText, extractMetadata, outputFormat, ...)"
- Changed
scrape2 fields changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema" - changed
Output schema / (root)Previous value: -nullNew value: +{ + "$schema": "https://json-schema.org/draft/2020-12/schema", + "additionalProperties": false, + "properties": { + "_cost": { + "additionalProperties": true, + "description": "Cost-transparency metadata (D3.5), present when injected into the text copy of the result", + "properties": { + "actual": { + "description": "Credits actually charged (0 in creator mode, half-rate on error)", + "type": "number" + }, + "projected": { + "description": "Credits projected for this call before execution", + "type": "number" + }, + "projection_note": { + "description": "Human-readable note about how the cost was projected", + "type": "string" + }, + "remaining_credits": { + "description": "Credits remaining on the account after this call, if known", + "type": [ + "number", + "null" + ] + } + }, + "type": "object" + }, + "content": { + "additionalProperties": true, + "description": "One key per requested format", + "properties": { + "branding": { + "additionalProperties": {}, + "description": "Static design tokens: colors, fonts, logo", + "type": "object" + }, + "html": { + "type": "string" + }, + "json": { + "description": "Result of the {type:\"json\"} format (LLM-structured extraction)" + }, + "links": { + "additionalProperties": true, + "properties": { + "external_count": { + "type": "number" + }, + "internal_count": { + "type": "number" + }, + "links": { + "items": { + "additionalProperties": true, + "properties": { + "href": { + "type": "string" + }, + "is_external": { + "type": "boolean" + }, + "original_href": { + "type": "string" + }, + "text": { + "type": "string" + } + }, + "type": "object" + }, + "type": "array" + }, + "total_count": { + "type": "number" + } + }, + "type": "object" + }, + "markdown": { + "type": "string" + }, + "metadata": { + "additionalProperties": true, + "properties": { + "author": { + "type": "string" + }, + "canonical_url": { + "type": "string" + }, + "description": { + "type": "string" + }, + "json_ld": { + "items": {}, + "type": "array" + }, + "keywords": { + "items": { + "type": "string" + }, + "type": "array" + }, + "microdata": { + "items": {}, + "type": "array" + }, + "og_tags": { + "additionalProperties": {}, + "type": "object" + }, + "robots": { + "type": "string" + }, + "title": { + "type": "string" + }, + "twitter_tags": { + "additionalProperties": {}, + "type": "object" + }, + "url": { + "type": "string" + }, + "viewport": { + "type": "string" + } + }, + "type": "object" + }, + "rawHtml": { + "type": "string" + }, + "screenshots": { + "description": "Present for the \"screenshot\" format; each item carries a resourceUri once published", + "items": { + "additionalProperties": true, + "properties": {}, + "type": "object" + }, + "type": "array" + }, + "text": { + "type": "string" + } + }, + "type": "object" + }, + "success": { + "description": "Whether the scrape completed", + "type": "boolean" + }, + "url": { + "description": "Final URL after redirects", + "type": "string" + }, + "warnings": { + "description": "Per-format warnings; partial success never fails the whole call", + "items": { + "type": "string" + }, + "type": "array" + } + }, + "type": "object" +}
- Changed
scrape_structured1 field changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema"
- Changed
scrape_template1 field changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema"
- Changed
scrape_with_actions3 fields changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema" - added
Input schema / properties / actions / items / properties / xAdded value: +{ + "description": "scroll: absolute X coordinate to scroll to (window.scrollTo; with y, takes precedence over direction/distance)", + "minimum": 0, + "type": "number" +} - added
Input schema / properties / actions / items / properties / yAdded value: +{ + "description": "scroll: absolute Y coordinate to scroll to (window.scrollTo; with x, takes precedence over direction/distance)", + "minimum": 0, + "type": "number" +}
- Changed
search_web2 fields changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema" - changed
Output schema / (root)Previous value: -nullNew value: +{ + "$schema": "https://json-schema.org/draft/2020-12/schema", + "additionalProperties": false, + "properties": { + "_cost": { + "additionalProperties": true, + "description": "Cost-transparency metadata (D3.5), present when injected into the text copy of the result", + "properties": { + "actual": { + "description": "Credits actually charged (0 in creator mode, half-rate on error)", + "type": "number" + }, + "projected": { + "description": "Credits projected for this call before execution", + "type": "number" + }, + "projection_note": { + "description": "Human-readable note about how the cost was projected", + "type": "string" + }, + "remaining_credits": { + "description": "Credits remaining on the account after this call, if known", + "type": [ + "number", + "null" + ] + } + }, + "type": "object" + }, + "cached": { + "type": "boolean" + }, + "effective_query": { + "description": "Present when query expansion changed the query actually used", + "type": "string" + }, + "expanded_queries": { + "items": { + "type": "string" + }, + "type": "array" + }, + "limit": { + "type": "number" + }, + "localization": { + "anyOf": [ + { + "additionalProperties": true, + "properties": { + "applied": { + "type": "boolean" + }, + "countryCode": { + "type": "string" + }, + "geoTargeting": { + "type": "boolean" + }, + "language": { + "type": "string" + }, + "searchDomain": { + "type": "string" + } + }, + "type": "object" + }, + { + "type": "null" + } + ] + }, + "offset": { + "type": "number" + }, + "processing": { + "additionalProperties": true, + "properties": { + "deduplication": { + "anyOf": [ + { + "additionalProperties": {}, + "type": "object" + }, + { + "type": "null" + } + ] + }, + "localization_applied": { + "type": "boolean" + }, + "query_expansion": { + "anyOf": [ + { + "additionalProperties": {}, + "type": "object" + }, + { + "type": "null" + } + ] + }, + "ranking": { + "anyOf": [ + { + "additionalProperties": {}, + "type": "object" + }, + { + "type": "null" + } + ] + } + }, + "type": "object" + }, + "provider": { + "additionalProperties": true, + "properties": { + "backend": { + "type": "string" + }, + "capabilities": { + "additionalProperties": {}, + "type": "object" + }, + "instanceUrl": { + "type": [ + "string", + "null" + ] + }, + "name": { + "type": "string" + }, + "note": { + "type": "string" + } + }, + "type": "object" + }, + "query": { + "type": "string" + }, + "results": { + "items": { + "additionalProperties": true, + "properties": { + "displayLink": { + "type": "string" + }, + "formattedUrl": { + "type": "string" + }, + "htmlSnippet": { + "type": "string" + }, + "link": { + "type": "string" + }, + "metadata": { + "additionalProperties": {}, + "type": "object" + }, + "pagemap": { + "additionalProperties": {}, + "type": "object" + }, + "snippet": { + "type": "string" + }, + "title": { + "type": "string" + } + }, + "type": "object" + }, + "type": "array" + }, + "search_time": { + "type": "number" + }, + "total_results": { + "type": [ + "string", + "number" + ] + } + }, + "type": "object" +}
- Changed
serp_rank2 fields changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema" - changed
Output schema / (root)Previous value: -nullNew value: +{ + "$schema": "https://json-schema.org/draft/2020-12/schema", + "additionalProperties": false, + "properties": { + "_cost": { + "additionalProperties": true, + "description": "Cost-transparency metadata (D3.5), present when injected into the text copy of the result", + "properties": { + "actual": { + "description": "Credits actually charged (0 in creator mode, half-rate on error)", + "type": "number" + }, + "projected": { + "description": "Credits projected for this call before execution", + "type": "number" + }, + "projection_note": { + "description": "Human-readable note about how the cost was projected", + "type": "string" + }, + "remaining_credits": { + "description": "Credits remaining on the account after this call, if known", + "type": [ + "number", + "null" + ] + } + }, + "type": "object" + }, + "allPositions": { + "description": "Every position the target holds on this SERP", + "items": { + "additionalProperties": true, + "properties": { + "domain": { + "type": "string" + }, + "position": { + "type": [ + "number", + "null" + ] + }, + "rankAbsolute": { + "type": [ + "number", + "null" + ] + }, + "snippet": { + "type": [ + "string", + "null" + ] + }, + "title": { + "type": [ + "string", + "null" + ] + }, + "url": { + "type": [ + "string", + "null" + ] + } + }, + "type": "object" + }, + "type": "array" + }, + "checkUrl": { + "description": "Link to view the real SERP on DataForSEO", + "type": "string" + }, + "checkedAt": { + "type": "string" + }, + "configured": { + "description": "False when DATAFORSEO_LOGIN/PASSWORD are unset — no rank was fabricated", + "type": "boolean" + }, + "cost": { + "description": "USD charged by DataForSEO for this lookup (separate from CrawlForge credits)", + "type": "number" + }, + "depthScanned": { + "type": "number" + }, + "device": { + "type": "string" + }, + "found": { + "description": "Whether the target appeared anywhere in the scanned SERP", + "type": "boolean" + }, + "keyword": { + "type": "string" + }, + "location": {}, + "note": { + "description": "Present when configured=false, explains how to enable", + "type": "string" + }, + "organicResults": { + "type": "number" + }, + "position": { + "description": "Best (lowest) organic rank; null = not within top `depth`", + "type": [ + "number", + "null" + ] + }, + "rankAbsolute": { + "type": [ + "number", + "null" + ] + }, + "results": { + "description": "Top organic competitors as Google actually ranks them (capped)", + "items": { + "$ref": "#/properties/allPositions/items" + }, + "type": "array" + }, + "seResultsCount": { + "type": "number" + }, + "target": { + "description": "Bare target domain, normalized", + "type": "string" + }, + "title": { + "type": [ + "string", + "null" + ] + }, + "url": { + "description": "URL of the target's best-ranking result", + "type": [ + "string", + "null" + ] + } + }, + "type": "object" +}
- Changed
stealth_mode1 field changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema"
- Changed
summarize_content2 fields changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema" - changed
Input schema / properties / options / additionalPropertiesPrevious value: -falseNew value: +true
- Changed
track_changes1 field changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema"
27 tool updates
v4.10.0- First observed
agent - First observed
analyze_content - First observed
batch_scrape - First observed
crawl_deep - First observed
deep_research - First observed
extract_content - First observed
extract_links - First observed
extract_metadata - First observed
extract_structured - First observed
extract_text - First observed
extract_with_llm - First observed
fetch_url - First observed
generate_llms_txt - First observed
get_batch_results - First observed
list_ollama_models - First observed
localization - First observed
map_site - First observed
process_document - First observed
scrape - First observed
scrape_structured - First observed
scrape_template - First observed
scrape_with_actions - First observed
search_web - First observed
serp_rank - First observed
stealth_mode - First observed
summarize_content - First observed
track_changes
TDQS
Scored across 31 tools
The descriptions draw explicit boundaries, using 'Not for...' guidance to separate scrape, extract_text, extract_content, fetch_url, scrape_with_actions, browser_session, and stealth_mode. Still, several sibling extraction and research tools overlap enough that an agent could plausibly misselect among them.
All tools use snake_case consistently, with no camelCase or chaotic style mixing. However, the pattern is not uniformly verb_noun: noun-phrase names like serp_rank, stealth_mode, localization, and deep_research break the otherwise predictable convention.
At 31 tools, the server exceeds the 25+ threshold for 'too many' in the rubric. Although many tools are specialized, the count makes the surface heavy and increases overlap among extraction, browser, and research paths.
Coverage is broad: single/batch/crawl/map scraping, interactive sessions, stealth, search, structured and embedded-state extraction, documents, Reddit, SERP, research, monitoring, and result paging. Minor lifecycle gaps remain, such as no explicit cancel/delete/list operations for batch jobs or scheduled monitors.
Maintenance
Related MCP Connectors
Web MCP: scrape/crawl sites, web search, brand assets, app stores, YouTube, Reddit, Hacker News.
Scrape, crawl and search the web for AI agents via MCP.
Unblocking and fresh web data for agents: URL to Markdown, YouTube, Maps, Amazon, jobs. Pay per call
Web scraping for AI agents: scrape, search, crawl, map any website to markdown + JSON. No browser.
Related MCP Servers
- FlicenseNot gradedqualityCmaintenanceProvides 42+ MCP tools for browser automation, web scraping, and search, enabling AI agents like Claude and Cursor to browse, extract data, and run research agents on the live web.9-
- AlicenseBqualityAmaintenanceEnables AI assistants to crawl websites, extract dynamic content, navigate links, and save structured Markdown files via the MCP protocol, with support for anti-bot bypass, CSS selectors, and custom JavaScript execution.146 PyPI45MIT
- AlicenseNot gradedqualityBmaintenanceEnables web search, page fetching and cleaning, OCR from images, and broken-link checking through MCP tools, with multi-provider aggregation, caching, and anti-detection features.50 npmMIT
- AlicenseNot gradedqualityBmaintenanceEnables agents to scrape, crawl, map, search, and extract web pages as clean markdown or structured JSON directly through MCP tools.AGPL 3.0