Skip to main content
Glama
mysleekdesigns

CrawlForge MCP Server

scrape_with_actions

Read-only

Scrape pages that require interaction: log in, click, type, scroll, and wait for dynamic content, then extract clean Markdown or JSON. Use for SPAs, login-gated pages, and multi-step flows.

Instructions

Use this when you must interact with a page before scraping - login, click buttons, fill forms, scroll, or wait for dynamic content to load - for SPAs, login-gated content, or multi-step flows. Actions: snapshot, wait, click, type, press, scroll, screenshot, executeJavaScript, select (dropdowns), hover, navigate. Start a chain with {type:"snapshot"} to list the page's interactive elements with stable refs (@e1, @e2 ...), then target those refs in later actions instead of guessing CSS selectors; navigation invalidates refs, so snapshot again after one. Set browserOptions.stealth:true to run the chain in the stealth browser, and browserOptions.engine to pick its engine ("auto" by default - camoufox when it is installed, Chromium otherwise, and the result says which ran). robots.txt is respected on every navigation. Screenshots from this tool are stored as crawlforge://screenshot/{actionId} resources. Not for pages that render without interaction (scrape) and not as the first attempt on a blocked site (stealth_mode operation:"scrape"). Cost: 5 credits. Example: scrape_with_actions({url: "https://app.com/dashboard", actions: [{type:"snapshot"},{type:"type",selector:"@e2",text:"user@a.com"},{type:"click",selector:"@e4"}]})

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
urlYesThe URL to scrape
actionsYesBrowser actions to perform before scraping
formatsNoOutput formats for scraped content
maxRetriesNoMaximum retry attempts on failure
redact_piiNoRedact personal data from the text this call returns, before it reaches your context window. true means the free regex pass over EMAIL, PHONE, FINANCIAL and SECRET. The result carries redaction:{entities,count}. Default: off
formAutoFillNoForm auto-fill configuration
browserOptionsNoBrowser configuration options
respect_robotsNoRespect the target site's robots.txt (default: true). Setting this to false is honoured, returns a warning in the response, and is recorded against your API key — it is your decision, not a silent default.
max_inline_charsNoLargest result to return inline, in characters of its JSON. Over it, the call returns a preview plus a result_handle for read_result instead of the whole result (default 40,000; env CRAWLFORGE_MAX_INLINE_CHARS)
extractionOptionsNoContent extraction options. selectors results are returned as content.json.extracted, so include "json" in formats when passing selectors — without it the extraction is not part of the response.
screenshotOnErrorNoCapture screenshot when an error occurs
captureScreenshotsNoTake screenshots during action execution
continueOnActionErrorNoContinue executing actions if one fails
captureIntermediateStatesNoCapture page state after each action

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed1 schema field changedv6.8.0
    • addedInput schema / properties / browserOptions / properties / engine
      Added value: +{
      +  "default": "auto",
      +  "description": "Stealth engine for the chain, with stealth:true. \"auto\" (default) runs camoufox when it is installed and Chromium otherwise; the result's `engine` says which ran. \"camoufox\" is Firefox-based with a higher anti-detect score; \"chromium\" (= \"playwright\") forces Chromium. Refused without stealth:true, where the browser is always Chromium.",
      +  "enum": [
      +    "auto",
      +    "chromium",
      +    "camoufox",
      +    "playwright"
      +  ],
      +  "type": "string"
      +}
  2. Changed4 schema fields changedv6.6.0
    • addedInput schema / properties / actions / items / properties / interactiveOnly
      Added value: +{
      +  "description": "snapshot: only interactive elements (default true)",
      +  "type": "boolean"
      +}
    • addedInput schema / properties / actions / items / properties / maxNodes
      Added value: +{
      +  "description": "snapshot: cap on nodes listed (default 200, max 1000); the result says truncated when the cap stopped the walk",
      +  "maximum": 1000,
      +  "minimum": 1,
      +  "type": "number"
      +}
    • addedInput schema / properties / actions / items / properties / selector / description
      Added value: +"A CSS selector, or a @e1 ref from an earlier snapshot action in this chain"
    • changedInput schema / properties / actions / items / properties / type / enum
      Previous value: -[
      -  "wait",
      -  "click",
      -  "type",
      -  "press",
      -  "scroll",
      -  "screenshot",
      -  "executeJavaScript",
      -  "select",
      -  "hover",
      -  "navigate"
      -]New value: +[
      +  "snapshot",
      +  "wait",
      +  "click",
      +  "type",
      +  "press",
      +  "scroll",
      +  "screenshot",
      +  "executeJavaScript",
      +  "select",
      +  "hover",
      +  "navigate"
      +]
  3. Changed11 schema fields changedv6.0.0
    • removedInput schema / additionalProperties
      Removed value: -false
    • removedInput schema / properties / actions / items / additionalProperties
      Removed value: -false
    • addedInput schema / properties / actions / items / properties / args / items
      Added value: +{}
    • removedInput schema / properties / actions / items / properties / position / additionalProperties
      Removed value: -false
    • removedInput schema / properties / browserOptions / additionalProperties
      Removed value: -false
    • removedInput schema / properties / extractionOptions / additionalProperties
      Removed value: -false
    • addedInput schema / properties / extractionOptions / properties / selectors / propertyNames
      Added value: +{
      +  "type": "string"
      +}
    • removedInput schema / properties / formAutoFill / additionalProperties
      Removed value: -false
    • removedInput schema / properties / formAutoFill / properties / fields / items / additionalProperties
      Removed value: -false
    • addedInput schema / properties / max_inline_chars
      Added value: +{
      +  "description": "Largest result to return inline, in characters of its JSON. Over it, the call returns a preview plus a result_handle for read_result instead of the whole result (default 40,000; env CRAWLFORGE_MAX_INLINE_CHARS)",
      +  "maximum": 10000000,
      +  "minimum": 1000,
      +  "type": "integer"
      +}
    • addedInput schema / properties / redact_pii
      Added value: +{
      +  "anyOf": [
      +    {
      +      "type": "boolean"
      +    },
      +    {
      +      "properties": {
      +        "entities": {
      +          "description": "Which classes to redact, case-insensitive: EMAIL, PHONE, FINANCIAL, SECRET, plus PERSON and LOCATION when mode is \"model\". Omitted or empty means all four regex classes (and both model classes in \"model\" mode). An unknown name, or a model-only name without mode:\"model\", is rejected",
      +          "items": {
      +            "type": "string"
      +          },
      +          "type": "array"
      +        },
      +        "mode": {
      +          "description": "\"fast\" (default) is regex only and free; \"model\" adds an Ollama NER pass for PERSON and LOCATION (+3 credits once per call)",
      +          "enum": [
      +            "fast",
      +            "model"
      +          ],
      +          "type": "string"
      +        },
      +        "replace_style": {
      +          "description": "\"tag\" (default) writes <EMAIL>, \"mask\" writes [REDACTED], \"remove\" deletes the value",
      +          "enum": [
      +            "tag",
      +            "mask",
      +            "remove"
      +          ],
      +          "type": "string"
      +        }
      +      },
      +      "type": "object"
      +    }
      +  ],
      +  "description": "Redact personal data from the text this call returns, before it reaches your context window. true means the free regex pass over EMAIL, PHONE, FINANCIAL and SECRET. The result carries redaction:{entities,count}. Default: off"
      +}
  4. Changed1 schema field changedv5.6.6
    • changedInput schema / properties / extractionOptions / description
      Previous value: -"Content extraction options"New value: +"Content extraction options. selectors results are returned as content.json.extracted, so include \"json\" in formats when passing selectors — without it the extraction is not part of the response."
  5. Changed7 schema fields changedv5.4.0
    • changedInput schema / properties / actions / items / properties / type / enum
      Previous value: -[
      -  "wait",
      -  "click",
      -  "type",
      -  "press",
      -  "scroll",
      -  "screenshot",
      -  "executeJavaScript"
      -]New value: +[
      +  "wait",
      +  "click",
      +  "type",
      +  "press",
      +  "scroll",
      +  "screenshot",
      +  "executeJavaScript",
      +  "select",
      +  "hover",
      +  "navigate"
      +]
    • addedInput schema / properties / actions / items / properties / url
      Added value: +{
      +  "description": "navigate: URL to navigate to — goes through the same SSRF and robots.txt gate as the initial URL",
      +  "format": "uri",
      +  "type": "string"
      +}
    • addedInput schema / properties / actions / items / properties / value
      Added value: +{
      +  "description": "select: option to choose, matched by value or label",
      +  "type": "string"
      +}
    • addedInput schema / properties / actions / items / properties / values
      Added value: +{
      +  "description": "select: options to choose in a multi-select, matched by value or label",
      +  "items": {
      +    "type": "string"
      +  },
      +  "type": "array"
      +}
    • addedInput schema / properties / actions / items / properties / waitUntil
      Added value: +{
      +  "description": "navigate: when to consider navigation complete",
      +  "enum": [
      +    "load",
      +    "domcontentloaded",
      +    "networkidle",
      +    "commit"
      +  ],
      +  "type": "string"
      +}
    • addedInput schema / properties / browserOptions / properties / stealth
      Added value: +{
      +  "default": false,
      +  "description": "Run the action chain in the stealth browser (randomized fingerprint, WebRTC/canvas spoofing) instead of the standard browser pool. Renders JavaScript; it does not solve challenges.",
      +  "type": "boolean"
      +}
    • addedInput schema / properties / respect_robots
      Added value: +{
      +  "description": "Respect the target site's robots.txt (default: true). Setting this to false is honoured, returns a warning in the response, and is recorded against your API key — it is your decision, not a silent default.",
      +  "type": "boolean"
      +}
  6. Changed3 schema fields changedv5.0.4
    • changedInput schema / $schema
      Previous value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema"
    • addedInput schema / properties / actions / items / properties / x
      Added value: +{
      +  "description": "scroll: absolute X coordinate to scroll to (window.scrollTo; with y, takes precedence over direction/distance)",
      +  "minimum": 0,
      +  "type": "number"
      +}
    • addedInput schema / properties / actions / items / properties / y
      Added value: +{
      +  "description": "scroll: absolute Y coordinate to scroll to (window.scrollTo; with x, takes precedence over direction/distance)",
      +  "minimum": 0,
      +  "type": "number"
      +}
  7. First observedv4.10.0

TDQS

A4.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description adds significant behavioral context beyond the annotations: the snapshot-ref workflow, navigation invalidating refs, robots.txt being respected, screenshot storage format (crawlforge://screenshot/{actionId}), the stealth engine auto-selection behavior, and the 5-credit cost. The annotations declare readOnlyHint=true and destructiveHint=false, which the description does not contradict. It doesn't mention that stealth doesn't solve challenges (that's in the schema), but the description is already rich.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but every sentence earns its place: usage trigger, action list, snapshot workflow, stealth/browserOptions semantics, robots.txt note, screenshot URI scheme, exclusions, cost, and a complete example. It front-loads the decision-relevant trigger and keeps the example last. No filler or redundancy with the schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex tool with 14 parametersaging the chain, the description covers the critical operational concepts (refs, navigation invalidation, stealth engine choice, robots gating, screenshot storage, cost, exclusions, example). The output schema is absent, but the description notes the result returns which engine ran and the truncation behavior is in the schema. Nothing needed to call it correctly is missing; the example ties it together.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description goes beyond the schema by explaining the snapshot-ref pattern (use @e1 refs instead of guessing CSS selectors), the stealth-engine selection semantics (auto = camoufox if installed, else Chromium, result says which ran), and the robots.txt behavior. The main action-list is recounted in the description, which overlaps with schema enums, but the ref workflow and cost info add real value.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb ('Use this when you must interact with a page before scraping') and enumerates concrete use cases (login, click buttons, fill forms, scroll, wait for dynamic content), then explicitly contrasts with the sibling tool 'scrape' for pages that render without interaction. The action list and example further clarify the resource and scope.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides explicit when-to-use guidance ('Use this when you must interact...'), lists the full action vocabulary, instructs to start with snapshot and re-snapshot after navigation, and explicitly states exclusion criteria: 'Not for pages that render without interaction (scrape) and not as the first attempt on a blocked site (stealth_mode operation:"scrape")'. This gives the agent clear decision rules and names the alternative tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.