Skip to main content
Glama

Spicrawl

Server Details

Web data for AI agents. Spicrawl MCP lets agents fetch public web pages as clean Markdown, HTML, text, screenshots, PDFs, or structured JSON. It handles JavaScript-heavy pages and browser actions like clicking, filling forms, and scrolling.

Ownership verified
Status
Healthy
Last Tested
Transport
Streamable HTTP · MCP 2025-11-25
URL

TDQS

A4.1/5.0

Scored across 25 tools

Disambiguation5/5

Each tool targets a clearly distinct resource or action, and overlapping areas (session_get vs session_context, batch_status vs results vs task_content, usage vs usage_summary/reconciliation) are explicitly differentiated in their descriptions. No two tools appear to do the same thing, so an agent can reliably pick the right one.

Naming Consistency4/5

All tools use the spicrawl_ prefix and snake_case, and most follow a domain_action pattern (batch_*, session_*, docs_*, usage_*). Minor deviation: spicrawl_request_get is singular while spicrawl_requests_list is plural, and spicrawl_browser_connect_url is less nested, but overall predictability is high.

Tool Count3/5

The 25 tools span seven functional areas and each batch/session lifecycle operation earns its place, but this is heavy for the stated scope and sits at the top of the borderline 16-25 range. Some consolidation (e.g., usage variants) could reduce the surface without losing capability.

Completeness4/5

The set covers single scrape, batch lifecycle (submit, add, close, cancel, retry, status, results, task content), sessions (create, get, list, context, release, delete), docs, usage, and request logs thoroughly. Minor gaps like no batch-job deletion or cache-management tool exist, but agents can work around them.

Available Tools

25 tools
spicrawl_batch_add_itemsAdd URLs to an open batch jobAInspect

Append URLs to a batch job that was submitted with open: true. Same item vocabulary as spicrawl_batch_submit (EITHER urls or items, not both, plus shared settings for these new items). Fails with CONFLICT if the job is closed or finished. Not idempotent: calling twice adds the URLs twice. Returns { job, items_added, items_dispatched }.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlsNoShorthand: URLs scraped with the shared settings and no per-item overrides. Mutually exclusive with `items`.
waitNoMilliseconds to wait after load (rendered scrapes).
itemsNoLong form: one /v1/scrape request object per item. Besides `url` and `external_id`, any /v1/scrape field (e.g. `js_render`, `wait_for`, `response_format`) overrides the shared setting for that item only. Mutually exclusive with `urls`.
engineNoPin the scrape engine for every item: `fetch` or `obscura` (must be allowed by the plan and served by the deployment); `camoufox` is coming soon, do not pin it yet. `chromium` cannot be pinned in a batch — use `render: true`, or spicrawl_scrape.
formatNoOutput format for every item (`response_format`). spicrawl_batch_submit defaults to 'markdown', the same as spicrawl_scrape; an item's own `response_format` overrides it.
job_idYesThe batch job id returned by spicrawl_batch_submit or spicrawl_batch_list.
renderNoRender every URL with a browser (`js_render`).
max_costNoCeiling in credits for ONE item (not the whole job — see credit_budget).
wait_forNoCSS selector to wait for before capturing (rendered scrapes).
premium_proxyNoComing soon: not available yet, do not send. Use residential managed-pool exits for every URL.
proxy_countryNoComing soon: not available yet, do not send. ISO-3166 alpha-2 exit country; requires premium_proxy.
custom_headersNoExtra request headers sent with every URL.
block_resourcesNoResource types to block while rendering, e.g. ["image","font","media"].
main_content_onlyNoStrip each page to its main article.

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare idempotentHint=false and destructiveHint=false; the description reinforces non-idempotence with a concrete consequence ('calling twice adds the URLs twice') and adds the CONFLICT failure mode plus the return shape `{ job, items_added, items_dispatched }`. That is real context beyond the annotation set.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four dense sentences, front-loaded with the precondition and vocabulary, then failure modes, then return value. Efficient overall; the mutual-exclusivity restatement slightly duplicates the schema but is defensible as a guardrail.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 14-parameter mutation tool with no output schema, the description supplies the missing pieces: the open-job precondition, the closed-job failure mode, non-idempotence, and the return shape. An agent has everything needed to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3, but the description adds cross-tool semantics the schema cannot express: the item vocabulary mirrors spicrawl_batch_submit and the shared settings apply only to the newly appended items, clarifying scope of the settings parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Specific verb (append) + resource (URLs to a batch job) with the exact precondition (`open: true`) that distinguishes it from spicrawl_batch_submit. An agent can tell what this does and which sibling it complements without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

States the precondition (`open: true`) and the failure condition (CONFLICT if closed/finished), and cross-references spicrawl_batch_submit for the shared item vocabulary. It stops short of naming an explicit alternative for the closed-job case, but the CONFLICT note makes the boundary clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

spicrawl_batch_cancelCancel a batch jobA
DestructiveIdempotent
Inspect

Cancel a batch job: items not yet started are cancelled and will never run; items already finished keep their results; items already running finish first, so the job comes back cancelling and moves to cancelled when they drain. Cannot be undone — resubmit to run the URLs again. Returns the job object.

ParametersJSON Schema
NameRequiredDescriptionDefault
job_idYesThe batch job id returned by spicrawl_batch_submit or spicrawl_batch_list.

TDQS

A4.3/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The annotations already declare destructiveHint=true and readOnlyHint=false, and the description meaningfully enriches this by explaining exactly what is destroyed or preserved: unstarted items never run, finished items keep results, running items finish first, and the job moves through `cancelling` to `cancelled`. It also states irreversibility and that the job object is returned.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with the core action and then spends its remaining length on high-value operational details. Every clause carries useful semantic information about item states or reversibility, with no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description correctly notes that the job object is returned. It covers mutation behavior, irreversibility, item-level outcomes, and job-state transitions, which is complete for a one-parameter destructive batch action.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the single job_id parameter is fully documented in the schema itself. The description adds no parameter syntax or format information beyond the schema, so the baseline of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: cancel a batch job. It also specifies the exact item-level effects and resulting job state, so an agent can distinguish it from sibling batch operations like retry, close, or add_items without opening their schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explains the consequences of cancelling but does not explicitly say when to use this tool versus alternatives such as spicrawl_batch_close or spicrawl_batch_retry. Usage is implied by the operation name and the note that resubmission is needed to run URLs again, but there is no direct routing guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

spicrawl_batch_closeClose an open batch jobA
Idempotent
Inspect

Stop an open batch job accepting items, so it completes once its queued work drains. Idempotent (closing a closed job succeeds). Returns the job object.

ParametersJSON Schema
NameRequiredDescriptionDefault
job_idYesThe batch job id returned by spicrawl_batch_submit or spicrawl_batch_list.

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare idempotentHint=true and destructiveHint=false, but the description concretizes the idempotency ('closing a closed job succeeds') and explains the graceful drain semantics, which is real context beyond the boolean hints. It does not mention auth requirements or rate limits, but for a job-state transition that is minor.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three compact sentences, no filler, with the core action and its completion semantics front-loaded and idempotency plus return value following. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter state-transition tool with no output schema, the description covers what it does, when the effect takes place, idempotency, and the return value ('Returns the job object'). Only the relationship to the sibling cancel tool is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the single job_id parameter is fully documented in the schema, including its provenance (returned by spicrawl_batch_submit or spicrawl_batch_list). The description adds no parameter detail, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Stop an open batch job accepting items') with the resulting behavior ('completes once its queued work drains'), which is far more informative than the title 'Close an open batch job'. However, it never explicitly contrasts itself with the sibling spicrawl_batch_cancel, leaving the close-vs-cancel distinction to inference.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'so it completes once its queued work drains' implies the graceful-drain use case and implicitly distinguishes it from cancellation, but no alternative tool is named and no when-not guidance (e.g. 'use batch_cancel to abandon queued work') is given.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

spicrawl_batch_listList batch jobsA
Read-onlyIdempotent
Inspect

List this API key's project's batch jobs, newest first, with status and progress. Returns { batches: [...], next_cursor? }; pass next_cursor back as cursor for the next page.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoPage size (default 50, max 200).
cursorNo`next_cursor` from a previous page, verbatim.
statusNoOnly jobs in this status.

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover the safety profile (readOnly, idempotent, non-destructive, closed-world), so the bar is lower; the description adds the ordering guarantee, the project-scoped visibility model, and the concrete return shape with a next_cursor field. That is genuine context beyond the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two compact sentences, front-loaded with scope/ordering and then the return and pagination contract. No filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description usefully supplies the response shape and pagination contract, and the schema covers all three parameters including the status enum. It omits any signal about how this differs from batch_status/batch_results, which is the remaining gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description earns above baseline by explicitly coupling the returned next_cursor to the cursor input ('pass next_cursor back as cursor'), which the schema cannot express on its own.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Specific verb+resource ('List ... batch jobs') with stated scope ('this API key's project's') and ordering ('newest first'). It does not contrast itself against the batch_status sibling, which is the natural ambiguity to resolve, so it stops short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use or when-not guidance is given, and no alternative sibling (e.g. spicrawl_batch_status for a single job) is named. The only procedural detail is the pagination loop, which is mechanics rather than tool selection.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

spicrawl_batch_resultsRead a batch job's resultsA
Read-onlyIdempotent
Inspect

Return a page of FINISHED items of a batch job (one record per item: seq, url, external_id, status, error, and content, or a result_url API path (GET /v1/batch/{id}/tasks/{seq}/content, same key) for large bodies). Unfinished items are not listed yet. Returns { results: [...], next_cursor? }; pass next_cursor back as cursor for the next page. Results expire after the job's retention window (410 Gone).

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoPage size (default 500, max 5000).
cursorNo`next_cursor` from a previous page, verbatim.
job_idYesThe batch job id returned by spicrawl_batch_submit or spicrawl_batch_list.
statusNoOnly items that finished with this status.

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover safety (readOnly, idempotent, non-destructive), and the description adds substantial extra behavior: only finished items appear, results expire with 410 Gone, large bodies are returned as a result_url path rather than inline content, and the page envelope shape. This is well beyond what the annotations provide.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single dense paragraph with no filler; the item shape and pagination mechanism are front-loaded, and the retention caveat is placed last as a secondary warning. Every sentence carries actionable information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description supplies the return envelope ({ results, next_cursor }) and the per-item fields, plus the large-body indirection and expiry behavior. An agent has everything needed to page correctly and interpret results.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents job_id, cursor, limit, and status. The description reinforces cursor usage and the status filter but adds little syntax or format detail beyond the schema. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: returns a page of FINISHED items of a batch job, with one record per item and the exact fields returned. It is clearly distinguishable from siblings like spicrawl_batch_status (job-level status) and spicrawl_batch_task_content (single item content).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear conditions: only finished items are listed, unfinished items are excluded, and results expire after the retention window (410 Gone). It also explains pagination flow via next_cursor. It does not explicitly name the sibling alternative for fetching a single large body, though it does give the raw API path.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

spicrawl_batch_retryRetry a batch job's failed itemsAInspect

Re-queue every FAILED item of a batch job so it runs again (succeeded items are untouched; a job with no failures is returned unchanged). Returns { job, items_reset, items_dispatched }.

ParametersJSON Schema
NameRequiredDescriptionDefault
job_idYesThe batch job id returned by spicrawl_batch_submit or spicrawl_batch_list.

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With annotations already declaring readOnlyHint=false, destructiveHint=false, and idempotentHint=false, the description adds useful behavioral details: only FAILED items are re-queued, succeeded items are untouched, and a job with no failures is returned unchanged. It also describes the return shape. It does not cover authentication or rate limits, but the key mutation semantics are well explained.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences, front-loaded with the core action and followed by critical edge-case clarifications and the return shape. Every sentence earns its place with no redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter mutation tool with no output schema, the description provides everything needed: the action, edge cases, and the return value shape. Annotations cover safety, so nothing important is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the single parameter job_id is fully documented in the schema. The description adds no further meaning beyond referring to the batch job, so the baseline of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Re-queue') and resource ('every FAILED item of a batch job'), and distinguishes itself from siblings by clarifying that succeeded items are untouched. An agent can immediately understand the tool's function without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear context for when to use this tool (to retry failed items of an existing batch job), but it does not explicitly name alternatives such as submitting a new batch or cancelling the job. No exclusions or when-not-to-use guidance is given.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

spicrawl_batch_statusCheck a batch job's statusA
Read-onlyIdempotent
Inspect

Return one batch job with its status and progress (total/completed/succeeded/failed/remaining, credits charged). Poll this after spicrawl_batch_submit; once status is completed, failed or cancelled, read spicrawl_batch_results.

ParametersJSON Schema
NameRequiredDescriptionDefault
job_idYesThe batch job id returned by spicrawl_batch_submit or spicrawl_batch_list.

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly, idempotent and non-destructive, so safety is covered. The description adds lifecycle context (terminal statuses, the two-call poll-then-read pattern) and names the returned progress fields, which is useful beyond the annotations, though it doesn't mention rate limits or polling cadence.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with what is returned, followed by the workflow routing. Every clause earns its place with no redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, but the description explicitly enumerates the return payload (status, progress counters, credits charged), which compensates. An agent has everything needed to call it and interpret the result.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% with a single job_id parameter whose description already names its source tools. The description adds no syntax or format detail beyond the schema, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Specific verb+resource: returns one batch job with status and progress. It also enumerates the exact progress fields (total/completed/succeeded/failed/remaining, credits charged), which distinguishes it cleanly from siblings like spicrawl_batch_list and spicrawl_batch_results.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit lifecycle routing: poll this after spicrawl_batch_submit, and switch to spicrawl_batch_results once status is completed, failed or cancelled. The when-to-use and when-to-move-on conditions are both stated, leaving nothing to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

spicrawl_batch_submitSubmit a batch scrape jobAInspect

Submit many URLs to be scraped as one asynchronous job, with shared settings applied to every URL. Give EITHER urls or items, not both (at least one URL, unless open is true). Returns the job object (id, status, estimated_credits, progress). Poll with spicrawl_batch_status, read with spicrawl_batch_results or spicrawl_batch_task_content. Use this instead of many spicrawl_scrape calls when you have tens or thousands of URLs. Set open: true to keep the job accepting more URLs via spicrawl_batch_add_items; an open job never finishes until you call spicrawl_batch_close.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameNoA human label for the job.
openNoKeep the job accepting items after submission (append with spicrawl_batch_add_items, finish with spicrawl_batch_close).
urlsNoShorthand: URLs scraped with the shared settings and no per-item overrides. Mutually exclusive with `items`.
waitNoMilliseconds to wait after load (rendered scrapes).
itemsNoLong form: one /v1/scrape request object per item. Besides `url` and `external_id`, any /v1/scrape field (e.g. `js_render`, `wait_for`, `response_format`) overrides the shared setting for that item only. Mutually exclusive with `urls`.
engineNoPin the scrape engine for every item: `fetch` or `obscura` (must be allowed by the plan and served by the deployment); `camoufox` is coming soon, do not pin it yet. `chromium` cannot be pinned in a batch — use `render: true`, or spicrawl_scrape.
formatNoOutput format for every item (`response_format`). spicrawl_batch_submit defaults to 'markdown', the same as spicrawl_scrape; an item's own `response_format` overrides it.
renderNoRender every URL with a browser (`js_render`).
max_costNoCeiling in credits for ONE item (not the whole job — see credit_budget).
priorityNoScheduling priority of the job relative to your other jobs.
wait_forNoCSS selector to wait for before capturing (rendered scrapes).
concurrencyNoMax items in flight at once for this job (may be capped by the server; see warnings).
max_attemptsNoPer-item retry attempts on failure.
credit_budgetNoCeiling on the WHOLE job, in credits; the run aborts when it would exceed this.
premium_proxyNoComing soon: not available yet, do not send. Use residential managed-pool exits for every URL.
proxy_countryNoComing soon: not available yet, do not send. ISO-3166 alpha-2 exit country; requires premium_proxy.
custom_headersNoExtra request headers sent with every URL.
block_resourcesNoResource types to block while rendering, e.g. ["image","font","media"].
failure_thresholdNoAbort the job once this many items have failed.
main_content_onlyNoStrip each page to its main article.
webhook_endpoint_idNoId of a registered webhook endpoint notified (with the job object) when the job completes.

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare non-destructive, non-idempotent, open-world behavior, and the description adds crucial async lifecycle detail beyond them: returns a job object with specific fields, requires polling, open jobs never finish until closed, and shared settings apply to every URL. This operational model is not in the annotations. Auth/rate-limit notes are absent but not critical here.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the purpose, followed by constraints, return behavior, workflow routing, and the open/close pattern. Every sentence contributes and no filler is present. Appropriate density for a 21-parameter asynchronous tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given 21 params, nested items, no output schema, and a complex async workflow, the description covers submission rules, return shape, polling, result retrieval, and open/close lifecycle. Nothing essential for correct invocation is missing; per-parameter details are handled by the schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% for 21 parameters, so the schema already documents all parameter meanings; baseline 3 applies. The description adds only the conditional 'at least one URL unless open is true' and repeats the urls/items mutual exclusivity already present in schema descriptions. Marginal additional value over structured fields.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Submit) and resource (batch scrape job) with scope (many URLs, one asynchronous job, shared settings). Explicitly distinguishes from siblings by naming spicrawl_batch_status, spicrawl_batch_results, spicrawl_batch_task_content, and spicrawl_scrape. An agent can route correctly without opening schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly says to use this instead of many spicrawl_scrape calls when you have tens or thousands of URLs, and points to status/results/task_content for follow-up. Also explains the open:true workflow with spicrawl_batch_add_items and spicrawl_batch_close. No ambiguity about when to choose this tool vs alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

spicrawl_batch_task_contentRead one batch item's contentA
Read-onlyIdempotent
Inspect

Return the full payload one item of a batch job produced, by its seq (items are numbered from 0 in submission order). Cheaper than paging results when you need one URL's output. A CONFLICT error means either the item has not finished yet (poll spicrawl_batch_status and ask again) or it finished as failed (it will never produce content; the message says which). NOT_FOUND means no such seq.

ParametersJSON Schema
NameRequiredDescriptionDefault
seqYesThe item's sequence number (0-based, submission order).
job_idYesThe batch job id returned by spicrawl_batch_submit or spicrawl_batch_list.

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover read-only, idempotent, and non-destructive behavior, so the bar is lower. The description adds crucial error semantics for CONFLICT (distinguishing unfinished vs. permanently failed) and NOT_FOUND, which are not captured by structured annotations. It stops short of describing payload format, size limits, or any auth/rate-limit context, but the added error behavior is genuinely useful.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences with zero waste: core purpose and seq semantics front-loaded, cost/alternative rationale next, and error handling last. Every sentence earns its place and the structure is easy to scan.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description carries the burden of explaining return values ('full payload') and does so only generically. The thorough error semantics compensate substantially, but an agent still lacks detail on payload format or whether partial content is possible. Given the tool's focused scope and existing annotations, it is nearly complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both parameters are already fully documented. The description repeats the 0-based submission-order note for `seq` rather than adding new meaning (e.g., allowed ranges, relationship to job_id). Baseline 3 is appropriate when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Return the full payload'), resource ('one item of a batch job'), and access key ('by its seq'). It explicitly contrasts with sibling `spicrawl_batch_results` by calling itself 'cheaper than paging results', making the distinction actionable without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It clearly states when to use it ('when you need one URL's output') and names the alternative (`poll spicrawl_batch_status`) for unfinished items. The conflict and not-found error conditions give concrete operational guidance on retrying or abandoning the request.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

spicrawl_browser_connect_urlBrowser (CDP) connect URLAInspect

Coming soon: remote browsers are not available yet, so minting fails; do not rely on this tool. For scripted page interaction today use actions on spicrawl_scrape. Mints a single-use connect URL for Spicrawl's remote browser (GET /v1/browser, Chrome DevTools Protocol) so YOUR code can drive it with puppeteer.connect({ browserWSEndpoint }) or playwright.chromium.connectOverCDP(url). It calls POST /v1/browser/token and returns a ws:// or wss:// URL carrying a short-lived token, never the API key: the URL works ONCE, within 60 seconds, and only with the options given here (editing them in the URL is refused), so call this right before connecting and again for every new session. Prefer sending the key as an Authorization header when the client supports headers on the upgrade (playwright connectOverCDP(url, { headers }), puppeteer.connect({ browserWSEndpoint, headers })): then connect to the plain /v1/browser URL with 'Authorization: Bearer ' and no token is needed. The endpoint exists only when the deployment was installed with --with-cdp, and the key must carry the browser scope; otherwise minting fails (404 or 403). Parameter errors are reported when minting.

ParametersJSON Schema
NameRequiredDescriptionDefault
proxyNoYour own proxy, as http://user:pass@host:port. Replaces our pool, so proxy_country, proxy_region and sticky_key cannot be combined with it. Kept out of the returned URL.
engineNoBrowser engine. Default: the deployment's default (chromium). 'obscura' cannot run headful.
headlessNofalse = headful browser with a display (desktop-sized screen). Omit for the deployment default.
sticky_keyNoKeep one exit IP for the session under this key (2-64 letters, digits, hyphens). Omit to rotate (the default).
session_ttlNoSession lifetime in seconds, 60-900. Default 180.
proxy_regionNoBroader exit region/pool name; 'global' means no constraint. Mutually exclusive with proxy_country.
proxy_countryNoExit country, ISO-3166 alpha-2 lowercase (e.g. 'de'). Mutually exclusive with proxy_region.

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the annotations (readOnly/openWorld/idempotent/destructive), the description discloses that the URL is single-use, expires in 60 seconds, carries a short-lived token and never the API key, that baked-in options cannot be edited without refusal, and that minting fails with 404/403 unless the deployment has --with-cdp and the key holds the `browser` scope. This is exactly the kind of behavioral context annotations cannot convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The most important fact (unavailable today, use an alternative) is front-loaded, which is the right structural choice. It is long and dense, but nearly every clause earns its place by covering expiry, token handling, prerequisites, and the header alternative.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 7-param, no-required, no-output-schema tool, the description covers the return value shape (ws:// or wss:// URL with token), lifetime, one-time-use constraint, auth requirements, and deployment/scope prerequisites. Nothing an agent needs to call it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the schema already documents all seven parameters — baseline 3. The description adds meaning beyond the schema by explaining that the options supplied at mint time are frozen into the URL and cannot be changed afterward, which tells the agent to choose parameters carefully rather than edit the returned URL.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource — mints a single-use CDP connect URL for Spicrawl's remote browser — and names both the endpoint (GET /v1/browser) and the sibling alternative (`actions` on spicrawl_scrape). An agent can distinguish it from the session/batch siblings without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives explicit when-not guidance up front ('remote browsers are not available yet... do not rely on this tool'), names the alternative for today (`actions` on spicrawl_scrape), and describes an alternate invocation mode (Authorization header on plain /v1/browser) with the condition that selects it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

spicrawl_docs_indexList the Spicrawl docsA
Read-onlyIdempotent
Inspect

The docs' llms.txt: every Spicrawl documentation page with its title, a one-line description and its Markdown URL. Use it to find the right page when a search does not, or to get an overview of what is documented; read a page with spicrawl_docs_read. No parameters.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly, idempotent, non-destructive, openWorld, so the safety profile is covered. The description adds useful content-level behavior: it is a static llms.txt-style index and precisely what fields each entry carries. It does not discuss size/pagination, but for a fixed index that is minor.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with the resource and payload, followed by usage and the sibling hand-off. No filler or restatement of the title.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description takes on the return-value burden and discharges it by naming the exact fields (title, one-line description, Markdown URL). For a no-parameter read-only index tool, an agent has everything needed to call it and use the result.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Zero parameters, so the 4 baseline applies. The description explicitly confirms 'No parameters', matching the empty schema with additionalProperties false, removing any doubt about how to invoke it.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific resource (the docs' llms.txt index) and exactly what it contains: every documentation page's title, one-line description, and Markdown URL. It is clearly distinguishable from siblings spicrawl_docs_search and spicrawl_docs_read, both of which are named in the text.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit when-to-use conditions: 'find the right page when a search does not' and 'get an overview of what is documented', and routes to the correct alternative for the follow-up step ('read a page with spicrawl_docs_read'). Nothing is left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

spicrawl_docs_readRead a Spicrawl docs pageA
Read-onlyIdempotent
Inspect

Fetches one Spicrawl documentation page as Markdown. path may be a page path (guides/anti-bot), a docs URL from spicrawl_docs_search (url or md_url), or either with a #anchor (the whole page is returned; the anchor names the section to look at). Only pages on the Spicrawl docs (https://docs.spicrawl.com/docs) can be read. Pages over 60000 characters are truncated, and the text says so. Don't know the path? Use spicrawl_docs_search, or spicrawl_docs_index (llms.txt) for the list of every page.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathYesPage path (`guides/anti-bot`, `errors#AUTH_INVALID_KEY`) or a full docs URL, as returned by spicrawl_docs_search.

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare read-only/idempotent/non-destructive, so the safety profile is covered; the description goes further by disclosing the 60000-character truncation with in-text notice and the domain restriction to https://docs.spicrawl.com/docs.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads what it does, then input forms, constraints, and alternatives. Every sentence carries distinct information with no padding.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter read tool with no output schema, the description covers the return format, truncation behavior, scope limits, and fallback tools. Nothing needed for correct invocation is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so baseline is 3, but the description adds meaning beyond the schema: accepted alternative forms (docs URL from search) and the anchor semantics (whole page is returned, the anchor only names the section to look at).

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (fetches), resource (one Spicrawl documentation page), and return format (Markdown). It is clearly distinguishable from sibling docs tools like spicrawl_docs_search and spicrawl_docs_index.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly routes the agent: when the path is unknown, use spicrawl_docs_search or spicrawl_docs_index (llms.txt). It also states the input forms accepted (page path, url, md_url, optional #anchor), leaving little to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

spicrawl_request_getGet one requestA
Read-onlyIdempotent
Inspect

The full log record of one past API call by its request id. Every response carries its id in the X-Request-Id header (and error bodies carry it too), so this is the tool for debugging a failed or slow scrape: it shows status, error_code/error_detail, attempts, the engine and proxy (tier, country, provider, sticky) actually used, whether and how the target blocked it (vendor/signal/rule), per-stage latency spans and credits charged. Returns not_found both for unknown ids and for ones older than your retention window, and for another project's request unless your key has the read scope.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesThe request id (a ULID, from X-Request-Id or a spicrawl_requests_list row's `id`).

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare read-only, idempotent, non-destructive, closed-world behavior, so the bar is lower — yet the description adds real context beyond them: the exact failure semantics ('not_found' for unknown ids, ids past the retention window, and other projects' requests) and the `read` scope requirement. That is substantive behavioral disclosure an agent cannot get from structured fields.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with the core purpose before the detailed enumeration of returned fields. It is dense and long, but nearly every clause earns its place by naming what the record contains; the parenthetical field list is the only slightly heavy portion.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description compensates by enumerating the returned fields (status, error_code/error_detail, attempts, engine/proxy details, blocking info, latency spans, credits) plus error and auth semantics. Nothing essential for correct invocation is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the single `id` parameter is already documented as a ULID sourced from X-Request-Id or a list row. The description restates that sourcing rather than adding new format or validation detail, so the schema carries the weight and baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('full log record of one past API call by its request id'), and the singular scope ('one') cleanly distinguishes it from the sibling spicrawl_requests_list. An agent can route between list and get without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly frames the tool's role: 'this is the tool for debugging a failed or slow scrape,' and tells the agent where to obtain the id (X-Request-Id header, error bodies, or a spicrawl_requests_list row). It stops short of stating when-not to use it or naming a competing tool, so it is clear context rather than full routing guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

spicrawl_requests_listList recent requestsA
Read-onlyIdempotent
Inspect

Recent API calls in your key's project, newest first, each with project_id, status, engine, http_status, error_code/error_detail, attempts, URL, proxy used, latency spans (queue/engine/upstream/total ms), block detection and credits. Use it to find failing calls (only_errors=true) or to spot patterns across a target. all_projects=true lists every project in your organization instead (needs a key with the read scope; 403 ERR::AUTH::INSUFFICIENT_SCOPE otherwise). Records only exist within the retention window reported in page.retention_hours / page.retained_since; an empty page means nothing in that window. Paginate by passing page.next_before and page.next_before_id back as before and before_id while page.has_more is true, keeping all_projects the same. Not billing-grade: use spicrawl_usage for totals.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoRows per page, 1-200. Default 50.
beforeNoCursor: RFC 3339 timestamp, exactly as returned in page.next_before. Pass together with before_id.
statusNoOnly requests with this outcome.
before_idNoCursor: request id, exactly as returned in page.next_before_id.
only_errorsNoOnly requests whose status is anything other than 'success'.
all_projectsNoEvery project in your organization, not just the key's own. Needs the `read` scope. Keep it the same across pages.

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly/idempotent/non-destructive, and the description still adds substantial behavioral context: retention-window semantics with page.retention_hours / page.retained_since, what an empty page means, scope-gated failure mode, and exact pagination contract. This is well beyond what the structured fields convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Dense but front-loaded: purpose and field list come first, then usage, then the pagination and retention mechanics. Every sentence carries operational information and none is filler, though the pagination sentence is long enough to slow scanning.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description compensates by naming the returned fields, the page envelope fields, and edge cases (empty page, retention window). An agent has everything needed to invoke and paginate correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds real meaning: it explains cursor round-tripping (page.next_before/page.next_before_id back as before/before_id), that all_projects must stay constant across pages, and that only_errors means anything other than 'success'. Only minor gaps remain (e.g., it doesn't restate limit/status semantics).

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a precise verb+resource+ordering ('Recent API calls in your key's project, newest first') and enumerates the returned fields, so the agent knows exactly what it gets. It also distinguishes itself from spicrawl_usage for billing totals, separating it from the nearest ambiguous sibling.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit when-to-use cases are given ('find failing calls (only_errors=true) or spot patterns across a target'), plus the alternative all_projects=true mode with its precondition (needs `read` scope, else 403 ERR::AUTH::INSUFFICIENT_SCOPE). It even routes to a sibling ('Not billing-grade: use spicrawl_usage for totals').

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

spicrawl_scrapeScrape a URLAInspect

Retrieve a web page and return its content as markdown (default), text, HTML, a JSON envelope, or a printed PDF. Handles the fetch vs. browser decision, charset decoding, and PDF-to-text for you. Optionally extract structured data with CSS selectors (extract) or autoparse (a natural-language ai_extract is coming soon), return discovered links, take a screenshot, or scope the DOM with include_tags/exclude_tags. Each successful call is billed in credits (a failure costs 0). Results are cached by default, and a cache hit is billed like the fetch that stored it (the cache saves time, not credits); set cache=false for time-sensitive pages. With the defaults it only reads the page. A method other than GET, HEAD or OPTIONS, and the actions that click, fill or run scripts on the page, can change state on the target site: set them only when the user asks for that. For markdown, text and HTML it returns { content }. With format: "json", or when extract, ai_extract, autoparse, links, screenshot or network_capture is set, it returns the JSON envelope: content, the site's status, credits, engine, warnings and data for extraction. Screenshots come back as image content blocks; each screenshots entry in the JSON keeps its metadata and says which block holds it. A PDF comes back as a resource content block, with its size and the engine and credits in the JSON.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYesThe absolute URL to retrieve.
modeNoRouting mode. `auto` lets the platform escalate fetch -> browser as needed, billing only the rung that worked.
waitNoMilliseconds to wait after load before capturing (render only).
cacheNoServe a recent cached result if fresh (on by default). A hit is billed at the same price as the fetch that stored it: the cache saves time, not credits. Set false only when you need live data.
linksNoAlso return the page's discovered links as an absolute, de-duplicated list.
proxyNoYour own proxy URL to egress through. Takes precedence over the pool; no surcharge.
engineNoPin the execution engine instead of letting the router choose. A pinned engine the deployment does not run is refused (ERR::ENGINE::UNAVAILABLE), never substituted. `camoufox` is coming soon: do not pin it yet.
formatNoOutput format. 'markdown' (default) is best for feeding to an LLM; 'html' returns the raw document; 'json' returns a structured envelope; 'pdf' prints the page in a browser and returns the file as a resource content block (needs `render` or `engine: "chromium"`).markdown
methodNoHTTP method sent to the target (default GET; no request body can be sent). Anything but GET, HEAD or OPTIONS can change state on the target site: use it only when the user asks.
renderNoRender with a real browser (executes JavaScript). Use for single-page apps or pages whose content is painted client-side. Costs more than the default fetch tier.
actionsNoBrowser steps run in order before capture (render only). Each step is {type, ...fields}. Examples: [{"type":"click","selector":"#load-more"},{"type":"wait_for","selector":".item:nth-child(20)","timeout_ms":5000}] · infinite scroll: [{"type":"scroll","to_bottom":true},{"type":"wait_for","ms":1000}] · login: [{"type":"fill","selector":"#user","value":"me"},{"type":"fill","selector":"#pass","value":"…","secret":true},{"type":"click","selector":"button[type=submit]","wait_for_navigation":true}]. A screenshot step switches the response to the JSON envelope (a FORMAT_COERCED warning says so).
extractNoSelector-based extraction: a map of field -> CSS/XPath selector, e.g. {"title":"h1","price":".amount"}. Returns structured `data`. Precise and cheap when you know the page structure.
stealthNoComing soon: not available yet, do not send. Stealth mode: render in the hardened Camoufox browser. Default false.
headlessNofalse runs Chromium on a real display (1920x1080 screen). chromium engine only.
max_costNoRefuse (rather than run) a request that would cost more than this many credits.
wait_forNoCSS selector to wait for before capturing (render only).
autoparseNoReturn the structured data the page publishes about itself (JSON-LD, OpenGraph, Twitter Card, microdata). No selectors; survives redesigns.
cache_ttlNoMax age in seconds of a cached copy to accept (0 forces a fresh fetch; capped at 48h). Freshness only: a hit is billed like a fetch.
parse_pdfNoTurn a PDF target into text/markdown (default true). false refuses PDFs.
ai_extractNoComing soon: not available yet, do not send. Model-driven extraction: describe what you want and let a model find it, no selectors. Refused with ERR::INTERNAL::UNAVAILABLE until it launches; use `extract` or `autoparse`.
screenshotNoCapture a screenshot (requires render).
session_idNoRun inside a session from spicrawl_session_create: same exit IP and browser state (cookies, localStorage) across calls.
sticky_keyNoComing soon: not available yet, do not send. Reuse the same managed-pool exit for every request carrying this key, without a session.
impersonateNoFetch tier only: present a real browser's TLS/JA3 + HTTP/2 fingerprint to clear passive bot gates. On by default; set false to send a plain client handshake.
exclude_tagsNoCSS selectors to remove (e.g. ['nav','.ad','#cookie']).
include_tagsNoCSS selectors to KEEP (everything else is dropped).
proxy_verifyNoVerify `proxy` is reachable before using it.
premium_proxyNoComing soon: not available yet, do not send. Use a residential exit from Spicrawl's managed pool instead of datacenter. Until then, pass your own `proxy`.
proxy_countryNoComing soon: not available yet, do not send. ISO-3166 alpha-2 country for a managed-pool exit, lowercase (e.g. 'us', 'de').
custom_headersNoExtra request headers sent to the target, e.g. {"Accept-Language":"de"}.
extract_presetNoComing soon: not available yet, do not send. A named extraction preset instead of an inline `extract` map; any value fails today. Put the rules in `extract`.
block_resourcesNoResource types the browser should not load, e.g. ["fonts","media","images"] (render only). Faster and cheaper.
network_captureNoRecord the API responses the PAGE fetched while rendering — often cleaner JSON than the DOM. Render only.
original_statusNoReturn the target's own HTTP status instead of 200 for a successful scrape.
wait_for_timeoutNoMax milliseconds to wait for `wait_for`.
main_content_onlyNoStrip the page to its main article, dropping nav/footer/aside. Defaults on for markdown.
screenshot_formatNoScreenshot image format.
screenshot_qualityNoScreenshot quality 1-100 (lossy formats).
screenshot_fullpageNoCapture the whole page, not just the viewport.
screenshot_selectorNoCapture only the element matching this CSS selector.
allowed_status_codesNoTarget statuses to treat as success instead of an error (e.g. [404] to capture a not-found page).

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Far exceeds the annotations: discloses credit billing (failures cost 0), the counterintuitive cache billing rule (a hit costs the same as the fetch), that pinning an unavailable engine is refused rather than substituted, and that non-GET methods or click/fill/script actions can mutate the target site. It also describes output shapes (image blocks for screenshots, resource block for PDF) that annotations cannot convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Dense and front-loaded: format/behavior first, then billing, then mutation warnings, then response shapes. Every sentence carries information, and the length is justified by 41 parameters and no output schema, though a few clauses (e.g. the screenshots metadata sentence) are harder to parse than necessary.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema and 41 parameters, the description carries the return-value burden and does so completely: it defines the { content } shape for text formats, the full JSON envelope fields (content, status, credits, engine, warnings, data), and how screenshots and PDFs are returned. Nothing an agent needs to call it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds genuine meaning beyond the schema: cache billing semantics, that screenshots switch the response to the JSON envelope, and that unset defaults keep the call read-only. It does not re-explain the many nested action fields, but those are well covered in the schema itself.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a precise verb (retrieve/scrape) and resource (web page) with the default output (markdown) and the full range of formats and extraction modes. It is unmistakably distinct from sibling tools like spicrawl_batch_submit or spicrawl_session_create, which manage sessions/batches rather than fetching a URL.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit when-to-use context (render for SPAs, cache=false for time-sensitive pages, method/actions only when the user asks) and clear when-not warnings about state-changing operations. It does not, however, route the agent to alternatives such as the batch tools for multi-URL work, so it falls short of full 5-level guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

spicrawl_session_contextDump a session's stored contextA
Read-onlyIdempotent
Inspect

Return the full stored state of a LIVE session — cookies, localStorage and other credentials — plus its fingerprint and domain scores. The session_context object can be passed to spicrawl_session_create to clone the session. The result contains secrets (login cookies): do not echo it to the user or logs unless asked.

ParametersJSON Schema
NameRequiredDescriptionDefault
session_idYesThe session id returned by spicrawl_session_create or spicrawl_session_list.

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly/ idempotent / non-destructive, so the safety profile is covered. The description adds meaningful non-obvious context: the operation only applies to LIVE sessions, the payload contains secrets (login cookies), and it should not be echoed to users or logs. That secret-handling warning is exactly the kind of behavioral caveat annotations cannot express.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences, front-loaded with what is returned, followed by the clone use case and the secret-handling caveat. No filler, every sentence carries distinct information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, so the description reasonably describes the return payload (context object, fingerprint, domain scores) and how to reuse it. It lacks any note on behavior for non-live/expired sessions or error handling, but for this shallow, single-param read tool it is close to complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and the single session_id parameter is fully documented in the schema, including where it comes from (create/list). The description adds nothing about the parameter, so the baseline 3 for schema-covered params is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Return) and resource (full stored state of a LIVE session), then enumerates contents — cookies, localStorage, credentials — plus fingerprint and domain scores. An agent can distinguish this from spicrawl_session_get/list, which the sibling names suggest provide metadata rather than full state.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives one concrete use case — passing the returned session_context to spicrawl_session_create to clone a session — which implies when to reach for it. However, it never states when NOT to use it or how it differs from spicrawl_session_get, leaving the agent to infer the distinction from sibling names alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

spicrawl_session_createCreate a browsing sessionAInspect

Create a session: a persistent browser identity (cookie jar, localStorage, pinned engine and, by default, a sticky exit IP) that survives across scrapes. Use it for multi-step flows — log in, then fetch pages behind the login — by passing the returned id as session_id to every spicrawl_scrape call that should share the state. The engine is pinned for the session's lifetime. Without engine, the server picks one this deployment runs: obscura, else chromium, else fetch. Release the session with spicrawl_session_release when done. Returns the session object (id, engine, status, sticky_key, proxy, expires_at).

ParametersJSON Schema
NameRequiredDescriptionDefault
engineNoEngine pinned for the session's lifetime. Default: `obscura` if deployed, else `chromium`, else `fetch`. `camoufox` is coming soon: do not pin it yet.
rotate_ipNoComing soon: not available yet, do not send. Keep the cookie jar stable but rotate the managed-pool exit IP per request. Cannot combine with sticky_key.
sticky_keyNoComing soon: not available yet, do not send. Pin the managed-pool exit IP to this key; sessions sharing a key share an exit. Cannot combine with rotate_ip.
fingerprintNoComing soon: not available yet, do not send. Camoufox fingerprint pin, stored verbatim. Only valid with engine `camoufox`.
region_poolNoComing soon: not available yet, do not send. Managed-pool region to draw the exit from.
ttl_secondsNoIdle lifetime in seconds (min 30, default 1800; the organization's maximum is enforced, not clamped).
domain_scoresNoSeed per-domain error scores (usually copied from a spicrawl_session_context dump).
premium_proxyNoComing soon: not available yet, do not send. Use a residential managed-pool exit, exactly as on spicrawl_scrape.
proxy_countryNoComing soon: not available yet, do not send. ISO-3166 alpha-2 exit country; requires premium_proxy.
session_contextNoSeed the session's state (cookies, storage) — e.g. the `session_context` object from spicrawl_session_context, to clone a session.

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare the mutation safety profile (readOnly=false, idempotent=false); the description adds the real behavioral context an agent needs — state survives across scrapes, the engine is pinned for the session's lifetime and cannot change, the default engine resolution order (obscura → chromium → fetch), and the fields returned. This is substantive disclosure beyond structured fields.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with purpose and usage, then defaults, release, and return shape. It is long but each sentence carries information; the closing field enumeration ('id, engine, status, sticky_key, proxy, expires_at') is slightly redundant padding since no output schema exists but the list is still useful.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 10-parameter, no-required-param, no-output-schema creation tool, the description covers purpose, sibling routing (scrape, release), engine-pinning semantics, defaults, and a field-level summary of the return value. Nothing essential for correct invocation is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the schema already flags the 'coming soon' parameters, so the baseline is 3. The description still adds value by explaining the engine default-selection fallback logic and the sticky exit IP default, though it does not go deeper into the remaining parameters (ttl_seconds, domain_scores, session_context) than the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Create a session') and immediately defines what a session is (persistent browser identity: cookie jar, localStorage, pinned engine, sticky exit IP), which cleanly separates it from sibling read/lifecycle tools like spicrawl_session_get or spicrawl_session_context.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly states when to use it ('multi-step flows — log in, then fetch pages behind the login'), how to consume it ('passing the returned id as session_id to every spicrawl_scrape call that should share state'), and how to end it ('Release the session with spicrawl_session_release'). Alternatives and lifecycle are named, not inferred.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

spicrawl_session_deleteDelete a sessionA
DestructiveIdempotent
Inspect

Permanently delete a session and its stored state, including the record itself (prefer spicrawl_session_release if you only want to stop using it). Idempotent: deleting an already-deleted session succeeds. Refused while a scrape holds the session's lease unless force is true.

ParametersJSON Schema
NameRequiredDescriptionDefault
forceNoBreak an active lease (a scrape currently using the session). Only set when you know it is safe.
session_idYesThe session id returned by spicrawl_session_create or spicrawl_session_list.

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true and idempotentHint=true, so the safety profile is covered. The description still adds real value beyond them: it specifies what is destroyed ('stored state, including the record itself') and the lease-refusal/force behavior, which annotations cannot express.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, all load-bearing: the core action and its scope, the sibling routing, then the idempotency and refusal semantics. Front-loaded with the verb and no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

A destructive mutation with no output schema and full annotation coverage needs precisely this: what gets destroyed, what happens on repeat calls, and the prerequisite for success. All present, so an agent can call it correctly without further context.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both parameters (session_id, force) are already documented in the schema, including what force does. The description references force's effect ('unless force is true') but adds no syntax or format detail beyond the schema. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (delete) plus resource (session and its stored state, including the record itself), and explicitly contrasts with the sibling spicrawl_session_release. An agent can distinguish it from release/get/list without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Names the alternative (spicrawl_session_release) with the exact selecting condition ('if you only want to stop using it'), and states the guard condition for refusal (active lease) plus the escape hatch (force). Nothing is left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

spicrawl_session_getGet a sessionA
Read-onlyIdempotent
Inspect

Return one session's metadata: status, engine, proxy/exit, usage count, expiry, lease holder, and a summary of its stored context. Never returns the credentials themselves (use spicrawl_session_context for that).

ParametersJSON Schema
NameRequiredDescriptionDefault
session_idYesThe session id returned by spicrawl_session_create or spicrawl_session_list.

TDQS

A4.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, non-destructive and closed-world, so the safety profile is covered. The description's genuinely additive contribution is the negative disclosure that credentials are never returned, which is useful but modest. It adds no information about pagination, caching, or error behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with the returned-field list and closing with the one important exclusion plus the alternative tool. No filler or redundant restatement of the title.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description carries the burden of describing return values and does so by enumerating the fields. It also closes the obvious security question by stating credentials are excluded, leaving nothing essential an agent needs before calling it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the single session_id parameter is fully documented in the schema, including where it comes from (spicrawl_session_create or spicrawl_session_list). The description adds no parameter-level meaning beyond that, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Return one session's metadata') and enumerates exactly which fields come back (status, engine, proxy/exit, usage count, expiry, lease holder, context summary). It also explicitly differentiates itself from the sibling spicrawl_session_context by naming what that tool is for instead.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description routes the agent away from this tool when credentials are needed, naming spicrawl_session_context as the alternative. It gives clear context for a read of session metadata but does not state when to prefer spicrawl_session_list versus this tool, so it stops short of full explicit when/when-not guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

spicrawl_session_listList sessionsA
Read-onlyIdempotent
Inspect

List this API key's project's sessions, newest first. Returns { sessions: [...], next_cursor? }; pass next_cursor back as cursor for the next page. Never includes the stored credentials.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoPage size (default 50, max 200).
cursorNo`next_cursor` from a previous page, verbatim.
engineNoOnly sessions pinned to this engine.
statusNoOnly sessions in this state.

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false, so the safety profile is covered. The description goes beyond them by disclosing result ordering, the return envelope shape, cursor semantics, and a useful negative guarantee ("Never includes the stored credentials") — meaningful behavioral context a caller would otherwise have to discover.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences with zero filler: ordering and scope first, then return shape and pagination. Every clause carries information the caller needs.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description compensates by stating the response envelope and pagination field. Combined with annotations covering the safety profile and a fully documented schema, an agent has everything needed to call and page through this tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description adds real value beyond the schema by documenting the end-to-end cursor contract (next_cursor out, cursor in for the next page), which is the non-obvious part of the call. It says nothing extra about the engine/status filters, which are already enum-documented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ("List this API key's project's sessions") plus scope and ordering ("newest first"), so the agent knows exactly what it returns. It does not, however, name or contrast itself with siblings like spicrawl_session_get or spicrawl_requests_list, so differentiation is left to inference.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explains the pagination workflow (pass next_cursor back as cursor) but gives no explicit when-to-use statement, no exclusions, and no pointer to alternative list tools. Usage is implied by the resource name rather than stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

spicrawl_session_releaseRelease a sessionA
DestructiveIdempotent
Inspect

End a session when you are done with it: frees its exit IP and browser state; it can no longer be used with spicrawl_scrape. Its stored cookies and storage are purged and cannot be recovered; start a new session and log in again to continue. The record stays visible (status released) so its history can still be inspected. Refused while a scrape holds the session's lease unless force is true. Returns the session object.

ParametersJSON Schema
NameRequiredDescriptionDefault
forceNoBreak an active lease (a scrape currently using the session). Only set when you know it is safe.
session_idYesThe session id returned by spicrawl_session_create or spicrawl_session_list.

TDQS

A4.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true and idempotentHint=true, so the safety profile is covered. The description goes further with concrete destructive effects (exit IP freed, cookies/storage purged irrecoverably), the retained record with `released` status, and the lease-refusal rule with `force` — all behavioral context the annotations do not carry.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The imperative trigger is front-loaded, and each subsequent clause earns its place by adding a distinct fact (destruction, irrecoverability, record retention, refusal condition, return value). No padding or restatement of the name.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Despite no output schema, the description states the return value ('Returns the session object'), and for a two-parameter destructive tool it covers preconditions, side effects, recovery, and post-state completely. An agent has everything needed to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds operational meaning to `force` by framing it as breaking a lease held by an active scrape ('only set when you know it is safe'), which enriches the schema's own wording rather than repeating it.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('End a session') and immediately scopes it against alternatives by noting the record stays visible as `released`, implicitly distinguishing it from spicrawl_session_delete. An agent can tell what this does without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit trigger ('when you are done with it'), consequences that steer away from misuse (cannot be used with spicrawl_scrape thereafter), recovery guidance (start a new session and log in again), and a precondition for refusal with the `force` escape hatch. Nothing about when to call this vs. siblings is left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

spicrawl_usageUsage breakdownA
Read-onlyIdempotent
Inspect

Billing-grade usage for your org over a date window, from the daily rollup (usage_daily), grouped by day, project, engine, feature or key. group_by=key returns per API key (key_id, name, prefix, project_id) counters per metric per day, from usage_key_daily; keyless (dashboard) usage is not in it, and a project-scoped key sees only its own project's keys. Use it to answer 'how many requests/credits did I use, and where'. The window defaults to the last 30 days including today; to is EXCLUSIVE (a single day 2026-08-01 is from=2026-08-01, to=2026-08-02) and the window may not exceed 400 days. The response carries stale_seconds (age of the newest rollup row; null when the window is empty). Requires the key to have the read scope.

ParametersJSON Schema
NameRequiredDescriptionDefault
toNoEXCLUSIVE end date, YYYY-MM-DD (UTC). Default: tomorrow, so today is included.
fromNoInclusive start date, YYYY-MM-DD (UTC). Default: 29 days before today.
metricsNoMetrics to return. Omit for all of them.
group_byNoHow to group the rows. Default 'day'.
project_idNoRestrict to one project id (must be a valid Spicrawl identifier).

TDQS

A4.3/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the readOnly/idempotent annotations, it discloses the data source (usage_daily vs usage_key_daily), that keyless/dashboard usage is excluded, that a project-scoped key only sees its own project's keys, the maximum window length, and the stale_seconds freshness field. It also states the required `read` scope, which is genuine authorization context the annotations do not carry.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Content is front-loaded with the core purpose and every clause carries information, but the first sentence is a long multi-clause construction rather than a scannable structure. It is dense without being wasteful, though a tighter lead sentence would improve it.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only analytics tool with no output schema, the description covers provenance, grouping, scoping limits, window rules, and even the one notable response field (stale_seconds). An agent has everything needed to call it correctly, including the scope requirement.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so baseline is 3, but the description adds meaning beyond the schema: the semantics of group_by=key (per-API-key counters with key_id/name/prefix/project_id), the keyless exclusion, and a worked example of the exclusive `to` bound. These clarifications exceed what the schema alone communicates.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a specific verb-and-resource with scope: retrieval of billing-grade usage over a date window, sourced from usage_daily and grouped by day/project/engine/feature/key. It clearly conveys what the tool is, but it never names or contrasts with the closest siblings spicrawl_usage_summary or spicrawl_usage_reconciliation, leaving the agent to infer the boundary.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives concrete context for use ('Use it to answer "how many requests/credits did I use, and where"') and operational conditions such as the default 30-day window and the 400-day cap. However, it offers no explicit when-not guidance or named alternatives, so the agent can't tell from this text when usage_summary would be the better choice.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

spicrawl_usage_reconciliationUsage reconciliationA
Read-onlyIdempotent
Inspect

Drift report for ONE day: re-derives usage from that day's request log and compares it with the billed daily rollup, so you can see whether the rollup missed or double-counted anything. Defaults to yesterday (UTC); today is still accumulating and would always show drift. Requires the read scope.

ParametersJSON Schema
NameRequiredDescriptionDefault
dayNoThe UTC day to reconcile, YYYY-MM-DD. Default: yesterday.

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already establish readOnly, idempotent, non-destructive. The description adds the auth requirement (`read` scope), the UTC default, and the rationale for excluding the current day—useful context beyond the annotations, though no return semantics are given.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core purpose, then defaults and prerequisite; every sentence carries information. Slightly dense but no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only, single-parameter diagnostic tool with no output schema, the description fully conveys what it does, when it applies, and the auth precondition. Only the returned drift-report shape is left implicit, which is a minor gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and the single `day` parameter is already documented with its pattern and default. The description's restatement of the default and why today is excluded adds rationale but little new parameter syntax, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (re-derives and compares usage) on a specific resource (one day's request log vs the billed daily rollup). The 'drift report' framing and the one-day scope clearly separate it from spicrawl_usage and spicrawl_usage_summary.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear context: use it to check whether the rollup missed or double-counted usage, and warns that today would always show drift, steering the agent to a completed day. It does not, however, explicitly name the aggregate sibling tools it complements.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

spicrawl_usage_summaryUsage summaryA
Read-onlyIdempotent
Inspect

A compact current-period summary of your org's usage (totals as of now, from the daily rollup). No parameters. Use for a quick 'how much have I used so far'; use spicrawl_usage for a windowed breakdown. Requires the read scope.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover readOnly/idempotent/non-destructive, but the description adds material context beyond them: the data provenance ('from the daily rollup', 'totals as of now') implies freshness lag, and it discloses the required `read` scope. It stops short of stating rate limits or response shape, so not a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, zero padding, and the leading sentence front-loads what the tool returns before moving to routing and auth prerequisites.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, yet the description gives only a loose sketch of the return ('totals', a compact summary) without field-level detail. Everything needed to invoke it correctly (no params, read scope, when to prefer it) is present, so it is nearly complete for a simple read tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Zero parameters, so the baseline is 4. The description correctly confirms 'No parameters', which matches the empty schema and prevents an agent from guessing at filters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource+scope: a current-period org usage summary sourced from the daily rollup. It explicitly contrasts with the sibling spicrawl_usage ('windowed breakdown'), so an agent can route between them without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit use case ('quick how much have I used so far') and names the alternative plus the condition that selects it ('use spicrawl_usage for a windowed breakdown'). Nothing is left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 25 tool updates
    • First observedspicrawl_batch_add_items
    • First observedspicrawl_batch_cancel
    • First observedspicrawl_batch_close
    • First observedspicrawl_batch_list
    • First observedspicrawl_batch_results
    • First observedspicrawl_batch_retry
    • First observedspicrawl_batch_status
    • First observedspicrawl_batch_submit
    • First observedspicrawl_batch_task_content
    • First observedspicrawl_browser_connect_url
    • First observedspicrawl_docs_index
    • First observedspicrawl_docs_read
    • First observedspicrawl_docs_search
    • First observedspicrawl_request_get
    • First observedspicrawl_requests_list
    • First observedspicrawl_scrape
    • First observedspicrawl_session_context
    • First observedspicrawl_session_create
    • First observedspicrawl_session_delete
    • First observedspicrawl_session_get
    • First observedspicrawl_session_list
    • First observedspicrawl_session_release
    • First observedspicrawl_usage
    • First observedspicrawl_usage_reconciliation
    • First observedspicrawl_usage_summary

Related MCP Servers

  • A
    license
    A
    quality
    C
    maintenance
    Enables brand visibility monitoring across major AI platforms like ChatGPT, Claude, Gemini, and Perplexity. It allows users to track visibility scores, analyze competitor data, and receive actionable insights to improve AI-generated brand recommendations.
    16
    22 npm
    1
    MIT
  • A
    license
    Not graded
    quality
    B
    maintenance
    Browse IndustryLens's published competitive-intelligence reports and head-to-head competitor comparisons from any AI agent — real, source-backed data.
    MIT
Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.

Resources