MeshArc
Server Details
Crawl and scrape websites to clean markdown, and track what changed between crawls.
- Status
- Healthy
- Last Tested
- Transport
- Streamable HTTP · MCP 2025-11-25
- URL
- Repository
- mesharc-org/mesharc-python
- GitHub Stars
- 0
TDQS
Scored across 17 tools
The descriptions explicitly distinguish overlapping capabilities: one-off crawl/scrape/map/extract, project creation and monitoring, run triggering, job polling, and stored-page reading. Each tool has a clear role despite the domain having many adjacent operations, so an agent can reliably select the right tool.
All tool names use snake_case and almost all follow a verb_noun or verb_phrase pattern, such as crawl_site, create_project, get_page, list_pages, and start_run. Minor variants like keep_crawl_as_project and describe_project_config are still predictable and readable.
With 17 tools, the surface is slightly above the ideal 3-15 range, but the server covers several distinct modes: discovery, one-off fetching, project lifecycle, run management, result inspection, and job polling. Each tool appears to earn its place, though the set is on the heavier side.
The server covers core crawling and monitoring workflows: discovery, fetching, project create/read/list/update, run triggering, change detection, page search, and job polling. The main gap is the absence of delete_project or delete_run for cleanup, which could matter when project limits are reached, but most agent workflows remain supported.
Available Tools
17 toolscrawl_siteCrawl a whole site onceAIdempotentInspect
Crawl a site once from a start URL -- following its links and reading its sitemap, up to limit pages -- and return what it found, without setting up a project. Use it to read a site or section whose page URLs you do not know; for known URLs use scrape_urls, and to watch a site over time use create_project (or keep_crawl_as_project on this crawl afterwards). Each page costs the credits of the engine that read it (usually 1 to 4) and refused pages are free. It waits for the crawl, minutes for a large limit; hosted, a crawl still going after the time budget comes back as a job for get_job, and repeating the call returns the same crawl. The answer is an index of the pages read plus excerpts inside 60,000 characters; get_job with url reads one page in full and with cursor the next window. The crawl is kept for a day; pass the result's crawl_id to keep_crawl_as_project to keep it for good.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | The absolute http(s) URL to start from, e.g. the home page. | |
| limit | No | The most pages to read, 1 to 5,000 (default 50); the plan's page cap also applies. | |
| config | No | Optional fetch settings: any subset of the keys describe_project_config lists, e.g. {"render_js": "always", "only_main_content": true}. Omit for the defaults. | |
| max_depth | No | How many links deep to follow from the start URL (0 = that page only; default 3). | |
| exclude_paths | No | Globs over the URL path to skip, e.g. ['/tag/*', '*.pdf']; an exclude wins over an include. | |
| include_paths | No | Globs over the URL path to keep, e.g. ['/blog/*'] for a section (its index included); a bare '/blog/' matches only that one page. Omit for the whole site. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations cover safety (idempotent, non-destructive) but the description adds substantial behavioral context beyond them: per-page credit costs (1-4, refused pages free), synchronous waiting with a time budget, hosted crawls returning a job retrievable via get_job, idempotent repeat-call behavior, and a one-day retention window. This is far past what annotations alone convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose and alternatives are front-loaded and every sentence carries information. It is, however, a dense run-on block with several clauses chained by semicolons, which slightly hurts scannability for an agent parsing it quickly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite no output schema, the description explains the return shape (index of pages plus excerpts capped at 60,000 characters), how to page through one page via get_job with url vs cursor, and the crawl_id retention handoff. An agent has enough to call and consume the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents url, limit, config, max_depth, and include/exclude paths thoroughly. The description reinforces `limit` and the crawl_id return but adds no parameter syntax or format detail beyond the schema. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (crawl) and resource (site from a start URL), including the mechanism (following links, reading sitemap, up to `limit` pages) and the key scoping outcome (returns an index without creating a project). It explicitly distinguishes itself from scrape_urls and create_project. An agent can select it correctly without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit when-to-use ('whose page URLs you do not know') and when-not, naming the alternatives: scrape_urls for known URLs, create_project to watch a site over time, keep_crawl_as_project to persist this crawl. Nothing is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_projectCreate a watched projectAInspect
Create a project that watches a site over time: it is crawled on its schedule and each run is compared with the last, recording pages added, modified and removed. Use it when the request is to monitor or track a site; for a one-off read use crawl_site (keep_crawl_as_project can turn that into a project later), and check list_projects first so the site is not added twice. Creating it reads the site's robots.txt and sitemaps but fetches no pages; every run then spends credits per page. Refused when the plan's project limit is reached. Returns the project, whose id start_run takes to crawl it now.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | A name for the project; omit to use the site's host name. | |
| seed | Yes | Where the crawl starts: a public site's URL or bare domain, e.g. 'https://example.com/blog/' or 'example.com'. The project covers that host. | |
| config | No | Any subset of the settings describe_project_config lists. For one section of a site pass include_paths, e.g. {"include_paths": ["/blog/*"]}; map_site shows the site's sections first. Omit for the defaults. | |
| schedule | No | How often it re-crawls on its own; 'manual' (default) runs only when start_run is called. | manual |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the annotations, it discloses important behavior: creation reads robots.txt and sitemaps but fetches no pages, each run spends credits per page, and creation is refused when the plan's project limit is reached. It also explains the return value enough to point the agent to start_run.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a compact paragraph that front-loads the core purpose and then adds usage, side effects, limits, and return behavior with no filler. Every sentence carries actionable information for an agent.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite having no output schema, the description explains what is returned (the project and its id for start_run). It also covers prerequisites, alternatives, operational cost, and plan limits, which is complete for a high-impact creation tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the input schema already documents all four parameters in detail. The description adds useful context about the project lifecycle but does not add syntax or per-parameter meaning beyond what the schema provides, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: creating a project that watches a site over time, crawling on a schedule and comparing runs. It also names the key siblings it is not (crawl_site, keep_crawl_as_project, list_projects), so an agent can distinguish it without opening schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly says when to use it (monitor or track a site) and when not to use it (one-off reads should use crawl_site, and keep_crawl_as_project can convert later). It also adds a concrete precondition: check list_projects first to avoid adding the same site twice.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
describe_project_configDescribe project settingsARead-onlyIdempotentInspect
List the project settings an assistant can set -- each key with its meaning and its default -- plus the valid schedules and how path globs work. Read it before create_project, update_project or a config argument whenever the request names a section, schedule, format, limit or behaviour: keys are exact and guessed names are refused. It describes settings in general; get_project shows the values one project has. Free, and it fetches nothing.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint/idempotentHint/destructiveHint=false, and the description adds context beyond them: 'Free, and it fetches nothing' (cost and no network/side effects) and 'keys are exact and guessed names are refused' (validation behavior that silently fails otherwise). It does not describe the return shape in depth, but the behavioral profile is well covered.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the purpose, then the when-to-use rule, then sibling disambiguation, then the cost note. Dense but every sentence earns its place; slightly long for a zero-parameter tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and no parameters, the description carries the full burden and does so: it says what the response contains (each key with meaning and default, valid schedules, path glob behavior). An agent knows both what the tool is for and what it will get back.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so the baseline is 4. The description does not need to explain argument semantics, and the empty schema is consistent with a pure reference-lookup tool.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb and resource: lists the project settings an assistant can set, plus valid schedules and path glob semantics. It explicitly separates itself from the sibling get_project ('describes settings in general; get_project shows the values one project has'), so an agent can route without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
States exactly when to read it: before create_project, update_project, or a `config` argument, and specifically whenever the request names a section, schedule, format, limit or behaviour. It also gives the reason (keys are exact and guessed names are refused), which is actionable gating guidance with a named alternative.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
extract_urlExtract one page in fullAInspect
Fetch one URL exactly as a project with the given settings would, and return the whole page: the bodies config's formats ask for plus the raw html (each capped at 12,000 characters), head and extracted fields, its images, and which engine read it -- without its link list. Use it for a single page that needs browser steps, structured fields or a settings trial before create_project; for plain content of one or many known URLs use scrape_urls, and get_page reads a page a project already stored without fetching. It climbs the fetch ladder -- plain http first, a browser only when the page needs one or render_js is 'always' -- and costs the credits of the rung that read it (1 to 4 for most pages, +1 when formats ask for a screenshot). Browser actions in config (click, type, select, press, wait, scroll; repeat 'until_gone' for Load-more buttons; each to act on every match) run before the page is read.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | The absolute http(s) URL of the page to read. | |
| config | No | Optional fetch settings: any subset of the keys describe_project_config lists, e.g. {"render_js": "always", "only_main_content": true}. Omit for the defaults. | |
| parse_documents | No | Read PDFs, Word files and spreadsheets as text (default true); false leaves their text out. A parse_documents key in config wins. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only cover safety flags (readOnly/openWorld/idempotent/destructive), so the description carries real weight and delivers: the fetch ladder (plain http first, browser only when needed or render_js is 'always'), credit cost 1-4 plus +1 for screenshots, the 12,000-char caps, and that browser actions run before the read. What it omits is failure/timeout behavior and auth requirements, but coverage is well above the annotation floor.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the return payload before the routing guidance, and every clause carries information (cap sizes, cost, ladder). The actions enumeration is dense but relevant; it is long without being padded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description compensates by describing the return shape in detail, and it supplies the cost, fallback, and browser-action behavior an agent needs to call correctly against a large sibling set.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds genuine semantic detail: it explains render_js 'always', the actions vocabulary (click, type, select, press, wait, scroll, 'until_gone', each), and that a parse_documents key in config overrides the top-level flag.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Fetch one URL ... return the whole page') and enumerates exactly what comes back (config formats, raw html capped at 12,000 chars, head, extracted fields, images, engine; no link list). It distinguishes itself from scrape_urls and get_page by name.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit when-to-use ('a single page that needs browser steps, structured fields or a settings trial before create_project') and routes two alternatives with their selecting conditions: scrape_urls for plain content of known URLs, get_page for an already-stored page without fetching.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_changesGet what changed in a runARead-onlyIdempotentInspect
Report what changed in a project run against the run before it -- the last finished run unless run_id is given: pages added, modified and removed, head-field changes (title, description, canonical...), and the run's coverage, with up to 200 entries per list plus the runs a record exists for. Use it to answer 'what changed on the site'; list_pages shows every page whatever changed, and get_page shows one page's text and versions. Removals are withheld when the crawl reached under 90% of the site, so a blocked crawl never reports the site as gone. A project's first run is a baseline with nothing to compare. Free and reads only what is stored.
| Name | Required | Description | Default |
|---|---|---|---|
| run_id | No | A run's id, from start_run or the runId a list_pages answer carries; leave empty for the last finished run. | |
| project_id | Yes | The project's id, as list_projects or create_project returns it. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnly, idempotent, non-destructive), yet the description adds substantial behavior beyond them: removals are suppressed when crawl coverage drops under 90% so a blocked crawl is never reported as site loss, the first run is a baseline with nothing to compare, and results are capped at 200 entries per list. These are non-obvious edge-case semantics an agent must know.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Dense but front-loaded: the core operation leads, then the returned lists, then routing to siblings, then edge-case caveats. Every clause carries information (limits, suppression rule, baseline rule, cost) with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description must characterize the response, and it does: entry types, per-list cap of 200, coverage, and the run-provenance caveat. Combined with the baseline and suppression rules, an agent has everything needed to call and interpret it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both parameters are already documented, including run_id's null default and provenance. The description restates the run_id fallback ('last finished run unless run_id is given'), which adds emphasis but little new syntax or format information. Baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('report what changed in a project run') and enumerates exactly what is reported: pages added/modified/removed, head-field changes, and run coverage. It distinguishes itself from list_pages and get_page by name, so an agent can pick it without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives the direct use case ('answer what changed on the site') and explicitly routes to alternatives with their conditions: list_pages shows every page regardless of change, get_page shows one page's text and versions. That is when-to-use plus when-not-to-use in one sentence.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_jobFollow a long-running jobARead-onlyIdempotentInspect
Check on a job a long tool handed back -- a crawl from crawl_site, a run from start_run or recrawl_pages, or a batch from scrape_urls -- and, once it has finished, return the same result the original tool would have given; while it is still going, its status and counts. Call it when a tool answered with status 'running' and a job; calling the original tool again is not needed and would not start a second job. For a crawl it is also how to read further: url returns one page in full (even mid-crawl) and cursor the next window of pages. For a project's stored pages use get_page instead. Free: it only reads stored results and never fetches.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | The job's id, from that 'job' object, or a crawl_site answer's crawl_id. | |
| url | No | Crawls only: one page's URL from the crawl's index, to read that page in full. | |
| kind | Yes | What the job is: 'crawl' (crawl_site), 'run' (start_run, recrawl_pages) or 'batch' (scrape_urls); the 'job' object a tool returned names it. | |
| cursor | No | Crawls only: the cursor a previous crawl answer returned, to read the next window of pages. | |
| project_id | No | Required for a run: the project it belongs to (the 'job' object carries it). Ignored for a crawl or a batch. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, non-destructive and closed-world, but the description adds real behavioral context: it is free, reads only stored results, never fetches, and returns the original tool's result once finished versus status/counts while running. It does not cover auth needs or any rate/expiry limits, so it stops short of a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core purpose and trigger, with the crawl-pagination and get_page routing details following. It is dense and runs long for one paragraph, but every sentence carries information an agent needs; no filler is present.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the return-value burden and does so by explaining the two states (running: status and counts; finished: same result as the originating tool). That is nearly complete, though the literal shape of a finished result is conveyed by analogy to the original tool rather than specified.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% so the baseline is 3, but the description adds semantics the schema does not: `url` reads one page in full even mid-crawl, and `cursor` reads the next window of pages, tying both to the crawl reading flow. `kind` and `project_id` semantics are largely already in the schema, so this is an increment rather than a full rewrite.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb and resource ('check on a job', 'return the same result the original tool would have given') and explicitly maps each job kind to the sibling that produced it (crawl_site, start_run, recrawl_pages, scrape_urls). An agent can distinguish it from every sibling without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It states the trigger condition precisely ('Call it when a tool answered with status "running" and a job'), rules out the wrong action ('calling the original tool again is not needed and would not start a second job'), and routes elsewhere when appropriate ('For a project's stored pages use get_page instead'). This is explicit when-to-use, when-not, and alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_pageGet one stored pageARead-onlyIdempotentInspect
Read one page a project run stored, in full: its markdown and other bodies (each capped at 12,000 characters), head fields (title, description, canonical, h1), extracted fields, and its versions across runs. Use it when you know the page's URL; list_pages or search_pages finds the URL first, and for a page from a crawl_site crawl use get_job with url instead. Free and reads only what is stored, so it never fetches the site -- use extract_url for a live copy. A URL the run did not store answers 404.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | The page's full URL as list_pages or search_pages shows it; the other scheme or trailing slash is matched too. | |
| run_id | No | A run's id, from start_run or a list_pages answer, to read that run's copy; leave empty for the newest copy across the project's recent runs. | |
| project_id | Yes | The project's id, as list_projects or create_project returns it. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnly, idempotent, non-destructive, closed-world), yet the description adds behavior not derivable from them: body truncation at 12,000 characters, per-run vs newest-copy selection, 404 semantics for unstored URLs, and the guarantee that it never hits the network. That is meaningful operational detail for an agent deciding how to source a page.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with what the tool returns before routing guidance, and every clause carries information (scope, alternatives, cost, failure mode). The opening sentence is dense with an inline parenthetical and a list, which slightly taxes readability, but there is no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description carries the return-value burden, and it does so by enumerating the returned sections and their truncation behavior. Combined with the routing rules and 404 case, an agent has everything needed to call this correctly without opening a sibling.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents url, run_id, and project_id with their formats and sourcing. The description reinforces the distinction between a run-specific copy and 'the newest copy across the project's recent runs' via 'its versions across runs', but adds no syntax or format detail the schema lacks.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Read one page a project run stored, in full') and immediately enumerates what the payload contains: markdown bodies (capped at 12,000 chars), head fields, extracted fields, and versions across runs. It distinguishes itself from siblings by naming search_pages/list_pages, get_job, and extract_url with the exact conditions that select each.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit when-to-use ('Use it when you know the page's URL'), explicit prerequisites (list_pages or search_pages finds the URL first), and explicit alternate routes for crawl_site crawls (get_job with url) and live copies (extract_url). It also states the failure case: a URL the run did not store answers 404.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_projectGet one projectARead-onlyIdempotentInspect
Read one project in full: its settings (config), schedule, retention, coverage and last run. Read it before update_project, to see the values you are about to change; list_projects is the lighter way to find projects, and list_pages or get_changes read what its runs found. Free, fetches nothing; an unknown or hidden project answers 404. The webhook signing secret is withheld from the answer.
| Name | Required | Description | Default |
|---|---|---|---|
| project_id | Yes | The project's id, as list_projects or create_project returns it. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly/idempotent/non-destructive, but the description goes further with operational traits: it is free, fetches nothing (no upstream crawl triggered), unknown or hidden projects return 404, and the webhook signing secret is withheld. These are exactly the behaviors an agent cannot infer from structured fields.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core purpose, then routed alternatives, then behavioral caveats, then the withheld-secret note. Dense but every clause carries distinct information and none is redundant with the schema or annotations.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description compensates by enumerating the returned fields (config, schedule, retention, coverage, last run) and disclosing the one field withheld. Error behavior (404) is also covered, so an agent has everything needed to call and interpret it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Only one parameter and schema description coverage is 100%, with the schema already explaining that project_id comes from list_projects or create_project. The description adds no format, syntax, or lookup guidance beyond the schema, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Read one project in full') plus the exact scope of what is returned: settings, schedule, retention, coverage, last run. It also names the siblings it differs from (list_projects, get_changes, list_pages), so an agent can disambiguate without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit when-to-use: read it before update_project to see the values being changed. Also names the lighter alternative (list_projects) for finding projects and points to list_pages/get_changes for run findings, covering both the positive and negative cases.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
keep_crawl_as_projectKeep a crawl as a projectAInspect
Turn a crawl_site crawl into a project, so the site is watched over time and each later run is compared against this one. Nothing is fetched again and it costs no credits: the crawl's pages become the project's first run. Use it after crawl_site when the site is worth watching; to start a watched site from scratch use create_project. Without it a crawl and its pages expire after a day. It counts toward the plan's project limit (refused with plan_limit when full). Returns the new project, whose id the other project tools take.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | A name for the project; omit to keep the crawl's own name. | |
| crawl_id | Yes | The crawl_id a crawl_site answer (or get_job on a crawl) returned. | |
| schedule | No | How often the project re-crawls on its own; 'manual' (default) runs only when start_run is called. Every run spends credits. | manual |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only cover the safety profile (non-read-only, non-idempotent, non-destructive), while the description adds the traits that actually matter: no refetching, no credit cost, the crawl's pages becoming the first run, the day-long expiry if unused, the plan project-limit interaction with a named plan_limit refusal, and the return value's id being consumed by other project tools.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the core transformation, then layers usage routing, cost/expiry caveats, and the return contract in roughly priority order. It is dense but every clause carries information; only the length keeps it short of a 5.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite having no output schema, the description states what is returned (the new project and its id, used by other project tools) and covers cost, expiry, and plan-limit failure. Nothing an agent needs to call this correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all three parameters (name, crawl_id, schedule) are already documented in the schema, including crawl_id's provenance and the credit cost of scheduled runs. The description adds no parameter-level detail beyond that, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb and resource ('Turn a crawl_site crawl into a project') with the scope of what that produces ('the site is watched over time'). It explicitly differentiates itself from create_project, so an agent can choose between them without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a clear trigger ('Use it after crawl_site when the site is worth watching') and names the alternative for the contrasting case ('to start a watched site from scratch use create_project'). Both when-to-use and when-to-use-something-else are stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_pagesList a run's pagesARead-onlyIdempotentInspect
List the pages one project run stored -- the last finished run unless run_id is given -- with each page's url, status, depth, word count and when it last changed, up to 500 rows, plus the runId. Use it to see what a crawl covered or to pick a URL; get_page reads one of them in full, search_pages finds pages by their text, and get_changes lists only what changed. Free and reads only what is stored: it never fetches the site. A project with no finished run answers with no pages.
| Name | Required | Description | Default |
|---|---|---|---|
| run_id | No | A run's id, from start_run or the runId a list_pages answer carries; leave empty for the last finished run. | |
| project_id | Yes | The project's id, as list_projects or create_project returns it. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint/idempotentHint/openWorldHint=false, so the safety profile is covered. The description adds genuinely new behavioral context: it is 'Free', it 'reads only what is stored: it never fetches the site', and it caps results at 500 rows. It stops short of describing ordering or truncation behavior beyond the row cap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action and scope, and every clause earns its place (fields returned, row cap, alternatives, edge case). The opening sentence is long and packs the returned-field list and default-run rule together, which costs a little readability.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, and the description compensates by enumerating the returned fields (url, status, depth, word count, last changed, runId) and the 500-row cap. Combined with the default-run resolution and the no-pages edge case, an agent has everything needed to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so run_id and project_id are already documented, including the 'leave empty for last finished run' default. The description restates the default resolution but adds no format or reference detail beyond the schema, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource (lists the pages a run stored) and immediately scopes it ('the last finished run unless run_id is given'). It also names the sibling alternatives (get_page, search_pages, get_changes) so the agent can tell it apart without opening other schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit when-to-use ('see what a crawl covered or to pick a URL') and explicit routing to alternatives with the condition that selects each: get_page for full content, search_pages for text lookup, get_changes for diffs only. Also states the empty-project edge case ('no finished run answers with no pages').
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_projectsList all projectsARead-onlyIdempotentInspect
List every project in the workspace -- the sites it watches over time -- with each one's id, name, seed URL, host, schedule, page count, coverage, last run and health. Call it first to find a project_id for the other project tools, or to check whether a site is already watched before create_project; get_project gives one project's full settings. Reads only what is stored: free, and it fetches nothing. A key limited to some projects sees only those.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly/idempotent/non-destructive, so the safety profile is covered. The description adds genuine context beyond them: it is free, fetches nothing (reads only stored data), and is scoped by API key permissions. It stops short of noting pagination or result-size behavior, so it is strong but not exhaustive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, front-loaded with what is listed, then routing guidance, then the read-only/permission caveat. No sentence is filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Zero-parameter read tool with annotations and an output schema; the description still covers routing, cost, data freshness ('reads only what is stored'), and key-scoping. Nothing an agent needs to invoke it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so the baseline is 4. The description's field list describes the response shape rather than inputs, and with an output schema present it adds no parameter meaning to compensate for.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb + resource ('List every project in the workspace') with the returned fields enumerated, and it explicitly distinguishes itself from get_project ('gives one project's full settings'). An agent can separate it from every sibling without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives concrete when-to-use conditions: call first to obtain a project_id for other project tools, or to check whether a site is already watched before create_project. It also names the alternative (get_project) and the condition that selects it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
map_siteMap a site's declared URLsAInspect
List every URL a site declares in its sitemaps -- found through robots.txt and the well-known paths, each index file walked to its children -- without fetching any of the pages. Use it first, to see how big a site is and what sections it has, before crawl_site or create_project; it cannot read page content (scrape_urls or crawl_site do) and misses pages a site links to but does not declare. Costs 1 credit per sitemap file read, usually 1 in total, never per URL. Returns the discovery method, totals, up to limit URLs and the section names.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | Any URL on the site, usually its home page. | |
| limit | No | The most URLs to return, 1 to 5,000 (default 1,000); totals still count them all. | |
| search | No | Keep only URLs containing this text, e.g. '/blog/'; omit for all. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Goes well beyond the annotations: it discloses the cost model (1 credit per sitemap file read, usually 1 total, never per URL), the non-fetching behavior, and the specific limitations (no content, misses undeclared pages). The credit side effect also explains why readOnlyHint is false despite the discovery-only behavior, so there is no contradiction.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core purpose and packed with useful detail in a small footprint; the em-dash-heavy run-on sentence is dense but every clause earns its place. Slightly harder to scan than two clean sentences.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description supplies the return shape (discovery method, totals, up to `limit` URLs, section names) as well as cost and coverage caveats. An agent has everything needed to decide and invoke correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all three parameters are already documented. The description restates `limit` ('up to `limit` URLs') and adds that totals still count all URLs, but does not mention the `search` filter or add syntax beyond the schema, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb+resource ('List every URL a site declares in its sitemaps') with the mechanism named (robots.txt, well-known paths, index files walked to children). It is immediately distinguishable from siblings like crawl_site and scrape_urls, which it explicitly says do something else.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit ordering guidance ('Use it first... before crawl_site or create_project') plus clear exclusions: it cannot read page content (scrape_urls or crawl_site do) and misses pages a site links to but does not declare. Both when-to-use and when-not-to-use are covered with named alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
recrawl_pagesRe-crawl chosen pagesAInspect
Fetch specific pages of a project again now, as a small run of their own whose record compares just those pages with the last full run. Use it to check a few pages you expect changed without crawling the whole site; start_run re-crawls everything, and extract_url reads a page outside any project. Each page costs credits (usually 1 to 4). The URLs must be on the project's site; refused with 409 while another run of the project is queued or running. Returns the queued run; follow it with get_job (kind 'run') and read results with get_changes.
| Name | Required | Description | Default |
|---|---|---|---|
| urls | Yes | 1 to 500 full URLs on the project's site to fetch again. | |
| project_id | Yes | The project's id, as list_projects or create_project returns it. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, openWorldHint=true, non-idempotent, non-destructive. The description adds valuable context beyond annotations: credit cost (1-4 per page), the 409 refusal condition (another run queued/running), URL host constraint, and the follow-up workflow (get_job with kind 'run', then get_changes). It does not explain what happens if some URLs fail or the exact shape of the returned run, but the coverage is strong.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the core action, then scenario, alternatives, cost, constraint, failure mode, and return/workflow. Every sentence earns its place with no redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with 2 well-documented params and annotations covering safety profile, the description fills remaining gaps: cost, failure condition, host constraint, and the follow-up workflow using sibling tools. An agent has everything needed to call this correctly and interpret the result.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the schema already defines both parameters well. The description reinforces the URL constraint ('must be on the project's site') and the credit-per-page cost, adding value beyond the schema's bare '1 to 500 full URLs on the project's site'.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource+scope: 'Fetch specific pages of a project again now, as a small run of their own...compares just those pages with the last full run.' Explicitly distinguishes from siblings start_run ('re-crawls everything') and extract_url ('reads a page outside any project').
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Names the intended scenario ('check a few pages you expect changed without crawling the whole site') and names two alternatives (start_run and extract_url) with the conditions that select them. Nothing is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
scrape_urlsScrape a list of URLsAIdempotentInspect
Fetch known URLs (1 to 500) once and return each page's content, markdown by default; no project is created. For one page with every format or browser steps use extract_url; to find pages by following links use crawl_site; to list a site's URLs without fetching them use map_site. Each page costs the credits of the engine that read it (1 plain fetch, 4 browser render), and a page the site refuses is free. A single URL answers in the same request; a list waits for the batch, and a large one comes back as an index with excerpts inside 60,000 characters. Hosted, a batch still going after the time budget comes back as a job for get_job, and repeating the same call returns that job instead of starting another.
| Name | Required | Description | Default |
|---|---|---|---|
| urls | Yes | The pages to fetch: 1 to 500 absolute http(s) URLs. | |
| config | No | Optional fetch settings: any subset of the keys describe_project_config lists, e.g. {"render_js": "always", "only_main_content": true}. Omit for the defaults. | |
| formats | No | Comma-separated bodies to return per page, from markdown (default), text, cleanHtml and rawHtml, e.g. 'markdown,text'. | markdown |
| parse_documents | No | Read PDFs, Word files and spreadsheets as text (default true); false leaves their text out. A parse_documents key in config wins. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Goes well beyond the annotations by disclosing credit economics (1 plain fetch, 4 browser render), that refused pages are free, that a single URL answers in-request while a list waits for the batch, the 60,000-character excerpt limit, and that a hosted batch still running returns a job for get_job with repeat-call idempotency. This is rich operational context that the annotations alone do not convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose and the sibling routing are front-loaded, and every sentence carries information about cost, batching or job behavior. It is dense and drifts into long run-on sentences at the end, but there is little pure filler to cut.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description still explains the return shape (index with excerpts capped at 60,000 characters) and the asynchronous job path for large hosted batches. Combined with full parameter coverage, an agent has everything needed to call and handle this tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents urls, config, formats and parse_documents in detail. The description repeats the 'markdown by default' default and echoes the batch/excerpt behavior but adds no new syntax or format guidance beyond the schema, making the baseline 3 appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource (fetch a list of known URLs, 1 to 500) plus the key scope fact that no project is created. It explicitly distinguishes itself from three named siblings (extract_url, crawl_site, map_site), so an agent can route correctly without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit when-to-use-which routing: extract_url for one page with every format or browser steps, crawl_site for finding pages via links, map_site for listing URLs without fetching. Nothing about alternative selection is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_pagesSearch a run's stored pagesARead-onlyIdempotentInspect
Find which pages of a project run contain some text or an element -- the last finished run unless run_id is given -- returning the matching URLs with how many pages were scanned. Use it to locate pages by what they say or contain; list_pages lists every page and get_page then reads one in full. It scans the stored copies and never fetches the site, so it is free and as fresh as the run. A selector search needs the run to have kept html (the rawHtml format); when none was kept the answer says so rather than reporting no matches.
| Name | Required | Description | Default |
|---|---|---|---|
| q | Yes | What to look for, 1 to 200 characters. In content mode every word must appear and "quoted phrases" match as written; in selector mode a CSS selector, or an XPath starting with / or (. | |
| mode | No | 'content' (default) searches the extracted markdown; 'selector' searches the stored html. | content |
| run_id | No | A run's id, from start_run or the runId a list_pages answer carries; leave empty for the last finished run. | |
| project_id | Yes | The project's id, as list_projects or create_project returns it. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnly, idempotent, non-destructive, closed-world), and the description adds genuinely new behavior: it scans stored copies rather than fetching the site, so it is free and only as fresh as the run. It also discloses the rawHtml/selector precondition and that a missing-html situation is reported explicitly rather than as zero matches.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences, front-loaded with the core purpose and then the disambiguation and the freshness/precondition caveats. No filler or repetition of what the schema or annotations already declare.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description still characterizes the result (matching URLs plus pages scanned) and the degenerate case where html was not retained. Combined with full parameter coverage and annotations, an agent has everything needed to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds cross-parameter semantics the schema does not: selector mode depends on the run having preserved rawHtml, and run_id defaults to the last finished run. The quoted-phrase and CSS/XPath syntax details remain schema territory.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('find which pages of a project run contain some text or an element') plus the return shape (matching URLs and scan count). It explicitly distinguishes itself from the siblings list_pages and get_page, so an agent can route without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit use case ('locate pages by what they say or contain') and names the two alternatives with their roles: list_pages lists every page, get_page reads one in full. It also states the default run scoping (last finished run unless run_id is given), so the when-to-use decision is fully specified.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
start_runStart a project run nowAInspect
Crawl a project's site now with its saved settings, outside its schedule; the run is then compared with the last one. Use it after create_project or update_project, or when fresh results are wanted; to re-read only a few pages use recrawl_pages, and for a site with no project use crawl_site. Every page read costs credits (usually 1 to 4) and a run stops when the credit budget runs out, keeping what it read; it is refused with 409 while a run of the project is already queued or running, and 402 when no credits are left. Without wait it returns the queued run at once; with wait it returns the finished run, except hosted, where a run still going after the time budget comes back as a job for get_job. Results are read with list_pages and get_changes.
| Name | Required | Description | Default |
|---|---|---|---|
| wait | No | true waits for the run to finish (minutes for a large site); false (default) returns as soon as it is queued. | |
| project_id | Yes | The project's id, as list_projects or create_project returns it. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations cover the safety profile (readOnlyHint=false, openWorldHint=true, non-idempotent, non-destructive), and the description layers on cost (1–4 credits per page), budget-exhaustion behavior (stops and keeps what it read), and the 409/402 refusal conditions. This is far beyond what the annotations alone convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads what the tool does, then usage routing, then cost/error semantics, then return behavior. Every clause carries information; no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description still explains the return shapes (queued run, finished run, or a job for hosted timeouts) and points to list_pages/get_changes for reading results. Nothing needed to call this correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds semantics the schema does not: 'with wait it returns the finished run, except hosted, where a run still going after the time budget comes back as a job for get_job.' That extra hosted-mode nuance materially changes how the wait parameter behaves.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Crawl a project's site now with its saved settings, outside its schedule') and immediately distinguishes itself from recrawl_pages and crawl_site by name. An agent can identify the exact operation without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit triggers ('after create_project or update_project, or when fresh results are wanted') and names the alternatives with the conditions that select them: recrawl_pages for a few pages, crawl_site for sites with no project. Nothing is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
update_projectUpdate a project's settingsADestructiveIdempotentInspect
Change an existing project's name, schedule or settings. Only what you pass changes: in config only the keys given are replaced, the rest stay; a replaced value is not kept. Call get_project first to see the current values, and describe_project_config for valid keys; to start a different site use create_project instead. Nothing is fetched and no credits are spent now; the next run uses the new settings, and a settings change that alters how pages are read starts a fresh comparison baseline. Returns the updated project.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | A new name; omit to leave it. | |
| config | No | Settings to change, e.g. {"max_pages": 200}; keys not given keep their values. Omit to leave them all. | |
| schedule | No | A new schedule; omit to leave it. | |
| project_id | Yes | The project's id, as list_projects or create_project returns it. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Goes well beyond the annotations (destructiveHint=true, idempotentHint=true): it discloses merge semantics ('only the keys given are replaced, the rest stay; a replaced value is not kept'), cost behavior ('no credits are spent now'), deferred effect ('the next run uses the new settings'), and a side effect ('starts a fresh comparison baseline'). These are exactly the traits an agent needs and the annotations do not convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the core action and then layers constraints in a single dense paragraph; every clause earns its place. It is slightly packed with stacked semicolons, which costs a little readability but not substance.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema, but the description states 'Returns the updated project.' Combined with the merge/cost/baseline behavior, everything needed to call this mutation tool correctly is present.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3. The description still adds the overwrite nuance ('a replaced value is not kept') and points to describe_project_config for valid config keys, which the schema does not say. That is real added meaning, though most per-parameter detail is already in the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Change) and resource (existing project's name, schedule or settings), and explicitly scopes it against siblings: 'to start a different site use create_project instead' and 'Call get_project first'. An agent can distinguish this from create_project and get_project without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit prerequisites ('Call get_project first to see the current values, and describe_project_config for valid keys') and names the alternative with its selecting condition ('to start a different site use create_project instead'). When-to-use and when-not-to-use are both covered.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
15 tool updates
- Changed
crawl_site6 fields changed- added
Input schema / properties / config / descriptionAdded value: +"Optional fetch settings: any subset of the keys describe_project_config lists, e.g. {\"render_js\": \"always\", \"only_main_content\": true}. Omit for the defaults." - added
Input schema / properties / exclude_paths / descriptionAdded value: +"Globs over the URL path to skip, e.g. ['/tag/*', '*.pdf']; an exclude wins over an include." - added
Input schema / properties / include_paths / descriptionAdded value: +"Globs over the URL path to keep, e.g. ['/blog/*'] for a section (its index included); a bare '/blog/' matches only that one page. Omit for the whole site." - added
Input schema / properties / limit / descriptionAdded value: +"The most pages to read, 1 to 5,000 (default 50); the plan's page cap also applies." - added
Input schema / properties / max_depth / descriptionAdded value: +"How many links deep to follow from the start URL (0 = that page only; default 3)." - added
Input schema / properties / url / descriptionAdded value: +"The absolute http(s) URL to start from, e.g. the home page."
- Changed
create_project5 fields changed- added
Input schema / properties / config / descriptionAdded value: +"Any subset of the settings describe_project_config lists. For one section of a site pass include_paths, e.g. {\"include_paths\": [\"/blog/*\"]}; map_site shows the site's sections first. Omit for the defaults." - added
Input schema / properties / name / descriptionAdded value: +"A name for the project; omit to use the site's host name." - added
Input schema / properties / schedule / descriptionAdded value: +"How often it re-crawls on its own; 'manual' (default) runs only when start_run is called." - added
Input schema / properties / schedule / enumAdded value: +[ + "manual", + "hourly", + "daily", + "weekly" +] - added
Input schema / properties / seed / descriptionAdded value: +"Where the crawl starts: a public site's URL or bare domain, e.g. 'https://example.com/blog/' or 'example.com'. The project covers that host."
- Changed
extract_url3 fields changed- added
Input schema / properties / config / descriptionAdded value: +"Optional fetch settings: any subset of the keys describe_project_config lists, e.g. {\"render_js\": \"always\", \"only_main_content\": true}. Omit for the defaults." - added
Input schema / properties / parse_documents / descriptionAdded value: +"Read PDFs, Word files and spreadsheets as text (default true); false leaves their text out. A parse_documents key in config wins." - added
Input schema / properties / url / descriptionAdded value: +"The absolute http(s) URL of the page to read."
- Changed
get_changes2 fields changed- added
Input schema / properties / project_id / descriptionAdded value: +"The project's id, as list_projects or create_project returns it." - added
Input schema / properties / run_id / descriptionAdded value: +"A run's id, from start_run or the runId a list_pages answer carries; leave empty for the last finished run."
- Changed
get_job6 fields changed- added
Input schema / properties / cursor / descriptionAdded value: +"Crawls only: the cursor a previous crawl answer returned, to read the next window of pages." - added
Input schema / properties / id / descriptionAdded value: +"The job's id, from that 'job' object, or a crawl_site answer's crawl_id." - added
Input schema / properties / kind / descriptionAdded value: +"What the job is: 'crawl' (crawl_site), 'run' (start_run, recrawl_pages) or 'batch' (scrape_urls); the 'job' object a tool returned names it." - added
Input schema / properties / kind / enumAdded value: +[ + "crawl", + "run", + "batch" +] - added
Input schema / properties / project_id / descriptionAdded value: +"Required for a run: the project it belongs to (the 'job' object carries it). Ignored for a crawl or a batch." - added
Input schema / properties / url / descriptionAdded value: +"Crawls only: one page's URL from the crawl's index, to read that page in full."
- Changed
get_page3 fields changed- added
Input schema / properties / project_id / descriptionAdded value: +"The project's id, as list_projects or create_project returns it." - added
Input schema / properties / run_id / descriptionAdded value: +"A run's id, from start_run or a list_pages answer, to read that run's copy; leave empty for the newest copy across the project's recent runs." - added
Input schema / properties / url / descriptionAdded value: +"The page's full URL as list_pages or search_pages shows it; the other scheme or trailing slash is matched too."
- Changed
get_project1 field changed- added
Input schema / properties / project_id / descriptionAdded value: +"The project's id, as list_projects or create_project returns it."
- Changed
keep_crawl_as_project4 fields changed- added
Input schema / properties / crawl_id / descriptionAdded value: +"The crawl_id a crawl_site answer (or get_job on a crawl) returned." - added
Input schema / properties / name / descriptionAdded value: +"A name for the project; omit to keep the crawl's own name." - added
Input schema / properties / schedule / descriptionAdded value: +"How often the project re-crawls on its own; 'manual' (default) runs only when start_run is called. Every run spends credits." - added
Input schema / properties / schedule / enumAdded value: +[ + "manual", + "hourly", + "daily", + "weekly" +]
- Changed
list_pages2 fields changed- added
Input schema / properties / project_id / descriptionAdded value: +"The project's id, as list_projects or create_project returns it." - added
Input schema / properties / run_id / descriptionAdded value: +"A run's id, from start_run or the runId a list_pages answer carries; leave empty for the last finished run."
- Changed
map_site3 fields changed- added
Input schema / properties / limit / descriptionAdded value: +"The most URLs to return, 1 to 5,000 (default 1,000); totals still count them all." - added
Input schema / properties / search / descriptionAdded value: +"Keep only URLs containing this text, e.g. '/blog/'; omit for all." - added
Input schema / properties / url / descriptionAdded value: +"Any URL on the site, usually its home page."
- Changed
recrawl_pages2 fields changed- added
Input schema / properties / project_id / descriptionAdded value: +"The project's id, as list_projects or create_project returns it." - added
Input schema / properties / urls / descriptionAdded value: +"1 to 500 full URLs on the project's site to fetch again."
- Changed
scrape_urls4 fields changed- added
Input schema / properties / config / descriptionAdded value: +"Optional fetch settings: any subset of the keys describe_project_config lists, e.g. {\"render_js\": \"always\", \"only_main_content\": true}. Omit for the defaults." - added
Input schema / properties / formats / descriptionAdded value: +"Comma-separated bodies to return per page, from markdown (default), text, cleanHtml and rawHtml, e.g. 'markdown,text'." - added
Input schema / properties / parse_documents / descriptionAdded value: +"Read PDFs, Word files and spreadsheets as text (default true); false leaves their text out. A parse_documents key in config wins." - added
Input schema / properties / urls / descriptionAdded value: +"The pages to fetch: 1 to 500 absolute http(s) URLs."
- Changed
search_pages5 fields changed- added
Input schema / properties / mode / descriptionAdded value: +"'content' (default) searches the extracted markdown; 'selector' searches the stored html." - added
Input schema / properties / mode / enumAdded value: +[ + "content", + "selector" +] - added
Input schema / properties / project_id / descriptionAdded value: +"The project's id, as list_projects or create_project returns it." - added
Input schema / properties / q / descriptionAdded value: +"What to look for, 1 to 200 characters. In content mode every word must appear and \"quoted phrases\" match as written; in selector mode a CSS selector, or an XPath starting with / or (." - added
Input schema / properties / run_id / descriptionAdded value: +"A run's id, from start_run or the runId a list_pages answer carries; leave empty for the last finished run."
- Changed
start_run2 fields changed- added
Input schema / properties / project_id / descriptionAdded value: +"The project's id, as list_projects or create_project returns it." - added
Input schema / properties / wait / descriptionAdded value: +"true waits for the run to finish (minutes for a large site); false (default) returns as soon as it is queued."
- Changed
update_project5 fields changed- added
Input schema / properties / config / descriptionAdded value: +"Settings to change, e.g. {\"max_pages\": 200}; keys not given keep their values. Omit to leave them all." - added
Input schema / properties / name / descriptionAdded value: +"A new name; omit to leave it." - added
Input schema / properties / project_id / descriptionAdded value: +"The project's id, as list_projects or create_project returns it." - changed
Input schema / properties / schedule / anyOfPrevious value: -[ - { - "type": "string" - }, - { - "type": "null" - } -]New value: +[ + { + "enum": [ + "manual", + "hourly", + "daily", + "weekly" + ], + "type": "string" + }, + { + "type": "null" + } +] - added
Input schema / properties / schedule / descriptionAdded value: +"A new schedule; omit to leave it."
17 tool updates
- First observed
crawl_site - First observed
create_project - First observed
describe_project_config - First observed
extract_url - First observed
get_changes - First observed
get_job - First observed
get_page - First observed
get_project - First observed
keep_crawl_as_project - First observed
list_pages - First observed
list_projects - First observed
map_site - First observed
recrawl_pages - First observed
scrape_urls - First observed
search_pages - First observed
start_run - First observed
update_project
Publisher details
- Operator
- MeshArc · Publisher source
- Operator website
- https://mesharc.dev
- Vendor relationship
- First-party · Publisher source
- Documentation
- https://mesharc.dev/mcp
- Trust center
- https://mesharc.dev/legal/security
- Restrictions
- A MeshArc account is needed; fetching pages uses credits (free plan available). · Publisher source
Related MCP Connectors
Web scraping to clean Markdown with JS rendering, multi-page crawl, structured extract, sitemaps.
Extract public webpages as Markdown, map sites, and run bounded crawl jobs. API key required.
Turn any URL into clean Markdown and structured data. Scrape, crawl, search and extract.
Scrape any URL into clean LLM-ready Markdown, or crawl a whole site.
Related MCP Servers
- AlicenseAqualityAmaintenanceAI-native web scraper: scrape, crawl and map any site to clean markdown over stdio. MIT-licensed.6149MIT
- AlicenseNot gradedqualityBmaintenanceProvides lightweight web crawling and scraping capabilities, enabling users to fetch pages as clean Markdown, recursively crawl domains, extract metadata, and perform search-and-crawl operations.1 npmMIT
- AlicenseAqualityBmaintenanceCrawl any website into clean Markdown, search through pages, read full content, and extract structured data using OpenAI, Claude, Gemini, or Grok — with auto-citation and resume support.55MIT
- FlicenseNot gradedqualityCmaintenanceEnables scraping single pages or crawling entire websites, converting content to markdown and optionally extracting structured data with Claude.-
Glama MCP Gateway
Add one secure layer between your agents and this server.