Skip to main content
Glama

writ-mcp connects a stdio MCP client — Claude Code, Claude Desktop, Cursor, Windsurf, Codex — to Writ, so the browser workflows you already recorded become tools your assistant can call: run them, read the data they collected, search past results, schedule them, expose them as REST endpoints, and kick off site crawls.

It is a transparent stdio↔HTTP proxy and nothing else. Every tool, schema and rule lives server-side, so what you get always matches what your instance can do — and a new tool never requires upgrading this package. Zero dependencies, one file, Node core only.

Clients that speak Streamable HTTP natively don't need this at all — point them straight at the /mcp endpoint with an Authorization: Bearer <key> header.

Quick start

1. Get an API key. In the Writ app: Settings → Developers → API keys. Keys look like wt_….

2. Add the server.

Against a self-hosted coordinator:

claude mcp add writ-selfhost -e WRIT_API_KEY=<YOUR_API_KEY> -- npx -y writ-mcp --url https://writ.example.com

Against a coordinator on this machine:

claude mcp add writ-selfhost -e WRIT_API_KEY=<YOUR_API_KEY> -- npx -y writ-mcp --url http://localhost:8000

Against a published per-workflow endpoint (an "Expose as MCP" slug URL is used verbatim — no path rewriting):

claude mcp add my-tools -e WRIT_API_KEY=<YOUR_API_KEY> -- npx -y writ-mcp --url https://mcp.example.com/mcp/my-tools

3. Check it.

claude mcp list

Writ Cloud is the default target when you pass no --url (https://api.usewrit.app). Pass --url to point the connector at your own coordinator instead. See Status.

Pass the key through the environment, not the command line. --api-key works, but it puts your key in the process's argument list where any local process can read it via ps, and your shell records it in history. The connector prints a note when you use it.

Your running coordinator also hands out these one-liners, pre-filled, on its Connect page and at GET /api/mcp/connect-info:

Claude Desktop / Cursor (config file)

Add to claude_desktop_config.json (Claude Desktop) or ~/.cursor/mcp.json (Cursor). The env form is recommended — it keeps the key out of the argument list:

{
  "mcpServers": {
    "writ-selfhost": {
      "command": "npx",
      "args": ["-y", "writ-mcp"],
      "env": {
        "WRIT_COORDINATOR_URL": "https://writ.example.com",
        "WRIT_API_KEY": "<YOUR_API_KEY>"
      }
    }
  }
}

Without npm

A self-host install bundles this connector at connectors/writ-mcp. There is no build step, so you can run it straight from disk:

node /path/to/writ/connectors/writ-mcp/index.js --url https://writ.example.com

It coexists with the other Writ servers. Each surface registers under its own slug on purpose — the desktop app is writ, Writ Cloud is writ-cloud, a self-hosted coordinator is writ-selfhost — and each identifies itself to the assistant with a distinct title. Keep any combination connected at once.

Related MCP server: Weavy MCP Server

Tools you get

Served by the coordinator, not by this package:

Tool

What it does

writ_list_workflows

Your saved workflows — plus a run_<name> tool per workflow

writ_run_workflow

Run one and wait for the extracted data

writ_workflow_data

Read a workflow's accumulated data table

writ_search_data

Search across everything already collected

writ_export_data

Export a workflow's data as CSV/JSON

writ_workflow_runs

Run history and status

writ_set_schedule

Schedule a workflow (interval / daily / weekly)

writ_expose_workflow_api

Publish a workflow as a callable REST endpoint

writ_crawl_site / writ_crawl_status

Start and poll a distributed site crawl

writ_create_automation

Event → run-workflow / notify chains

writ_create_monitor / writ_wire_monitor

Watch a page and react to changes

Every target additionally exposes a build family — writ_browser_use, writ_record_website, writ_build, writ_website_to_api, then writ_browser_act / _context / _network / _save / _cancel. Your assistant opens a real browser, drives it turn by turn, and saves the session as a reusable workflow that afterwards replays with no model in the loop. Writ Cloud runs it on Writ's fleet; a self-hosted coordinator runs it on your own fleet agent. Either way your assistant is the brain — no second model key is involved.

Reusing a recent result (max_age)

Running a workflow drives a real browser, so asking the same question twice in one session costs two full runs and two waits. Every workflow tool takes an optional max_age (seconds) meaning a recent answer is good enough:

{ "name": "run_price_check", "arguments": { "sku": "B0C123", "max_age": 300 } }
  • omitted or 0 — always run fresh (the default; nothing goes stale on its own).

  • N — reuse a result younger than N seconds, otherwise run.

A reused answer carries _cache: {hit: true, age_seconds: N} so the assistant can tell how current it is.

max_age works identically against Writ Cloud and a self-hosted coordinator, and on a workflow's own generated tool as well as writ_run_workflow.

Calling a saved crawl (writ_run_saved_crawl)

A whole-site crawl is slow and metered, so re-crawling to answer the same question is the most expensive mistake an assistant can make. A saved crawl is a stored crawl configuration with a stable name, and the same max_age contract applies to it:

{ "name": "writ_run_saved_crawl", "arguments": { "crawl": "docs", "max_age": 86400 } }
  • hit — the pages that crawl already collected come back inline, instantly, with nothing crawled and nothing metered.

  • miss — the site is crawled again with the saved settings, and you get a crawl id to poll with writ_crawl_status (a crawl outlives a single tool call).

Three tools cover the surface:

Tool

What it does

writ_saved_crawls

List saved crawls. Check here before crawling a site again.

writ_run_saved_crawl

Run one, reusing recent data when max_age allows.

writ_saved_crawl_data

Read what one already collected, at any age. Never crawls.

To create one, pass save_as to writ_crawl_site — that saves the settings and runs them, so the crawl becomes callable by REST as well. Re-using the same save_as updates that saved crawl instead of piling up duplicates. Saving needs an admin-scoped credential (it creates reusable, callable configuration); running one needs only run.

If a tool call times out

It comes back as status: "running" with retryable: true. The run was not cancelled — calling the tool again starts a second run. Wait, then retry with a max_age wide enough to pick up the first run's result once it lands.

Configuration

Flags take precedence over environment variables.

Flag

Env

Default

Purpose

--url

WRIT_COORDINATOR_URL / WRIT_URL

https://api.usewrit.app

Target base URL. A URL whose path is already /mcp or /mcp/<slug> is used verbatim.

--api-key

WRIT_API_KEY

— (required)

Key for Authorization: Bearer. Prefer the env var.

--insecure

WRIT_INSECURE_TLS=1

off

Accept a self-signed local-CA cert. Private networks only.

--timeout

WRIT_MCP_TIMEOUT_MS

600000

Per-request timeout (ms); covers long writ_run_workflow waits.

--help / --version

—

—

Print usage or version and exit.

HTTPS with a local CA (self-host): trust the CA (recommended) and use the https:// address, or set NODE_EXTRA_CA_CERTS=/path/to/ca.pem. Use --insecure only for localhost testing — it disables certificate verification entirely, which exposes your key to a man-in-the-middle.

Security

This process exists to carry a credential, so everything that could expose one is made loud rather than convenient. Full detail in SECURITY.md.

  • The key goes to --url and nowhere else. No telemetry, no analytics, no update check. Zero dependencies means there is no transitive code in the process that could phone home — and CI fails the build if that ever changes.

  • Redirects are never followed. Replaying your Authorization header to whatever origin a Location header names would hand your key to a host you didn't choose. A 3xx becomes an error telling you to point --url at the final URL.

  • Credentials in a URL are redacted from every diagnostic. MCP clients write a server's stderr to a log file on disk; https://user:pass@host/ would otherwise be written down in plaintext.

  • Loud warnings, on stderr, for every way a key leaks: TLS verification disabled, plaintext http:// to a non-loopback host, or a key passed on the command line.

  • Retries never double-run a workflow. Only read-only methods are retried; tools/call is sent exactly once, because a retry could re-execute a side effect the connector cannot see.

  • Responses are bounded at 32 MB, so a broken endpoint can't grow this process until the OS kills your session.

  • No request is ever left unanswered — a hung MCP client is a denial of service on your assistant, and that is the failure this connector works hardest to make impossible.

Scope your keys. Give a key only workflows:read / workflows:execute unless a tool you actually use needs more.

Verifying what you install

Releases are published from CI with npm provenance, so the tarball is cryptographically linked to the commit and workflow that built it:

npm audit signatures

Troubleshooting

Symptom

Cause and fix

Unauthorized: … rejected the API key

The key is wrong, disabled, or lacks scope. Recreate it under Settings → Developers with workflows:read / workflows:execute.

Cannot reach …

Wrong --url, or the target is down. For a self-signed cert see the local-CA note above.

… redirected (HTTP 301)

Your reverse proxy redirects (usually http → https). Point --url at the final URL.

The API key contains characters that cannot be sent in an HTTP header

A newline or control character got into the key — usually a copy-paste artifact. Re-copy it.

No tools listed

You have no saved workflows yet, or the key can't read them. The static writ_* tools appear regardless.

Client won't connect, no error

Read the connector's stderr — your client logs it. Claude Code: ~/Library/Caches/claude-cli-nodejs/<project>/mcp-logs-<name>/.

Status

Self-hosted coordinator

Supported and verified end to end.

Published /mcp/<slug> endpoints

Supported.

Writ Cloud (https://api.usewrit.app, the no---url default)

Supported. Used automatically when no --url is passed.

Development

npm test

No dev dependencies — the suite uses Node's built-in node:test and drives the real index.js as a subprocess against a mock MCP server, exercising the same stdio path an MCP client uses. It runs in about six seconds. npm publish runs it automatically via prepublishOnly.

See CONTRIBUTING.md — note the two hard rules: zero dependencies, permanently, and no tool logic here.

The rest of Writ

usewrit/writ

The self-host coordinator — web UI, API, your data. Start here.

usewrit/writ-agent

The Rust fleet worker that does the actual browsing.

writ-mcp (this repo)

The MCP connector.

License

MIT — see LICENSE.

This package is deliberately permissive because it runs inside your MCP client, not inside the coordinator, so it has to be embeddable anywhere. The coordinator it talks to is AGPL-3.0-only; the two licenses are not interchangeable.

Available Tools

39 tools
writ_browser_actAct in a browser sessionA
Destructive
Inspect

Run one batch of actions on an open browser session and get the fresh page back. YOU are the brain: every navigation, click, fill, sign-in, capture, probe and script is yours to decide, one batch at a time. No writ_browser_compose in this client? This tool composes too: actions=[{action:'define_function', name:'feed.list', from_index:3, ...}] or [{action:'compose', operation, payload}], never mixed with clicks in one batch. After a navigate, a click or a select that changes the page, END the batch and look at the new page before acting on it. RECORDING RULES: (1) a caller INPUT — write the value as {{name}} in the action (select/fill/type_text value, navigate url) and pass the real value in inputs ({"name": "real value"}); the page gets the real value, the recorded step keeps {{name}}, and name becomes a workflow input by itself. (2) DATA — record an extract {variable,script} at the position that shows it (a read-only JS IIFE returning rows/fields); its result comes back in this answer, so check it before saving. (3) a SECRET — fill with data_key, never a literal. Interactions (navigate/click/fill/select/press_key/...) are recorded as steps; SEE/HEAR/NETWORK probes and wait never are — replay waits for each step's selector by itself, and a wait the task truly needs is an explicit wait_for step (writ_browser_compose add_steps). ACTIONS: DRIVE: navigate {url} · click {selector | field_index | button_index} · fill {selector,value,data_key?} · type_text {selector,value} · select {selector,value} · check {selector} · hover {selector} · submit {selector} · press_key {key} · scroll {direction,amount} · back · wait {seconds} · wait_for {selector,timeout}. SEE (granular first): query_dom {selector,limit,offset,attrs?,text_chars?,html_chars?} (every match as compact records with a css path to target next) · count {selector} · find_text {text,selector?,exact?,limit?} (the deepest elements showing that text, with paths) · get_attributes {selector,index?} (one element: all attrs, value, box, options) · read_text {selector,all?,limit?,max_chars?} · inspect {selector,limit?,max_chars?} (match count + outerHTML) · list_candidates (the page's repeating row shapes — start here for any list/table) · list_frames · get_dom {selector?,depth?,max_chars?} (the real cleaned HTML — the expensive last resort) · get_screenshot {x?,y?,width?,height?}. TABS / FILES / 2FA: list_tabs / switch_tab {index} · upload {selector,mode,file_slot} · wait_for_download {trigger_selector,output_key} · twofa {challenge_method,selector?,submit_selector?} (the persona's one-time code, minted server-side — see 2FA RULES). HEAR: get_console {level?,since?,query?,limit?} (console messages, uncaught JS errors with stack, failed/blocked requests since your last read — the page's console_since_last_read counts tell you when it is worth a call; read it BEFORE guessing why a sign-in, click or extraction did nothing) · page_errors (only the uncaught exceptions). NETWORK: capture_network {reload?} (the backend calls the page makes — how you find the site's real API; then search/read them with writ_browser_network) · get_request {url substring} (one call in full). RUN CODE: evaluate_js {script} (any JS on the live page, returns JSON — your main probing tool; read-only). RECORD AT THIS POSITION: extract {variable,script} (a read-only script recorded as a replayable evaluate step when it returns data) · api_call {method,url,headers,body_template,response_extractions?,variable} for one request, or api_call {flow:{version:1,steps:[...]},inputs:{...},variable} for a multi-request bootstrap/pagination/transform program (both execute NOW inside the session with its cookies; the flow uses the same interpreter as browserless replay and returns a bounded result sample) · login_post {method,url,headers,body_template} (replay a sign-in as one request) · probe_write {selector} (learn a create/update/delete request WITHOUT sending it) · confirm_write {selector} (perform it ONCE for its real confirmation — changes real data, only when authorized). A sensitive fill MUST carry data_key so the value is held server-side and the saved step keeps a {{secret:...}} placeholder. ASK the user for any credential, 2FA code or decision — never invent one. 2FA RULES: the one-time code is minted server-side from the attached persona and never shown to you. BEFORE twofa, READ the challenge and IDENTIFY the method the page is using — a phone number / 'text message' = sms, an email address = email, 'authentication app' = authenticator, approve-on-phone / passkey / QR / WhatsApp = other — and pass it as challenge_method. A persona receives exactly ONE method (see twofa_method in writ_personas) and a site picks its own default: Facebook texts an SMS even when the account has email. If the page's method is not the persona's, do not emit twofa: click the page's 'Try another way' / 'Use another method' / 'More options' / 'Didn't get a code?' control (in the page's own language), choose the persona's method, confirm, THEN twofa. On twofa_method_required, twofa_method_mismatch or twofa_verify_method do exactly what the message says — verify the method on the page and switch or resend — and do NOT ask the user yet. Only twofa_mint_failed / twofa_no_persona mean: call writ_browser_ask_user kind='twofa' so the Writ user supplies it.

ParametersJSON Schema
NameRequiredDescriptionDefault
inputsNoValues held server-side for {{placeholder}} substitution, e.g. {"city":"Paris"}. Secrets belong here or on a fill's data_key — never hardcoded into a step.
actionsYesOrdered action objects, e.g. [{"action":"click","selector":"#login"}].
max_charsNoClip of the returned page_dom / page_text (default 40000, ≤200000). A get_dom probe on a real app is 500KB — prefer evaluate_js / inspect / read_text to target what you need.
session_idYesSession id from the start tool.

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations flag destructiveHint=true/openWorldHint=true, and the description adds rich context beyond them: confirm_write 'changes real data, only when authorized', sensitive fills must carry data_key, secrets are held server-side, the 2FA code is minted server-side and never shown, and which actions are/aren't recorded as replay steps.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core purpose, and given the tool's complexity most sentences carry real information. However it is a very long, dense wall of text with heavy inline action enumeration that could be tightened; structure is functional but not economical.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a high-complexity, open-world, destructive browser automation tool with no output schema, the description covers recording semantics, probe vs drive actions, error recovery, tabs/files/2FA, and credential handling — essentially everything an agent needs to invoke it safely.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description materially enriches the opaque `actions` array by illustrating action shapes, the {{placeholder}}/inputs contract, and data_key secret handling — more than the schema's terse 'ordered action objects' conveys.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource+scope: 'Run one batch of actions on an open browser session and get the fresh page back.' It immediately frames the agent's role ('YOU are the brain') and distinguishes batch interaction from compose/network siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit when-to-use rules abound: end the batch after page-changing actions, use compose actions when writ_browser_compose is absent, choose granular SEE probes before get_dom, and a clear decision tree for twofa errors and when to call writ_browser_ask_user instead of guessing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

writ_browser_ask_userAsk the user to step inAInspect

Ask the WRIT USER (the person who owns this Writ account) to step in on an open browser session — complete a security check (CAPTCHA, 'confirm it's you', Arkose/hCaptcha/reCAPTCHA puzzle) in the live browser, supply a one-time 2FA code, or answer a question you cannot decide yourself. Use it when writ_browser_act returns security_check with auto_solved false, when a twofa action fails with twofa_mint_failed or twofa_no_persona (kind='twofa'), or whenever only a human can proceed. A one-time code is NOT a first-resort ask: when twofa answers twofa_method_mismatch or twofa_verify_method, first verify on the page which method it is using and switch it to the persona's (or resend) as the message says; interrupt the user only once that has failed. Never try to click through a CAPTCHA yourself, and never ask for a one-time code through kind='question' or type one into the page: with kind='twofa' the user pastes the code in the Writ app and Writ enters it server-side, so it never reaches you. The session pauses (the user is notified in the app and by email and controls the live page); this call holds up to 60s and returns status 'answered' (with solved / answer, or entered for twofa), 'waiting_for_user' (call again with the same session_id to keep waiting — do not act meanwhile), or 'expired'.

ParametersJSON Schema
NameRequiredDescriptionDefault
kindNocaptcha: the user completes a check in the live browser. question: the user answers in text. twofa: the user supplies the one-time code the page is asking for; Writ types it server-side and returns `entered`, never the code.
questionNoWhat you need, in one short sentence (required for kind 'question').
session_idYes
wait_secondsNoHold up to this long (1-60, default 60).

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Even with annotations present, the description adds rich behavior the annotations cannot convey: the session pauses, the user is notified in-app and by email and controls the live page, the call holds up to 60s, the three possible statuses (answered / waiting_for_user / expired), and that with kind='twofa' the code is entered server-side and never reaches the agent. This is well beyond the readOnly/idempotent hints.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose, triggers, and prohibitions are front-loaded in the first two sentences, and each subsequent clause carries actionable information. The description is dense and long (roughly a paragraph and a half), which is justified by the branching logic, though it could be trimmed slightly without loss.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description supplies the return contract itself (the three statuses and their payloads, including 'entered' for twofa) and the timing/polling behavior. Combined with the error-driven usage triggers, an agent has everything needed to invoke and re-invoke this tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 75%, so the schema already documents most parameters (including the kind enum and wait_seconds bounds). The description nonetheless adds real meaning: it warns not to use kind='question' to obtain a one-time code, and explains that kind='twofa' returns 'entered' rather than the code. session_id re-use on 'waiting_for_user' is also clarified.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb and resource ('Ask the WRIT USER ... to step in on an open browser session') and enumerates the concrete cases it covers: CAPTCHA/security check, one-time 2FA code, or an undecidable question. It clearly separates this human-in-the-loop tool from the autonomous siblings (writ_browser_act), so an agent can route without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit triggers are given (writ_browser_act returns security_check with auto_solved false; twofa_mint_failed / twofa_no_persona) plus an explicit when-not: a 2FA code 'is NOT a first-resort ask' until the agent has tried switching/resending the method. Alternatives (verify on the page and switch the twofa method) are named, and the prohibition on clicking CAPTCHAs yourself is stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

writ_browser_cancelClose a browser sessionA
DestructiveIdempotent
Inspect

Close an open browser session. Call this when the task is done and the user does not want to reuse it, or when abandoning a session — an open cloud browser keeps consuming execution time until it is closed. Work you never saved is NOT lost: a session with defined functions, or a recording that did more than visit pages, is auto-saved as a workflow first (the reply names it; inactive draft when it would not replay). Pass discard=true to close without keeping anything. A session bound to a build is settled by that auto-save, or marked cancelled.

ParametersJSON Schema
NameRequiredDescriptionDefault
discardNotrue = throw the session's unsaved steps and functions away instead of auto-saving them. Default false.
session_idYes

TDQS

A4.3/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already flag destructiveHint=true and idempotentHint=true, and the description goes well beyond them: unsaved work is auto-saved as a workflow first, the reply names it, it may be an inactive draft if it would not replay, and build-bound sessions are settled or marked cancelled. This is exactly the mutation-side effect an agent could not infer from structured fields.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three dense sentences, front-loaded with the action and reason, and each clause adds real information (auto-save, discard, build-bound handling). Slightly compressed in the final sentence, but nothing is padding.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, yet the description still explains what the reply contains (the workflow name) and how the session is finalized. For a destructive, non-idempotent-looking close operation with no annotation-level detail on data retention, this is complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 50%: discard is documented in the schema and the description reinforces it meaningfully ('Pass discard=true to close without keeping anything', default false). session_id is undocumented in both, but its meaning is unambiguous. The description adds the consequence of discard that the schema only hints at.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource — 'Close an open browser session' — with clear scope (session is not being reused, or is being abandoned). An agent can distinguish it from session-management siblings, though no sibling is named to draw an explicit contrast.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives concrete when-to-use conditions: task finished and no reuse intended, or abandoning a session, with the rationale that an open cloud browser keeps consuming execution time. It does not name an alternative tool for cases where the user does want to keep going, so context is clear but exclusion logic is left implicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

writ_browser_composeCompose a workflow in a sessionA
Destructive
Inspect

AUTHOR the workflow being built in an open browser session — the power to turn what you drove into a real, complex, callable workflow rather than a replay of clicks. YOU decide its shape. WHEN YOU NEED IT: not for a plain recording — there a caller input is a {{name}} value + inputs on writ_browser_act and the data is an extract action. Use this to expose NAMED FUNCTIONS (an API), to give an input a description/default, or to add a step the recorder cannot see. OPERATIONS: define_function {name, fn_type api|list|script|extraction, ...} · compile_function {name, from_index} — DETERMINISTIC (no-LLM) capture->function: traces session tokens to an is_auth bootstrap, generates per-call ids ({{uuid()}}), and marks a write so the BUILD never sends it (a real run does) · test_function {name, sample_inputs} · remove_function {name} · set_inputs {inputs:{name:{default?,description?,required?,example?}}} · add_steps {steps:[{type,...}], at?} · remove_step {id} · list (the draft: steps, data_steps, functions, inputs). FASTEST PATHS: (a) a list / table / search-results page → writ_browser_context section=lists returns a live-tested define_function payload; pass it here with then_save:{name} — ONE call defines, tests and saves. (b) a site endpoint → capture_network, find the call with writ_browser_network, then define_function {name:'quotes.list', from_index:, request:{url:'https://site/api/quotes?page={{page}}'}, input_variables:[{name:'page',example:'1'}], response_extractions:{quotes:{from:'json',path:'quotes'}, has_next:{from:'json',path:'has_next'}}, then_save:{name:'...'}}. Every function is LIVE-TESTED as you define it; a failed test keeps NOTHING — fix it and define it again with the SAME name. Saved functions are called with writ_run_workflow function_name. DETAILS:

  • define_function: a NAMED callable the saved workflow exposes. fn_type api (backed by one of the site's endpoints — pass from_index=<a captured call's index from writ_browser_network> and Writ seeds method/url/headers/body from the capture; override request fields to parameterize them with {{name}} placeholders; secrets as {{secret:name}}, anti-CSRF echoes as {{cookie:NAME}}), script (a read-only JS IIFE returning the data from the page), list (PREFERRED for any list/table: row_selector + fields {name: sub-selector | {selector, attr}}; the JS is generated for you, and writ_browser_context section=lists hands you this payload ready-made), or extraction (one selector's text). A list/script/extraction function reads the page it was defined on: pass page_url as a template (https://site/search?q={{query}}) or give each input_variable an example and the URL is templated from it. Add then_save:true (or {name, description}) to SAVE the workflow the moment the function passes its live test: one call instead of compose then save. Name it . (orders.list, orders.create) so functions group by surface. Declare input_variables=[{name,description,required,example}], output_fields, and response_extractions for the fields callers get back. Supported specs: JSON {from:'json',path:'data.items'}, embedded JSON {from:'embedded_json',kind:'array',has:['id']}, server HTML {from:'html_css',selector:'.row',attribute:'data-id',all:true} — with fields it returns ROW OBJECTS, which is how a server-rendered list becomes a BROWSERLESS function: {from:'html_css',selector:'tr.athing',all:true,base_url:'',fields:{title:{selector:'.titleline > a'},url:{selector:'.titleline > a',attribute:'href'}}} on an api function that GETs the page (no browser at replay — prefer this over a list/script function whenever the rows are in the served HTML), regex {from:'regex',pattern:'...',group:1}, header {from:'header',name:'x-next'}, body {from:'body'}, or legacy '$.json.path'. Default to an Auphan-style named graph: ordered is_auth functions publish dynamic token/id/origin values consumed as {{extracted:name}}, while each data function remains independently callable. Pass flow={version:1,steps:[...]} only for request loops, recursive mapping, cross-page dedupe, cursor pagination or a composite return. The flow is schema-validated and live-tested immediately in the current browser session with its cookies, persona and egress, using the same interpreter as the saved HTTP lane. The later saved run remains the final engine=http parity proof. is_auth=true marks the sign-in function: it runs first on every replay and its response_extractions publish values the others consume as {{extracted:}}. The function is LIVE-TESTED the moment you define it (an in-session request, or a DOM read) with sample_inputs={name: value}; a failed test returns feedback and keeps NOTHING — fix it and define it again with the SAME name. test=false skips the proof (a real run proves it later).

  • test_function {name, sample_inputs}: prove a defined function again.

  • remove_function {name}.

  • add_steps {steps:[...], at?}: explicit replayable steps the DOM recorder cannot see — navigate, click, fill, select, press, wait, wait_for, extract, evaluate, api_call, login_post, return, upload, wait_for_download — inserted at a position (default: append).

  • remove_step {id}.

  • set_inputs {inputs:{name:{default?,description?,required?,example?}}}: the parameters a caller passes at run time; every {{name}} in a step or function must be a declared input, a credential, a {{cookie:}}/{{extracted:}} runtime reference, or produced by an earlier step, or the save is refused. Credentials are never inputs — they come from the persona or a data_key fill.

  • list: the draft so far (steps, functions, inputs, build). Then writ_browser_save: api functions become api_call steps (auth first), the workflow becomes api_recorded when every step is a call, and each function is callable by name (writ_run_workflow function_name) and documented at GET /api/v1/workflows/{id}/api-docs.

ParametersJSON Schema
NameRequiredDescriptionDefault
payloadNoThe operation's arguments. define_function: {name, fn_type?, description?, surface?, from_index?, request?{method,url,headers,body_template}, flow?{version,steps}, script?, selector?, input_variables?, output_fields?, response_extractions?, is_auth?, order?, sample_inputs?, test?}. compile_function {name, from_index, sibling_index?, input_variables?, response_extractions?, page_url?}: the DETERMINISTIC (no-LLM) way to turn a captured authenticated request — a GraphQL/RPC POST, a form submit — into a callable function. Writ decodes the body, keeps the static parameters, TRACES each session-minted token (csrf/xsrf/dtsg/lsd/etc.) to where a fresh session re-reads it and emits an is_auth bootstrap that publishes it as {{extracted:}}/{{cookie:}}, replaces per-call client values (idempotence token, session id, timestamp) with runtime GENERATORS ({{uuid()}}, {{uuid(session)}}, {{timestamp_ms()}}, {{counter()}}), templates your caller inputs, and CLASSIFIES a write (create/post/send/delete): the BUILD never sends it, a real run does (pass mutation_mode='dry_run' to preview). Prefer this over a hand-built api function for any authenticated mutation or token-bound endpoint; pass sibling_index=<another capture of the same request> so the pagination inputs are told apart from the minted noise. test_function: {name, sample_inputs?}. remove_function: {name}. add_steps: {steps:[{type, config|flat fields, description?}], at?}. remove_step: {id}. set_inputs: {inputs:{name:{default?, description?, required?, example?}}}.
operationYes
session_idYesSession id from the start tool.

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Well beyond the annotations (destructiveHint/openWorldHint/idempotentHint), the description discloses that every function is live-tested on definition, that a failed test 'keeps NOTHING', that the build never sends writes while a real run does, that a save is refused if a {{name}} is undeclared, and that credentials are never inputs. These are exactly the behavioral traits an agent needs and are not inferable from the hints.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with purpose, then WHEN YOU NEED IT, OPERATIONS, FASTEST PATHS and DETAILS, which is strong structure for an eight-operation tool. It is nonetheless extremely long, and the 'live-tested / failed test keeps NOTHING / redefine with the SAME name' point is repeated in both FASTEST PATHS and DETAILS, costing some conciseness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description still covers return shape (functions callable by name, api-docs endpoint), save behavior, draft inspection via `list`, and per-operation argument semantics for all eight operations. Given the complexity and nested payload, almost nothing an agent needs to call it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 67% and the payload schema is already richly documented, so the baseline is high. The description nonetheless adds real semantic depth the schema lacks: {{name}} placeholder semantics, {{secret:}}/{{cookie:}}/{{extracted:}} references, response_extractions spec formats, and then_save testing behavior. Some of this overlaps the existing payload description, keeping it just shy of a 5.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The opening line states a specific verb and resource ('AUTHOR the workflow being built in an open browser session') and immediately frames the tool's role versus a plain recording. It explicitly names sibling tools (writ_browser_act, writ_browser_context, writ_browser_network, writ_browser_save) and distinguishes their roles, so an agent can separate this from neighbors without opening schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The 'WHEN YOU NEED IT' block gives explicit when-to-use (expose named functions, add descriptions/defaults, add steps the recorder cannot see) and when-not ('not for a plain recording'), plus the alternative pattern for that case. FASTEST PATHS routes to specific siblings with concrete conditions, covering both selection and exclusion.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

writ_browser_contextRead a browser sessionA
Read-onlyIdempotent
Inspect

Read context for an open browser session. section=page (default) re-reads the LIVE page — url, form fields, buttons, links, and the cleaned DOM. section=map reads the BUILD this session is bound to: the endpoints, specs and candidate functions Writ's crawl rungs found (evidence to verify), plus what you have composed so far. For ONE list/search API: open the guided session ON the results URL, read section=lists, pass its define_function to writ_browser_compose, save. section=lists SCANS the live page for you at no AI cost: the repeating rows, their field selectors, which captured request carries them, and a live-tested define_function payload to hand to writ_browser_compose. Read it BEFORE probing a results page by hand. section=explorer pages through Writ's full recording policy; section=concierge_api pages through the API-builder policy. The policy sections are reference for unusual flows (logins, multi-request chains), not a prerequisite.

ParametersJSON Schema
NameRequiredDescriptionDefault
offsetNoPaging offset for the policy sections.
sectionNopage (default) | lists | map | explorer | concierge_api
max_charsNoCharacters per page (1000–10000, default 8000).
session_idYes

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already establish readOnly, non-openWorld, idempotent, and non-destructive behavior, so the safety profile is covered. The description adds that section=lists scans at no AI cost and is live-tested, and that map output is evidence to verify — useful context beyond annotations. It does not describe pagination behavior for non-policy sections, output format, or rate limits, leaving some behavioral gaps.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Information-dense but front-loaded with the section definitions before procedures. However, the middle section is a run-on that jumps between the 'for ONE list/search API' workflow and section=lists scanning, making it harder to parse. Could be tighter without losing essential details.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no output schema and a complex tool with five distinct sections, the description is substantially complete: it covers all section modes, their purposes, and the integration path with writ_browser_compose. It lacks detail on offset/max_chars paging behavior and exact output shape, but for a multi-modal read tool it is mostly sufficient.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 75%, so most parameters are already documented in the schema (section enum, offset, max_chars). The description adds meaning to section values (what page vs map vs lists vs explorer vs concierge_api returns) that goes beyond the schema's terse enum listing. Baseline 3 is appropriate given the schema does the heavy lifting and the description adds some section semantics.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Read) and resource (browser session context), then enumerates the sections and their distinct outputs — live page state, build map, list scanning, and policy reference. This is far more precise than the sibling read tools like Writ_browser_network or writ_browser_sessions, which read different facets of a session.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly tells when to use each section: section=lists is for scanning results pages and must be read BEFORE probing by hand; section=map is for the build; policy sections are for unusual flows and not a prerequisite. Directs the agent to pass define_function to writ_browser_compose, naming a sibling and the handoff.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

writ_browser_networkRead a session's network callsA
Read-onlyIdempotent
Inspect

Search or read the requests the live page has made — how you find a site's real backend API instead of scraping its HTML. operation=search lists matching calls (filter with query / method); operation=detail returns one call in full by index (method, url, request headers, request body, response body). Indices are stable for the session, so one from an earlier search still resolves later — only the oldest calls age out of the retained window, and asking for one of those says so rather than returning a different call. Nothing captured yet? Run the capture_network action with writ_browser_act first — it reloads the page with capture armed; to catch a POST (login/search/submit), perform the action that triggers it, then capture. A call you want as a callable function goes straight into writ_browser_compose define_function via from_index=. Held credential values are replaced with their placeholder in the output.

ParametersJSON Schema
NameRequiredDescriptionDefault
indexNoWhich call to read, for operation=detail.
queryNoSubstring filter across method, url, status, and bodies.
methodNoFilter by HTTP method.
offsetNo
max_charsNoWindow size, 1000-10000 (default 8000; larger values are clamped). Page with offset.
operationNosearch (default) | detail. list/get are aliases.
session_idYes

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly/idempotent/non-destructive, and the description adds real behavioral context beyond them: indices are stable for the session, only oldest calls age out, and a stale index errors rather than silently returning a different call. It also discloses that held credential values are replaced with placeholders in output, which is important for interpretation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Information-dense and front-loaded, with the core purpose stated first and operational detail following. It is on the long side and slightly paragraph-heavy, but nearly every sentence carries load; a minor trim would make it tighter.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 7-parameter, two-operation tool with no output schema, the description covers both operations' return shapes (detail returns method, url, request/response bodies), the staleness edge case, the capture prerequisite, and credential redaction. Nothing an agent needs to call it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 71% schema coverage the schema already documents most parameters, but the description adds meaning: operation semantics, that query/method filter the search results, and that index refers to the stable call index from a prior search. It does not mention the offset/max_chars pagination behavior of the search window, leaving a small gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource (search/read the network requests the live page made) and frames the payoff against scraping HTML, which clearly separates it from siblings like writ_scrape or writ_crawl_site. It also names the two operations the tool supports, so an agent knows the tool's surface without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit when-to-use workflow: if nothing is captured, run capture_network via writ_browser_act first, and to catch a POST, trigger the action then capture. It also routes output to writ_browser_compose's define_function via from_index, so the agent knows where this tool sits in the broader flow.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

writ_browser_saveSave a session as a workflowA
Destructive
Inspect

Save the open browser session as a clean, replayable workflow and close the browser. Everything you composed (writ_browser_compose) is materialized: named api functions become api_call steps with the auth function first, declared inputs become the workflow's parameters, explicit steps land at their position. The saved workflow is ACTIVE immediately: it runs on demand with writ_run_workflow (function_name calls one function) or its own run_ tool (writ_pin_workflow_tool) at zero AI cost, can be scheduled with writ_set_schedule, exposed as a REST endpoint with writ_expose_workflow_api, and documented at GET /api/v1/workflows/{id}/api-docs. A session bound to a build (writ_website_to_api guided) settles that build. The save is REFUSED, with the reasons, when a step or function references something no run could resolve — fix it in the session and save again. Only save once the task actually worked on the live page — verify first.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameNoShort workflow name (defaults to the goal).
keep_openNoLeave the browser open after saving (default false — saving closes it).
session_idYes
descriptionNo
allow_no_dataNoSave a workflow that yields NO data (navigation/actions only) on purpose. Off by default: a save with no data step and no defined function is REFUSED for an API build (define_function or an extract first) and WARNED for a task recording — a workflow like that runs green and returns nothing.

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well beyond the annotations (destructiveHint=true, openWorldHint=true) by disclosing that saving closes the browser, that the saved workflow is ACTIVE at zero AI cost, and that the save is REFUSED with reasons when a step cannot resolve at run time. The allow_no_data refusal/warning semantics add real behavioral detail. It does not explicitly restate the destructive/irreversibility angle beyond 'closes the browser,' so it lands short of a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core action and outcome in the first sentence, with consequences and refusal conditions following in a logical order. It is dense with parenthetical tool cross-references, which is useful routing information but pushes the prose near the limit of what one tool description should carry.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description still covers the mutation's side effects, activation state, refusal behavior, and follow-on tooling — enough for an agent to call it correctly. The one gap is the return value: it never says what a successful save yields (e.g. a workflow id) or what the refusal response looks like.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 60%, and the description compensates on the ambiguous parameters: it clarifies keep_open means the browser stays open (default closes it), that name defaults to the goal, and it expands the allow_no_data refusal-vs-warning semantics well past the schema text. session_id and description remain undocumented in both places, which caps it below 5.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Specific verb (save) plus resource (open browser session) plus the exact transformation: composed session materializes into a replayable workflow and the browser closes. It distinguishes itself from siblings by naming writ_browser_compose as the producer and writ_run_workflow / writ_pin_workflow_tool as the consumers, so an agent can place it in the lifecycle without reading other schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit gate: 'Only save once the task actually worked on the live page — verify first,' plus refusal conditions tied to unresolvable references. It also routes to the correct downstream tools for running, scheduling, exposing, and documenting the saved workflow. What is missing is a direct when-not-to-save rule versus sibling tools like writ_export_data or writ_record_website.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

writ_browser_sessionsList open browser sessionsA
Read-onlyIdempotent
Inspect

List the cloud browser sessions this account has open, so you can RESUME one instead of opening a second browser beside it. A session you already opened is warm, parked on its current page, and keeps billing while it stays open — so when you need a browser, check here first and continue an open one by passing its session_id to writ_browser_act / writ_browser_context, rather than calling writ_browser_use again. Returns each session's id, status, resumable flag, current url and goal.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMax sessions to return (default 20).
include_closedNoAlso list recently-closed sessions (not resumable) for reference. Default false — only open, resumable sessions.

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover read-only/idempotent/non-destructive safety, so the bar is lower; the description adds real value by disclosing that sessions stay warm, parked on a page, and keep billing while open, plus the resumable flag concept. It stops short of describing pagination or ordering of results, which keeps it from a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the purpose and the resume rationale before naming alternatives; every clause carries instruction. Slightly dense with the billing aside, but nothing is filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, so the description usefully enumerates the returned fields (id, status, resumable flag, current url, goal), and it supplies the workflow context needed to act on the result. Nothing an agent needs to call or use the tool is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and both parameters (limit, include_closed) are documented there, so the schema does the heavy lifting. The description never references limit or include_closed, adding nothing beyond the schema — the baseline 3 when coverage is high.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource (list open cloud browser sessions) with account scope, and explicitly differentiates itself from siblings by naming writ_browser_use, writ_browser_act and writ_browser_context. An agent can select it vs. those tools without reading any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit when-to-use ('when you need a browser, check here first'), the when-not ('rather than calling writ_browser_use again'), and the follow-on action (pass session_id to writ_browser_act / writ_browser_context). The alternative path is fully spelled out.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

writ_browser_useOpen a browserA
Destructive
Inspect

A REAL CLOUD BROWSER FOR A TASK ON A WEBSITE: the user's own signed-in account (email, social, shop, bank or work portal — pass a persona_id from writ_personas; the password stays sealed in Writ, so never ask for a password), a click, form, submit, search inside an app, setting change or buy/book/post, or a page a plain fetch cannot open (login wall, 403, CAPTCHA). OPENS a real cloud browser and returns the FIRST live page observation. It is NOT an autonomous agent — YOU are the brain and the driver (Writ runs no model here): do every step with writ_browser_act(session_id), turn by turn, until the task is done. Call writ_browser_use ONCE per task; never again for the same task, and never wait for it to finish anything. RECORDING IS ALWAYS ON, SAVING IS ON DEMAND: every interaction you drive is recorded. If the user wants to REUSE the task ("record it", "so I can re-run it"), drive it the recordable way from the first action — a value they will want to change goes in as {{name}} with the real value in writ_browser_act inputs, the data goes out through an extract action — then writ_browser_save(name): its answer lists the steps, inputs and a run_example, and the workflow replays at zero AI cost (writ_run_workflow). Otherwise just finish the task and writ_browser_cancel — an open browser bills until it is closed. WHAT YOU CAN DO IN IT: navigate, click, fill, type, select, press keys, scroll, switch tabs, upload a file, sign in (a persona's 2FA code is minted server-side); SEE the page (read_text, get_dom for the real HTML, inspect a selector, list_candidates for repeating rows, get_screenshot); read EVERY backend call the page makes (capture_network, then search/read them with writ_browser_network); run ANY JavaScript on the live page (evaluate_js) and call the site's backend from inside the session with its cookies (api_call); work a page whose content only appears after interaction. NOT FOR: just READING a page, a few pages, or the top N items of a listing — that is writ_scrape (one call, 2-10s, no browser); collecting a site into a dataset (writ_crawl_site); turning a site into an API (writ_website_to_api, which opens the same browser bound to a build). A browser costs execution time for as long as it is open, so open one only when the task needs interaction.

The page comes back after every batch and on demand via writ_browser_context(section=page). FOLLOW THE USER'S DIRECTIONS and ASK the user directly in chat whenever you need a decision, a value to type, a credential, or a 2FA/OTP code — never guess or invent secrets (for a sensitive fill set data_key so the saved step keeps a placeholder, never the raw value). writ_browser_compose adds what the recorder cannot see (named functions, explicit steps, an input's description/default). Prefer replaying an existing saved workflow (writ_list_workflows -> writ_run_workflow) when one already does the task. BLOCKED BY A BOT WALL / CAPTCHA? A browser's exit IP is fixed once it opens, so cancel it (writ_browser_cancel) and reopen with use_residential=true. And before opening a browser at all, call writ_browser_sessions: an open one is warm and cheaper to continue than a new one.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYesStarting URL to open (required).
goalNoOptional label describing the task, for the run log and your own reference. Passing it does NOT make the tool carry out the task — you still drive every step yourself with writ_browser_act after this returns.
persona_idNoSaved identity to sign in with (list them with writ_personas). Required for sites behind a login with 2FA. A desktop persona ('device:…') opens the browser on its own desktop, which fills {{secret:username}} / {{secret:password}} / 2FA itself — you never see its values.
use_residentialNoOpen on the platform residential network (premium) for a site that blocks datacenter IPs or shows a bot wall / captcha. Default off (free); turn on when a plain open is blocked.
execution_targetNoWhere the session runs: 'cloud' (managed cloud fleet), 'auto' (prefer the user's OWN linked Writ desktop app when it is online, else cloud), or 'local' (REQUIRE their own desktop app). Running on the user's own machine keeps the session on their computer + IP and exposes nothing local to the cloud. Omit to use the account's default set in the Writ app. If a 'local' run reports local_agent_offline, their app is closed — tell them, and only pass execution_target='cloud' to re-run in the cloud if they agree.
residential_countryNoTwo-letter ISO country the residential exit should be in (e.g. 'us', 'fr') — for a site that serves a different page per country, or throttles foreign traffic. Omit for an automatic exit. Ignored unless the session egresses residential.

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare destructive/openWorld/non-idempotent, and the description goes well beyond them: recording is always on, browsers bill until closed, the exit IP is fixed once open (so cancel and reopen with use_residential), and the tool is explicitly non-autonomous (the agent drives every step). This is unusually rich operational disclosure for a mutating, open-world tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core action and the key one-shot-per-task constraint, then organized under clear labels (RECORDING, WHAT YOU CAN DO, NOT FOR). It is very long, but nearly every clause carries operational information the agent must act on, so the length is mostly earned rather than padding.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description carries the return burden and does so: it states it returns the FIRST live page observation and that the page comes back after every batch via writ_browser_context. It also covers the full task lifecycle (use -> act -> save/cancel) and the credential/2FA policy.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is already 100%, giving a baseline of 3, but the description adds real meaning beyond it: persona_id ties a sealed credential and server-minted 2FA, execution_target explains the local_agent_offline recovery path, and use_residential is tied to the cancel-and-reopen bot-wall flow.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (opens a cloud browser) and its scope (real signed-in account, interactive site work) and explicitly contrasts itself against named siblings writ_scrape, writ_crawl_site and writ_website_to_api. An agent can tell exactly what this tool is and is not without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Contains an explicit NOT FOR block routing reading, listing and dataset collection to the correct alternatives, plus positive conditions (login wall, 403, CAPTCHA, form/click/submit). It also gives procedural gates: call writ_browser_sessions first, call this ONCE per task, and cancel to stop billing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

writ_crawl_filesList a crawl's captured filesA
Read-onlyIdempotent
Inspect

The ORIGINAL documents a crawl captured (PDFs, office docs, images, CSVs) as stored files — filename, size, version, source_url, and a short-TTL download_url fetchable with no further auth. The crawl's DATASET holds the extracted text; use this when you want the actual files. Pass crawl_id for one run, or crawl (saved crawl slug/name/id) for its most recent completed run(s).

ParametersJSON Schema
NameRequiredDescriptionDefault
runsNoWith `crawl`: how many recent completed runs to aggregate (default 1 — the current version of every document).
crawlNoSaved crawl slug, name, or id (alternative to crawl_id).
limitNoMax files to return (default 100, cap 200).
crawl_idNoCrawl id from writ_crawl_site.

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover safety (readOnly, idempotent, non-destructive, closed-world), so the bar is lower. The description adds real behavioral value beyond that: the download_url is short-TTL and fetchable with no further auth, and the returned field set is enumerated. It does not mention pagination beyond the schema's limit.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences, front-loaded with the resource and its differentiator, and no filler. The contrast with the DATASET is placed immediately after the identification, which is the right order.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description carries the burden of describing the return shape, and it does so (filename, size, version, source_url, download_url). Combined with 100% schema coverage for inputs and annotations covering safety, an agent has everything needed to call this correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the parameter baseline is 3. The description goes beyond that by clarifying the crawl_id-vs-crawl selection semantics ('for one run, or its most recent completed run(s)'), which the per-parameter schema text does not fully spell out.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific resource ('the ORIGINAL documents a crawl captured') and enumerates the file types and returned fields, so the agent knows exactly what comes back. It also explicitly contrasts itself with the crawl's DATASET (extracted text), which is the nearest confusion point among siblings like writ_saved_crawl_data.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It states the selection condition clearly ('use this when you want the actual files') and names the alternative artifact (the DATASET with extracted text). It also explains the crawl_id vs. crawl routing choice. It stops short of naming a sibling tool by name or listing exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

writ_crawl_siteCrawl a site into a datasetA
Destructive
Inspect

COLLECT A SITE (or a section of it) INTO A DATASET — a distributed Dragnet crawl that discovers pages and stores every one as a queryable, change-tracked row. This is the tool for 'crawl ', 'get every page of the docs', 'all products in this category', 'build a dataset of ', or anything that will be queried, exported, monitored or re-run later. NOT FOR: reading a page or a handful of pages right now — that is writ_scrape (url / urls / top_n answers in one call, no dataset); acting on a page (writ_browser_use).

CHOOSE THE MODE — all three fetch pages the same way; they differ in who READS each page:

  • CLASSIC (default: extract_mode='markdown', executor='regular') — every page becomes clean markdown, no AI spent, fastest. Right for content, docs, articles, discussions (threads keep [top-level]/[reply · depth N] tags). Pick this unless a rule below applies.

  • SCHEMA (extract_mode='schema' + extract_schema) — every page holds the SAME structured record (a product, a listing row) and you want rows, not prose. Deterministic CSS extraction, no AI.

  • AI-ASSISTED (executor='ai' + extract_prompt) — the wanted fields need understanding and vary per page (sentiment, pros/cons, a classification, free-form values with no stable selector) across MANY pages. Each page waits on a model call (~10s) and bills 5x the page rate — never use it for a few pages you could read yourself, and never to 'be sure'.

SCOPE IT — an unscoped crawl of a real site collects hundreds of nav, tag and pagination pages and bills for every one. Match the ask to a shape:

  • A SECTION ('the docs', 'the pricing and blog pages'): pass intent in plain language — the server derives include/exclude paths and depth from the site's real URLs — and relevance_threshold ≈0.3 to drop off-goal pages.

  • KNOWN PAGES as a dataset: seed_urls (no discovery). For an immediate answer use writ_scrape(urls) instead.

  • TOP-N of a listing as a dataset (re-run later, monitored): rank_cap=N. For an immediate answer use writ_scrape(url, top_n) instead.

  • WHOLE SITE ('every page'): the defaults; set page_budget to cap the spend.

DELIVERY: a bounded crawl (rank_cap / seed_urls) waits and returns its pages in data.rows in this call; an open site crawl returns a crawl id to poll with writ_crawl_status — results land as a workflow dataset (writ_workflow_data, writ_search_data, writ_export_data). If the ask mentions comments or discussion, set content_spec {"preset": "full", "include_comments": true} (rank_cap crawls do this already). Behind a login: persona_id — never sign in yourself. save_as ONLY when the user will re-run it; writ_saved_crawls lists those — re-run one only when its scope matches the ask.

YOU OWN THE RESPONSE SHAPE. Every answer is Writ's envelope by default (definition + crawl status + a data table whose rows wrap fields in run bookkeeping, and whose records carry page metadata like content_kind/depth). When you are BUILDING AN API on a crawl — anything a program or the user will consume — set output so the answer is THEIR shape: {shape:'record'} for one entity (a usage meter, a dashboard), {shape:'records'} for a list, fields to pick/rename ('percent_used as pct', dotted paths), exclude to drop, key to wrap. Page metadata is stripped unless include_meta=true. Saved with save_as, it becomes the API's default shape (override per call on writ_run_saved_crawl / writ_saved_crawl_data). Do NOT try to prompt the metadata away in extract_prompt — it is added after the model answers; output is the fix.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYesSeed URL (required).
nameNo
waitNoBlock until the crawl converges and return the collected pages IN THIS CALL (`data.rows`). Default TRUE for a bounded crawl (rank_cap or seed_urls — a few pages, seconds) and false for an open site crawl (returns a crawl id to poll with writ_crawl_status). Past the 75s ceiling you get a 504 that still carries the crawl id.
limitNoRows of collected data to return when wait=true (default 50).
speedNoThroughput tier: slow | normal (default) | fast — how much of your parallel-agent allowance the crawl uses. slow ≈ ¼ at a discounted page rate, normal ≈ ½ at standard rate, fast = all of it at a premium.
deviceNoA linked Writ desktop's agent_id (writ_devices): act ON it. Omit to use the desktop this connection chose with writ_devices action='use' (if any).
intentNoPlain-English goal. The server derives include/exclude paths and a depth from it against a sample of the site's real URLs, and ranks the frontier by relevance — so on an unfamiliar site this beats guessing path regexes yourself.
outputNoRESPONSE SHAPE — set this whenever the answer is for a program or an API you are building, not for you to read. {shape: 'envelope' (default: Writ's full answer, projected) | 'table' ({columns, rows, total}) | 'records' (bare list of records) | 'record' (the newest record alone — one entity, a usage meter, a dashboard), fields: ['used', 'percent_used as pct', 'items.0.price as first_price'] (ordered pick, renames, dotted paths; missing → null so keys are stable), exclude: ['depth'], include_meta: false (page metadata content_kind/depth/thumbnails are STRIPPED unless true), key: 'usage' (wrap)}. On writ_crawl_site with save_as it is SAVED as the API's default shape.
max_ageNoOnly meaningful with `save_as`: if that saved crawl already completed within this many seconds, return its collected data instead of crawling again. 0 always crawls.
save_asNoONLY when the user will want to re-run this crawl later (a recurring pull, an API they asked for): saves these settings as a named, callable crawl. A one-off question does NOT need one — saved crawls are listed to every future session as 'already collected', so a saved one-off misleads the next agent. Reusing the same name updates that saved crawl instead of creating a duplicate.
delay_msNoPoliteness delay between fetches per host (default 250).
executorNoregular (default) = deterministic crawl, no AI. ai = a fleet of AI agents reads every page against `extract_prompt` and returns structured records — for data with no clean CSS selector. Bills at 5x the page rate.
ocr_modeNoauto (default) | off | force
rank_capNoTOP-N ASK — set this whenever the user wants the top/first N items from a listing (a front page, search results, a category). The server reads the seed page's link order (which IS the ranking), seeds exactly those N item pages, and pins the crawl to them — one page per agent, in parallel. WITHOUT it the same request becomes a breadth crawl that mostly collects nav and pagination and does not answer the question. Pair with include_paths when you know the item-link shape.
max_depthNo
seed_urlsNoExact pages to start from, when you already know them — the crawl collects these instead of discovering its own. Cheapest way to scrape a known set.
persona_idNoSaved identity to crawl AS (list them with writ_personas) — for pages behind a login. Every shard shares the persona's signed-in session and one sticky exit IP. 2FA is minted server-side. A desktop persona ('device:…') crawls on its own desktop, from that machine.
shard_sizeNoURLs fetched per shard batch (default 20).
page_budgetNo
render_modeNoHow each page is FETCHED — independent of `executor`, which decides who READS it. auto (default) = plain HTTP first, warm browser only for JS-challenge or near-empty pages; http = never open a browser (fastest, static HTML); browser = warm-render every page (JS/SPA sites). executor=ai works on either lane.
same_domainNo
content_specNoWhich ELEMENTS of each page to keep: {preset: 'full'|'main', include_comments: bool, exclude_selectors: [css], include_selectors: [css], keep: {images: bool}}. 'main' = article body only; 'full' = the whole page INCLUDING comment and discussion threads — use 'full' with include_comments when the ask mentions comments, replies or discussion, or they will be stripped out.
extract_modeNomarkdown (default) | schema (uniform records via extract_schema) | html (each page's RAW HTML, for selectors or embedded JSON)
exclude_pathsNo
include_pathsNo
preview_charsNoCut each inline page's text cells to this many characters (default 12000; 0 = full pages). Cut rows list the fields under `_truncated`; fetch a full page with writ_workflow_data(workflow_id=<data_workflow_id>, refs=['<run_id>:<record_index>']).
extract_promptNoRequired with executor=ai: what each agent should extract from each page, in plain language (e.g. 'the product name, price and SKU').
extract_schemaNo
respect_robotsNoHonor robots.txt (default true).
timeout_secondsNoMax seconds to hold when wait=true (≤75).
use_residentialNoRoute every shard through the platform residential network (premium). Turn on for sites that block datacenter IPs / show a bot wall — a persona crawl forces it on automatically. Costs residential bandwidth; default off.
allow_subdomainsNo
relevance_thresholdNo0-1. Score every discovered page against `intent` and SKIP anything below the bar, so a broad crawl collects only what the goal needs (≈0.3 for 'the pricing and docs pages'). Leave unset for a whole-site sweep.
residential_countryNoTwo-letter ISO country the residential exit should be in (e.g. 'us', 'fr') — pins the exit pool's geo for every shard. Omit for an automatic exit. Ignored unless the session egresses residential.
max_concurrent_shardsNoExplicit parallel-shard cap; overrides the `speed` allocation.

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare a write, non-idempotent, open-world, destructive operation, and the description adds substantial context beyond that: cost multipliers (AI executor 5x page rate), spend-capping via page_budget, wait/504 behavior, dataset persistence, login handling through persona_id, and re-run semantics for saved crawls. No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long but tightly organized with front-loaded caps headings (NOT FOR, CHOOSE THE MODE, SCOPE IT, DELIVERY) that let an agent navigate quickly. Given the tool's 35-parameter complexity the length is largely justified, though some expanded mode explanations could be trimmed.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex tool with no output schema, the description explains the return envelope, data.rows behavior for bounded vs. open crawls, polling with writ_crawl_status, and how to reshape output for API consumers. This covers what an agent needs to call the tool correctly and interpret results.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 35 parameters at ~77% schema coverage, the description adds decision-level meaning for the most consequential knobs: intent, rank_cap, seed_urls, extract_mode/executor interplay, output shape configuration, save_as, and content_spec. This meaning is not derivable from the schema descriptions alone and directly shapes correct invocation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource ('COLLECT A SITE ... INTO A DATASET') with clear scope, and explicitly contrasts itself against writ_scrape and writ_browser_use. An agent can identify this as the crawl-to-dataset tool without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It provides explicit when-to-use guidance (crawl <site>, build a dataset), when-not-to-use (one-off page reads, page actions), and names the alternatives (writ_scrape, writ_browser_use). It further routes between three execution modes and several scoping shapes with concrete conditions for each.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

writ_crawl_statusCheck a crawlA
Read-onlyIdempotent
Inspect

Status of a crawl by id — page counts, status, the dataset workflow id. With wait=true it is ONE held call (≤75s) that returns when the crawl converges, WITH the collected rows inline (data, shaped by output): the answer to a crawl tool's 504 / crawl-id handle. Never poll this in a loop — pass wait=true and, if it is still running at the ceiling, call it once more the same way.

ParametersJSON Schema
NameRequiredDescriptionDefault
waitNoHold until the crawl is terminal and inline its rows (default false).
limitNoRows to inline when it converged (default 50).
outputNoRESPONSE SHAPE — set this whenever the answer is for a program or an API you are building, not for you to read. {shape: 'envelope' (default: Writ's full answer, projected) | 'table' ({columns, rows, total}) | 'records' (bare list of records) | 'record' (the newest record alone — one entity, a usage meter, a dashboard), fields: ['used', 'percent_used as pct', 'items.0.price as first_price'] (ordered pick, renames, dotted paths; missing → null so keys are stable), exclude: ['depth'], include_meta: false (page metadata content_kind/depth/thumbnails are STRIPPED unless true), key: 'usage' (wrap)}. On writ_crawl_site with save_as it is SAVED as the API's default shape.
crawl_idYesCrawl id from writ_crawl_site.
preview_charsNoCut inline text cells to this many chars (default 12000; 0 = full).
timeout_secondsNoCeiling for wait=true (≤75).

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover the safety profile (readOnly, idempotent, non-destructive, closed-world), and the description goes well beyond them: it discloses that wait=true becomes a single held call bounded at ≤75s, that rows are returned inline when it converges, and it explicitly forbids loop-polling. Latency ceiling and retry semantics are exactly the traits annotations cannot express.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core purpose, and the operational rule (never poll, use wait=true) lands early. Dense with em-dashes and parentheticals, and the phrase 'the answer to a crawl tool's 504 / crawl-id handle' takes a second read, but every clause carries information an agent needs.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description does the work of naming what a non-wait status call returns and what an inline `data` payload is shaped by. It stops short of describing pagination or the exact envelope for the plain status case, which is a minor gap for a 6-parameter tool with a nested object.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and the nested `output` object is documented in exhaustive detail in the schema itself, so the baseline is 3. The description adds real meaning on top: it ties `wait` to a held call and `data` shaping, and frames the return fields (counts, status, workflow id) that the schema cannot. It does not touch `limit`, `preview_chars`, or `timeout_seconds`, which is why it stays at 4 rather than 5.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Names a specific verb+resource ('Status of a crawl by id') and enumerates what comes back (page counts, status, dataset workflow id). It also situates itself relative to the crawl-creating sibling by explaining it answers 'a crawl tool's 504 / crawl-id handle', so an agent can place it without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit and prescriptive: use wait=true rather than polling, and if the crawl is still running at the ceiling, re-call once the same way. It also states the condition for setting `output` (when the answer is for a program rather than for the agent to read). No inference required.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

writ_create_automationCreate an automationAInspect

Create an automation: on an EVENT, run a workflow, send a notification, and/or wake an AI agent. Chain workflows (when workflow A completes → run workflow B), alert on completion, or have an agent act on the event (ai_prompt). Give a source workflow via on_workflow for workflow_* events, and at least one of run_workflow / notify / ai_prompt. EMAIL THE DATA, not just that it ran — notify is a template over the event: after a WORKFLOW, {{result.extracted_data.0.title}} / {{result.extracted_data..0.url}} (the run's own rows); after a CRAWL, {{rows.0.}} … {{row_count}} (the records it collected, schema fields included) plus {{seed_host}} {{pages_done}}; after a monitor change, {{extracted.price}}. A missing path renders empty, so a digest of N rows is N numbered lines. THE DIGEST PATTERN: writ_set_schedule on the workflow (or a saved crawl), then this with when=workflow_completed / crawl_completed. ON A CLOCK: when='scheduled' + schedule fires the actions at that time (a notify after run_workflow waits for that run, so {{result.extracted_data...}} is filled). run_functions + inputs call chosen functions of a multi-function workflow (what they need runs too). Anything else (conditions, scrape, extract, branches): a raw blocks tree. ON A WEBHOOK: when='webhook_received' MINTS a signed inbound URL; the answer's webhook holds the url, its signing_secret (shown only then), the two headers every call signs, curl and Python examples and an example_body. Each top-level JSON field of a call becomes the run input of the same name; inputs, notify and ai_prompt templates read any field as {{payload.}}. A call is acknowledged at once; with ?wait=true it is held until the workflow runs the automation starts finish and answers their data. webhook_trigger_id reuses an existing URL (writ_list_webhooks).

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYesName for the automation (required).
whenNoEvent: workflow_completed | workflow_started | crawl_completed | crawl_failed | ai_session_completed | ai_session_started | change_detected (give `target_id`) | webhook_received (mints a signed URL; `webhook_trigger_id` reuses one) | scheduled (a clock: give `schedule`).
titleNoNotification title (with `notify`).
blocksNoRAW flow tree instead of the action arguments (max 50): each {id, type: event|condition|action, blockType, config, parentId}. The FIRST block is the root event (its blockType is the event: one of `when`; scheduled config {mode, interval_ms | time, days, tz}); every other block names an EARLIER block as parentId. Actions: workflow {workflow_id, function_name | function_names, input_mapping}; notification {template, title, channels, recipients}; ai_session {goal, entry_url}; scrape {urls:[<=5, templates ok], format: markdown|html|both, on_error} -> {{scraped.content}} {{scraped.pages}}; extract {source:'{{scraped.content}}', fields:[{key, from: html_css|json|regex|embedded_json|body, selector, attribute, all, path, pattern, group, number, required}]} -> {{extracted.<key>}}; condition {field, operator, value}. Wait for a run: an event block workflow_completed {linked_to_block:<workflow block id>}. Every string setting takes {{placeholders}}.
inputsNoWith run_workflow: its inputs {input_name: value}, each a literal or a {{template}} over the event (e.g. {{extracted.price}}, or {{payload.order.id}} from a webhook call's body); saved values fill the rest. Never secrets.
notifyNoSend a notification with this message — a template over the event's data (see the tool description: {{result.extracted_data.0.title}} after a workflow, {{rows.0.title}} / {{row_count}} after a crawl, {{extracted.price}} on a change).
enabledNo
channelsNoNotification channels for `notify`, e.g. ["pushover","email"] (required for delivery).
priorityNo
scheduleNoWith when='scheduled': {kind:'interval', interval_minutes:N} or {kind:'daily', time:'HH:MM', tz:'<IANA zone>'} or {kind:'weekly', time, days:[1..7] (1=Mon .. 7=Sun), tz}.
ai_promptNoWake an AI agent with this task when the event fires. The agent gets the event context (page URL, diff, extracted values) and works the task in a cloud browser. Supports {{placeholders}}.
target_idNoWith when='change_detected' (required there): the monitor whose changes fire this, i.e. the monitor_id writ_create_monitor returned.
recipientsNoNotification recipients, e.g. ["email:3"]. OMIT to reach EVERY enabled recipient on the channel — the answer names who the alert actually reaches, and warns when nobody is configured.
descriptionNo
on_workflowNoSource workflow name whose event fires this (required for workflow_* events).
ai_entry_urlNoPage the woken agent starts on. Defaults to the event's page (the monitored URL on change_detected); required in practice for workflow_*/webhook events.
run_functionNoOne function name (alias of run_functions).
run_workflowNoWorkflow to RUN when the event fires (by name).
ai_session_idNoWith when='ai_session_completed' / 'ai_session_started': only this AI session.
run_functionsNoWith run_workflow: call these functions of a multi-function workflow instead of the whole workflow — one run of the selection plus what it needs (sign-in, a token, the search whose ids another reads). Omit for the whole workflow.
on_workflow_idNo
run_workflow_idNo
cooldown_minutesNoMinimum minutes between AI wakes for `ai_prompt` (default 10; 0 disables).
target_selector_idNoWith target_id: only changes of this one selector of that monitor.
webhook_trigger_idNoWith when='webhook_received': fire on this EXISTING inbound webhook (its id from writ_list_webhooks) instead of minting a new URL. Its senders keep signing with its secret; one that has no secret yet gets one, shown once in the answer.

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With annotations already declaring readOnlyHint=false, openWorldHint=true, and destructiveHint=false, the description adds rich behavioral context: webhook URLs are minted with a signing secret shown only once, calls are acknowledged immediately unless ?wait=true is used, and notification templates render empty on missing paths. These details go well beyond the safety profile provided by annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with a clear purpose statement and organized into thematic paragraphs (notify, clock, webhook), which helps navigation. However, it is quite lengthy and uses dense all-caps emphasis, making it less concise than ideal, though most sentences carry useful, non-redundant information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (25 parameters, nested objects, no output schema), the description is highly complete: it covers event triggers, action types, template syntax, webhook lifecycle, scheduling, and references to sibling tools. Annotations already cover the safety profile, so the description's depth fills all remaining gaps an agent would need to invoke the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Although schema description coverage is high (80%), the description adds substantial meaning beyond the schema, especially for notify templates (e.g., {{result.extracted_data.0.title}}, {{rows.0.<field>}}, {{extracted.price}}), webhook payload substitution ({{payload.<path>}}), and the blocks tree structure. It also clarifies parameter interactions like on_workflow being required for workflow_* events and run_functions calling selected workflow functions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a precise verb and resource: 'Create an automation: on an EVENT, run a workflow, send a notification, and/or wake an AI agent.' It clearly distinguishes the tool from siblings by naming related tools (writ_set_schedule, writ_list_webhooks) and explaining the unique combination of event triggers, actions, and scheduling.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly covers when to use each trigger mode (workflow_completed, scheduled, webhook_received) and provides concrete patterns like 'THE DIGEST PATTERN' that route the agent to specific sibling tools for complementary tasks. It also specifies required combinations ('at least one of run_workflow / notify / ai_prompt') and references alternatives such as writ_set_schedule and writ_list_webhooks.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

writ_create_http_extractionBuild an HTTP extractionA
Destructive
Inspect

Create or revise an advanced browserless HTTP extraction from plain language and real browser-network evidence. Use AFTER a browser experiment/capture_network when one simple request or an Auphan-style named auth/function graph is insufficient (request loops, GraphQL descriptor discovery, recursive JSON, cross-page dedupe, sorting, cursor pagination). Ordinary login/bootstrap/data chains belong in writ_browser_compose define_function with is_auth/order/typed response_extractions. This tool generates a universal api_call.config.flow; site behavior stays inside the workflow. First call with apply=false to review validation, then apply=true. After applying, run the workflow and prove engine=http with writ_diagnose_http_workflow before exposing it. Never put fetch in evaluate_js and never send Writ-only controls to the site.

ParametersJSON Schema
NameRequiredDescriptionDefault
goalYesExact inputs, output fields, filters, ordering and pagination behavior wanted.
applyNofalse (default) returns a reviewable draft; true writes a valid draft to the workflow.
workflowNoExisting workflow name (or use workflow_id).
step_indexNoapi_call step to replace; defaults to the first, or appends one.
workflow_idNo
requirementsNoExtra mapping, dedupe, filtering or cursor requirements.
desired_inputsNoCaller parameters such as query, min_price, max_price, limit and cursor.
request_samplesNoRelevant calls returned by writ_browser_network/capture_network, including representative response bodies when available.
response_sampleNoOptional representative JSON/HTML response when it is not in request_samples.

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the annotations, the description discloses the draft-versus-write behavior of apply, the generated artifact as universal api_call.config.flow, and the operational requirement to run and prove engine=http with writ_diagnose_http_workflow before exposing the workflow. It also adds safety constraints such as never putting fetch in evaluate_js and never sending Writ-only controls to the site. These details materially help an agent avoid destructive or invalid actions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense and front-loads purpose, usage conditions, workflow steps, and safety constraints. Although long for a tool definition, each sentence carries routing, behavioral, or validation guidance rather than filler. Minor density in the middle section keeps it from a perfect score.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex mutation tool with rich annotations but no output schema, the description covers the full lifecycle: preconditions, draft review, apply behavior, post-apply workflow execution, and diagnostic proof. It does not need to explain return values because no output schema exists, and the annotations already cover open-world and destructive traits. Nothing critical for correct invocation is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 89%, so the input schema already carries most parameter meaning, including apply, goal, workflow, step_index, request_samples, and desired_inputs. The description adds sequencing for apply but does not meaningfully extend the semantics of the remaining parameters beyond what the schema provides. Thus baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: create or revise an advanced browserless HTTP extraction from plain language and browser-network evidence. It clearly distinguishes this from siblings such as writ_browser_compose for ordinary login/bootstrap chains and writ_diagnose_http_workflow for post-apply validation. An agent can identify the tool's scope without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives explicit when-to-use guidance: use AFTER a browser experiment/capture_network when simple requests or Auphan-style auth/function graphs are insufficient, and lists concrete triggers like request loops, GraphQL discovery, recursive JSON, dedupe, sorting, and cursor pagination. It also routes ordinary chains to writ_browser_compose and prescribes apply=false first, then apply=true, followed by workflow proof.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

writ_create_monitorWatch a page for changesAInspect

Create a MONITOR — a target Writ checks on a schedule and that fires a change_detected event when the page, a CSS selector's text, or a visual ZONE of the page changes. Use when the user wants to WATCH a URL for changes/updates. Returns the monitor id for writ_wire_monitor. PROVE THE SELECTOR FIRST: open the page with writ_browser_use and pass its session_id — the selector is checked on the live page before saving. One that matches SEVERAL elements (Amazon '.a-price' matches a dozen; the check would join them into one blob) is pinned to the one shown; one that matches nothing or only an image becomes a VISUAL ZONE watch. NO SELECTOR FOUND AT ALL? mode='visual' + zone_text=<the value exactly as the page prints it, e.g. '51,77 EUR'> watches that area's pixels (digits are compared, so 51.77 finds '51,77 EUR'). A zone fires on ANY visual change — wire a change alert, not a threshold. selector_check in the answer says what was done. HOW OFTEN + ACCOUNT: no interval → needs_input (this plan's options + a tell_user to relay; nothing created); one the plan refuses is refused with the allowed ones. The page is read first: behind a sign-in → needs_persona; a bot check → it offers use_residential.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYesURL to monitor (required).
modeNo'visual' watches the on-screen ZONE of `selector`'s element (or of `zone_text`) and diffs its pixels — for charts, images, badges, or a value no selector can read.
watchNo'price' also rejects a selector whose text holds no number (it becomes a zone). Default 'content'.
deviceNoA linked Writ desktop's agent_id (writ_devices): act ON it. Omit to use the desktop this connection chose with writ_devices action='use' (if any).
enabledNoStart the monitor enabled (default true).
extractNoBROWSERLESS alternative to `selector`: a response-extraction spec the check applies to a plain HTTP response — {"from":"json","path":"data.price"} (a JSON/XHR endpoint, with request_url), {"from":"html_css","selector":".a-offscreen"} (the page markup), {"from":"regex","pattern":"..."} (a value in a script/JSON blob). PREFER it over `selector`/requires_browser whenever the value is readable without JavaScript — most prices and stock lines are: no browser at any check, cheaper, and harder to wall. Best first: structured data (an endpoint, JSON-LD, a JSON blob) survives a redesign. Same grammar as api_call response_extractions.
intervalNoHow often to check: an option id from the needs_input answer ('5m', '15m', '1h', '6h', '24h') or a number of seconds (3600). Omit it and the answer is needs_input: this account's options (checks a day, how long the check allowance lasts, allowed or which plan) and a tell_user to relay — ask the user, then call again with their pick. An interval the plan refuses is refused with the allowed ones; one that runs the allowance out before it renews is created with a `warning`.
selectorNoCSS selector for content-change monitoring; omit for uptime/status monitoring.
zone_textNoThe text the page PRINTS where the value is (e.g. '51,77 EUR', 'Currently unavailable'). Locates the zone when no selector exists.
persona_idNoEvery check carries this persona's LIVE session (kept fresh by the persona's own sign-in), so a login or a bot wall it passed stays passed (writ_personas). Only for a page behind a sign-in: a public page is watched without one, and the answer says when one is needed.
session_idNoAn open writ_browser_use session on this page. The selector is proved there before saving (pinned / switched to a zone); REQUIRED for mode='visual' or zone_text.
try_anywayNoAfter a bot-check answer: create it on Writ's servers anyway (a check that meets the bot check reads nothing).
request_urlNoWith `extract`: the endpoint the value comes from (an XHR the page calls), when it is not `url` itself.
use_residentialNoCheck through a residential exit — for sites that wall datacentre traffic (Amazon, marketplaces). A bot check found on the page is answered with this offer, or with the alternatives when the plan cannot pay for one.
interval_minutesNoLegacy: how often to check, in minutes. Prefer `interval`.
requires_browserNoRender with a real browser (JS) instead of plain HTTP. Set it for JS-rendered/SPA pages and framed pages (framesets/iframes): the check matches the rendered, frame-flattened DOM, and selector validation is deferred to the first browser render instead of a raw-HTML fetch. Omitted, a page the access check could only read in a browser is checked in one.
residential_countryNoISO-2 exit country for use_residential (e.g. 'ca'); implies use_residential.

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, destructiveHint=false, openWorldHint=true, but the description goes well beyond: it explains the needs_input flow when interval is omitted, the needs_persona path for sign-in pages, the bot-check offer of use_residential, the refusal/warning outcomes for disallowed intervals, and the `selector_check` field returned. This is deep behavioral disclosure for a mutating, open-world tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads purpose, then the selector-proof procedure, then frequency/account behavior; every section maps to a real prerequisite for a 17-parameter tool. It is dense and caps-heavy, and some interval semantics restate the schema, so it is efficient but not spartan.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 17-parameter, open-world, non-idempotent creation tool with no output schema, the description covers the full lifecycle an agent needs: prerequisite session, selector validation, browserless alternative, scheduling/needs_input, auth/bot-wall handling, and the returned monitor id. Nothing required to call it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the per-parameter text already carries most semantics (baseline 3). The description adds cross-parameter workflow meaning the schema cannot convey: session_id must be supplied before saving, selector matching several elements gets pinned, mode='visual' requires session_id/zone_text, and extract/request_url may replace selector/requires_browser.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Opens with a specific verb+resource ('Create a MONITOR') and immediately defines the resource's scope: a scheduled target that fires a change_detected event on content, selector-text, or visual-zone change. It also names the sibling it hands off to (writ_wire_monitor) for alerting, so the agent can distinguish it from the rest of the browser/workflow family.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

States explicitly when to use it ('when the user wants to WATCH a URL for changes/updates'), and routes specific situations to alternatives: writ_browser_use to prove the selector, extract over selector when the value is readable without JS, mode='visual' when no selector matches at all. The when-not guidance (extract preferred for browserless, no interval → needs_input) is unusually complete.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

writ_devicesChoose a linked desktopA
Idempotent
Inspect

The user's LINKED WRIT DESKTOPS (they can have several): which are connected right now, and how many local workflows and personas each offers. action='list' (default) shows them; action='use' device= scopes THIS connection to one — runs (writ_run_workflow), browser sessions (writ_browser_use), desktop workflows (writ_list_workflows, runs_on='desktop') and desktop personas (writ_personas, source='device') then target it — and its run history (writ_workflow_runs), data (writ_workflow_data, writ_search_data) and monitors (writ_create_monitor, writ_wire_monitor; list them with action='monitors') are read from and created ON it. action='clear' removes the scope. A desktop's personas sign in from its own vault: the credentials never leave it. Use this when the user says 'on my laptop', 'on my work computer', or wants their own browser and logins.

ParametersJSON Schema
NameRequiredDescriptionDefault
actionNolist (default) | use | clear | monitors (that desktop's monitors) | datasets (what its exposed workflows collected).
deviceNouse: the desktop's agent_id (or its exact name) from list.

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Adds real behavioral context beyond the annotations: which downstream tools get re-scoped, that run history/data/monitors are read from and created on the chosen desktop, and that persona credentials never leave the desktop's vault. The annotations already cover the safety profile (idempotent, non-destructive), so a 4 reflects solid added value without full return-format disclosure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the resource and the default, then the actions in order. Sentences are dense but each carries a distinct consequence; slightly long parenthetical tool lists, though none are pure filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema, so the description does the work of describing what list returns (connection status, workflow/persona counts), which it handles. The one gap is that the 'datasets' action appears in the enum but is never addressed in the description.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description goes further by explaining the downstream effect of each action (use scopes runs/browser/workflows/personas; clear removes scope) and that device accepts an agent_id or exact name from list. Only action='datasets' is left unexplained in prose.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource (list linked WRIT desktops, or scope the connection to one) with a clear scope distinction. The tool is instantly separable from siblings like writ_list_workflows or writ_browser_use because it is explicitly the connector they route through.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly names when to use it: 'on my laptop', 'on my work computer', or when the user wants their own browser and logins. It also states the default action and the clear action, so when-not is covered by the list/use/clear triad.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

writ_diagnose_http_workflowDiagnose a workflow's HTTP laneA
Read-onlyIdempotent
Inspect

Diagnose HTTP-lane readiness for a saved workflow. Reports browser-only dependencies, invalid flow actions/expressions, unsafe internal query parameters, eligibility versus actual proof, and the next repair action. Pass task_id after a representative run to confirm it really used engine=http, did not fall back, and returned RECORDS (an answer such as {ok:true,count:0,listings:[]} is not proof). With task_id it also returns the run's steps and, for a flow, what EVERY request it made was answered with (status, sizes, a response sample) — read that FIRST when a run returned nothing or failed: a 200 with an empty feed means the site stopped serving this session (signed out, blocked, or a null written over a working request variable), not 'no matches'.

ParametersJSON Schema
NameRequiredDescriptionDefault
task_idNoOptional run to verify actual HTTP execution and output.
workflowNo
workflow_idNo

TDQS

A4.3/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly/idempotent/non-destructive, and the description adds substantial behavior beyond that: it discloses what the report contains, how task_id changes the output (run steps plus every request's status, sizes, and a response sample), and how to interpret a 200 with an empty feed versus genuinely 'no matches'. That interpretation guidance is real diagnostic value not present in structured fields.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with purpose, then usage, then diagnostic interpretation, so the ordering is sound and little is filler. The second sentence is very long and packs several ideas together, which costs some scannability, but the density is largely substantive rather than redundant.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description must explain returns and it does so thoroughly, including what a successful proof looks like and what the task_id run payload contains. The remaining gap is the unresolved workflow/workflow_id parameter pair, which an agent still has to guess about.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 33% (task_id is the sole documented parameter), so the description carries extra burden. It explains task_id well, including what its presence changes, but says nothing to disambiguate 'workflow' (string) from 'workflow_id' (integer), leaving the weaker two-thirds of the surface unresolved.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource with scope: 'Diagnose HTTP-lane readiness for a saved workflow.' It then enumerates exactly what it reports (browser-only dependencies, invalid actions/expressions, unsafe query params, proof vs eligibility, next repair action), which distinguishes it from siblings like writ_workflow_runs or writ_run_workflow without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear conditional usage: pass task_id after a representative run to confirm engine=http, no fallback, and RECORDS returned, and says to read the returned step/request answer data FIRST when a run returned nothing or failed. It does not name an alternative tool or state when not to use this one, so it stops short of explicit routing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

writ_discovery_statusCheck an API buildA
Read-onlyIdempotent
Inspect

Poll a build started by writ_website_to_api. Terminal states are succeeded, failed and cancelled. RESTING state: needs_guidance — Writ's mechanical rungs are done and the build is YOURS to finish: it carries map (endpoints seen, specs, candidate functions) and you continue it with writ_website_to_api mode=guided build_id=. A rung the ladder replaced reports superseded=true with fallback_build_id; with wait=true this tool FOLLOWS that pointer for you and answers with the rung now running (followed_from lists the ids it walked; poll the returned build_id from then on). escalations lists every earlier rung with the real reason it handed over (robots.txt refused the crawl, no pages fetched, nothing matched the goal, no structured list...) — quote it when you explain why a browser was needed. On success it returns the workflow_id, the surfaces mapped, and verified — FALSE for a fast/browser map, whose functions are candidates until a real run proves them; say so when you report them. Run any function with writ_run_workflow (function_name), and fetch the generated API docs (OpenAPI 3, Markdown or a Postman collection, all pointing at the real Writ endpoint) from GET /api/v1/workflows/{workflow_id}/api-docs.

ParametersJSON Schema
NameRequiredDescriptionDefault
waitNoONE held call (≤75s) that follows escalations to the newest rung and returns when that build is terminal or parked as needs_guidance — instead of polling. If still building at the ceiling, call again with wait=true and the build_id this answer carries.
build_idYesThe build id from writ_website_to_api.
timeout_secondsNoCeiling for wait=true (≤75).

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover safety (readOnly, idempotent, non-destructive), yet the description adds substantial behavioral context: terminal states, the needs_guidance 'resting' state and its ownership semantics, superseded/fallback_build_id pointer following, escalations as historical hand-off reasons, and the meaning of verified=false. This is far beyond what the annotations convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with purpose and terminal states, and nearly every sentence carries operating information. It is dense and some sentences are long, but there is little filler given the amount of behavior being disclosed.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description must describe returns, and it does: workflow_id, mapped surfaces, verified flag, escalations, map contents, and followed_from. An agent has everything needed to call it and interpret the response.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds real meaning on top: it explains the wait=true continuation contract ('call again with wait=true and the build_id this answer carries') and the follow-pointer behavior, which the schema's terse wait description only gestures at.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Opens with a specific verb+resource ('Poll a build started by writ_website_to_api') and names the sibling that creates the build. An agent can immediately tell this is the status/polling counterpart to writ_website_to_api.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly states when to poll vs. use wait=true, what to do in the needs_guidance state, how to continue a build, and routes downstream actions to writ_run_workflow and the api-docs endpoint. It names alternatives and the conditions that select them.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

writ_export_dataExport collected dataB
Read-onlyIdempotent
Inspect

Export a workflow's full extracted-data table as CSV or JSON (search/filter applied, un-paginated).

ParametersJSON Schema
NameRequiredDescriptionDefault
qNo
viewNo
formatNocsv (default) or json
workflowNo
workflow_idNo

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly, idempotent and non-destructive, so the safety profile is covered. The description usefully adds that the export is un-paginated and returns the FULL table, which warns the agent about potentially large payloads, but it says nothing about auth requirements, size limits, or whether filters must be set to avoid huge results.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler; the key behavioral fact (un-paginated, filter-applied full table) is surfaced immediately.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With five mostly-undescribed parameters, no output schema, and no indication of which identifiers are required, the description leaves an agent guessing how to target the right workflow and what the returned payload looks like. Given the low schema coverage, it needed to compensate more than it does.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 20% — just 'format' is documented, and the description only restates csv/json. The meaning of q, view, workflow and workflow_id (and which of the redundant workflow/workflow_id to use) is left entirely undocumented despite 5 parameters and zero required fields.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Export) and resource (a workflow's full extracted-data table) plus the output formats, so an agent knows exactly what it produces. It doesn't explicitly distinguish itself from close siblings like writ_workflow_data or writ_search_data, which is the only reason it isn't a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'search/filter applied, un-paginated' implies this is the bulk-export counterpart to a paginated/search tool, but no alternative is named and no when-to-use or when-not-to-use condition is given. Usage must be inferred from the parenthetical.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

writ_expose_workflow_apiPublish a workflow as a REST APIA
Idempotent
Inspect

Publish a saved workflow through Writ's managed REST gateway, using the same resource as the frontend's REST endpoint switch. The returned POST URL waits for the workflow and returns its JSON result. Repeated calls reuse the existing endpoint.

ParametersJSON Schema
NameRequiredDescriptionDefault
labelNoOptional name for the endpoint.
workflowNo
workflow_idNo
wait_timeoutNoDeprecated alias of timeout_seconds.
timeout_secondsNoManaged run timeout. Defaults to 120, or 300 for AI navigation workflows.

TDQS

A3.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations cover safety (readOnlyHint=false, destructiveHint=false, idempotentHint=true), and the description adds real context: the returned POST URL blocks until the workflow finishes and returns its JSON result, and repeat calls are idempotent (reuse the endpoint). It does not mention auth/permission requirements, which is the remaining gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences, front-loaded with the core action, then the return behavior, then the idempotency note. No filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description appropriately explains the return (a POST URL that returns the workflow's JSON result) and the idempotency behavior. It is complete for invoking, though it omits permission/prerequisite context and gives no parameter guidance.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Five parameters with only 60% schema coverage; 'workflow' and 'workflow_id' have no description in either schema or description text, and the description adds nothing about how to specify which workflow, the label, or timeout behavior. It does not compensate for the coverage gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource: publishing a saved workflow as a managed REST endpoint, and clarifies it uses the same resource as the frontend's REST switch. It is distinguishable from siblings like writ_run_workflow or writ_website_to_api, though it never names an alternative explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage context is implied (expose a saved workflow as an API endpoint) and it notes that repeated calls reuse the existing endpoint, but there is no explicit when-to-use/when-not guidance or comparison against alternatives such as running the workflow directly.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

writ_list_webhooksList inbound webhook URLsA
Read-onlyIdempotent
Inspect

The account's inbound webhook URLs: for each, its webhook_trigger_id, the callable url, whether it is signed, the automations it fires and how often it was called. Signing secrets are stored encrypted and never shown. An automation reuses one with writ_create_automation webhook_trigger_id=; when='webhook_received' without it mints a new URL.

ParametersJSON Schema
NameRequiredDescriptionDefault
webhook_trigger_idNoOnly this webhook.

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false, so the safety profile is covered. The description adds value beyond them: signing secrets are stored encrypted and never returned, and it clarifies that reuse vs. minting of webhook URLs depends on supplying webhook_trigger_id.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, front-loaded with the resource and its returned fields, then the security caveat, then the integration rule. No filler or repetition; every clause carries information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description usefully enumerates the returned fields and the security constraint, and annotations cover the safety profile. What is missing is minor: no guidance on empty results, pagination, or ordering of the listed URLs.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and the single optional webhook_trigger_id is documented in the schema as 'Only this webhook.' The description adds no filter syntax or behavior beyond that, so the baseline 3 for a fully documented parameter set applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Specific verb+resource: it lists the account's inbound webhook URLs and enumerates exactly what is returned for each (webhook_trigger_id, callable url, signed status, firing automations, call count). An agent immediately knows this is a read of webhook configuration, not a crawler, workflow, or browser tool like its siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explains the relationship to writ_create_automation (reuse an existing webhook_trigger_id; when='webhook_received' without it mints a new URL), which supplies real context for the trigger id. However it never states when to call this listing tool rather than alternatives, nor any exclusion or precondition for the optional filter.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

writ_list_workflowsList saved workflowsA
Read-onlyIdempotent
Inspect

List the workflows saved in your Writ account — each runs on demand without live browsing. Returns id, name, declared inputs, schedule, and whether it is pinned as its own run_ tool.

ParametersJSON Schema
NameRequiredDescriptionDefault
searchNoOptional name/description filter.

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, openWorldHint=false and destructiveHint=false, so the safety profile is covered. The description still adds genuinely useful behavior: each workflow "runs on demand without live browsing" and the response enumerates id, name, declared inputs, schedule, and pinned-tool status.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tightly written clauses: purpose/scope first, then the return contract. No filler, and the most decision-relevant information is front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description usefully enumerates the returned fields and the execution model (on-demand, no live browsing), which an agent needs to decide follow-up actions. It omits only routing guidance to alternative list tools.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% for the single optional `search` parameter, so the schema already carries the semantics. The description adds nothing about the filter, making the baseline 3 appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ("List the workflows saved in your Writ account") and clarifies the scope relative to live-browsing siblings. It does not name a specific alternative sibling (e.g. writ_workflow_runs vs writ_list_workflows), so differentiation is contextual rather than explicit.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is only implied: the tool lists saved workflows, so an agent can infer this is the discovery step before running or pinning one. There is no explicit when-to-use, when-not-to-use, or named alternative, leaving the agent to infer routing against siblings like writ_workflow_runs and writ_saved_crawls.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

writ_paymentPay with a card the user approvesA
Destructive
Inspect

Pay on a website with one of the user's cards, without the card number ever reaching this conversation: the AI never sees or asks for card numbers. action='request' (site, max_amount, purpose) returns a tell_user and an open_url: a Writ window where the user picks or adds a card (or makes a virtual card) and approves it. A purchase needs that approval. action='wait' grant_id= is ONE held call (up to 60s) that answers when they approve, with the card's brand and last four digits only. PLACING THE ORDER: action='checkout' (session on the final checkout page, commit_selector = the place-order button, total_selector = the order total, grant_id, or max_amount when the store uses its own saved card) pauses the session and the user confirms (emailed link, the Writ app, or an auto-confirm rule they turned on); then Writ types the card, checks the total and clicks the order button itself: a real purchase. You never click an order button yourself (refused). action='wait' confirmation_id= returns the outcome. action='fill' types an approved virtual card early (multi-page checkouts). Card use is limited to the site and amount granted; an unused grant expires after 15 minutes. action='list' shows the user's payment methods as handles (kind, brand, last4, id), never numbers; a handle goes into writ_wire_monitor buy.payment {kind, ref}.

ParametersJSON Schema
NameRequiredDescriptionDefault
siteNorequest: the store's domain, e.g. 'store.example.com'. The card is typed only on this site and its payment frames.
stepsNocheckout: steps Writ runs after confirmation, before the order button, e.g. [{"type":"fill_card"}, {"type":"click","selector":"#continue"}, {"type":"wait","seconds":2}] for a card page followed by a review page. Default: fill the granted card, then the order button.
actionYesrequest = ask the user to approve a card for one purchase; wait = hold until they answer (grant_id) or until a checkout is settled (confirmation_id); checkout = pause at the order button for the user's confirmation, then Writ places the order; fill = type an approved virtual card early; list = the user's payment methods as handles.
fieldsNorequest, for live browsing: card field -> CSS selector on the checkout page (read them with writ_browser_context), or {selector, frame_url} for a field inside a payment provider's frame (frame_url = part of the frame's URL, e.g. 'js.stripe.com'). Keys: number, exp (MM/YY), exp_month, exp_year, exp_year2, cvc, name, zip. Writ types into exactly these.
purposeNorequest: one line the user reads when approving, e.g. 'Buy Nike Dunk Low, size 10'.
sessionNocheckout / fill (required) / request: the session_id of the writ_browser_use session on the checkout page.
summaryNocheckout: one line the user reads when confirming, e.g. '1 x Nike Dunk Low, size 10, shipped to home'.
currencyNorequest: three-letter ISO code (default 'usd').
grant_idNowait / fill / checkout: the grant_id action='request' returned.
max_amountNorequest / checkout: the most this purchase may charge, tax and shipping included (e.g. 129.99). checkout without grant_id needs it.
automation_idNorequest: the automation the card is for, when the purchase is a saved automation.
total_selectorNorequest / checkout: CSS selector of the order total on the checkout page. Writ reads it and never clicks on when it is above the maximum. Needed for an auto-confirm rule to apply.
commit_selectorNocheckout (required): CSS selector of the button that places the order. Writ clicks it after the user confirms.
confirmation_idNowait: the confirmation_id action='checkout' returned.
payment_method_idNorequest: a handle id from action='list' to pre-select that card; the user still approves.

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the annotations (destructiveHint=true, openWorldHint=true), the description discloses critical behavior: card numbers never reach the conversation, card use is scoped to site and amount granted, grants expire after 15 minutes, wait is a single held call up to 60s, and checkout pauses for user confirmation before Writ places a real purchase. This is unusually rich context that annotations cannot convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The opening clause front-loads the key constraint (no card numbers) and the request flow, and every sentence carries substantive information. It is dense and long, but for a 15-parameter, five-action tool this density is largely earned; slightly tighter per-action grouping would improve scanning.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex mutation tool with 15 parameters, nested objects, and no output schema, the description covers the entire workflow from approval to order placement and explains the return shape (brand and last four digits, confirmation outcome). Nothing an agent needs to invoke it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3, but the description adds cross-parameter semantics: which action consumes grant_id versus confirmation_id, that checkout without grant_id requires max_amount, and how commit_selector/total_selector gate the auto-confirm rule. It enriches schema-level info rather than merely repeating it.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb and resource (pay on a website using the user's cards) and enumerates all five actions with their distinct functions. It clearly differentiates from siblings by routing card handles into writ_wire_monitor and by referencing writ_browser_context for selector reading.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Each action's use case is stated explicitly: request to obtain approval, wait to hold for a grant/confirmation, checkout to place the order, fill for multi-page checkouts, list to enumerate payment methods. It also names alternatives (writ_wire_monitor for buy.payment, writ_browser_context for reading selectors) and an exclusion ('You never click an order button yourself (refused)').

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

writ_personasUse the user's saved sign-insAInspect

The user's OWN accounts on websites. When a task needs them signed in (their email, social, shop, bank or work portal), a persona is how Writ signs in as them: no password passes through this tool or the conversation, so never ask for one. A persona is a saved sign-in identity: a site's username plus credentials sealed server-side (never readable here), optional 2FA whose codes are minted server-side, and a warm signed-in session. USE one by passing its persona_id to writ_browser_use, writ_crawl_site, writ_scrape or writ_run_workflow. BEFORE asking the user for credentials for a site, call action='list' (filter by domain). action='get' inspects one (include_runs adds its recent runs); action='sign_in' runs its login workflow NOW (force=true re-logs-in even when the session looks usable); action='record_login' has a server-side AI sign in as it once and RECORD the flow as its login workflow, so it can always sign itself back in. NONE FITS? This tool can NOT create a persona and no credential ever passes through it. list with a domain answers persona_needed: a tell_user, a create_url (a MINTED link: a small Writ window with only the persona form, pre-filled for that site) and its link_id; action='request' domain= mints one on purpose (another account for a site). Relay it BEFORE starting the task, then action='wait' link_id=: ONE held call that answers the moment the user saves it, with the persona_id — continue on your own. ALSO LISTED: the personas of the user's linked Writ DESKTOP (source='device', id device:<agent>:<id>, name and site only). Pass that id as persona_id to writ_run_workflow, writ_browser_use / writ_record_website, writ_scrape or writ_crawl_site: the work goes to that desktop, which signs in from its own vault and only on the persona's own site - the credentials never leave it.

ParametersJSON Schema
NameRequiredDescriptionDefault
whyNorequest: one line the user sees in the window — what you need the account for.
waitNolist + domain: hold up to 75s for a persona for that site to appear. With a link_id prefer action='wait'. Never poll in a loop instead.
forceNosign_in: re-run the login even when the current session still looks usable.
actionYesWhat to do (default list). request = mint the persona link for a site (domain); wait = hold until the user saves it (link_id).
domainNolist: only personas usable on this host (suffix match), e.g. 'github.com'. With no match the answer is `persona_needed` — the ask to relay to the user. request: the site the new persona is for.
link_idNowait: the `link_id` of a persona_needed / request answer. Holds up to 75s and answers the moment the user saves the persona (persona_id), declines, or closes the window; `still_waiting` means call it again.
login_urlNorecord_login: exact sign-in page URL when known; defaults to the persona's domain root (the AI finds the form from there). list / request: the site's sign-in page, carried into `create_url` so the new persona can record it.
persona_idNoWhich persona — required for get / sign_in / record_login.
include_runsNoget: include the persona's recent runs (which workflows acted as it, and whether they succeeded).

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well beyond annotations: no password passes through the tool or conversation, credentials are sealed server-side and never readable, 2FA codes are minted server-side, sessions are warmed, wait holds up to 75s and answers on save/decline/close, and desktop personas sign in from their own vault only on their own site. This is rich operational context that annotations (openWorld, not read-only) do not convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The content is valuable and the action-routing is front-loaded, but it is a dense wall of run-on, heavily-bolded sentences that mix definitions, workflows, and return shapes. Some of the length is justified by six actions, yet it is harder to parse than it needs to be.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 9-parameter, six-action tool with no output schema, the description does the heavy lifting by describing return keys (persona_id, link_id, still_waiting, persona_needed with tell_user/create_url). It is nearly complete, though the response shapes could be stated slightly more cleanly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds cross-parameter semantics the schema does not: the request→link_id→wait→persona_id workflow, that domain does suffix matching and yields persona_needed, and that link_id answers with persona_id/still_waiting. This meaningfully enriches how the parameters relate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description defines the resource precisely (a 'saved sign-in identity' with sealed credentials) and enumerates the concrete actions (list, get, sign_in, record_login, request, wait). It also distinguishes itself from siblings by naming the tools it feeds (writ_browser_use, writ_crawl_site, writ_scrape, writ_run_workflow), so an agent can tell what it is and how it connects to the ecosystem.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit routing is given: use persona_id in the sibling tools, call action='list' (filtered by domain) BEFORE asking the user for credentials, use 'sign_in' to run the login now, 'record_login' to capture a workflow, and 'request'/'wait' to mint and hold for a new persona. It even states what the tool cannot do ('can NOT create a persona') and warns against polling loops.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

writ_pin_workflow_toolPin a workflow as a toolA
Idempotent
Inspect

Pin (or unpin) a saved workflow as its own run_ tool on this server. Workflows are NOT exposed as individual tools by default — every one is always callable via writ_run_workflow — so pin only the few the user runs often enough to deserve a first-class tool (the derived list is capped). Do this when the user asks for it, or after saving a workflow the user clearly intends to call as a tool from here.

ParametersJSON Schema
NameRequiredDescriptionDefault
pinnedNotrue (default) pins; false unpins.
workflowNoWorkflow name (or use workflow_id).
workflow_idNo

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover the safety profile (readOnlyHint=false, idempotentHint=true, destructiveHint=false), so the description's job is to add context. It does: workflows are never exposed by default, unpinning reverses the effect, and the derived tool list is capped. Return-shape detail is absent, but that is minor here.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, front-loaded with the action and effect, then the constraint, then the trigger condition. Slightly dense with parentheticals, but every clause carries routing or constraint information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema and full annotation coverage, the description supplies the missing piece — why and when this tool exists relative to writ_run_workflow. Only the workflow vs workflow_id choice and any cap/pagination behavior remain unstated.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 67% and the description mirrors rather than extends it, restating the pin/unpin toggle and naming the workflow as the target. It adds no disambiguation between workflow and workflow_id, and the undocumented workflow_id param is left untouched.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a precise verb pair (pin/unpin) plus the concrete resource and effect: creating a derived run_<name> tool on this server. It explicitly contrasts itself with writ_run_workflow, so an agent can distinguish it from the default invocation path.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit when-to-use ('when the user asks for it, or after saving a workflow the user clearly intends to call as a tool') and an implicit when-not ('pin only the few… the derived list is capped') plus the named alternative writ_run_workflow for non-pinned use.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

writ_record_websiteRecord a website taskA
Destructive
Inspect

Record a repeatable website TASK as a workflow that replays on demand. Use whenever the user asks to record, capture, teach, automate or repeat actions on a site and the point is the TASK, not an API surface. Writ opens a real cloud browser and returns an observation; YOU are the brain: drive it with writ_browser_act, author what the recorder cannot see with writ_browser_compose (inputs a caller passes, explicit steps, named functions), and call writ_browser_save when the goal is complete — the saved workflow then replays at zero AI cost (writ_run_workflow, or its own run_ tool once pinned with writ_pin_workflow_tool) and can be scheduled. A goal that asks for an API ("turn into an API", "an endpoint for") is routed to writ_website_to_api's intelligent ladder automatically (static crawl, rendered crawl, then Writ's AI browser rung), so the recording is a real build. Before anything runs it proposes the user's OWN matching workflows (existing_workflows) and ready-made marketplace APIs (marketplace_candidates); skip_existing / skip_marketplace bypass those.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYesWebsite URL to start on (required).
goalYesWhat should be recorded on the website, in plain language
deviceNoA linked Writ desktop's agent_id (writ_devices): act ON it. Omit to use the desktop this connection chose with writ_devices action='use' (if any).
persona_idNoSaved identity to sign in with (list them with writ_personas). Required for sites behind a login with 2FA — the one-time code is then minted server-side and never shown to you. A desktop persona ('device:…') records on its own desktop, which fills its values itself.
skip_existingNoAPI builds first propose the user's OWN matching workflows (replaying is instant and free); set true after the user declined those.
use_residentialNoOpen on the platform residential network (premium) for a site that blocks datacenter IPs or shows a bot wall. Default off (free datacenter egress).
skip_marketplaceNoAPI builds then propose compatible ready-made marketplace APIs; set true to skip that and record fresh.
residential_countryNoTwo-letter ISO country the residential exit should be in (e.g. 'us', 'fr') — for a site that serves a different page per country, or throttles foreign traffic. Omit for an automatic exit. Ignored unless the session egresses residential.

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Adds substantial context beyond the annotations: it opens a real cloud browser and returns an observation, the agent must drive it via writ_browser_act/compose/save, the saved workflow replays at zero AI cost and can be scheduled, and it proposes existing_workflows/marketplace_candidates before running. It does not explain the destructiveHint=true implications or permission/auth needs, which keeps it short of a 5 against already-rich annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads purpose in sentence one and each following sentence carries routing or workflow information. It is dense and long with heavy parenthetical content referencing many sibling tools, which costs some readability, but little of it is padding.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex 8-parameter, open-world, non-idempotent recording tool with no output schema, the description covers the trigger, the alternatives, the multi-tool authoring flow and the bypass flags. It stops short of describing the return/observation shape or failure behavior, but is complete enough to invoke correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all 8 parameters, making 3 the baseline. The description adds only marginal meaning (the skip_existing/skip_marketplace bypass rationale) and no extra syntax or format detail beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Opens with a specific verb+resource: 'Record a repeatable website TASK as a workflow that replays on demand.' It explicitly distinguishes this tool from the sibling writ_website_to_api by stating API-shaped goals are routed there automatically, so an agent can tell them apart without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit trigger ('whenever the user asks to record, capture, teach, automate or repeat actions on a site and the point is the TASK'), names the alternative and the condition that selects it (API goals -> writ_website_to_api), and sketches the in-tool flow (act/compose/save) plus skip_existing/skip_marketplace bypass semantics.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

writ_run_saved_crawlRun a saved crawlAInspect

Run a saved crawl with its stored settings. Pass max_age to get the data it already collected if that run is recent enough — the cheap path. Otherwise it re-crawls. The response carries _cache.hit and _cache.age_seconds so you can tell which happened.

ParametersJSON Schema
NameRequiredDescriptionDefault
waitNoBlock until the crawl converges (default false — a crawl is slow).
crawlYesSaved crawl slug, name, or id (from writ_saved_crawls).
limitNoRows of collected data to include (default 50).
outputNoRESPONSE SHAPE — set this whenever the answer is for a program or an API you are building, not for you to read. {shape: 'envelope' (default: Writ's full answer, projected) | 'table' ({columns, rows, total}) | 'records' (bare list of records) | 'record' (the newest record alone — one entity, a usage meter, a dashboard), fields: ['used', 'percent_used as pct', 'items.0.price as first_price'] (ordered pick, renames, dotted paths; missing → null so keys are stable), exclude: ['depth'], include_meta: false (page metadata content_kind/depth/thumbnails are STRIPPED unless true), key: 'usage' (wrap)}. On writ_crawl_site with save_as it is SAVED as the API's default shape.
max_ageNoReuse the last completed crawl if it finished within this many seconds. 0 (default) always re-crawls.
preview_charsNoCut each inline page's text cells to this many characters (default 12000; 0 = full). Full page: writ_workflow_data(workflow_id=<data_workflow_id>, refs=[...]).
timeout_secondsNoMax seconds to wait when wait=true.

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare non-readOnly, non-idempotent, open-world, and the description adds genuinely new behavioral context on top: the cache-vs-recrawl decision, the cheap path, and the `_cache.hit` / `_cache.age_seconds` response fields. It still doesn't note that a run may be slow or that repeated calls without max_age duplicate work, but the added caching semantics are substantive.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, zero filler, front-loaded with purpose followed by the caching decision and the means to verify it. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description discloses the most decision-relevant return fields (`_cache.hit`/`age_seconds`) and the caching model, which is enough to call it correctly across 7 params with full schema coverage. It stops short of describing the general response envelope or the non-idempotent side effects of an actual re-crawl.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so baseline is 3, and the description goes beyond the schema by framing max_age as the 'cheap path' shortcut and by documenting the `_cache` response fields that appear in no parameter description. The wait/timeout/output parameters are left entirely to the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (run) and resource (a saved crawl) and immediately scopes it with 'with its stored settings', which distinguishes it from writ_crawl_site (new crawl) and writ_update_saved_crawl (mutation of the definition). An agent can pick it out from the sibling list without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives one useful conditional: pass max_age for the cheap cached path, otherwise it re-crawls. But it never says when to choose this tool over writ_crawl_site, writ_scrape, or writ_saved_crawl_data, nor does it state prerequisites for the saved crawl existing. Usage is implied rather than routed.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

writ_run_workflowRun a saved workflowA
Destructive
Inspect

Run a saved workflow by id or name and (by default) wait for it to finish, returning the extracted data. Pass workflow inputs as top-level fields or under inputs, and any file inputs under files.

ParametersJSON Schema
NameRequiredDescriptionDefault
waitNoWait for completion and return the data (default true).
filesNoOptional file inputs for this run, as {slot: file_id}. Slot names come from the workflow's `file_slots` (writ_list_workflows); file_ids come from the account's file library. A workflow whose upload step already has a file pinned runs fine with no `files` at all — pass it only to swap the file for THIS run.
deviceNoRun on this linked Writ desktop (an agent_id from writ_devices) — overrides the desktop this connection chose with writ_devices action='use'.
inputsNoRun inputs (or pass them as top-level fields).
outputNoRESPONSE SHAPE — set this whenever the answer is for a program or an API you are building, not for you to read. {shape: 'envelope' (default: Writ's full answer, projected) | 'table' ({columns, rows, total}) | 'records' (bare list of records) | 'record' (the newest record alone — one entity, a usage meter, a dashboard), fields: ['used', 'percent_used as pct', 'items.0.price as first_price'] (ordered pick, renames, dotted paths; missing → null so keys are stable), exclude: ['depth'], include_meta: false (page metadata content_kind/depth/thumbnails are STRIPPED unless true), key: 'usage' (wrap)}. On writ_crawl_site with save_as it is SAVED as the API's default shape.
max_ageNoOptional. Reuse a previous result if it is younger than this many seconds, instead of running the workflow again. 0 (the default) always runs fresh. Use it when a recent answer is good enough — much faster and cheaper.
workflowNoWorkflow name (or use workflow_id).
persona_idNoRun AS this saved identity (see writ_personas) — the run signs in with the persona's warm session. Omit to use the workflow's default persona, if it has one. A persona of the user's linked DESKTOP (`device:<agent>:<id>`, source='device' in writ_personas) sends the run to that desktop, which signs in from its own vault: the credentials never leave it.
workflow_idNoA number, or `local:<id>` for a workflow that lives on the user's linked Writ desktop (writ_list_workflows, runs_on='desktop') - it runs there, signed in as one of that desktop's personas when persona_id is `device:...`.
function_nameNoCall ONE named function of a multi-function workflow (an API built with writ_website_to_api / writ_browser_compose define_function): only that function and the sign-in functions it depends on run. Omit to run the whole workflow. It is a control, never a workflow input.
mutation_modeNoHow a WRITE function (one that creates/posts/sends/deletes) runs on THIS run. Default 'live': you invoked the workflow, so its writes ARE sent. Pass 'dry_run' to preview the request without sending, or 'private_test' to send it with the function's safe overrides. (Separately, BUILDING a function — define/compile/test — never sends a write, whatever this is.) Reads ignore this.
function_namesNoCall SEVERAL functions in ONE run instead of `function_name`: the union of their steps runs once, in recorded order (a prerequisite they share runs once). Not for a desktop (`local:`) workflow. A control, never a workflow input.
timeout_secondsNoMax seconds to wait for completion (default 120).
use_residentialNoPer-call network override for an owned workflow: true uses the platform residential network, false disables the workflow's residential default. Use it for geo-sensitive or datacenter-blocking sites.
execution_targetNoWhere this run executes: 'cloud' (managed fleet), 'auto' (prefer the user's OWN linked Writ desktop app when online, else cloud), or 'local' (require their own desktop app — keeps the run on their machine + IP). Omit to keep the workflow's own configured target. If a 'local' run fails because the app is offline, tell the user and only fall back to 'cloud' with their agreement.
residential_countryNoTwo-letter ISO-3166 exit country for this call (for example ca, us, fr). Keep it aligned with the requested storefront, coordinates or delivery market. It is applied when the run uses residential egress.

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, idempotentHint=false, and openWorldHint=true, so the mutation profile is partly covered. The description adds the default synchronous wait-and-return-extracted-data behavior, but never mentions the live-write default or the dry_run/private_test escape hatch (that lives only in the schema), leaving key behavioral context on the param descriptions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences, front-loaded with the core action and the default behavior; no filler or repetition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 16-parameter, nested, open-world execution tool with no output schema, the description is thin: it omits mention of function_name/function_names selection, mutation_mode, execution_target, persona/device routing, and max_age caching. The schema covers these well, but the description does little to orient an agent on this complexity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description does add the convention that inputs may be passed as top-level fields or under `inputs` and file inputs under `files`, but much of this is also restated in the schema, so marginal added value.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Run a saved workflow') and adds useful scope: identified by id or name, defaults to waiting for completion and returning extracted data. It does not name or distinguish itself from siblings like writ_workflow_runs or writ_website_to_api, so it stops short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage through the default-wait behavior and input-passing conventions, but gives no explicit when-to-use/when-not guidance versus the many sibling run/inspect tools (writ_workflow_runs, writ_replay, etc.).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

writ_saved_crawl_dataRead a saved crawl's dataA
Read-onlyIdempotent
Inspect

Read the data a saved crawl already collected on its most recent completed run. Never starts a crawl — use this when you want what is already there, at any age.

ParametersJSON Schema
NameRequiredDescriptionDefault
crawlYesSaved crawl slug, name, or id.
limitNoRows to return (default 25).
outputNoRESPONSE SHAPE — set this whenever the answer is for a program or an API you are building, not for you to read. {shape: 'envelope' (default: Writ's full answer, projected) | 'table' ({columns, rows, total}) | 'records' (bare list of records) | 'record' (the newest record alone — one entity, a usage meter, a dashboard), fields: ['used', 'percent_used as pct', 'items.0.price as first_price'] (ordered pick, renames, dotted paths; missing → null so keys are stable), exclude: ['depth'], include_meta: false (page metadata content_kind/depth/thumbnails are STRIPPED unless true), key: 'usage' (wrap)}. On writ_crawl_site with save_as it is SAVED as the API's default shape.
preview_charsNoCut string cells (page markdown) to this many characters (default 2000; 0 = full cells). Full single pages: writ_workflow_data(refs=...) per the response hint.

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false, so the safety profile is covered. The description adds real behavior beyond that: the data is a cached snapshot of the most recent completed run only, and calling it never triggers a crawl — an important expectation-setting detail.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, both load-bearing, with the read-only, no-run-started constraint front-loaded before the usage hint. Nothing is repeated from the title or schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description usefully states which run's data is returned and that it is never mutated, which compensates for the missing return-value documentation. Pagination and response-shape behavior are left entirely to the schema, which documents the output param but no return format.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the nested output object, limit, and preview_chars are all documented in the schema with defaults and semantics. The description adds no parameter-level guidance, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Read the data a saved crawl already collected') and pins the exact data set ('most recent completed run'). The phrase 'Never starts a crawl' cleanly separates it from writ_run_saved_crawl and writ_crawl_site without needing to open either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'use this when you want what is already there, at any age' gives a clear condition for selection, and the negative clause rules out triggering a fresh run. It stops short of naming the alternative tool (writ_run_saved_crawl) for the fresh-run case, leaving that inference to the agent.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

writ_saved_crawlsList saved crawlsA
Read-onlyIdempotent
Inspect

List the crawls the user SAVED for re-running (callable by API, each with a scope: seed_url, rank_cap, include_paths, extract_mode, executor). Use one via writ_run_saved_crawl(max_age=…) when its scope matches the ask — recent data then comes back instantly and costs nothing. A saved crawl of a DIFFERENT page, or one using executor=ai, is not a shortcut for a fresh question: start writ_crawl_site(rank_cap=N) instead.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMax saved crawls to return (default 50).

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint and destructiveHint=false, so the safety profile is covered. The description adds real value beyond that: it discloses the API-callable nature of saved crawls, the per-item `scope` structure, and the behavioral caveat that executor=ai crawls are not reusable for fresh questions. It does not discuss result ordering or pagination behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, front-loaded with the purpose before the routing advice, and each sentence carries information. It is dense with parenthetical detail, but nothing is filler or restated boilerplate.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, and the description compensates by describing the shape of each returned crawl and its scope fields, plus the correct follow-up tool. The only gaps are result count/ordering and whether limit truncates silently, which are minor for a read-only list tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There is a single parameter (limit) at 100% schema description coverage, so the schema already carries the semantics; the description never mentions limit and adds nothing parameter-level. Baseline 3 is appropriate when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('List the crawls the user SAVED for re-running') and immediately characterizes what a saved crawl contains (scope with seed_url, rank_cap, include_paths, extract_mode, executor). It is unambiguously distinguishable from sibling write tools like writ_update_saved_crawl or writ_run_saved_crawl.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly routes between alternatives: use writ_run_saved_crawl(max_age=…) when the scope matches, otherwise start writ_crawl_site(rank_cap=N). It even names the exclusion case (different page, or executor=ai) that disqualifies a saved crawl as a shortcut, which is unusually precise guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

writ_scrapeRead web pagesAInspect

READ PAGE CONTENT NOW — one page, a list of pages, or the top N items of a listing — returned as clean markdown IN THIS CALL (2-10s). This is the tool for 'what does say', 'summarize ', 'the top N posts/products/results of and what's on each', 'fetch these 3 links' — including a page a plain fetch cannot read: blocked or empty (403, bot wall), rendered by JavaScript, or behind the user's sign-in (render_mode, use_residential, persona_id). NOT FOR: collecting a whole site or section into a dataset (writ_crawl_site); clicking, typing, signing in or any action on a page (writ_browser_use).

THREE SHAPES, ONE CALL EACH:

  • url → that page.

  • urls=[...] (≤20) → all of them, fetched in parallel, pages in the order given.

  • url= + top_n=N (≤20) → the listing (listing) AND the N top-ranked item pages it links to (pages, in rank order) — e.g. url='https://news.ycombinator.com/', top_n=3 returns the front page and the 3 top stories' discussion pages. Add include_paths=['item\?id='] when you know the item-link shape; the server otherwise detects it.

Discussion pages keep their comment threads; each comment is tagged [top-level] or [reply · depth N], so 'the top-level comments' is answerable from the text. Long pages are preview-cut at 12000 chars (_truncated); the hint tells you how to fetch a full page. YOU read the markdown — no AI is spent here. Behind a login: persona_id. Bot wall: use_residential=true.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlNoA page to read — or, with `top_n`, the LISTING page whose top items to read.
urlsNoKnown pages to read together (max 20), fetched in parallel in one call.
top_nNoWith `url` = a listing/front/search/category page: also read its N top-ranked item pages (the page's link order IS the ranking). One call, parallel.
deviceNoA linked Writ desktop's agent_id (writ_devices): act ON it. Omit to use the desktop this connection chose with writ_devices action='use' (if any).
formatNomarkdown (default): the page's clean main content. html: the RAW HTML as fetched (same egress/persona/render) — for selectors, embedded JSON, anything the cleaned text drops; clipped to preview_chars (default 40000, 0 = whole page). both: html plus the markdown derived from it, one fetch. Single `url` only.
persona_idNoSaved identity to read AS (list them with writ_personas) — for pages behind a login. Forces the identity's own residential exit. A desktop persona ('device:…') reads on its own desktop, from that machine.
render_modeNoauto (default: plain HTTP, browser only if the page needs JS) | http | browser.
include_pathsNoWith top_n: regex(es) the item links match (e.g. 'item\\?id=', '/products/'). Optional — the server detects detail links when omitted.
preview_charsNoCut each page's text to this many characters (default 12000; 0 = full pages).
respect_robotsNoApply robots.txt to the explicitly requested page(s). Default false for scrape; writ_crawl_site defaults true for autonomous discovery.
use_residentialNoFetch through the platform residential network (premium) for a site that blocks datacenter IPs or shows a bot wall. Default off.
residential_countryNoTwo-letter ISO country the residential exit should be in (e.g. 'us', 'fr') — also used by the automatic residential retry on a blocked page. Omit for an automatic exit. Ignored unless the session egresses residential.

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Discloses behavior well beyond the annotations: 2-10s latency, parallel fetching, 12000-char preview cut with a `_truncated` flag, a `hint` for full retrieval, 'YOU read the markdown — no AI is spent here', and the escalation paths for bot walls (use_residential) and logins (persona_id). Nothing here contradicts readOnlyHint=false/openWorldHint=true; a network egress that may consume premium residential is consistent with non-read-only.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded and sectioned ('THREE SHAPES, ONE CALL EACH'), so an agent can scan the operative rules quickly. It is nonetheless long and stylistically dense (all-caps opener, heavy em-dash nesting), so a little more restraint would help; every section still carries non-redundant information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema and 12 parameters, the description carries the return-shape burden itself, naming `pages`, `listing`, `_truncated`, `hint`, and the comment depth tags, plus failure modes (403/bot wall/JS/empty) and their remedies. Nothing an agent needs to call or interpret this tool is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Although schema coverage is 100%, the description adds combinatorial meaning the schema cannot express: the three argument shapes (single url, urls≤20, url+top_n≤20) and exactly what each returns (`pages` in given order, `listing` + ranked `pages`, comment depth tagging). It also clarifies defaults and cross-tool differences for respect_robots, include_paths auto-detection, and residential_country.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('READ PAGE CONTENT') and differentiates immediately from siblings: 'NOT FOR: collecting a whole site ... (writ_crawl_site); clicking, typing, signing in ... (writ_browser_use)'. An agent can pick this tool without opening a sibling schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit when-to-use phrasing mapped to user intents ('what does <page> say', 'summarize <url>', 'top N posts') plus a dedicated NOT FOR section naming the two alternative tools and their domains. Exclusions and alternatives are both covered.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

writ_search_dataSearch collected dataA
Read-onlyIdempotent
Inspect

Search across everything already collected by your workflows for a term — answers data questions from past runs without running anything. Scopes to one workflow when given, else fans out. Matches come back preview-sized; fetch full records with writ_workflow_data(refs=...).

ParametersJSON Schema
NameRequiredDescriptionDefault
qYesSearch term (required).
limitNoRows per workflow (default 10).
deviceNoA linked Writ desktop's agent_id (writ_devices): act ON it. Omit to use the desktop this connection chose with writ_devices action='use' (if any).
workflowNo
workflow_idNo
preview_charsNoCut matched string cells to this many characters (default 300; 0 = full cells).

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true and destructiveHint=false, so the safety profile is covered. The description adds value beyond that: it discloses the fan-out-across-workflows behavior and that matches return preview-sized (truncated) rather than full records, which is important for interpreting results. It does not state rate limits or result-count caps beyond the per-workflow limit.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three compact sentences: purpose and scope first, scoping behavior second, and the follow-up retrieval path last. Every sentence carries information and none repeats the name or schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description appropriately covers what results look like (preview-sized matches) and how to get complete records. For a read-only search with 6 params, the remaining gap is the untold interaction between the workflow and workflow_id params and any result caps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 67%, so the schema already documents q, limit, device, and preview_chars. The description adds conceptual meaning for preview truncation and for workflow scoping, but does not explain how workflow vs workflow_id interact or what default preview size it uses. This is adequate-but-marginal value on top of the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (search) and resource (data collected by workflows), and clarifies the scope ('everything already collected') versus running something new. It explicitly distinguishes itself from the sibling writ_workflow_data, which is used to fetch full records, so an agent can route without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Clearly frames the use case ('answers data questions from past runs without running anything') and names the follow-up tool with the exact trigger ('fetch full records with writ_workflow_data'). It also explains scoping ('Scopes to one workflow when given, else fans out'). It lacks explicit when-not-to-use guidance versus other data siblings like writ_saved_crawl_data or writ_export_data.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

writ_set_scheduleSchedule a workflowA
DestructiveIdempotent
Inspect

Schedule a saved workflow to run automatically. Use every_minutes for an interval, or kind='daily'/'weekly' with time (HH:MM) and days. On a multi-function workflow (an API from writ_website_to_api) functions schedules ONE OR SEVERAL of its functions instead of the whole workflow, with saved inputs for them: each run executes those functions plus whatever they need (sign-in, a token, the search whose ids another reads) exactly once. The answer names also_runs and any missing_inputs to fill.

ParametersJSON Schema
NameRequiredDescriptionDefault
tzNoIANA timezone for daily/weekly.
daysNoISO weekdays for weekly: 1=Mon .. 7=Sun.
kindNointerval | daily | weekly
timeNoHH:MM local time for daily/weekly.
inputsNoScheduled-run inputs {input_name: value} (text, number, true/false), used over the workflow's saved values for scheduled runs only. {} clears them; omit to keep them. Never secrets: those are workflow secrets or a persona.
enabledNoTurn the schedule on/off (default on).
functionNoOne function name (alias of `functions`); "" or "all" = the whole workflow.
workflowNo
functionsNoFunction names a scheduled run calls (from writ_list_workflows / writ_update_workflow). [] or ["all"] = the whole workflow. Omit to keep the current target.
workflow_idNo
every_minutesNoInterval schedule: minutes between runs.

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, idempotentHint=true, and readOnlyHint=false. The description adds behavioral context beyond that: each run executes targeted functions plus dependencies exactly once, and the response names also_runs and missing_inputs. No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core purpose, then parameter modes, then advanced multi-function behavior, then response details. It is dense but each sentence earns its place. Slightly complex phrasing in the third sentence, but overall efficient.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given 11 parameters, nested objects, no output schema, and annotations that already cover safety, the description provides enough context to invoke the tool correctly. It explains scheduling modes, function targeting, dependency resolution, and response hints. Minor details like tz defaults and workflow vs workflow_id selection are left to the schema, which is reasonable at 82% coverage.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is high at 82%, so the baseline is 3. The description adds semantic meaning beyond the schema by clarifying the interplay between functions, inputs, and whole-workflow targeting on multi-function APIs, and by explaining the dependency execution model. It does not fully cover every parameter, but it meaningfully supplements the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Schedule') and resource ('a saved workflow') with automatic recurrence, distinguishing it from writ_run_workflow (manual run) and writ_update_workflow (modify). The added multi-function targeting explanation further narrows its scope. An agent can identify the tool's purpose without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides mode selection guidance ('Use every_minutes for an interval, or kind="daily"/"weekly" with time and days'), which is parameter-level usage. However, it never explicitly says when to choose this tool over alternatives like writ_run_workflow, writ_create_automation, or writ_update_workflow. Usage is implied rather than fully contextualized.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

writ_update_saved_crawlEdit a saved crawlA
DestructiveIdempotent
Inspect

Inspect or customize a saved crawl in place. With no changes, returns its complete definition. settings recursively merges any crawl option into the stored config (scope, paths, budgets, extraction, rendering, persona, residential egress/country, speed, output shape and other flags) without dropping unrelated settings.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameNo
crawlYesSaved crawl slug, name, or id.
settingsNoSparse crawl config patch merged recursively into the complete saved settings.
descriptionNo
default_max_age_secondsNo

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, idempotentHint=true, and readOnlyHint=false. The description adds genuinely useful context beyond them: the settings patch is a recursive merge that preserves unrelated settings, which materially softens how an agent should interpret the destructive hint. It omits auth requirements and whether the merge can be reverted.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with purpose, then the non-obvious merge semantics. No filler or repetition of the title.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation tool with annotations covering safety and no output schema, the description explains the read-vs-edit behavior and merge semantics adequately. It could say more about whether changes require a separate apply/save step or what happens on invalid patch keys.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 40%, but the description compensates for the most complex parameter by explaining that `settings` is a sparse recursive patch over the full stored config. `name`, `description`, and `default_max_age_seconds` are self-evident and need no further prose.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('inspect or customize') and resource ('a saved crawl') and the in-place nature, distinguishing it from writ_saved_crawls (listing) and writ_run_saved_crawl (executing). It is clear but does not explicitly name a sibling to contrast against.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The dual behavior is implied ('With no changes, returns its complete definition' vs. merging changes), which tells an agent it can be used as a read. However, there is no explicit when-to-use guidance, prerequisites, or named alternatives among the many sibling tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

writ_update_workflowEdit a saved workflowA
Destructive
Inspect

Inspect or completely customize an existing saved workflow in place. An edit answers with a compact outline of the steps (verbose=true for the full definition). An api_call step is {type:'api_call', config:{function_name, method, url, headers, body_template, response_extractions}} — function_name binds the step to the function of that name, which is what writ_run_workflow function_name selects. With no edit arguments, returns the full workflow and its zero-based step indexes. patch changes workflow settings including persona, residential egress/country, human behavior, headless/fast mode, device routing, timeouts, retries, schedules, auth/browser config, functions, streaming, sessions and raw replay. step_updates recursively edits one or more recorded steps (selectors, scripts, URLs, waits, flags, extraction config, and generic api_call.config.flow programs); replace_steps deliberately replaces the complete step list. AGENT SKILL: skill_md (in the read) is the SKILL.md that teaches an agent to call this workflow over MCP; regenerate_skill=true rebuilds it from the current functions, then patch.skill_md writes your edited version (YAML frontmatter name + description, then Markdown); patch.skill_md=null removes it.

ParametersJSON Schema
NameRequiredDescriptionDefault
patchNoSparse settings patch. Supports every WorkflowUpdate field except credentials and captured recorded_session material; use vault/persona/session management for those.
verboseNoAnswer an edit with the full workflow definition instead of the compact step outline (default false).
workflowNoWorkflow name (or use workflow_id).
workflow_idNo
step_updatesNo
replace_stepsNoComplete recorded-step replacement. Prefer step_updates for a targeted edit.
regenerate_skillNoRebuild the workflow's agent skill (SKILL.md) from its current functions, inputs and sign-in, replacing any edited version; the answer carries the new skill_md. Runs after any patch in the same call.

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already mark destructiveHint=true and readOnlyHint=false, so the description correctly implies mutation and the ability to replace steps. The description adds critical context: 'replace_steps deliberately replaces the complete step list', and details on patch fields and skill regeneration. It doesn't elaborate on permissions or side effects like data loss, but with annotations covering the safety profile, this is strong.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense and front-loads the core purpose, but it is quite long and includes detailed technical examples (api_call structure) that could be moved to a separate section. Some sentences are lengthy and could be split for clarity, but overall it is informative without excessive fluff.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (7 parameters, nested objects, many patch fields), the description covers key behaviors: read vs edit, verbose mode, step_updates vs replace_steps, skill regeneration, and patch scope. It lacks explicit prerequisites (e.g., authentication needs) and error handling, but with no output schema and schema at 71%, it is largely complete for an agent to call correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 71%, so there is some gap. The description adds meaning beyond schema: it explains that patch changes settings including a long list of specific fields; that step_updates recursively edits steps with examples (selectors, scripts, URLs, waits, flags, extraction config); and clarifies the api_call step structure with config fields. This compensates for the schema gap effectively, though not exhaustively.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Inspect or completely customize') and resource ('existing saved workflow in place'), and differentiates itself from siblings by naming writ_run_workflow function_name binding. An agent can tell this is the edit tool, distinct from writ_list_workflows or writ_run_workflow.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides clear context: 'With no edit arguments, returns the full workflow', 'Prefer step_updates for a targeted edit', and explains replace_steps vs step_updates. However, it doesn't explicitly state when-not-to-use (e.g., for creation, use a different tool) or list alternatives beyond replace_steps.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

writ_website_to_apiTurn a website into an APIAInspect

TURN A WEBSITE INTO A CALLABLE API — the one tool for this, every lane. Use it whenever a service has no official/practical API but the user wants its data or actions programmatically: "turn into an API", "map the API of ", "expose every feature", "give me an endpoint for ". THE WHOLE JOB IS 3 CALLS: (1) this tool with url + goal. START ON THE PAGE THAT ALREADY SHOWS THE ROWS (the search-results / category / listing URL, e.g. https://www.google.com/maps/search/bakeries+Montreal/ — NOT the app's home page: a build seeded at an empty shell spent 6 minutes over three rungs and produced no function), and name the inputs and the fields wanted ("page number in; quotes with text/author/tags and has_next out") — + save_as; (2) writ_discovery_status(build_id, wait=true): ONE held call that follows every rung; (3) on succeeded, run it exactly as the answer's run_example shows (writ_run_workflow: workflow_id + function_name + inputs), with TWO different inputs, and check the answers differ — then report. Do not open a browser or start a second build for the site meanwhile. An answer of existing_workflows / marketplace_candidates is a PROPOSAL: run the match, or call again with skip_existing / skip_marketplace for a fresh build. status needs_guidance = the build is YOURS: call again with mode=guided build_id= (you are the brain of that browser). An empty run → writ_diagnose_http_workflow(workflow_id, task_id). LOGIN: if the app is behind a sign-in, ASK THE USER which saved identity to use (writ_personas) and pass its persona_id — never guess or type credentials; a persona also carries 2FA. NOT FOR: reading a page's content (writ_scrape), collecting a site as a dataset (writ_crawl_site), or a task that is not an API surface (writ_record_website). DEFAULT = intelligent: Writ runs the WHOLE cost ladder for you, cheapest rung first, and you only start it and wait. The ladder: the user's OWN matching workflows (answered as existing_workflows — propose replaying those; skip_existing=true to bypass), ready-made MARKETPLACE APIs (marketplace_candidates; skip_marketplace=true), a STATIC HTTP crawl (forms, search boxes and query links become functions with inputs; inline JS; OpenAPI/Swagger specs; server-rendered listings), then a RENDERED crawl for JS/SPA pages, then Writ's AI BROWSER rung (its discovery brain drives a browser, ranks data-bearing traffic, promotes the site's own HTTP requests, tests inputs and pagination, saves the workflow). A rung that proves enough ENDS the build there; one that does not escalates, and the status of the newer rung carries escalations: why each cheaper rung handed over (robots.txt refused the crawl, no pages fetched, nothing matched the goal, no structured list...). Read it before telling the user why a browser was needed. A crawl-rung result is UNVERIFIED (verified:false) until a real run proves it — say so. ROBOTS: the crawl rungs obey the site's robots.txt by default. A site that disallows the target (many search paths; some whole hosts) admits zero pages, so both crawl rungs end at once and the build goes to the AI browser. When escalations names robots.txt and the user vouches for the target, call again with respect_robots=false. mode=auto is the same ladder with YOU as the last rung: the crawl rungs run, and when they do not prove enough the build PARKS as status=needs_guidance with map (every endpoint seen, specs, candidate functions) instead of spending Writ's agent. Continue it — or start directly — with mode=guided (and build_id=). That opens a real browser BOUND TO THE BUILD that you drive turn by turn (writ_browser_act: navigate, sign in, capture_network, evaluate_js, read calls with writ_browser_network), on which you DEFINE the API (writ_browser_compose define_function — api functions from captured calls via from_index, or proven scripts/extractions; each is live-tested as you define it; set_inputs for parameters; is_auth for a sign-in function whose response_extractions feed the others). writ_browser_save settles the build: the workflow is callable at once (writ_run_workflow with function_name), pinnable, schedulable, exposable as REST (writ_expose_workflow_api), and its API docs are at GET /api/v1/workflows/{workflow_id}/api-docs. HTTP-FIRST GATE: this browser is the experiment bench. Capture a representative search/filter and next-page request, then define direct API functions. Use the Auphan-style named function graph by default: ordered is_auth functions publish tokens/ids/origins through response_extractions and data functions consume {{extracted:name}}. Use config.flow only for loops, recursive mapping, cross-page dedupe or composite returns. Typed extraction sources are json, embedded_json, html_css, regex, header and body. Expose search/filter/limit/page/offset/cursor as declared inputs and return next_cursor/next_offset/has_more. After save, run with the intended persona and call writ_diagnose_http_workflow(task_id=...). Do not expose until engine=http returns non-empty data and pagination matches the browser baseline, unless you can name a measured browser-only dependency. Pass mode=guided to drive the browser yourself; mode=fast / mode=browser START on that crawl rung with you as the driver (it parks as needs_guidance when it falls short, exactly like auto).

ParametersJSON Schema
NameRequiredDescriptionDefault
urlNoThe site's entry/home URL (required unless build_id continues a parked build).
goalNoWhat the API should return or do, in plain language. Matches your own workflows and marketplace listings first, and NARROWS the crawl to what was asked (equivalent to `scope`) instead of mapping the whole app.
modeNointelligent (DEFAULT): the full ladder driven by Writ — static crawl, then rendered crawl, then Writ's own AI browser rung, stopping at the first rung that proves enough. Start it and wait. guided: open a browser bound to a build that YOU drive and compose — sign in, capture_network, define_function, save; with build_id it continues a parked build and inherits its map. 'auto': the same crawl rungs with YOU as the last rung — own workflows and marketplace proposed first, then a static crawl escalating to a rendered crawl; not enough => the build parks as needs_guidance for you to finish guided. 'fast' = start on the static crawl (seconds, no browser, UNVERIFIED). 'browser' = start on the rendered crawl. For compatibility, regular/deep remain aliases of guided.
levelNoWrite policy for the crawl rungs. 'light' (default) captures every endpoint and payload but NEVER performs a real create/update/delete. 'deep' performs each write once to capture its real confirmation response — it CHANGES real data, so only use it when the user explicitly asks.
scopeNoCrawl lanes: map ONLY this surface (e.g. 'employees') and what it depends on. Omit to map the whole app.
deviceNoA linked Writ desktop's agent_id (writ_devices): act ON it. Omit to use the desktop this connection chose with writ_devices action='use' (if any).
save_asNoName for the workflow the build saves.
build_idNoContinue a parked build (status needs_guidance from writ_discovery_status) on the guided rung: opens the browser bound to it, seeded with its map.
anonymousNoBuild from what is visible WITHOUT an account even though the site shows a sign-in page — only when the user said the public part is enough.
persona_idNoSaved identity to sign in with (writ_personas). Without one, a site whose entry page IS a sign-in wall is not built: the answer names the persona to pass, or returns `persona_needed` — relay its tell_user + create_url to the user, then call again with the new persona_id. Required for 2FA — the code is minted server-side and never shown to you.
ai_superviseNoAI-supervised crawl rungs (default true): after the crawl mines forms, POSTs, query links, scripts and listings, one bounded Writ AI call authors the API from them — which functions serve the goal, their names, inputs and example values; the live verify call then measures each response shape. false = the purely mechanical surface map (no AI spend on the crawl rungs).
skip_existingNoSkip the proposal of the user's OWN matching workflows (set after they declined).
respect_robotsNoCrawl rungs obey the site's robots.txt (default true). Pass false ONLY when the user vouches for the target and a rung reported that robots.txt refused the crawl (`escalations` / `message` name the rule) — otherwise the crawl rungs fetch nothing and the build goes straight to a browser. Does not apply to the AI or guided browser rungs.
use_residentialNoRun on the platform residential network (premium) for a site that blocks datacenter IPs. Default off. Continuing a build (build_id) keeps the build's own persona, residential exit and country — pass these only to change them.
execution_targetNo'cloud' (the fleet) or a linked desktop's agent_id: build there, in its own browser and connection. Omitted = the desktop chosen with writ_devices, else the cloud.
skip_marketplaceNoSkip the ready-made marketplace proposals and build fresh.
residential_countryNoTwo-letter ISO country the residential exit should be in (e.g. 'us', 'fr') — applies to every rung of the build, the guided browser included. Omit for an automatic exit. Ignored unless the session egresses residential.

TDQS

A3.6/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description adds extensive behavioral context beyond annotations, such as the cost ladder, escalations, robots handling, verification status, persona login, and parked builds. However, it states that level=deep performs real create/update/delete writes and 'CHANGES real data,' while the annotations declare destructiveHint=false. That is a direct annotation contradiction, so per the rubric this dimension scores 1.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The opening is front-loaded and identifies the core job quickly, but the description becomes an all-caps wall of text that repeats mode and parameter details already covered by the schema. For a tool whose parameter documentation is already complete, this level of redundancy and density is not concise.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex 17-parameter orchestration tool with no output schema, the description covers selection, the full call flow, modes, authentication, robots restrictions, verification caveats, and follow-up tools. Despite the verbosity, an agent has enough context to start, continue, and recover a build.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all 17 parameters, making 3 the baseline. The description adds useful cross-parameter context—build_id continues a parked build, persona_id handles sign-in/2FA, respect_robots ties to escalations—but much of the mode and level explanation duplicates the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource: turning a website into a callable API. It explicitly distinguishes itself from sibling tools by naming writ_scrape, writ_crawl_site, and writ_record_website as NOT FOR cases. An agent can identify the tool's core job without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit when-to-use guidance, the intended three-call workflow, the follow-up status/run/diagnose tools, and clear exclusions. It also explains which mode to choose and when to escalate, leaving little to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

writ_wire_monitorConnect a monitor to an actionAInspect

Wire a monitor's change_detected event to an action. action='workflow' runs a saved workflow when the monitored page changes; action='notify' sends a notification (provide channels + recipients the account has configured); action='ai_task' WAKES AN AI AGENT with a task prompt — the agent opens the monitored page in a cloud browser, sees what changed (diff + extracted values) and works the prompt autonomously (add channels/recipients to also get notified when it finishes). Use after writ_create_monitor to make the monitor DO something on change. A PRICE watch takes threshold: the action then runs ONCE, when a check reads a price at or below it (threshold_op picks the side) — not on every change. AUTO-BUY: action='workflow' with a checkout workflow and buy (payment {kind, ref} from writ_payment, approval, spend cap) makes it a purchase, always saved as a rehearsal (dry run) that stops before the order is placed; the user turns real purchases on in the Writ app.

ParametersJSON Schema
NameRequiredDescriptionDefault
buyNoaction='workflow' only: the workflow is a CHECKOUT and runs as a purchase. Always saved with dry_run on (a rehearsal that stops before placing the order), whatever is sent; the user arms real purchases in the Writ app. Cloud monitors only.
nameNoOptional automation name.
titleNo
actionYesWhat to do on a detected change.
deviceNoA linked Writ desktop's agent_id (writ_devices): act ON it. Omit to use the desktop this connection chose with writ_devices action='use' (if any).
promptNoaction='ai_task': what the agent should do when the monitor fires, e.g. "Check whether the price dropped below $500 and summarize what changed". Supports {{placeholders}} like {{diff_snippet}} and {{extracted.price}}.
enabledNo
messageNoNotification body template (supports {{event.url}}).
channelsNoNotification channels, e.g. ["pushover","email"] — required for action='notify'; optional with action='ai_task' (finish alert).
workflowNoWorkflow to run (name) — required for action='workflow'.
entry_urlNoaction='ai_task': page the agent starts on (defaults to the monitored URL).
max_stepsNoaction='ai_task': cap on agent steps per wake (default 20, max 100).
thresholdNoPrice watch: act only when the watched price reaches this number (e.g. 477.04), once; by default at or below it (see `threshold_op`), and a price on the other side, or one the page no longer shows, does nothing. The monitor must read the price (a selector or extract, not a visual zone). Omit it for an alert on any change. Works for a desktop monitor too (`device`).
monitor_idYesMonitor id from writ_create_monitor.
recipientsNoNotification recipients, e.g. ["pushover:1"]. OMIT to reach EVERY enabled recipient on the channel — the answer names who the alert actually reaches, and warns when nobody is configured.
workflow_idNo
threshold_opNoWhich side of `threshold` fires: lte = at or below (default: "drops to 477.04" fires at 477.04), lt = strictly below ("below / under / less than"), gte = at or above, gt = strictly above (a rise alert). Pass what the user's words say.
ai_session_idNoaction='ai_task': re-run this saved AI session instead of (or as well as) a prompt.
cooldown_minutesNoaction='ai_task': minimum minutes between wakes (default 10; 0 disables).

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare the safety profile (readOnly=false, destructive=false, openWorld=true); the description goes well beyond, disclosing that auto-buy is always saved as a dry-run rehearsal until armed in the app, that price thresholds fire once and require a selector/extract, and that ai_task wakes an autonomous agent. These are material behavioral facts not available in annotations or schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core purpose, then mode behaviors, price watch, and auto-buy in a logical order with UPPERCASE emphasis on the riskiest mode. Dense but every clause is load-bearing; slightly long for a single description but not padded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 19-parameter, nested-object tool with no output schema, the description covers the routing logic across action modes, the price-watch constraint, and the auto-buy safety model, which is what an agent needs to invoke it correctly. Nothing critical is left unexplained.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 84% schema coverage the baseline is 3, but the description adds real semantics: how action modes partition which params apply (channels/recipients for notify, prompt/workflow for others), the threshold single-fire semantics tied to threshold_op, and the buy/spend_cap defaulting behavior. This meaningfully exceeds the schema text.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Wire a monitor's change_detected event to an action') and immediately enumerates the three concrete action modes with their distinct behaviors. It is clearly distinguished from writ_create_monitor, which is named as the prerequisite sibling.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly says 'Use after writ_create_monitor to make the monitor DO something on change,' and gives per-mode usage conditions (workflow vs notify vs ai_task), plus the price-watch single-fire rule and the auto-buy rehearsal rule. Alternatives and exclusions are stated, not inferred.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

writ_workflow_dataRead a workflow's collected dataA
Read-onlyIdempotent
Inspect

Read the accumulated extracted data for a saved WORKFLOW as a table (columns + rows). Filter with q, or inspect one run with run_id. Long text cells arrive preview-cut (_truncated lists the fields) — hydrate full records via refs. NOTE: crawl ids are a different namespace — a writ_crawl_site run's data lives behind writ_crawl_status / writ_saved_crawl_data, not here.

ParametersJSON Schema
NameRequiredDescriptionDefault
qNoSubstring filter across fields.
refsNoHydration: fetch FULL untruncated records by ref '<run_id>:<record_index>' (both fields are on every row). Max 100.
viewNoall | latest | run
limitNo
deviceNoA linked Writ desktop's agent_id (writ_devices): act ON it. Omit to use the desktop this connection chose with writ_devices action='use' (if any).
outputNoRESPONSE SHAPE — set this whenever the answer is for a program or an API you are building, not for you to read. {shape: 'envelope' (default: Writ's full answer, projected) | 'table' ({columns, rows, total}) | 'records' (bare list of records) | 'record' (the newest record alone — one entity, a usage meter, a dashboard), fields: ['used', 'percent_used as pct', 'items.0.price as first_price'] (ordered pick, renames, dotted paths; missing → null so keys are stable), exclude: ['depth'], include_meta: false (page metadata content_kind/depth/thumbnails are STRIPPED unless true), key: 'usage' (wrap)}. On writ_crawl_site with save_as it is SAVED as the API's default shape.
run_idNo
workflowNo
workflow_idNo
preview_charsNoCut string cells to this many characters (default 2000; 0 = full cells).

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly/idempotent/non-destructive, and the description adds non-obvious behavior beyond them: text cells arrive preview-cut, `_truncated` lists the cut fields, and full records are hydrated via refs. It omits pagination/limit defaults and what `view` modes do, but the truncation-and-hydration disclosure is genuinely additive.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core action, then filtering, then truncation/hydration behavior, closing with the namespace caveat. Every sentence carries load-bearing information and there is no restating of the title or schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 10-param tool with no output schema and a nested output object, the description covers the essential read path, truncation semantics, and cross-namespace disambiguation well. Gaps remain around the view/limit/device parameters and the saved-workflow concept, but nothing blocks a correct call.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 60%, and the description meaningfully explains q, run_id and the refs hydration path (equivalent to '<run_id>:<record_index>'). However several parameters (view, limit, device, workflow/workflow_id) get no clarification beyond the schema, so it only partially compensates for the coverage gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Read the accumulated extracted data for a saved WORKFLOW as a table') and describes the output shape (columns + rows). The NOTE explicitly differentiates this from the crawl-data siblings (writ_crawl_status / writ_saved_crawl_data), so an agent can route without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives concrete routing guidance: 'Filter with `q`, or inspect one run with `run_id`', and the NOTE names the alternatives to use for crawl data with the condition that selects them. It stops short of contrasting with other plausible siblings like writ_search_data or writ_export_data.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

writ_workflow_runsList workflow runsB
Read-onlyIdempotent
Inspect

Inspect run history — status, timing, errors — for one workflow or across all of them.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
deviceNoA linked Writ desktop's agent_id (writ_devices): act ON it. Omit to use the desktop this connection chose with writ_devices action='use' (if any).
statusNopending|running|success|failed|cancelled|skipped
workflowNo
workflow_idNo

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false, so the safety profile is clear. The description adds that it surfaces status, timing, and errors, which is useful context beyond the annotations. However, it doesn't disclose pagination, default limits, or ordering, which are relevant for a list endpoint.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Single sentence, front-loaded with the action and resource, with a clear scope clause. Efficient, though it could be more informative without waste.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given 5 parameters, no output schema, and moderate annotation richness, the description is minimally adequate. It tells what the tool returns conceptually but doesn't cover parameter usage or behavioral details like pagination, which an agent needs to invoke correctly. It's not misleading, but it leaves gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 40%, so the description should compensate but does not: it mentions status, timing, and errors but not the parameters limit, device, status, workflow, or workflow_id. The schema itself provides descriptions for device and status, leaving limit, workflow, and workflow_id undocumented in both schema and description. Baseline 3 when schema partially covers; here it's slightly below because description adds almost no parameter meaning.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Clearly states the verb and resource: inspecting run history for workflows. It specifies scope ('one workflow or across all of them') and names the returned attributes (status, timing, errors). It doesn't explicitly differentiate from siblings like writ_list_workflows or writ_workflow_data, but the resource (runs vs workflows vs data) is distinct enough.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this vs alternatives. The sibling writ_list_workflows lists workflows and writ_workflow_data retrieves workflow data; it's unclear when an agent should call this tool instead. The description implies it's for inspecting run history but doesn't mention exclusions or companions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 39 tool updatesv1.1.0
    • First observedwrit_browser_act
    • First observedwrit_browser_ask_user
    • First observedwrit_browser_cancel
    • First observedwrit_browser_compose
    • First observedwrit_browser_context
    • First observedwrit_browser_network
    • First observedwrit_browser_save
    • First observedwrit_browser_sessions
    • First observedwrit_browser_use
    • First observedwrit_crawl_files
    • First observedwrit_crawl_site
    • First observedwrit_crawl_status
    • First observedwrit_create_automation
    • First observedwrit_create_http_extraction
    • First observedwrit_create_monitor
    • First observedwrit_devices
    • First observedwrit_diagnose_http_workflow
    • First observedwrit_discovery_status
    • First observedwrit_export_data
    • First observedwrit_expose_workflow_api
    • First observedwrit_list_webhooks
    • First observedwrit_list_workflows
    • First observedwrit_payment
    • First observedwrit_personas
    • First observedwrit_pin_workflow_tool
    • First observedwrit_record_website
    • First observedwrit_run_saved_crawl
    • First observedwrit_run_workflow
    • First observedwrit_saved_crawl_data
    • First observedwrit_saved_crawls
    • First observedwrit_scrape
    • First observedwrit_search_data
    • First observedwrit_set_schedule
    • First observedwrit_update_saved_crawl
    • First observedwrit_update_workflow
    • First observedwrit_website_to_api
    • First observedwrit_wire_monitor
    • First observedwrit_workflow_data
    • First observedwrit_workflow_runs

TDQS

A3.7/5.0

Scored across 39 tools

Disambiguation3/5

The core domains (scrape, crawl, browser task, API build, workflows, data, monitors, automation) are mostly distinct and the descriptions use explicit NOT FOR cues. However, several authoring/build/debug tools overlap in practice: writ_website_to_api vs writ_record_website, writ_browser_compose vs writ_browser_act define_function/compose, and writ_create_http_extraction vs writ_update_workflow can be confused, forcing careful reading.

Naming Consistency4/5

All tools use the writ_ snake_case prefix, and most follow a predictable verb_noun pattern (list_workflows, run_workflow, create_monitor, crawl_site). Minor deviations are noun-phrase resources (writ_personas, writ_payment, writ_devices, writ_saved_crawls) and the prepositional writ_website_to_api.

Tool Count2/5

39 tools is very large for an MCP surface; many are sub-actions of longer browser/crawl/workflow workflows that could be grouped. While the platform is broad, the count far exceeds the typical 3-15 well-scoped range and increases selection and context burden.

Completeness4/5

Coverage is strong across scraping, crawling, browser automation, API building, saved workflows, data access, scheduling, monitors, automations, personas and payments. Some lifecycle gaps remain, such as no obvious delete/unwire paths for several resources and monitor listing only indirectly via writ_devices action='monitors', but most gaps are workaroundable.

Maintenance

ActivitySlowing
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers