Turn a website into an API
writ_website_to_apiTurn a website that lacks an official API into callable functions: map its pages, search and forms into endpoints with typed inputs, so you can fetch data or automate actions programmatically.
Instructions
TURN A WEBSITE INTO A CALLABLE API — the one tool for this, every lane. Use it whenever a service has no official/practical API but the user wants its data or actions programmatically: "turn into an API", "map the API of ", "expose every feature", "give me an endpoint for ".
THE WHOLE JOB IS 3 CALLS: (1) this tool with url + goal. START ON THE PAGE THAT ALREADY SHOWS THE ROWS (the search-results / category / listing URL, e.g. https://www.google.com/maps/search/bakeries+Montreal/ — NOT the app's home page: a build seeded at an empty shell spent 6 minutes over three rungs and produced no function), and name the inputs and the fields wanted ("page number in; quotes with text/author/tags and has_next out") — + save_as; (2) writ_discovery_status(build_id, wait=true): ONE held call that follows every rung; (3) on succeeded, run it exactly as the answer's run_example shows (writ_run_workflow: workflow_id + function_name + inputs), with TWO different inputs, and check the answers differ — then report. Do not open a browser or start a second build for the site meanwhile. An answer of existing_workflows / marketplace_candidates is a PROPOSAL: run the match, or call again with skip_existing / skip_marketplace for a fresh build. status needs_guidance = the build is YOURS: call again with mode=guided build_id= (you are the brain of that browser). An empty run → writ_diagnose_http_workflow(workflow_id, task_id).
LOGIN: if the app is behind a sign-in, ASK THE USER which saved identity to use (writ_personas) and pass its persona_id — never guess or type credentials; a persona also carries 2FA.
NOT FOR: reading a page's content (writ_scrape), collecting a site as a dataset (writ_crawl_site), or a task that is not an API surface (writ_record_website).
DEFAULT = intelligent: Writ runs the WHOLE cost ladder for you, cheapest rung first, and you only start it and wait. The ladder: the user's OWN matching workflows (answered as existing_workflows — propose replaying those; skip_existing=true to bypass), ready-made MARKETPLACE APIs (marketplace_candidates; skip_marketplace=true), a STATIC HTTP crawl (forms, search boxes and query links become functions with inputs; inline JS; OpenAPI/Swagger specs; server-rendered listings), then a RENDERED crawl for JS/SPA pages, then Writ's AI BROWSER rung (its discovery brain drives a browser, ranks data-bearing traffic, promotes the site's own HTTP requests, tests inputs and pagination, saves the workflow). A rung that proves enough ENDS the build there; one that does not escalates, and the status of the newer rung carries escalations: why each cheaper rung handed over (robots.txt refused the crawl, no pages fetched, nothing matched the goal, no structured list...). Read it before telling the user why a browser was needed. A crawl-rung result is UNVERIFIED (verified:false) until a real run proves it — say so.
ROBOTS: the crawl rungs obey the site's robots.txt by default. A site that disallows the target (many search paths; some whole hosts) admits zero pages, so both crawl rungs end at once and the build goes to the AI browser. When escalations names robots.txt and the user vouches for the target, call again with respect_robots=false.
mode=auto is the same ladder with YOU as the last rung: the crawl rungs run, and when they do not prove enough the build PARKS as status=needs_guidance with map (every endpoint seen, specs, candidate functions) instead of spending Writ's agent. Continue it — or start directly — with mode=guided (and build_id=). That opens a real browser BOUND TO THE BUILD that you drive turn by turn (writ_browser_act: navigate, sign in, capture_network, evaluate_js, read calls with writ_browser_network), on which you DEFINE the API (writ_browser_compose define_function — api functions from captured calls via from_index, or proven scripts/extractions; each is live-tested as you define it; set_inputs for parameters; is_auth for a sign-in function whose response_extractions feed the others). writ_browser_save settles the build: the workflow is callable at once (writ_run_workflow with function_name), pinnable, schedulable, exposable as REST (writ_expose_workflow_api), and its API docs are at GET /api/v1/workflows/{workflow_id}/api-docs.
HTTP-FIRST GATE: this browser is the experiment bench. Capture a representative search/filter and next-page request, then define direct API functions. Use the Auphan-style named function graph by default: ordered is_auth functions publish tokens/ids/origins through response_extractions and data functions consume {{extracted:name}}. Use config.flow only for loops, recursive mapping, cross-page dedupe or composite returns. Typed extraction sources are json, embedded_json, html_css, regex, header and body. Expose search/filter/limit/page/offset/cursor as declared inputs and return next_cursor/next_offset/has_more. After save, run with the intended persona and call writ_diagnose_http_workflow(task_id=...). Do not expose until engine=http returns non-empty data and pagination matches the browser baseline, unless you can name a measured browser-only dependency.
Pass mode=guided to drive the browser yourself; mode=fast / mode=browser START on that crawl rung with you as the driver (it parks as needs_guidance when it falls short, exactly like auto).
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| url | No | The site's entry/home URL (required unless build_id continues a parked build). | |
| goal | No | What the API should return or do, in plain language. Matches your own workflows and marketplace listings first, and NARROWS the crawl to what was asked (equivalent to `scope`) instead of mapping the whole app. | |
| mode | No | intelligent (DEFAULT): the full ladder driven by Writ — static crawl, then rendered crawl, then Writ's own AI browser rung, stopping at the first rung that proves enough. Start it and wait. guided: open a browser bound to a build that YOU drive and compose — sign in, capture_network, define_function, save; with build_id it continues a parked build and inherits its map. 'auto': the same crawl rungs with YOU as the last rung — own workflows and marketplace proposed first, then a static crawl escalating to a rendered crawl; not enough => the build parks as needs_guidance for you to finish guided. 'fast' = start on the static crawl (seconds, no browser, UNVERIFIED). 'browser' = start on the rendered crawl. For compatibility, regular/deep remain aliases of guided. | |
| level | No | Write policy for the crawl rungs. 'light' (default) captures every endpoint and payload but NEVER performs a real create/update/delete. 'deep' performs each write once to capture its real confirmation response — it CHANGES real data, so only use it when the user explicitly asks. | |
| scope | No | Crawl lanes: map ONLY this surface (e.g. 'employees') and what it depends on. Omit to map the whole app. | |
| device | No | A linked Writ desktop's agent_id (writ_devices): act ON it. Omit to use the desktop this connection chose with writ_devices action='use' (if any). | |
| save_as | No | Name for the workflow the build saves. | |
| build_id | No | Continue a parked build (status needs_guidance from writ_discovery_status) on the guided rung: opens the browser bound to it, seeded with its map. | |
| anonymous | No | Build from what is visible WITHOUT an account even though the site shows a sign-in page — only when the user said the public part is enough. | |
| persona_id | No | Saved identity to sign in with (writ_personas). Without one, a site whose entry page IS a sign-in wall is not built: the answer names the persona to pass, or returns `persona_needed` — relay its tell_user + create_url to the user, then call again with the new persona_id. Required for 2FA — the code is minted server-side and never shown to you. | |
| ai_supervise | No | AI-supervised crawl rungs (default true): after the crawl mines forms, POSTs, query links, scripts and listings, one bounded Writ AI call authors the API from them — which functions serve the goal, their names, inputs and example values; the live verify call then measures each response shape. false = the purely mechanical surface map (no AI spend on the crawl rungs). | |
| skip_existing | No | Skip the proposal of the user's OWN matching workflows (set after they declined). | |
| respect_robots | No | Crawl rungs obey the site's robots.txt (default true). Pass false ONLY when the user vouches for the target and a rung reported that robots.txt refused the crawl (`escalations` / `message` name the rule) — otherwise the crawl rungs fetch nothing and the build goes straight to a browser. Does not apply to the AI or guided browser rungs. | |
| use_residential | No | Run on the platform residential network (premium) for a site that blocks datacenter IPs. Default off. Continuing a build (build_id) keeps the build's own persona, residential exit and country — pass these only to change them. | |
| execution_target | No | 'cloud' (the fleet) or a linked desktop's agent_id: build there, in its own browser and connection. Omitted = the desktop chosen with writ_devices, else the cloud. | |
| skip_marketplace | No | Skip the ready-made marketplace proposals and build fresh. | |
| residential_country | No | Two-letter ISO country the residential exit should be in (e.g. 'us', 'fr') — applies to every rung of the build, the guided browser included. Omit for an automatic exit. Ignored unless the session egresses residential. |