Skip to main content
Glama

Turn a website into an API

writ_website_to_api

Turn a website that lacks an official API into callable functions: map its pages, search and forms into endpoints with typed inputs, so you can fetch data or automate actions programmatically.

Instructions

TURN A WEBSITE INTO A CALLABLE API — the one tool for this, every lane. Use it whenever a service has no official/practical API but the user wants its data or actions programmatically: "turn into an API", "map the API of ", "expose every feature", "give me an endpoint for ". THE WHOLE JOB IS 3 CALLS: (1) this tool with url + goal. START ON THE PAGE THAT ALREADY SHOWS THE ROWS (the search-results / category / listing URL, e.g. https://www.google.com/maps/search/bakeries+Montreal/ — NOT the app's home page: a build seeded at an empty shell spent 6 minutes over three rungs and produced no function), and name the inputs and the fields wanted ("page number in; quotes with text/author/tags and has_next out") — + save_as; (2) writ_discovery_status(build_id, wait=true): ONE held call that follows every rung; (3) on succeeded, run it exactly as the answer's run_example shows (writ_run_workflow: workflow_id + function_name + inputs), with TWO different inputs, and check the answers differ — then report. Do not open a browser or start a second build for the site meanwhile. An answer of existing_workflows / marketplace_candidates is a PROPOSAL: run the match, or call again with skip_existing / skip_marketplace for a fresh build. status needs_guidance = the build is YOURS: call again with mode=guided build_id= (you are the brain of that browser). An empty run → writ_diagnose_http_workflow(workflow_id, task_id). LOGIN: if the app is behind a sign-in, ASK THE USER which saved identity to use (writ_personas) and pass its persona_id — never guess or type credentials; a persona also carries 2FA. NOT FOR: reading a page's content (writ_scrape), collecting a site as a dataset (writ_crawl_site), or a task that is not an API surface (writ_record_website). DEFAULT = intelligent: Writ runs the WHOLE cost ladder for you, cheapest rung first, and you only start it and wait. The ladder: the user's OWN matching workflows (answered as existing_workflows — propose replaying those; skip_existing=true to bypass), ready-made MARKETPLACE APIs (marketplace_candidates; skip_marketplace=true), a STATIC HTTP crawl (forms, search boxes and query links become functions with inputs; inline JS; OpenAPI/Swagger specs; server-rendered listings), then a RENDERED crawl for JS/SPA pages, then Writ's AI BROWSER rung (its discovery brain drives a browser, ranks data-bearing traffic, promotes the site's own HTTP requests, tests inputs and pagination, saves the workflow). A rung that proves enough ENDS the build there; one that does not escalates, and the status of the newer rung carries escalations: why each cheaper rung handed over (robots.txt refused the crawl, no pages fetched, nothing matched the goal, no structured list...). Read it before telling the user why a browser was needed. A crawl-rung result is UNVERIFIED (verified:false) until a real run proves it — say so. ROBOTS: the crawl rungs obey the site's robots.txt by default. A site that disallows the target (many search paths; some whole hosts) admits zero pages, so both crawl rungs end at once and the build goes to the AI browser. When escalations names robots.txt and the user vouches for the target, call again with respect_robots=false. mode=auto is the same ladder with YOU as the last rung: the crawl rungs run, and when they do not prove enough the build PARKS as status=needs_guidance with map (every endpoint seen, specs, candidate functions) instead of spending Writ's agent. Continue it — or start directly — with mode=guided (and build_id=). That opens a real browser BOUND TO THE BUILD that you drive turn by turn (writ_browser_act: navigate, sign in, capture_network, evaluate_js, read calls with writ_browser_network), on which you DEFINE the API (writ_browser_compose define_function — api functions from captured calls via from_index, or proven scripts/extractions; each is live-tested as you define it; set_inputs for parameters; is_auth for a sign-in function whose response_extractions feed the others). writ_browser_save settles the build: the workflow is callable at once (writ_run_workflow with function_name), pinnable, schedulable, exposable as REST (writ_expose_workflow_api), and its API docs are at GET /api/v1/workflows/{workflow_id}/api-docs. HTTP-FIRST GATE: this browser is the experiment bench. Capture a representative search/filter and next-page request, then define direct API functions. Use the Auphan-style named function graph by default: ordered is_auth functions publish tokens/ids/origins through response_extractions and data functions consume {{extracted:name}}. Use config.flow only for loops, recursive mapping, cross-page dedupe or composite returns. Typed extraction sources are json, embedded_json, html_css, regex, header and body. Expose search/filter/limit/page/offset/cursor as declared inputs and return next_cursor/next_offset/has_more. After save, run with the intended persona and call writ_diagnose_http_workflow(task_id=...). Do not expose until engine=http returns non-empty data and pagination matches the browser baseline, unless you can name a measured browser-only dependency. Pass mode=guided to drive the browser yourself; mode=fast / mode=browser START on that crawl rung with you as the driver (it parks as needs_guidance when it falls short, exactly like auto).

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
urlNoThe site's entry/home URL (required unless build_id continues a parked build).
goalNoWhat the API should return or do, in plain language. Matches your own workflows and marketplace listings first, and NARROWS the crawl to what was asked (equivalent to `scope`) instead of mapping the whole app.
modeNointelligent (DEFAULT): the full ladder driven by Writ — static crawl, then rendered crawl, then Writ's own AI browser rung, stopping at the first rung that proves enough. Start it and wait. guided: open a browser bound to a build that YOU drive and compose — sign in, capture_network, define_function, save; with build_id it continues a parked build and inherits its map. 'auto': the same crawl rungs with YOU as the last rung — own workflows and marketplace proposed first, then a static crawl escalating to a rendered crawl; not enough => the build parks as needs_guidance for you to finish guided. 'fast' = start on the static crawl (seconds, no browser, UNVERIFIED). 'browser' = start on the rendered crawl. For compatibility, regular/deep remain aliases of guided.
levelNoWrite policy for the crawl rungs. 'light' (default) captures every endpoint and payload but NEVER performs a real create/update/delete. 'deep' performs each write once to capture its real confirmation response — it CHANGES real data, so only use it when the user explicitly asks.
scopeNoCrawl lanes: map ONLY this surface (e.g. 'employees') and what it depends on. Omit to map the whole app.
deviceNoA linked Writ desktop's agent_id (writ_devices): act ON it. Omit to use the desktop this connection chose with writ_devices action='use' (if any).
save_asNoName for the workflow the build saves.
build_idNoContinue a parked build (status needs_guidance from writ_discovery_status) on the guided rung: opens the browser bound to it, seeded with its map.
anonymousNoBuild from what is visible WITHOUT an account even though the site shows a sign-in page — only when the user said the public part is enough.
persona_idNoSaved identity to sign in with (writ_personas). Without one, a site whose entry page IS a sign-in wall is not built: the answer names the persona to pass, or returns `persona_needed` — relay its tell_user + create_url to the user, then call again with the new persona_id. Required for 2FA — the code is minted server-side and never shown to you.
ai_superviseNoAI-supervised crawl rungs (default true): after the crawl mines forms, POSTs, query links, scripts and listings, one bounded Writ AI call authors the API from them — which functions serve the goal, their names, inputs and example values; the live verify call then measures each response shape. false = the purely mechanical surface map (no AI spend on the crawl rungs).
skip_existingNoSkip the proposal of the user's OWN matching workflows (set after they declined).
respect_robotsNoCrawl rungs obey the site's robots.txt (default true). Pass false ONLY when the user vouches for the target and a rung reported that robots.txt refused the crawl (`escalations` / `message` name the rule) — otherwise the crawl rungs fetch nothing and the build goes straight to a browser. Does not apply to the AI or guided browser rungs.
use_residentialNoRun on the platform residential network (premium) for a site that blocks datacenter IPs. Default off. Continuing a build (build_id) keeps the build's own persona, residential exit and country — pass these only to change them.
execution_targetNo'cloud' (the fleet) or a linked desktop's agent_id: build there, in its own browser and connection. Omitted = the desktop chosen with writ_devices, else the cloud.
skip_marketplaceNoSkip the ready-made marketplace proposals and build fresh.
residential_countryNoTwo-letter ISO country the residential exit should be in (e.g. 'us', 'fr') — applies to every rung of the build, the guided browser included. Omit for an automatic exit. Ignored unless the session egresses residential.

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv1.1.0

TDQS

A3.6/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description adds extensive behavioral context beyond annotations, such as the cost ladder, escalations, robots handling, verification status, persona login, and parked builds. However, it states that level=deep performs real create/update/delete writes and 'CHANGES real data,' while the annotations declare destructiveHint=false. That is a direct annotation contradiction, so per the rubric this dimension scores 1.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The opening is front-loaded and identifies the core job quickly, but the description becomes an all-caps wall of text that repeats mode and parameter details already covered by the schema. For a tool whose parameter documentation is already complete, this level of redundancy and density is not concise.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex 17-parameter orchestration tool with no output schema, the description covers selection, the full call flow, modes, authentication, robots restrictions, verification caveats, and follow-up tools. Despite the verbosity, an agent has enough context to start, continue, and recover a build.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all 17 parameters, making 3 the baseline. The description adds useful cross-parameter context—build_id continues a parked build, persona_id handles sign-in/2FA, respect_robots ties to escalations—but much of the mode and level explanation duplicates the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource: turning a website into a callable API. It explicitly distinguishes itself from sibling tools by naming writ_scrape, writ_crawl_site, and writ_record_website as NOT FOR cases. An agent can identify the tool's core job without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit when-to-use guidance, the intended three-call workflow, the follow-up status/run/diagnose tools, and clear exclusions. It also explains which mode to choose and when to escalate, leaving little to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.