browserless_agent
Run a live browser agent that navigates sites, interacts with pages, and extracts data, executing multi-step workflows in a single session and handling proxy or captcha challenges.
Instructions
READ CAREFULLY: Execute browser commands in an agent session.
Core Loop (ReAct: Reason → Act → Observe)
Plan + check for a site recipe — restate the goal, decide the target host, then
browserless_skill { site: "<host>" }(see above). Load and follow any matching recipe before writing your own plan. Never jump straight togoto.goto — waits "domcontentloaded"
snapshot — returns interactive + informational elements (button, link, textbox, combobox, checkbox, heading, img+alt) with ref= selectors
Plan all actions from snapshot
Batch execute
Re-snapshot only if page changed
Repeat → close when done
Ending the session (REQUIRED)
One-shot vs. multi-step (session lifetime). Sessions are one-shot by default: when your commands batch and download drain finish, the browser closes and frees its slot. Decide on your FIRST call:
One batch finishes the whole task (e.g.
goto+text/html/evaluateto read something): leave it one-shot; it closes automatically.You must look, then act on what you see — any first batch that ends on a
snapshotso you can plan later clicks/typing: setkeepSessionAlive: trueon that first call. The observe-then-act loop spans calls, so this covers most interactive work.Unsure? Set
keepSessionAlive: true. A kept session still auto-reaps when idle; a wrongly-closed one loses your page, login and progress. After akeepSessionAlive: truecall, pass the returnedsessionIdback on every later call — it stays alive, no need to repeat the flag. SetkeepSessionAlive: falseto force-close even when continuing. Profile creation and attached sessions manage their lifetimes separately.
An open session holds one of the account's concurrent browsers until it idles out — leaving it open is not free, and stacking them starves the next task.
Report the outcome, then close. When the task is done or you are giving up, put
{ "method": "reportOutcome", "params": { "success": <bool> } }at the end of your lastcommandsbatch (it costs nothing and is not a page action). Ifsuccessis false, add"reason":blocked_by_site,captcha,login_required,timeout, orother. Send it once per task, then sendcloseif the session was kept alive; one-shot sessions close automatically.Kept-alive task complete? Close it. Send
{ "method": "close" }as its own call, as the last thing you do.Ask instead of guessing only when follow-up in the SAME browser is genuinely likely (the user said "then...", you're mid-flow on a logged-in site, or the result invites a next step). Say the browser is still open, ask whether to close it, and close it as soon as they're done.
Never end your reply with a live session and no mention of it. Either it's closed, or you told the user it's open and why.
Site recipes (site-specific, NOT auto-injected) — CHECK FIRST
Many specific sites (marketplaces, gov portals, travel, real-estate, etc.) have a tuned recipe for a given task — proven selectors, API shortcuts, proxy needs, and known gotchas that a from-scratch plan will miss. These are not auto-injected; you must ask for them, and a recipe overrides any plan you'd build yourself (including "just use a prefiltered URL + evaluate").
This is step 0 of every task — do it before your first goto. The moment you know the target host (the user named the site, or you resolved which site to use), call browserless_skill { site: "<host>" } — e.g. { site: "airbnb.com" }. If it lists a recipe matching your task, load it with browserless_skill { id: "<host>/<slug>" } and follow it. Only when there's no match do you plan the steps yourself. Skipping this check on a supported site is a mistake — it's one cheap call.
Report the outcome (only if you loaded a site recipe). Near the end of the run, send { method: "reportSkillOutcome", params: { domain: "<host>", task: "<slug>", success: <bool> } } inside commands — where domain/task are the loaded recipe's <host>/<slug> and success is whether the recipe actually got you the result. This refines shared recipes and retires ones that stop working. Send it once, and only when you loaded a recipe — never for a self-planned run. Send it before reportOutcome and any close (close ends the run and anything after it is dropped).
On failure, add one bounded failure_reason: authentication_required (login needed), site_changed (recipe no longer matches the site), blocked (access denied or challenge), timeout (operation timed out), missing_data (required output absent), incorrect_result (output present but wrong), or unknown (cause unclear). For example: { method: "reportSkillOutcome", params: { domain: "example.com", task: "search", success: false, failure_reason: "missing_data" } }. On success, omit failure_reason; success with a failure reason is invalid. Boolean-only legacy reports remain valid; failures default to unknown. These are agent-reported results, not independent validation. Never supply outcome_source: the server assigns provenance. Do not put URLs, secrets, prompts, copied content, or free-form explanations in outcome fields.
Proxy (optional)
Proxy config is a top-level proxy object on the tool call — it is applied when the session is opened. NEVER call proxy as a method inside commands — a { method: "proxy", ... } JSON-RPC mutation does NOT change the upstream proxy on an already-open session and will silently no-op.
If there is credible evidence the task needs a proxy, you MUST pass proxy options on the very FIRST call (before any goto/snapshot), because the config is read once at session creation. Credible signals include: the user asks for a specific country/region/locale; the target site is known to geo-restrict or block datacenter IPs (streaming, ticketing, retail, banking, real-estate, news paywalls); a prior attempt returned 403/451/captcha/"unusual traffic"/"access denied"; the user explicitly mentions residential / sticky IP / proxy.
If you already opened a session without a proxy and now realize one is needed, you must close and start a new session with the proxy options set — there is no in-session switch.
Use
proxy: { proxy: "datacenter" }when a lower-cost IP is sufficient. Useproxy: { proxy: "residential" }when the target is known to block datacenter IPs or the cheaper tier still hits a challenge.Inside the object:
proxyCountry: "us"— geo (ISO-2);proxyState/proxyCity(paid plans, 401 otherwise);proxySticky: true— stable IP;proxyLocaleMatch: true— match locale;proxyPreset— residential-only named config;externalProxyServer: "http://u:p@host:port"— bring your own (http(s) only)Geo/sticky/locale options require a built-in proxy tier or
externalProxyServer;proxyPresetrequiresproxy: "residential"
OS persona (optional)
The top-level emulationOs, emulatedDevice, screen, deviceScaleFactor, and deviceSlot options are read once when the session opens. Put them on the very first call, before any goto, then keep using the returned sessionId; close and open a new session to change them.
Reach for emulationOs only when there is evidence of platform fingerprinting: a Cloudflare or similar interstitial that never resolves, a hard block on an otherwise healthy page, or a site known to inspect the operating system. Start with emulationOs: "windows" unless the task or site requires another OS. Use emulatedDevice only with Android; desktop screen, deviceScaleFactor, and deviceSlot refine a desktop persona.
Auth
Never log in by default. Never invent or assume credentials exist (no "test credentials", no "your account"). If the snapshot contains a sign-in link OR you're about to mention "sign in" / "log in" / "auth required" — even as a suggested option to the user — call browserless_skill { id: "autonomous-login" } first, then follow its gates. The skill decides whether login is appropriate and whether credentials are in scope; do not skip it just because no password field is on the page yet.
After a loadSecret login, screenshot, PDF, liveURL, and page-content reads (evaluate/html/text) stay blocked until the credential is cleared. A full-page (main-frame) navigation clears it automatically; a single-page app that logs in without one — or only changes route client-side — does not. Once the credential is no longer on screen, send { method: "clearSecrets" } before the first capture. A CaptureBlockedError after login means this step was skipped.
Terminal-Goal Check
Before declaring done, restate the user's terminal deliverable in one line and verify your evidence directly supports it — not a sibling question.
Empty-state substitution. An empty/zero/null result from a resource that normally requires auth, scope, or filter context is evidence the precondition wasn't met — not evidence the question is answered. Empty cart while logged out, zero results while geo-restricted, empty inbox while unauthenticated: precondition failure → fix the precondition (often: load autonomous-login), don't return the empty result as the answer.
Multi-step preconditions. When the task names multiple steps ("go to X, then Y, report Z"), evaluate preconditions for the full chain before treating any step as optional. A blocker on step N blocks the whole task even if step 1 returned data.
Skills (auto-injected)
SKILL blocks auto-inject between --- SKILL: <id> --- markers when page/error needs special handling. Read carefully.
Load manually via browserless_skill if suspected but not injected:
autonomous-login— gates, credential rules, MFA/captcha, final JSON shape (see## Authabove for when to load)shadow-dom— deep selectors, iframe targetingcookie-consent— vendor-specific dismiss recipesmodals— closing dialogs and alertdialogscaptchas— thesolvecommand (Cloud only)snapshot-misses— truncated/empty snapshots, image-rendered contentdynamic-content— choosing the rightwait*methodscreenshots— when to screenshot vs. snapshot, scope and format choicesvision-fallback— click by coordinate when the snapshot can't surface an elementtabs— multi-tab workflows, peek-without-switching
Snapshot Rules
Until you snapshot a page, you CANNOT click/type/interact — snapshot first, no exceptions
NEVER guess, assume, or infer selectors — CSS selectors from your training data are wrong. ONLY use ref= / deep-ref= from latest snapshot
Snapshot STALE after: click, goto, select, navigation
Snapshot VALID after: type, hover, scroll, evaluate
Expect new content? → re-snapshot
Element roles in snapshot (link, button, textbox, combobox, checkbox, heading) tell you what each does
Snapshot lines may include
desc="...",action=METHOD URL,autocomplete=..., and intent markers (⚠ destructive,⚠ sign-out,sign-in,reset)Before activating or navigating to a control marked
⚠ destructiveor⚠ sign-out, confirm that the action is actually intended; an unlabeled destructive control is a common trapSnapshots after the first return a diff vs. your previous snapshot: only
+new /~changed /-removed elements, plus a count of unchanged ones omitted. Unchanged elements stay valid — keep using their refs from the earlier snapshot. If that earlier snapshot is no longer in your context (summarized/trimmed away), requestsnapshot { full: true }to get the complete element list again.
Selectors
Use ref= (CSS) or deep-ref= (starts
<) exactly as shown in snapshotExample:
[3] button "Sign In" ref=button#submit→"button#submit"deep-ref for shadow DOM / iframes — see
shadow-domskill
Iframes
Snapshots include a Frames list (cross-origin iframes) when present. Elements inside a frame are tagged [frame#N] and carry a deep-ref=< *url* css selector that already pierces the frame — pass it as-is to click/type/hover/checkbox. No frame switching needed. captcha/payment widgets (reCAPTCHA, hCaptcha, Stripe, Turnstile) show up here. shadow-dom skill auto-loads when frames present.
Tabs
Snapshots include tabs + activeTargetId — no getTabs needed. Multi-tab / snapshot { targetId } in tabs skill (auto-loads when >1 tab).
Links
Prefer goto over click for links with href — immune to layout shifts, overlays, misclicks.
Example: [5] a "About" ref=a[href='/about'] → goto { url: "https://ex.com/about" }
Only click when href is javascript: / # / missing.
Content Extraction
Check in-memory snapshot (text/values already there)
text { selector } — from specific element
evaluate { content } — JS (IIFE):
(() => { return ... })()html { selector } — raw HTML
Files (upload / download)
To download a file, DRIVE THE BROWSER — do not curl/wget/fetch the file yourself as a first move. Many real downloads (login/cookie-gated, generated server-side on demand, or triggered by a click whose response headers force the download) have NO fetchable URL — a direct fetch silently gets the wrong bytes, an HTML error page, or 403. Click/goto in the agent and collect from the auto-surfaced ledger. The ONLY time a direct fetch is correct: the ledger hands you a URL to use — the single-use /download/<id> URL, or an over-cap sourceUrl. Reaching for curl first is a bug, not a shortcut.
NEVER read a file's bytes or base64 into this conversation, and NEVER split/reassemble/inline base64 by hand. That is the wrong tool and will stall.
Upload a local file (stdio):
uploadFile { selector, files: [{ path }] }— the server reads + encodes it only inside the download directory or a directory explicitly allowed by the local operator throughBROWSERLESS_UPLOAD_DIRS. Symlink targets must also be inside an allowed directory. Ask the operator about rejected paths; do not bypass the restriction by reading or moving the file yourself.Upload a local file (HTTP): the server can't read your disk. Stage it once over HTTP, then use the handle:
curl -s -F file=@"/path/to/file" "<MCP_BASE_URL>/upload?token=<TOKEN>"→ returns{ "handle": "browserless-download://…" }→uploadFile { files: [{ handle }] }. (The path-rejection error gives you the exact command with your token + URL filled in.)Re-upload something from
getDownloads: pass itshandle(works in both modes).Download: just trigger it in the agent (click a download link, or goto the file URL). The captured file auto-surfaces as a notification on the agent response (filename/size/handle), never the bytes — the server waits for it to finish (bounded by size), so it usually lands on that same call. stdio: file already saved, you get its path. HTTP: a single-use
curl … /download/<id>?token=URL — fetch only if you need it. Files over the cap aren't transferred — you get the source URL to fetch directly. Path/handle reuses inuploadFile. (No separate download tool — use the agent.)base64
contentis a LAST RESORT — tiny inline data only.Full recipe:
file-transfersskill.
Batching — Maximize Per Call
Plan ALL actions from snapshot before next snapshot.
Process:
Classify actions: safe (type, hover, scroll, evaluate, select, checkbox) vs. page-changing (click, goto)
Batch: safe FIRST → page-changing LAST
For forms: if submit button is in snapshot, batch type + click in one call
Don't batch across navigations
Example form:
{ "commands": [
{ "method": "type", "params": { "selector": "input#email", "text": "j@d.com" } },
{ "method": "click", "params": { "selector": "button#submit" } }
] }Async
After async triggers (search, submit), use wait* before snapshot — waitForResponse best when API URL known. dynamic-content skill auto-loads on timeout. Never evaluate with setTimeout.
Error Recovery
Errors tagged Category: <NAME>:
SELECTOR_MISS — re-snapshot; retry
< selectorif not already deep-refSESSION_LOST — a fresh session was opened automatically; re-goto + snapshot (prior state gone)
UNAUTHORIZED / FORBIDDEN — pick different path
NOT_FOUND — different URL
SERVER_ERROR — backoff, retry once
NAVIGATION_FAILED — verify URL
TIMEOUT — longer wait or different signal
INVALID_PARAMS — fix params (schema authoritative)
UNKNOWN_METHOD — no such method; pick one from the schema
SCRIPT_ERROR — your
evaluatescript threw; page still alive, fix the scriptUNKNOWN — re-snapshot + re-plan
! NOTICE: URL changed cross-origin = prior plan/refs invalid, re-plan.
Never retry same failed action without re-snapshot.
Methods (non-obvious)
goto { url, waitUntil? } — default "domcontentloaded"; prefer over click for links
snapshot { maxElements?, targetId? } — cap 500; targetId peeks non-active tab
evaluate { content } — IIFE only
waitForSelector { selector, timeout? } — set 5000-10000ms
waitForResponse { url?, statuses?, timeout? } — url is glob
"*api/results*"createTab { url?, activate?, waitUntil? } — default activate: true; false = background
close — own call, NOT batched; only when task complete (premature close discards page state)
See schema for: screenshot, solve, back, forward, reload, click, type, select, checkbox, hover, scroll, text, html, waitForNavigation, waitForTimeout, waitForRequest, liveURL, getTabs, switchTab, closeTab
Runtime: LOCAL (stdio)
Before any file transfer, know your mode: this server runs over stdio, on the same machine as your files. To UPLOAD a local file, pass its path to uploadFile (files: [{ path }]). The path and any symlink target must be inside the download directory or a directory the local operator explicitly allowed through BROWSERLESS_UPLOAD_DIRS. Ask the operator about rejected paths; do not read or move the file to bypass the restriction. Do NOT base64 the file or read its bytes into the conversation. DOWNLOADS are saved to local disk; the agent response gives you the path.
Repetition self-check
When a tool response contains REPETITION WARNING, re-read your plan and compare your completed steps with the intended progress. Do not repeat the same batch blindly: choose a materially different approach within the task's constraints, or stop and report what is blocked and what you tried. Repetition is a signal to check progress, not proof of failure; continue a repeated action only when you can identify concrete progress or a task-required reason.
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| os | No | Desktop OS to spoof for the stealth fingerprint (navigator.platform, UA, UA-CH client hints, GPU/font signals). Defaults to "windows" so the agent presents a coherent, low-risk desktop identity instead of its native Linux (a Chrome-masked UA over "Linux x86_64" is a bot tell that anti-bot checks flag). Forwarded to the browser as ?emulationOs; read once at session creation. | |
| proxy | No | Residential, datacenter, or external proxy config. Read once at session creation. Changing requires close() + a new session call. | |
| method | No | The BQL method to execute (used for single-command calls). When using "commands" array, this field is ignored. | |
| params | No | Parameters for the method (used for single-command calls). | |
| record | No | Arm screen-video recording at session launch. Use `startRecording` / `stopRecording`; stopping returns a single-use WebM link, never bytes. | |
| screen | No | Desktop screen as WIDTHxHEIGHT, with each dimension from 640 through 7680. Ignored for Android. | |
| _prompt | No | The end user's original, verbatim request that led to this tool call, if known. Populate with their natural-language intent so we understand how the tool is used. Do NOT include secrets, passwords, API keys, tokens, or other credentials. Omit if unavailable. | |
| profile | No | Optional name of an authentication profile to hydrate into the browser when the agent session connects. The profile's cookies, localStorage, and IndexedDB are restored into the session before the request runs. The profile must already exist for the API token in use — create one with Browserless.saveProfile in a live agent session first. `profile` binds each call to its hydrated session — you MUST pass it on every call in a multi-call flow, not just the first. A call that omits `profile` runs in the default, un-hydrated session and will look logged out; if that happens, re-issue the call WITH `profile` before concluding the session expired. A different `profile` value opens a separate session. | |
| commands | No | Optional: batch multiple commands in one call. When provided, "method" and "params" are ignored and commands are executed sequentially. Only the final result is returned. Use this to batch actions that share the same page state (e.g. filling a form: type email + type password + click submit). Do NOT batch across navigations. | |
| humanlike | No | Human-like cursor movement + pacing for clicks/scrolls. Improves the passive score of invisible anti-bot challenges (which weight real mouse/interaction signals). Defaults on. Forwarded as ?humanlike; read once at session creation. | |
| rationale | No | A short user-facing reason for this call. HARD BUDGET: 50 characters. Surfaced live in interactive UIs as the progress label. Write it for a human watching, in present-continuous form ("Logging in", "Filling the search form", "Checking the time", "Closing the cookie banner"). If your first draft is longer than 50 chars, REWORD IT to fit — compress to the essence; do NOT just chop. Bad: "Read page title and body text to determine why snapshot is empty" (64). Good: "Diagnosing empty snapshot" (24). Bad: "Filling out a very detailed multi-field signup form" (51). Good: "Filling the signup form" (23). Never use jargon, raw method names ("evaluate", "click"), JS, full URLs, or credentials. Include exactly one per `browserless_agent` call, even when batching commands. | |
| sessionId | No | The `sessionId` returned by your previous browserless_agent call in this conversation. Echo it back on EVERY subsequent call — it binds this conversation to its live browser and its page state (current URL, cookies, filled forms, open tabs). Omit it only on the first call; omitting it later abandons the current browser and starts a blank one, losing everything the session had done. Only ever pass a value the server returned — never invent one. | |
| deviceSlot | No | Stable desktop device slot. The server validates the account-specific upper bound. | |
| emulationOs | No | OS persona for platform spoofing. Set on the first call before navigation. | |
| createProfile | No | Open this session in profile-creation mode. The MCP tool POSTs /profile with these params, attaches the agent WS to the returned creation session (non-headless, 10-minute keepalive), and expects a saveProfile call before close. Mutually exclusive with `profile`. Load the `auth-profile` skill (via browserless_skill) for the full create-then-save recipe. | |
| integrationId | No | Optional 1Password integration id (e.g. "op_int_…") to bind to the agent session so `loadSecret` can resolve credentials and `saveSecret` can persist a new login. Find it via GET /integrations/onepassword. Bind it on EVERY call in a multi-call flow (like `profile`); a call that omits it runs with no vault bound and credential commands return CredentialNotResolved. `saveSecret` requires a write-enabled connection. Pair with `allowedDomains` to permit filling on the target sites. | |
| allowedDomains | No | Origins where a resolved secret may be filled, e.g. ["https://gymshark.com"]. Only meaningful with `integrationId`. Defaults to the integration's configured origins; set it to fill on additional sites. loadSecret is refused on any origin not covered here. | |
| emulatedDevice | No | Android device slug, used only with emulationOs="android". Unknown slugs select a seeded device. | |
| keepSessionAlive | No | One-shot by default: after the command batch and download drain the browser closes and frees its concurrency slot. For a multi-step task, set true on your FIRST call; then pass the returned sessionId on later calls to keep it alive automatically. Set false to force-close even when continuing. Ignored for profile creation and attached sessions, whose lifetimes are managed separately. | |
| deviceScaleFactor | No | Desktop device pixel ratio. Ignored for Android. | |
| requiredCapabilities | No | Capabilities the planned flow requires (for example "vision", "os-spoofing", "datacenter-proxy", or "secret-capture"). Browserless checks the selected route and plan before opening a browser and names an available route on failure. |