Skip to main content
Glama

Act in a browser session

writ_browser_act
Destructive

Run a batch of browser actions—navigate, click, fill, extract—on an open session and get the updated page back to decide the next step.

Instructions

Run one batch of actions on an open browser session and get the fresh page back. YOU are the brain: every navigation, click, fill, sign-in, capture, probe and script is yours to decide, one batch at a time. No writ_browser_compose in this client? This tool composes too: actions=[{action:'define_function', name:'feed.list', from_index:3, ...}] or [{action:'compose', operation, payload}], never mixed with clicks in one batch. After a navigate, a click or a select that changes the page, END the batch and look at the new page before acting on it. RECORDING RULES: (1) a caller INPUT — write the value as {{name}} in the action (select/fill/type_text value, navigate url) and pass the real value in inputs ({"name": "real value"}); the page gets the real value, the recorded step keeps {{name}}, and name becomes a workflow input by itself. (2) DATA — record an extract {variable,script} at the position that shows it (a read-only JS IIFE returning rows/fields); its result comes back in this answer, so check it before saving. (3) a SECRET — fill with data_key, never a literal. Interactions (navigate/click/fill/select/press_key/...) are recorded as steps; SEE/HEAR/NETWORK probes and wait never are — replay waits for each step's selector by itself, and a wait the task truly needs is an explicit wait_for step (writ_browser_compose add_steps). ACTIONS: DRIVE: navigate {url} · click {selector | field_index | button_index} · fill {selector,value,data_key?} · type_text {selector,value} · select {selector,value} · check {selector} · hover {selector} · submit {selector} · press_key {key} · scroll {direction,amount} · back · wait {seconds} · wait_for {selector,timeout}. SEE (granular first): query_dom {selector,limit,offset,attrs?,text_chars?,html_chars?} (every match as compact records with a css path to target next) · count {selector} · find_text {text,selector?,exact?,limit?} (the deepest elements showing that text, with paths) · get_attributes {selector,index?} (one element: all attrs, value, box, options) · read_text {selector,all?,limit?,max_chars?} · inspect {selector,limit?,max_chars?} (match count + outerHTML) · list_candidates (the page's repeating row shapes — start here for any list/table) · list_frames · get_dom {selector?,depth?,max_chars?} (the real cleaned HTML — the expensive last resort) · get_screenshot {x?,y?,width?,height?}. TABS / FILES / 2FA: list_tabs / switch_tab {index} · upload {selector,mode,file_slot} · wait_for_download {trigger_selector,output_key} · twofa {challenge_method,selector?,submit_selector?} (the persona's one-time code, minted server-side — see 2FA RULES). HEAR: get_console {level?,since?,query?,limit?} (console messages, uncaught JS errors with stack, failed/blocked requests since your last read — the page's console_since_last_read counts tell you when it is worth a call; read it BEFORE guessing why a sign-in, click or extraction did nothing) · page_errors (only the uncaught exceptions). NETWORK: capture_network {reload?} (the backend calls the page makes — how you find the site's real API; then search/read them with writ_browser_network) · get_request {url substring} (one call in full). RUN CODE: evaluate_js {script} (any JS on the live page, returns JSON — your main probing tool; read-only). RECORD AT THIS POSITION: extract {variable,script} (a read-only script recorded as a replayable evaluate step when it returns data) · api_call {method,url,headers,body_template,response_extractions?,variable} for one request, or api_call {flow:{version:1,steps:[...]},inputs:{...},variable} for a multi-request bootstrap/pagination/transform program (both execute NOW inside the session with its cookies; the flow uses the same interpreter as browserless replay and returns a bounded result sample) · login_post {method,url,headers,body_template} (replay a sign-in as one request) · probe_write {selector} (learn a create/update/delete request WITHOUT sending it) · confirm_write {selector} (perform it ONCE for its real confirmation — changes real data, only when authorized). A sensitive fill MUST carry data_key so the value is held server-side and the saved step keeps a {{secret:...}} placeholder. ASK the user for any credential, 2FA code or decision — never invent one. 2FA RULES: the one-time code is minted server-side from the attached persona and never shown to you. BEFORE twofa, READ the challenge and IDENTIFY the method the page is using — a phone number / 'text message' = sms, an email address = email, 'authentication app' = authenticator, approve-on-phone / passkey / QR / WhatsApp = other — and pass it as challenge_method. A persona receives exactly ONE method (see twofa_method in writ_personas) and a site picks its own default: Facebook texts an SMS even when the account has email. If the page's method is not the persona's, do not emit twofa: click the page's 'Try another way' / 'Use another method' / 'More options' / 'Didn't get a code?' control (in the page's own language), choose the persona's method, confirm, THEN twofa. On twofa_method_required, twofa_method_mismatch or twofa_verify_method do exactly what the message says — verify the method on the page and switch or resend — and do NOT ask the user yet. Only twofa_mint_failed / twofa_no_persona mean: call writ_browser_ask_user kind='twofa' so the Writ user supplies it.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
inputsNoValues held server-side for {{placeholder}} substitution, e.g. {"city":"Paris"}. Secrets belong here or on a fill's data_key — never hardcoded into a step.
actionsYesOrdered action objects, e.g. [{"action":"click","selector":"#login"}].
max_charsNoClip of the returned page_dom / page_text (default 40000, ≤200000). A get_dom probe on a real app is 500KB — prefer evaluate_js / inspect / read_text to target what you need.
session_idYesSession id from the start tool.

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv1.1.0

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations flag destructiveHint=true/openWorldHint=true, and the description adds rich context beyond them: confirm_write 'changes real data, only when authorized', sensitive fills must carry data_key, secrets are held server-side, the 2FA code is minted server-side and never shown, and which actions are/aren't recorded as replay steps.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core purpose, and given the tool's complexity most sentences carry real information. However it is a very long, dense wall of text with heavy inline action enumeration that could be tightened; structure is functional but not economical.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a high-complexity, open-world, destructive browser automation tool with no output schema, the description covers recording semantics, probe vs drive actions, error recovery, tabs/files/2FA, and credential handling — essentially everything an agent needs to invoke it safely.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description materially enriches the opaque `actions` array by illustrating action shapes, the {{placeholder}}/inputs contract, and data_key secret handling — more than the schema's terse 'ordered action objects' conveys.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource+scope: 'Run one batch of actions on an open browser session and get the fresh page back.' It immediately frames the agent's role ('YOU are the brain') and distinguishes batch interaction from compose/network siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit when-to-use rules abound: end the batch after page-changing actions, use compose actions when writ_browser_compose is absent, choose granular SEE probes before get_dom, and a clear decision tree for twofa errors and when to call writ_browser_ask_user instead of guessing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.