Skip to main content
Glama

Server Configuration

Describes the environment variables required to run the server.

NameRequiredDescriptionDefault

No arguments

Instructions

Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.

This server publishes no instructions, or was last inspected before Glama recorded them.

Capabilities

Features and capabilities supported by this server

Protocol revision2025-11-25

CapabilityDetails
tools
{
  "listChanged": false
}
experimental
{}

Tools

Functions exposed to the LLM to take actions

NameDescription
browser_openA

Wake the workspace and open or navigate your current browser tab. Pass either url or a bounded urls list (maximum 5). To check several sites, open them in ONE call with urls: it opens independent tabs in the same authenticated browser context and returns one compact, ordered result per URL; that results list is the authoritative source for each tab_id. A URL that fails (DNS, block) is reported on its own row and the other tabs stay open. Batch results omit page bodies: next, call browser_extract/observe for every returned tab_id together in one response, since calls on different tabs run in parallel. Returns once the page has settled. By default it reuses your current actor-owned tab, including across a resumed task; set reuse=false only for a genuinely independent second tab (the result carries reused=true when reused). The single-URL result includes a settled page snapshot; use it directly instead of immediately calling browser_observe. Only a returned attention_required=true represents an explicit handoff. A blocked=true page result does not suspend the run: decide whether to take another route or tell the user plainly what could not be completed. For broad public discovery use web_search rather than navigating this tab to a search engine; browser search remains valid when explicitly requested or interaction/personalization matters.

browser_tabsA

List, close, or sleep workspace tabs. Closing takes tab_id for one tab or tab_ids for several in a single call (at most 20); each tab's outcome is reported separately, so one unknown id does not strand the rest.

browser_observeA

Observe a live tab and return generation-scoped element refs, on-screen controls first, current viewport content_blocks and measured scroll_containers. Treat auth_state as page evidence for your next decision; a visible login form does not itself stop the run. diagnostics lists page errors and failed requests since your previous result. When the result says elements were not shown, narrow the view instead of reading the DOM with browser_evaluate: query finds controls by their words, within scopes to a region, frame or element, and cursor continues where the last response stopped. Every ref returned is valid for browser_act.

browser_actA

Perform one verified action in a leased tab. action.kind must be navigate|click|fill|fill_form|sequence|select|check|date|press|upload|scroll|wait (action.type is accepted as an alias for kind; type=type means fill). ref must be an exact ref=e… token from the newest page/observe/extract result for this tab — labels and row numbers are rejected; CSS/XPath/text selectors are not an agent-facing escape hatch. Reuse the fresh refs returned in an action's page; do not observe again after every field. A successful act returns either page (the page it produced, with fresh refs) or page_unchanged: true, meaning the URL and visible text are what your last observation showed and its refs still apply — observe again only for a different view (query, selector, region). Batch stable fields with fill_form, which reports each field separately. For anything that picks from a list (native selects, custom ARIA comboboxes and dropdowns, type-to-search boxes) use kind=select with option= and optional query=; this performs and verifies open/filter/choose in one call. For anything with a checked or selected state, including list items, use kind=check with checked=true|false. Independent calls for distinct tab_ids may be issued together; calls targeting one tab remain ordered. Use sequence for a short reversible multi-step workflow when later-page targets can be named semantically. Every step is re-resolved from current page state, authority is rechecked, and the sequence stops before dispatch when an exact expect_before guard or a unique target is missing; use expect_after to guard a planned transition. Do not place a consequential final submit in a sequence. until waits inside the call for a late effect (redirect, async save, hot reload); kind=wait only waits. The same select fields can be included in fill_form. For a reversible draft/save action, a click receipt with state=network_effect_observed, a successful same-origin non-read response, matching control readback, and the page's save status is sufficient browser evidence. Do not activate unrelated controls, inspect hidden databases, or use shell/page scripts solely to obtain stronger private persistence proof. upload attaches files to an input[type=file] ref (or file_chooser=true with an observed attachment button ref) and is the ONLY way to do so — fill rejects file inputs. For a request such as 'keep scrolling until I stop you', use one scroll action with duration_seconds instead of spending repeated model rounds; cancelling/replacing the run stops it.

browser_extractA

Extract the complete current visible content and structured field/page state from a tab. Pass target_ref to scope a fresh read to an observed content block, control or the ARIA region it owns. This is a fresh high-level read, not a change-only observation. If a large result returns evidence_ref and next_cursor from browser_extract, continue it with this same tool. Do not continue a managed-output reference from browser_open, browser_observe, or browser_act; start a fresh extraction with tab_id and instruction instead. Pass selector for an instant CSS query instead of a full read: it counts and lists matching elements (tag, text, requested attributes, and the ref of any match already in the current observation) without building a snapshot, e.g. every product link's href, how many rows a table has, or each card's price. read=console|network reads the tab's logs (with same-site error bodies); read=inspect with target_ref says why a click is refused or a control will not take a value: what receives the click there, what covers or clips it, disabled state, styles. read=design returns how the page looks in CSS terms: variables, colors by use, type scale, radii, shadows, spacing and layout regions with sizes; target_ref limits it to one component. read=audit audits the page you are on (a deployed or localhost site): accessibility violations (axe-core) with the ref of each element, load timings and weight, title/lang/description/headings/alt gaps, broken same-origin links and console errors, worst first.

browser_screenshotA

Take a picture of an agent-owned tab. This is a last resort, not a way to find or operate controls: browser_observe returns the refs browser_act needs, and a picture returns none. Use it only for what an observation cannot describe (canvas, charts, images, visual layout). Password and payment fields are masked. The result states the viewport size at capture; a person may resize the live window, so trust it over sizes you saw earlier. state captures a ref's :hover/:focus styling; compare_with diffs against an earlier picture or another tab and boxes what changed. To check a page you are building against an original, open both and take one with compare_with='tab:': one picture of both.

browser_viewportA

Check a responsive layout in the current agent-owned tab. action=set with preset phone (390x844), tablet (768x1024) or desktop (1365x768), or width and height in CSS pixels, resizes the real browser window so the page reflows as it would on that screen (Linux; elsewhere the window keeps its size). action=restore returns to the launch size and clears emulation; call it when the check is done. action=get reads the current size. action=emulate sets color_scheme (light or dark), reduced_motion, forced_colors or offline. A successful set or restore includes a fresh page snapshot; earlier element positions are stale after a resize, so use the new refs.

browser_evaluateA

Run policy-gated read-only JavaScript in a tab only when browser_extract or browser_observe cannot return the required fact. For visible page text, labels, controls, or a region owned by aria-controls, use browser_extract instead; for computed styles, colors, fonts, sizes, icons or backgrounds, browser_extract read=design (the page, or one component with target_ref) or read=inspect (one element's box and styles).

browser_flowA

Record and replay browser work. Every verified browser_act on a tab is recorded (targets by what they are, not by ref; passwords and other secrets never). After a task succeeds once, action=save turns that tab's recorded steps into a flow; action=run replays it on a tab with no model call between steps, re-finding each element on the live page, so repeating the task (the next job application on the same site, the next record in the same form) takes seconds. Pass start_url for the new page and variables for the values that change; save reports the detected variables. A run stops at the first step it cannot find unambiguously or verify, and says which steps are done and which remain; finish that step yourself, then run again with from_step. The same review and authority apply to a replayed submit as to one you click yourself: run only a flow whose steps the user's request covers.

wait_for_bot_wallA

Wait out a bot wall on a tab, then attempt to solve it if it does not clear on its own. Polls the tab while the wall is auto-verifying (e.g. a Cloudflare 'Just a moment…' check that clears itself). If it becomes a user-required challenge it runs the configured solver, and if that fails it returns a structured verdict. Also call it before submitting a form that carries an embedded Turnstile, reCAPTCHA or hCaptcha checkbox: it waits for the widget to pass, presses the checkbox if needed, and reports whether a token was issued (image puzzles are never solved). When the page is required for the user's goal, stop and tell them to solve it (or hand off if you are a browser sub-agent); otherwise continue to the next task.

Prompts

Interactive templates invoked by user choice

NameDescription

No prompts

Resources

Contextual data attached and managed by the client

NameDescription

No resources

TDQS

A4.4/5.0

Scored across 10 tools

Disambiguation4/5

Most tools target clearly different browser operations, but the read-side tools (browser_observe, browser_extract, browser_evaluate, browser_screenshot) overlap in purpose and require the long descriptions to separate them. The descriptions do provide strong guidance, so an agent can usually pick correctly, but some ambiguity remains.

Naming Consistency4/5

Nine tools use a consistent snake_case browser_ prefix, and the naming is readable and predictable. The lone exception is wait_for_bot_wall, and some browser_ tools are noun-based rather than verb-based, which are minor deviations.

Tool Count5/5

Ten tools is well-scoped for a browser automation server, with each tool covering a distinct capability such as opening, acting, observing, extracting, screenshotting, viewport control, flow replay, and bot-wall handling. No tool feels redundant or missing at the count level.

Completeness5/5

The surface covers the browser lifecycle: opening/navigating, tab management, interaction, observation, extraction, evaluation, screenshots, viewport emulation, workflow replay, and bot-wall handling. There are no obvious dead ends for typical browser automation tasks.

Maintenance

ActivityNo data
ResponsivenessNo issues