ascended-browser
Server Configuration
Describes the environment variables required to run the server.
| Name | Required | Description | Default |
|---|---|---|---|
No arguments | |||
Instructions
Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.
This server publishes no instructions, or was last inspected before Glama recorded them.
Capabilities
Features and capabilities supported by this server
Protocol revision2025-11-25
| Capability | Details |
|---|---|
| tools | {
"listChanged": false
} |
| experimental | {} |
Tools
Functions exposed to the LLM to take actions
| Name | Description |
|---|---|
| browser_openA | Wake the workspace and open or navigate your current browser tab. Pass either url or a bounded urls list (maximum 5). To check several sites, open them in ONE call with urls: it opens independent tabs in the same authenticated browser context and returns one compact, ordered result per URL; that |
| browser_tabsA | List, close, or sleep workspace tabs. Closing takes tab_id for one tab or tab_ids for several in a single call (at most 20); each tab's outcome is reported separately, so one unknown id does not strand the rest. |
| browser_observeA | Observe a live tab and return generation-scoped element refs, on-screen controls first, current viewport content_blocks and measured scroll_containers. Treat auth_state as page evidence for your next decision; a visible login form does not itself stop the run. |
| browser_actA | Perform one verified action in a leased tab. action.kind must be navigate|click|fill|fill_form|sequence|select|check|date|press|upload|scroll|wait (action.type is accepted as an alias for kind; type=type means fill). |
| browser_extractA | Extract the complete current visible content and structured field/page state from a tab. Pass target_ref to scope a fresh read to an observed content block, control or the ARIA region it owns. This is a fresh high-level read, not a change-only observation. If a large result returns evidence_ref and next_cursor from browser_extract, continue it with this same tool. Do not continue a managed-output reference from browser_open, browser_observe, or browser_act; start a fresh extraction with tab_id and instruction instead. Pass selector for an instant CSS query instead of a full read: it counts and lists matching elements (tag, text, requested attributes, and the ref of any match already in the current observation) without building a snapshot, e.g. every product link's href, how many rows a table has, or each card's price. read=console|network reads the tab's logs (with same-site error bodies); read=inspect with target_ref says why a click is refused or a control will not take a value: what receives the click there, what covers or clips it, disabled state, styles. read=design returns how the page looks in CSS terms: variables, colors by use, type scale, radii, shadows, spacing and layout regions with sizes; target_ref limits it to one component. read=audit audits the page you are on (a deployed or localhost site): accessibility violations (axe-core) with the ref of each element, load timings and weight, title/lang/description/headings/alt gaps, broken same-origin links and console errors, worst first. |
| browser_screenshotA | Take a picture of an agent-owned tab. This is a last resort, not a way to find or operate controls: browser_observe returns the refs browser_act needs, and a picture returns none. Use it only for what an observation cannot describe (canvas, charts, images, visual layout). Password and payment fields are masked. The result states the viewport size at capture; a person may resize the live window, so trust it over sizes you saw earlier. state captures a ref's :hover/:focus styling; compare_with diffs against an earlier picture or another tab and boxes what changed. To check a page you are building against an original, open both and take one with compare_with='tab:': one picture of both. |
| browser_viewportA | Check a responsive layout in the current agent-owned tab. action=set with preset phone (390x844), tablet (768x1024) or desktop (1365x768), or width and height in CSS pixels, resizes the real browser window so the page reflows as it would on that screen (Linux; elsewhere the window keeps its size). action=restore returns to the launch size and clears emulation; call it when the check is done. action=get reads the current size. action=emulate sets color_scheme (light or dark), reduced_motion, forced_colors or offline. A successful set or restore includes a fresh page snapshot; earlier element positions are stale after a resize, so use the new refs. |
| browser_evaluateA | Run policy-gated read-only JavaScript in a tab only when browser_extract or browser_observe cannot return the required fact. For visible page text, labels, controls, or a region owned by aria-controls, use browser_extract instead; for computed styles, colors, fonts, sizes, icons or backgrounds, browser_extract read=design (the page, or one component with target_ref) or read=inspect (one element's box and styles). |
| browser_flowA | Record and replay browser work. Every verified browser_act on a tab is recorded (targets by what they are, not by ref; passwords and other secrets never). After a task succeeds once, action=save turns that tab's recorded steps into a flow; action=run replays it on a tab with no model call between steps, re-finding each element on the live page, so repeating the task (the next job application on the same site, the next record in the same form) takes seconds. Pass start_url for the new page and variables for the values that change; save reports the detected variables. A run stops at the first step it cannot find unambiguously or verify, and says which steps are done and which remain; finish that step yourself, then run again with from_step. The same review and authority apply to a replayed submit as to one you click yourself: run only a flow whose steps the user's request covers. |
| wait_for_bot_wallA | Wait out a bot wall on a tab, then attempt to solve it if it does not clear on its own. Polls the tab while the wall is auto-verifying (e.g. a Cloudflare 'Just a moment…' check that clears itself). If it becomes a user-required challenge it runs the configured solver, and if that fails it returns a structured verdict. Also call it before submitting a form that carries an embedded Turnstile, reCAPTCHA or hCaptcha checkbox: it waits for the widget to pass, presses the checkbox if needed, and reports whether a token was issued (image puzzles are never solved). When the page is required for the user's goal, stop and tell them to solve it (or hand off if you are a browser sub-agent); otherwise continue to the next task. |
Prompts
Interactive templates invoked by user choice
| Name | Description |
|---|---|
No prompts | |
Resources
Contextual data attached and managed by the client
| Name | Description |
|---|---|
No resources | |
TDQS
Scored across 10 tools
Most tools target clearly different browser operations, but the read-side tools (browser_observe, browser_extract, browser_evaluate, browser_screenshot) overlap in purpose and require the long descriptions to separate them. The descriptions do provide strong guidance, so an agent can usually pick correctly, but some ambiguity remains.
Nine tools use a consistent snake_case browser_ prefix, and the naming is readable and predictable. The lone exception is wait_for_bot_wall, and some browser_ tools are noun-based rather than verb-based, which are minor deviations.
Ten tools is well-scoped for a browser automation server, with each tool covering a distinct capability such as opening, acting, observing, extracting, screenshotting, viewport control, flow replay, and bot-wall handling. No tool feels redundant or missing at the count level.
The surface covers the browser lifecycle: opening/navigating, tab management, interaction, observation, extraction, evaluation, screenshots, viewport emulation, workflow replay, and bot-wall handling. There are no obvious dead ends for typical browser automation tasks.