Skip to main content
Glama
yogesh-joshi-0333

browser-control-mcp-server

Server Configuration

Describes the environment variables required to run the server.

NameRequiredDescriptionDefault

No arguments

Instructions

Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.

This server publishes no instructions, or was last inspected before Glama recorded them.

Capabilities

Features and capabilities supported by this server

Protocol revision2025-11-25

CapabilityDetails
tools
{
  "listChanged": true
}
logging
{}

Tools

Functions exposed to the LLM to take actions

NameDescription
browser_statusA

Check browser readiness: reports whether a Chrome debug port is available (for connect mode) and lists all active Puppeteer session IDs. No parameters needed. Use this to verify the server is running and see which browser sessions are available before interacting with pages.

browser_select_modeA

Check available browser modes and set the session default. CALL THIS FIRST before using any other browser tool. Two modes: "connect" (connects to user's Chrome via debug port) or "headless" (invisible background Puppeteer browser, always available). Call without params to see what is available. Call with mode="headless" or mode="connect" to set the default for all subsequent calls. If only headless is available, it is auto-selected.

browser_screenshotA

Capture a screenshot of the current page as a PNG image. Returns the actual rendered visual — you can see and describe layout, design, text, images, errors, and any visual content. Use this after browser_navigate to see a page, after browser_click to verify what happened, or anytime you need to visually inspect the browser. Pass sessionId to screenshot a specific headless session.

browser_get_urlA

Get the current URL of the browser page. Useful to verify navigation succeeded, check the current location after redirects, or confirm which page the browser is on.

browser_navigateA

Navigate the browser to any URL — websites, localhost, web apps. Opens the page and waits for it to fully load. In headless mode, you can set custom viewport dimensions (width/height in pixels) to simulate desktop (1920x1080), tablet (768x1024), or mobile (375x812) screen sizes. Returns the final URL and sessionId. Pass the returned sessionId to all subsequent tool calls to reuse this browser session. Pass proxyServer or profile (only when NOT reusing an existing sessionId) — both apply at browser launch and cannot be changed on an existing session.

browser_clickA

Click any element on the page by CSS selector — buttons, links, menus, dropdowns, checkboxes, tabs, etc. Examples: "#submit-btn", "button[type=submit]", ".nav-link", "a.login". Waits for the element to be visible, clicks it, then waits for the page to stabilize (handles AJAX, animations, re-renders). Use browser_get_dom first if you need to find the right selector. To click inside an iframe, pass frameIndex from browser_frames (this always uses the simple DOM-click path, since humanClick's mouse simulation only works on the main page). To reach into shadow DOM, prefix the selector with "pierce/" instead.

browser_scrollA

Scroll the page by pixel amount. Use positive y to scroll down, negative y to scroll up. Examples: scroll down 500px → {y: 500}, scroll up → {y: -500}, scroll to bottom → {y: 99999}, scroll right → {x: 500}. Returns the final scroll position. Useful for viewing content below the fold, triggering lazy-loaded images, or reaching elements further down the page.

browser_typeA

Type text into any input field, textarea, search bar, or contenteditable element by CSS selector. Use this to fill out forms, enter search queries, write messages, type login credentials, etc. Examples: browser_type({selector: "input[name=email]", text: "user@example.com"}), browser_type({selector: "#search", text: "search query"}). Combine with browser_click to submit forms after filling them. To type into an element inside an iframe, pass frameIndex from browser_frames. To reach into shadow DOM, prefix the selector with "pierce/" (e.g. "pierce/input#name") — no frameIndex needed for that.

browser_get_domA

Get the HTML source code of the current page (or a scoped subtree). Returns the rendered DOM including dynamically loaded content. Use this to: find CSS selectors for browser_click/browser_type, understand page structure, check element attributes and classes, inspect form fields, or analyze the page content as HTML or plain text. For large pages, prefer browser_snapshot (accessibility-tree view) or scope with selector — a full-page dump can be tens of thousands of characters.

browser_snapshotA

Get a compact accessibility-tree view of the page: interactive and semantic elements (links, buttons, inputs, headings, landmarks) as a small indented text list, each tagged with a stable ref like [ref=e12]. Far cheaper than browser_get_dom for finding what to interact with — use this first when exploring an unfamiliar page. Every ref is also a valid CSS selector — pass [data-mcp-ref="e12"] directly as the selector argument to browser_click, browser_type, browser_hover, or browser_select_option. Pass diff:true to get only what changed since your last browser_snapshot call in this session — much cheaper than a full snapshot for checking the result of a click/type.

browser_console_logsA

Read JavaScript console output from the current page — includes console.log, console.warn, console.error, and console.info messages with timestamps. Use this to: debug JavaScript errors, check for failed API calls, find runtime exceptions, see application logs, or diagnose why something is not working on the page.

browser_executeA

Execute arbitrary JavaScript code on the current page and return the result. Use this for complex interactions that other tools cannot handle: jQuery/Select2 dropdowns (e.g. jQuery("#select").val("value").trigger("change")), reading JavaScript variables, calling page functions, manipulating complex widgets, dispatching custom events, or any DOM manipulation. The code runs in the page context with full access to window, document, jQuery, etc. Return a value from your code and it will be sent back as the result. Examples: Set Select2 dropdown: 'jQuery("#project").val("123").trigger("change")' | Read a value: 'document.querySelector("#total").textContent' | Fill multiple fields: 'jQuery("#name").val("John"); jQuery("#email").val("john@example.com"); "done"' | Trigger form submit: 'document.querySelector("form").submit()' | Check if element exists: '!!document.querySelector(".success-message")'

browser_keyboardA

Press keyboard keys like Enter, Tab, Escape, arrow keys, or key combinations with Ctrl/Shift/Alt modifiers. Use to: submit forms (Enter), navigate between fields (Tab), close modals (Escape), select all text (Ctrl+A), copy/paste, or navigate autocomplete menus (ArrowDown/ArrowUp). If selector is provided, focuses that element first.

browser_hoverA

Hover over an element to trigger dropdown menus, tooltips, hover effects, or any mouseover-activated content. The mouse moves to the element center and stays there. Use this before browser_click when a menu only appears on hover, or to reveal hidden UI elements.

browser_select_optionA

Select an option from a native HTML dropdown by value or visible label text. Also works with Select2/jQuery dropdowns — automatically detects and triggers the correct change events. Examples: browser_select_option({selector: "#country", value: "US"}) or browser_select_option({selector: "#country", label: "United States"}). For complex custom dropdowns that don't use , use browser_execute with jQuery instead.

browser_wait_forA

Wait for a condition before proceeding — an element to appear, text to become visible, or text to disappear. Use this to handle dynamic content, AJAX loading, animations, or any async page updates. Examples: wait for a success message after form submit, wait for a loading spinner to disappear, wait for a modal to open. Timeout defaults to 10 seconds.

browser_handle_dialogA

Handle JavaScript alert(), confirm(), and prompt() dialogs. Call this BEFORE triggering an action that shows a dialog (e.g. before clicking a delete button that shows a confirm). The handler will auto-accept or auto-dismiss the next dialog that appears. For prompt() dialogs, provide promptText to enter text before accepting. If a dialog is already showing, it will be handled immediately.

browser_file_uploadB

Upload one or more files to a file input element. Provide the CSS selector of the element and an array of absolute file paths to upload. Works with single and multiple file inputs. Examples: browser_file_upload({selector: "input[type=file]", paths: ["/home/user/document.pdf"]}) or multiple files: browser_file_upload({selector: "#photos", paths: ["/tmp/img1.png", "/tmp/img2.png"]})

browser_drag_dropA

Drag an element and drop it onto another element. Use for kanban boards, sortable lists, reordering items, or any drag-and-drop interface. Simulates a full human drag: mousedown on source → mousemove to target → mouseup on target, with proper drag events dispatched.

browser_tabsA

Manage browser tabs — list open tabs, create new tabs, switch between tabs, close tabs, or screenshot a specific tab WITHOUT switching to it. Use this for: OAuth flows that open popups, links that open in new tabs, multi-page workflows, comparing two open tabs side by side, or cleaning up tabs. Manages pages within the Puppeteer browser session.

browser_navigate_backA

Navigate back to the previous page in browser history — equivalent to clicking the browser back button. Use after navigating to a page and wanting to return to the previous one.

browser_cookiesA

Get, set, delete, or clear cookies on the current page. Use this to inspect auth/session cookies, inject a saved session to skip a login flow, or clean up before testing a fresh session.

browser_storageA

Get, set, or clear localStorage/sessionStorage on the current page. Use this alongside browser_cookies to save and restore full session state (auth tokens, app state) and skip repeated login flows across calls.

browser_extractA

Extract structured text or attributes from the page without dumping raw HTML. With no selector, returns the page's visible text content (like a reader view). With a selector, returns the text (or a given attribute) of the matching element(s). Much smaller and more directly useful than browser_get_dom when you just need the content, not the markup.

browser_pdfA

Render the current page to a PDF file on disk (not returned inline — PDFs are too large for tool-call context). Returns the saved path and file size. Headless mode only.

browser_networkA

Inspect or control network traffic for a headless session: list captured requests (url, method, resourceType, status), clear the log, or block requests by resource type (e.g. "image", "font", "stylesheet", "media") to speed up pages and reduce noise. Blocking is only available in headless mode — connect mode (the user's real Chrome tab) never intercepts requests, to avoid adding latency or risk to their live browsing.

browser_downloadsA

Configure where downloads are saved for this session, and list what has landed there. Call action "configure" once with a directory before triggering a download (e.g. clicking a download link), then action "list" afterward to see the files and their sizes.

browser_statsA

Cheap pre-check for page size and complexity: DOM node count, approximate HTML length, and approximate visible-text length. Call this BEFORE browser_get_dom on an unfamiliar page to decide whether a full dump is safe, or whether you should scope it (selector/maxLength) or use browser_snapshot/browser_extract instead.

browser_macroB

Save a named sequence of tool calls once, then replay it by name later against any session — collapses a repeated multi-step flow (e.g. "login") into a single tool call instead of N round trips.

browser_authA

Set or clear HTTP Basic/Digest Auth credentials for the current page — the kind of login prompt the browser itself shows (a native dialog), not a login form on the page. Call this BEFORE browser_navigate to a protected URL (e.g. a password-protected staging site). Use clear:true to remove credentials afterward.

browser_emulateA

Emulate device/environment conditions on the current page — device presets, dark mode, reduced motion, timezone, locale, geolocation, permissions, network throttling, and CPU throttling. Pass only the options you need; each is applied independently and the response lists what was actually changed. Geolocation and permissions are scoped to the page's current origin, so navigate first.

browser_framesA

List all frames (iframes) on the current page, each with an index, url, and name. Pass that index as frameIndex to browser_click, browser_type, or browser_get_dom to target content inside that iframe — the main document is not part of this list navigation-wise, but IS included as index 0 if it has children. Note: shadow DOM (unlike iframes) needs no special tool — prefix any selector with "pierce/" (e.g. "pierce/.my-shadow-element") in browser_click/type/get_dom/hover/select_option to reach into shadow roots.

browser_form_fillA

Fill multiple form fields in one call: { "#email": "user@example.com", "#country": "IN" }. Automatically uses the right interaction per element — types into text inputs/textareas, uses native selection, and clicks checkboxes/radios (interpreting a truthy value as "check it"). Optionally click a submit button afterward.

browser_reloadA

Reload the current page. Equivalent to pressing the browser's refresh button. Use ignoreCache:true for a hard refresh that bypasses the browser cache (equivalent to Ctrl+Shift+R) — useful when testing a deployed change that the page might otherwise serve stale/cached.

browser_navigate_forwardA

Navigate forward to the next page in browser history — equivalent to clicking the browser forward button. Only works after browser_navigate_back (or a real back navigation) has happened; if there is no forward history, the page simply stays put.

browser_page_errorsA

Read uncaught JavaScript exceptions thrown by the page (window "error" events / unhandled errors in page scripts). Different from browser_console_logs: console_logs only captures explicit console.log/warn/error calls, this captures real crashes — a script that threw and stopped running, even if it never called console.error itself. Check this whenever a page seems "stuck" or a feature silently does nothing.

browser_clipboardA

Read or write the system clipboard from the page context. Automatically grants the clipboard-read/clipboard-write permissions for the current origin first — you do NOT need to call browser_emulate separately for this. Use to verify a "Copy to clipboard" button worked, or to paste text a page doesn't expose an input for.

browser_get_elementA

Get full detail on ONE element by selector: tag name, all attributes, bounding box, visibility, text content, and a curated set of computed CSS properties. Complements browser_extract (text/attribute only) and browser_snapshot (a whole page of elements) — use this when you need to inspect a single specific element closely, e.g. checking why something looks wrong, or confirming an element is actually visible/enabled before interacting.

browser_findA

Search for element(s) by visible text and/or ARIA role, without needing a CSS selector or a full browser_snapshot dump. e.g. find({text: "sign in"}) or find({role: "button", text: "submit"}). Matches are tagged with a ref, usable as '[data-mcp-ref="f1"]' in browser_click/type/hover/select_option — same mechanism as browser_snapshot, but with an "f" prefix so the two never collide.

browser_profilesA

List or delete named persistent browser profiles. A profile is a Chrome user-data directory whose cookies, localStorage, and login state survive across separate MCP server runs — unlike a normal session, which is wiped once it ends. Create one implicitly by passing profile: "name" to browser_navigate on a brand-new session (letters, digits, dash, underscore only in the name).

Prompts

Interactive templates invoked by user choice

NameDescription

No prompts

Resources

Contextual data attached and managed by the client

NameDescription

No resources

TDQS

A3.6/5.0

Scored across 40 tools

Disambiguation3/5

Several inspection tools (browser_get_dom, browser_snapshot, browser_extract, browser_get_element, browser_find, browser_stats) all retrieve page content with subtle differences, making misselection likely. Interaction and navigation tools are mostly distinct, but the large surface adds ambiguity. Descriptions help, but the overlap is still notable.

Naming Consistency5/5

All 40 tools use the browser_ prefix with snake_case verb/noun naming (e.g., browser_navigate, browser_click, browser_get_dom). There are no deviations in case or structure. The pattern is highly predictable throughout.

Tool Count2/5

40 tools is well above the 15-tool guideline and feels excessive for a browser control server. While the domain is broad, many capabilities could be consolidated (e.g., the inspection tools). The large count forces agents to navigate a heavy menu for common actions.

Completeness4/5

The surface covers navigation, interaction, inspection, network, storage, cookies, emulation, dialogs, uploads/downloads, tabs, frames, and more, with no obvious major gaps. A minor missing piece is an explicit session-close tool, though closing all tabs is a workaround. Overall coverage is strong.

Maintenance

ActivityMaintained
ResponsivenessNo issues