mcp-textbrowser
Server Configuration
Describes the environment variables required to run the server.
| Name | Required | Description | Default |
|---|---|---|---|
No arguments | |||
Instructions
Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.
This server publishes no instructions, or was last inspected before Glama recorded them.
Capabilities
Features and capabilities supported by this server
Protocol revision2025-11-25
| Capability | Details |
|---|---|
| tools | {
"listChanged": true
} |
Tools
Functions exposed to the LLM to take actions
| Name | Description |
|---|---|
| browser_navigateA | Open a URL in the browser and return the page context (DOM elements + OCR text). Default text-only mode — screenshot is captured only for OCR then discarded, zero image tokens. Set visual=true to also receive the PNG. |
| browser_clickA | Click an element on the current page. Identify the target by CSS selector, XPath, or visible text. Returns updated page context after the click. |
| browser_typeA | Type text into an input field on the current page. By default clears the field first. Returns updated page context. |
| browser_scrollA | Scroll the current page. Use direction (up/down/left/right) with an optional pixel amount, or provide a selector to scroll that element into view. |
| browser_screenshotA | Capture the current page and return an OCR-based text map. The screenshot itself is discarded unless visual=true. |
| browser_readA | Read the current page context (DOM elements + OCR text) without navigating or clicking anything. |
| browser_evaluateA | Execute JavaScript in the current page context and return the result. Restricted to safe DOM read operations — eval, Function, fetch, WebSocket, and document.write are blocked. |
Prompts
Interactive templates invoked by user choice
| Name | Description |
|---|---|
No prompts | |
Resources
Contextual data attached and managed by the client
| Name | Description |
|---|---|
No resources | |
TDQS
Scored across 7 tools
Each tool targets a distinct browser action (navigate, click, type, scroll, read, screenshot, evaluate) with no overlap in purpose.
All tool names follow the consistent 'browser_verb' pattern, making the set predictable and easy to understand.
7 tools is well-scoped for a browser automation server, covering essential interactions without being excessive or insufficient.
Covers core browsing actions well, but lacks explicit back/forward navigation or page refresh, which are minor gaps for typical workflows.