webcontrol-mcp
Server Configuration
Describes the environment variables required to run the server.
| Name | Required | Description | Default |
|---|---|---|---|
No arguments | |||
Instructions
Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.
This server publishes no instructions, or was last inspected before Glama recorded them.
Capabilities
Features and capabilities supported by this server
Protocol revision2025-11-25
| Capability | Details |
|---|---|
| tools | {} |
Tools
Functions exposed to the LLM to take actions
| Name | Description |
|---|---|
| open_pageA | Open a URL in the controlled browser and wait for it to load. Returns the page title. |
| clickB | Click an element on the current page, located by a CSS selector. |
| fillC | Type text into an input field on the current page. |
| screenshotB | Take a PNG screenshot of the current page. Returns the image as base64. |
| get_textB | Get the visible text content of the page, or of an element matching a CSS selector. |
Prompts
Interactive templates invoked by user choice
| Name | Description |
|---|---|
No prompts | |
Resources
Contextual data attached and managed by the client
| Name | Description |
|---|---|
No resources | |
TDQS
Scored across 5 tools
Each tool maps to a single, clearly distinct browser action (navigate, click, type, capture, read), with no semantic overlap between them. An agent can unambiguously pick the right tool for any given step.
Most tools follow a verb_noun/verb-object pattern (open_page, get_text) or a standard single-action verb (click, fill), which is readable and predictable. The lone noun-form 'screenshot' is a minor deviation from the dominant verb style but still self-explanatory.
Five tools is a reasonable, well-scoped surface for a minimal browser controller and each tool earns its place. It leans slightly thin, since common interactions like navigation or key presses have no dedicated tool.
The core read/interact loop (open, click, fill, screenshot, read text) is covered, but notable gaps exist: no back/forward/reload navigation, no key press (e.g. Enter/Submit), no select/dropdown handling, scroll, or hover. These omissions will force agents to work around common form and multi-page flows.