MCP-Midscene
Server Configuration
Describes the environment variables required to run the server.
| Name | Required | Description | Default |
|---|---|---|---|
No arguments | |||
Instructions
Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.
This server publishes no instructions, or was last inspected before Glama recorded them.
Capabilities
Server capabilities have not been inspected yet.
Tools
Functions exposed to the LLM to take actions
| Name | Description |
|---|---|
| midscene_playwright_exampleC | Provides Playwright code examples for Midscene. If users need to generate Midscene test cases, they can call this method to get sample Midscene Playwright test cases for generating end-user test cases. Each step must first be verified using the mcp method, and then the final test case is generated based on the playwright example according to the steps executed by mcp |
| midscene_navigateA | Navigates the browser to the specified URL. Always opens in the current tab. |
| midscene_get_tabsA | Retrieves a list of all open browser tabs, including their ID, title, and URL. |
| midscene_set_active_tabA | Switches the browser's focus to the tab specified by its ID. Use midscene_get_tabs first to find the correct tab ID. |
| midscene_aiWaitForA | Waits until a specified condition, described in natural language, becomes true on the page. Polls the condition using AI. |
| midscene_aiAssertA | Asserts that a specified condition, described in natural language, is true on the page. Polls the condition using AI. |
| midscene_aiKeyboardPressC | Presses a specific key on the keyboard. |
| midscene_screenshotB | Captures a screenshot of the currently active browser tab and saves it with the given name. |
| midscene_aiTapC | Locates and clicks an element on the current page based on a natural language description (selector). |
| midscene_aiScrollC | Scrolls the page or a specified element. Can scroll by a fixed amount or until an edge is reached. |
| midscene_aiInputC | Inputs text into a specified form field or element identified by a natural language selector. |
| midscene_aiHoverA | Moves the mouse cursor to hover over an element identified by a natural language selector. |
Prompts
Interactive templates invoked by user choice
| Name | Description |
|---|---|
No prompts | |
Resources
Contextual data attached and managed by the client
| Name | Description |
|---|---|
No resources | |
TDQS
Scored across 12 tools
Each tool has a clearly distinct purpose with no overlap. For example, midscene_aiAssert checks conditions, midscene_aiHover hovers over elements, and midscene_aiInput inputs text, all targeting different actions in browser automation. The tools are well-defined and unlikely to cause misselection.
All tools follow a consistent 'midscene_' prefix and snake_case pattern, with clear verb_noun combinations like aiAssert, aiHover, and get_tabs. This predictability makes it easy for agents to understand and use the toolset without confusion.
With 12 tools, the count is well-scoped for browser automation and testing. It covers core actions like navigation, interaction, waiting, and tab management, with each tool earning its place without being excessive or insufficient for the domain.
The toolset provides comprehensive coverage for browser automation, including navigation, element interaction, condition checking, and tab management. A minor gap is the lack of tools for handling browser contexts or windows, but core workflows are fully supported, allowing agents to perform most common tasks effectively.