MCP Browser Automation Server
# MCP Browser Automation
This is demo project to practice Model Context Protocol based server implemenation for automating browsing with Playwright. It interacts with a Claude Desktop client to accept user prompts and use server to control browser.
<a href="https://glama.ai/mcp/servers/hokppvk1dy"><img width="380" height="200" src="https://glama.ai/mcp/servers/hokppvk1dy/badge" alt="Browser Automation Server MCP server" /></a>
## Pre-requisites
- [Playwright](https://playwright.dev/)
- [Claude Desktop](https://claude.ai/download)
- [Node.js](https://nodejs.org/en/download/)
## Building
1. Clone the repository: `git clone https://github.com/hrmeetsingh/mcp-browser-automation.git`
2. Install dependencies: `npm install`
3. Verify the output executables are present in `dist` folder
## Integration
1. Create a configuration file in `~/Application\ Support/Claude/claude_desktop_config.json` (This is for macOS)
2. Copy the following to the file:
```json
{
"mcpServers": {
"mcp-browser-automation": {
"command": "node",
"args": ["/path/to/mcp-browser-automation/dist/index.js"]
}
}
}
```
3. Start Claude Desktop
## Usage
1. Open Claude Desktop
2. Start a new conversation to open a browser and navigate to a URL
## Example
- Added MCP Server options

- Navigating to a URL and doing actions with playwright

TDQS
Scored across 12 tools
Most tools have distinct purposes, but there is some potential confusion between playwright_click and playwright_select, as both involve interacting with page elements. The HTTP methods (GET, POST, PUT, PATCH, DELETE) are clearly differentiated, and other tools like navigate, screenshot, and evaluate have unique functions.
All tools follow a consistent 'playwright_verb' naming pattern, using snake_case throughout. This predictability makes it easy for agents to understand and use the toolset without confusion about naming conventions.
With 12 tools, this server is well-scoped for browser automation, covering essential actions like navigation, interaction, HTTP requests, and debugging. Each tool serves a clear purpose, and the count aligns with the domain's complexity without being overwhelming.
The toolset provides strong coverage for core browser automation tasks, including navigation, element interaction, HTTP methods, and screenshots. A minor gap is the lack of tools for handling browser contexts, tabs, or more advanced JavaScript execution, but agents can still accomplish most workflows effectively.