Browser MCP Server
# Browser MCP Server
A Model Context Protocol (MCP) server that enables AI agents to understand web page structure and content without screenshots. Built with Playwright for reliable browser automation.
## šÆ Purpose
Solves the problem where AI agents give "wrong and random suggestions" about web pages by providing them with structured page data instead of requiring visual screenshots.
## ⨠Features
- š **Page Structure Analysis** - Understand layout, forms, and interactive elements
- š **Smart Content Extraction** - Extract text using CSS selectors
- šÆ **Element Discovery** - Find and analyze specific page elements with accessibility info
- š **Fast & Local** - Uses Playwright's accessibility tree (no screenshots needed)
- š¤ **AI-Optimized** - Designed specifically for AI agent integration
## š¦ Installation
```bash
npm install @your-org/browser-mcp-server
```
## š Usage
### Claude Desktop Integration
Add to your Claude Desktop configuration (`~/Library/Application Support/Claude/claude_desktop_config.json`):
```json
{
"mcpServers": {
"browser": {
"command": "npx",
"args": ["@your-org/browser-mcp-server"]
}
}
}
```
### VS Code + Copilot Chat
Configure in your VS Code MCP settings to use with Copilot Chat.
### Windsurf Integration
Add to your Windsurf MCP configuration.
### Local Development
```bash
# Link for local development
npm link @your-org/browser-mcp-server
# Use in your project
const browserMCP = require('@your-org/browser-mcp-server');
```
## š ļø Available Tools
### `browser_navigate`
Navigate to a URL and get page information.
```json
{
"url": "https://example.com"
}
```
### `browser_analyze_page`
Get comprehensive page structure analysis.
```json
{
"includeContent": true,
"includeForms": true,
"includeLinks": true,
"maxDepth": 2
}
```
### `browser_extract_text`
Extract text content from page or specific elements.
```json
{
"selector": "h1, h2, p",
"includeHidden": false
}
```
### `browser_get_elements`
Find elements with detailed attributes and accessibility info.
```json
{
"selector": "button, input, a[href]",
"includeAttributes": true,
"maxResults": 20
}
```
## š Example AI Conversations
### Understanding Page Flow
```
Human: Go to https://github.com/login and help me understand the login process
AI: I'll navigate to the GitHub login page and analyze its structure.
[Uses browser_navigate and browser_analyze_page]
The login page has:
- 2 input fields: username/email and password (both required)
- 1 "Sign in" submit button
- Links for "Forgot password" and "Create account"
- Form submits to /session with POST method
- Additional 2FA field may appear after initial submission
The typical flow is: enter credentials ā click Sign in ā (optional 2FA) ā redirect to dashboard.
```
### Content Extraction
```
Human: Extract all the main article titles from https://news.ycombinator.com
AI: [Uses browser_navigate and browser_extract_text with CSS selectors]
Found 30 article titles including:
1. "Show HN: My new startup idea"
2. "Why distributed systems are hard"
3. "The future of web development"
[etc...]
```
## š§ Development
```bash
# Clone and setup
git clone <your-repo>
cd browser-mcp-server
npm install
# Install browser
npx playwright install chromium
# Build
npm run build
# Test locally
npm start
```
## šÆ Perfect For
- **AI Agents** that need to understand web page workflows
- **Automated Analysis** of website structure
- **Content Extraction** without visual parsing
- **Form Discovery** for automation planning
- **Accessibility Analysis** of web pages
## šļø Architecture
```
āāā src/
ā āāā index.ts # MCP server entry point
ā āāā server.ts # MCP server implementation
ā āāā browser/ # Browser management
ā ā āāā manager.ts # Playwright browser lifecycle
ā āāā tools/ # MCP tools
ā ā āāā navigate.ts # Page navigation
ā ā āāā analyze.ts # Page structure analysis
ā ā āāā extract.ts # Content extraction
ā ā āāā elements.ts # Element discovery
ā āāā types.ts # TypeScript definitions
```
## š Requirements
- Node.js 18+
- Chromium (auto-installed via Playwright)
- MCP-compatible AI agent (Claude Desktop, VS Code, Windsurf, etc.)
## š License
MIT
---
**Built for AI agents to understand the web, not just see it.** š¤šTDQS
Scored across 4 tools
Each tool has a distinct purpose: navigation, high-level page analysis, text extraction, and element inspection. There is slight overlap between analyze_page and get_elements, but the descriptions clarify that one provides an overview while the other targets specific selectors.
All tools follow a consistent 'browser_verb_noun' pattern (navigate, analyze_page, extract_text, get_elements). This makes the set predictable and easy to understand.
With only 4 tools, the server is well-scoped for its apparent read-only browser analysis purpose. Each tool is necessary and there is no redundancy.
The tool set covers the core workflow of navigating to a page and extracting both text and structural information. It lacks interaction capabilities like clicking or typing, but for a read-only analysis server this is a minor gap rather than a fatal omission.