Skip to main content
Glama
README.md
# Smart DOM MCP Server

Token-efficient browser automation for AI agents. Drop-in alternative to Playwright MCP that filters the DOM before returning it — only interactive elements, grouped by page section.

## Why?

Playwright MCP sends the **full accessibility tree** on every interaction. On complex pages, that's 10-40K tokens per snapshot. Your 128K context fills up after ~3-12 actions.

Smart DOM filters to **only interactive elements** (buttons, links, inputs, selects, dialogs) and caps output at 30 elements per section with dedup. Result: **3-13x fewer tokens**.

| Site | Playwright MCP | Smart DOM | Reduction |
|---|---|---|---|
| Amazon | ~37K tokens | ~3.2K tokens | **11x** |
| Hacker News | ~10K tokens | ~757 tokens | **13x** |
| GitHub | ~10K tokens | ~2.5K tokens | **4x** |
| Wikipedia | ~8.4K tokens | ~2.6K tokens | **3.3x** |

**Actions in 128K context window:**
- Amazon: 3 (Playwright) → **39** (Smart DOM)
- Hacker News: 12 → **169**

## Install

```bash
git clone https://github.com/Frontrunner0x/smart-dom-mcp.git
cd smart-dom-mcp
npm install
npx playwright install chromium
```

## Setup

### Claude Code
```bash
claude mcp add smart-dom -s user -- node /path/to/smart-dom-mcp/src/index.js
```

### Claude Desktop / Cursor / etc.
Add to your MCP config:
```json
{
  "mcpServers": {
    "smart-dom": {
      "command": "node",
      "args": ["/path/to/smart-dom-mcp/src/index.js"]
    }
  }
}
```

## Tools

### `navigate`
Open a URL and get filtered interactive elements grouped by page section.
```
navigate({ url: "https://github.com" })
→ { sections: { header: [...], main: [...] }, totalShown: 90 }
```

### `dom_summary`
Re-scan current page for interactive elements. Use `focus` to get only one section.
```
dom_summary({ focus: "main" })
→ { section: "main", elements: [...], count: 28 }
```

### `dom_act`
Click, type, select, check/uncheck elements by their ref (from navigate/dom_summary).
```
dom_act({ ref: "@e5", action: "click" })
→ { ok: true, action: "clicked", label: "Sign in" }

dom_act({ ref: "@e12", action: "type", value: "hello world" })
→ { ok: true, action: "typed", value: "hello world" }
```

### `dom_read`
Extract readable content: tables, headings, alerts, form values, or plain text.
```
dom_read({ selector: "main" })
→ { headings: [...], tables: [...], text: "..." }
```

### `screenshot`
Visual fallback. Capture viewport, full page, or a specific element.
```
screenshot({ fullPage: true })
screenshot({ selector: "#results" })
```

### `wait`
Wait for text to appear, disappear, or a fixed delay.
```
wait({ text: "Results loaded" })
wait({ seconds: 2 })
```

## How it works

1. **Persistent Chromium session** — browser stays open across tool calls (no cold start per action)
2. **DOM filter** — `querySelectorAll` for interactive elements only, grouped by ARIA landmarks (header, nav, main, sidebar, footer, dialog)
3. **Dedup + cap** — identical label+role combos are deduplicated, max 30 elements per section
4. **Ref system** — each element gets a stable ref (`@e0`, `@e1`, ...) for use with `dom_act`
5. **Playwright actions** — `dom_act` uses Playwright's native click/fill/selectOption for reliability with SPAs

## Limitations

- **Headless only** — no access to your logged-in browser sessions (use Chrome DevTools MCP for that)
- **No iframe support** — elements inside iframes are not captured (planned)
- **No shadow DOM** — web components with shadow roots are not traversed (planned)
- **Single tab** — one page at a time

## License

MIT