Skip to main content
Glama

Tactab šŸŒšŸ¤–

Tactile Browser Control & Multimodal Vision Bridge for AI Agents

Glama Score MCP Standard Manifest V3 License: MIT Node.js

Tactab (Tactile + Tab) is a high-performance, local, and open-source bridge that gives any MCP-compliant AI client (Cursor, Claude Desktop, Antigravity, Windsurf, Cline, etc.) live visual eyes and tactile hands in your Google Chrome browser using the Model Context Protocol (MCP), WebSockets, and a Chrome Extension (Manifest V3).


⚔ Highlights

  • 🌐 Works with Any Plan (Free or Paid): Seamlessly connects whether you are on free tiers or paid Pro/Enterprise plans. Zero subscriptions or paid cloud automation platforms (like Browserbase or MultiOn) required.

  • šŸ‘ļø Visual Multimodal Vision & DOM Control: Enables your AI to capture high-res PNG screenshots for visual layout inspection, alongside full DOM scraping, button clicks, form filling, and navigation.

  • šŸ”’ Private & 100% Local: Operates strictly over local loopback (127.0.0.1). Your session cookies, authenticated tabs (AWS, Jira, GitHub), and local dev servers (localhost:3000) never leave your machine.

  • 🧩 Zero Extension Bloat (One Extension for All Agents): Eliminates the need to install separate, heavy browser extensions for each AI tool. Tactab acts as a single, lightweight gateway that connects Claude, Cursor, Antigravity, or custom agents through standard MCP.

  • šŸ”„ Universal MCP Standard: Plug-and-play with Claude Desktop, Cursor, Antigravity, Windsurf, or custom AI agents over standard I/O (stdio).


Related MCP server: real-browser-mcp

šŸ—ļø Architecture

ā”Œā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”         STDIO (JSON-RPC)         ā”Œā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”
│     AI Agent / IDE Client       │  ◄─────────────────────────────► │    Tactab MCP Server (Node)   │
│ (Cursor, Claude, Antigravity)   │                                  │         tactab/server         │
ā””ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”˜                                  ā””ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”¬ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”˜
                                                                                     │
                                                                            WebSocket (ws://127.0.0.1:8765)
                                                                                     │
ā”Œā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”     chrome.tabs.sendMessage     ā”Œā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā–¼ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”
│        Webpage DOM (Tab)        │  ◄─────────────────────────────► │     Tactab Extension (MV3)    │
│          (content.js)           │                                  │    (background.js + badge)    │
ā””ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”˜                                  ā””ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”˜
  1. AI Agent invokes the automate_website tool over standard input/output (stdio).

  2. Tactab Bridge Server (bridge.js) listens on 127.0.0.1:8765 and translates tool calls into JSON WebSocket messages.

  3. Tactab Chrome Extension Service Worker (background.js) receives commands, manages connection status, and routes to active tabs.

  4. Content Script (content.js) executes actions inside the active web page and returns structured results back up the pipeline.


šŸ“‚ Project Structure

tactab/
│
ā”œā”€ā”€ extension/                      # Chrome Extension (Manifest V3)
│   ā”œā”€ā”€ manifest.json               # Extension manifest declaration
│   ā”œā”€ā”€ background.js               # Service worker WebSocket client & badge manager
│   └── content.js                  # DOM interaction & scraping execution engine
│
ā”œā”€ā”€ server/                         # Local Node.js MCP Server
│   ā”œā”€ā”€ package.json                # Dependencies (@modelcontextprotocol/sdk, ws)
│   └── bridge.js                   # MCP Stdio transport + WebSocket bridge
│
ā”œā”€ā”€ AGENTS.md                       # Universal Multi-Agent Handoff Guardrail
ā”œā”€ā”€ PLAN.md                         # Active milestone tracker & Decision Log
ā”œā”€ā”€ mcp_config.example.json         # Universal MCP configuration template
ā”œā”€ā”€ LICENSE                         # MIT License
ā”œā”€ā”€ .gitignore                      # Git ignore file (excludes plane.md, node_modules)
└── README.md                       # Documentation & setup guide

šŸš€ Quick Start Guide

1. Install Server Dependencies

Open your terminal in the server/ directory:

cd server
npm install

(On Windows systems where script execution policies restrict npm, run npm.cmd install)


2. Load the Chrome Extension

  1. Open Google Chrome and navigate to chrome://extensions/.

  2. Toggle on Developer mode in the top-right corner.

  3. Click the Load unpacked button in the top-left corner.

  4. Select the extension/ folder in this repository.

  5. Tactab — Browser MCP Bridge will now appear in your extensions list.

    • When connected to the local MCP server, the badge displays green ON.

    • When disconnected, it displays red OFF and automatically retries every 5 seconds.


3. Configure Your AI Client

Add Tactab to your AI client's MCP configuration using your absolute path to server/bridge.js:

A. Claude Desktop

Config file location:

  • Windows: %APPDATA%\Claude\claude_desktop_config.json

  • macOS: ~/Library/Application Support/Claude/claude_desktop_config.json

  • Linux: ~/.config/Claude/claude_desktop_config.json

{
  "mcpServers": {
    "tactab": {
      "command": "node",
      "args": [
        "/ABSOLUTE/PATH/TO/tactab/server/bridge.js"
      ]
    }
  }
}

Windows Note: Use escaped backslashes in paths, e.g.: "C:\\path\\to\\tactab\\server\\bridge.js"

B. Cursor IDE

Open Cursor Settings āž” Features āž” MCP Servers (or edit .cursor/mcp.json):

{
  "mcpServers": {
    "tactab": {
      "command": "node",
      "args": [
        "/ABSOLUTE/PATH/TO/tactab/server/bridge.js"
      ]
    }
  }
}

C. Antigravity / Windsurf / Cline

Add the exact same JSON block into your environment's MCP server configuration file.


šŸ› ļø Supported Browser Actions

The automate_website tool supports the following actions on any standard webpage:

Action

Required Payload

Optional Payload

Description

take_screenshot

(none)

(none)

Captures a high-resolution visual PNG screenshot of the active tab for AI multimodal vision analysis.

scrape_data

(none)

maxLength (number)

Extracts page title, URL, meta description, and clean inner text content.

click_button

selector (string)

text (string)

Clicks an element by CSS selector with fallback matching by inner text.

fill_form

selector (string), value (string)

(none)

Sets input/textarea values and dispatches standard input/change events for modern frameworks (React, Vue, Angular).

get_elements

(none)

selector (default "a"), limit (default 20)

Scrapes multiple elements matching a selector, returning text, links, classes, and IDs.

scroll

(none)

direction ("down", "up", "top", "bottom"), amount (pixels)

Scrolls the active tab smoothly.

navigate

url (string)

(none)

Navigates the active tab to a new URL.

get_html

(none)

selector (default "body"), maxLength (number)

Extracts outer HTML of a specific element or the entire page.


šŸ’” Example Prompts to Ask Your AI

Once configured, simply instruct your AI in natural language:

  • šŸ“ø "Take a screenshot of my current tab and inspect the layout design."

  • šŸ“Š "Look at the charts on my active tab and summarize the visual data."

  • šŸ” "Look at my active Chrome tab and summarize what you see."

  • āœļø "Fill in the login form with test credentials and submit."

  • šŸ”— "Extract all product titles and links visible on the page."

  • šŸ“œ "Scroll down 800 pixels and extract the pricing table."


šŸ¤– Multi-Agent Handoff Protocol (The "Triple-Anchor" Architecture)

Whether you are switching between specialized models (e.g. Cursor for rapid coding, Claude for system architecture), managing token budgets, or navigating provider rate limits, this repository includes the Triple-Anchor Architecture for seamless context transfer with zero loss of progress or hallucination.

ā”Œā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”
│ Anchor 1: AGENTS.md (Universal Agent Guardrail)               │
│ -> Tells any newly opened agent: "Read PLAN.md and Git first" │
ā””ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”¬ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”˜
                               │
       ā”Œā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”“ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”
       ā–¼                                               ā–¼
ā”Œā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”   ā”Œā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”
│ Anchor 2: PLAN.md             │   │ Anchor 3: Git Status & Diff   │
│ (The Intent & Architecture)   │   │ (The Ground Truth of Code)    │
│ • Completed tasks             │   │ • Exact lines of code changed │
│ • Next pending tasks          │   │ • Clean syntax state          │
│ • "Decision Log" & Gotchas    │   │ • Zero hallucinated changes   │
ā””ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”˜   ā””ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”˜

How to Use in Any Project

Copy AGENTS.md and PLAN.md into the root of any codebase:

  1. Anchor 1 (AGENTS.md): Automatically read by modern AI tools (Cursor, Antigravity, Claude, ChatGPT CLI). It instructs the model to inspect PLAN.md and git status before modifying any files.

  2. Anchor 2 (PLAN.md): Contains the live checklist and Decision Log. The Decision Log is critical: it prevents the incoming agent from accidentally reverting intentional architecture choices (such as why a heartbeat ping was added).

  3. Anchor 3 (Git Milestones): When your current agent approaches its token limit, tell it:

    "Commit your progress to Git and update PLAN.md."

  4. Instant Continuation: Open your next AI agent and simply prompt:

    "Continue." The new model reads AGENTS.md āž” PLAN.md āž” git status, and picks up immediately with zero lost context.


šŸ”§ Troubleshooting

  • "Chrome extension is disconnected": Make sure Google Chrome is open, the extension is loaded, and you have an active standard website open (not a restricted internal page like chrome://extensions or about:blank).

  • Port Customization: By default, the bridge uses WebSocket port 8765 (avoiding conflict with standard HTTP 8080 proxies). To use a different port, set the WS_PORT environment variable before launching (e.g. set WS_PORT=8090 / export WS_PORT=8090) and update the port in extension/background.js.


IMPORTANT NOTICE: This software is provided for personal workflow automation, educational research, and authorized testing purposes only.

  1. Compliance with Terms of Service: Users are solely responsible for ensuring that their automated actions, scraping requests, and web interactions comply with all applicable local, national, and international laws, as well as the Terms of Service, Acceptable Use Policies, and robots.txt guidelines of any websites visited.

  2. No Unauthorized Access: This software must not be used to bypass authentication barriers, paywalls, CAPTCHAs, rate limits, or security controls, nor to access or extract proprietary, copyrighted, or sensitive personal data without explicit permission.

  3. Limitation of Liability: The author(s) and contributor(s) of this project assume no liability or responsibility for any misuse, website bans, legal disputes, data loss, damages, or consequences resulting from the installation or execution of this software. By using this project, you agree to assume all associated risks and responsibilities.


šŸ“„ License

Released under the MIT License. Free for open source and commercial use.

Available Tools

1 tool
automate_websiteB

Executes custom browser actions on the currently active webpage via the local Chrome extension bridge.

ParametersJSON Schema
NameRequiredDescriptionDefault
actionYesThe targeted script action. Supported actions: 'take_screenshot' (capture visual PNG screenshot of the tab for multimodal vision), 'scrape_data' (get text content), 'click_button' (click element by selector), 'fill_form' (type into input/textarea), 'get_elements' (query elements), 'scroll' (scroll down/up), 'navigate' (open URL), 'get_html' (extract HTML).
payloadNoKey-value parameters needed for the action (e.g. { selector: 'button.submit' }, { selector: 'input[name=q]', value: 'search text' }, { direction: 'down', amount: 500 }, { url: 'https://example.com' })

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It only discloses the transport (local Chrome extension bridge), omitting that several actions are mutating (click_button, fill_form, navigate), whether they require user-visible permission, what happens on failure, or what the result looks like. For a tool that can drive page state, that is a significant omission.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler. It conveys verb, resource, and mechanism without redundancy, which is exactly the right size for this tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The parameters are fully documented via schema, but with no annotations and no output schema the description should still flag operational requirements (extension installed, tab focused) and the mutating nature of several actions. It leaves an agent to infer those, so it is minimally adequate rather than complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the action enum-like list plus payload examples are fully documented in the schema, so the baseline of 3 applies. The description adds no syntax, format, or default information beyond what the schema already conveys.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: executes browser actions on the currently active webpage. It also names the execution channel (local Chrome extension bridge), which tells an agent what this tool is. It does not enumerate the action types, but the schema covers that, so the gap is minor and there are no siblings to differentiate from.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use or when-not-to-use guidance, no alternatives, and no stated prerequisite such as an active tab or an installed extension. The phrase 'currently active webpage' gestures at context but does not say when an agent should reach for this tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 1 tool updatev1.0.0
    • First observedautomate_website

TDQS

B3.3/5.0

Scored across 1 tool

Disambiguation5/5

Only one tool exists, so there is no possibility of selecting the wrong tool. Its purpose is clearly stated as executing browser actions via a Chrome extension bridge.

Naming Consistency5/5

The single tool follows a clear verb_noun snake_case pattern (automate_website). There are no mixed conventions or naming conflicts to penalize.

Tool Count3/5

A single tool is borderline thin for a browser automation server. It could be intentional as a generic executor, but the set lacks distinct operations that agents would typically expect.

Completeness2/5

The surface provides one generic action executor with no explicit operations, parameters, or lifecycle coverage. Agents lack discoverable building blocks for common browser automation tasks, creating likely gaps.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    D
    maintenance
    An MCP server and Chrome extension that allows users to select browser DOM elements via a keyboard shortcut to provide detailed technical context to AI coding tools. It captures HTML attributes, CSS styles, and React component metadata, enabling agents to analyze and modify web elements directly.
    598
    MIT
  • A
    license
    Not graded
    quality
    D
    maintenance
    An extension-based MCP server that enables AI assistants to control your browser, leveraging existing sessions and login states for automation and content analysis. It provides over 20 tools for semantic tab search, interactive element manipulation, and network monitoring directly within your daily Chrome environment.
    MIT