tactab
Tactab is a local MCP server that gives any MCP-compliant AI client tactile control and visual access to your active Google Chrome tab.
Take high-resolution PNG screenshots for multimodal vision analysis.
Scrape page text, title, URL, and meta description.
Click buttons/elements by CSS selector or inner text.
Fill input/textarea fields and dispatch input/change events.
Query multiple elements (links, text, classes, IDs) by selector.
Scroll the page down, up, to top, or bottom by pixels.
Navigate the active tab to a specified URL.
Extract outer HTML of a specific element or the whole page.
Connect via standard stdio to clients like Claude Desktop, Cursor, Windsurf, Cline, etc.
Operates 100% locally over loopback for privacy.
Enables AI agents to control and inspect Google Chrome tabs through the Tactab extension, including taking screenshots, scraping page data, clicking elements, filling forms, scrolling, navigating, and extracting HTML.
Tactab šš¤
Tactile Browser Control & Multimodal Vision Bridge for AI Agents
Tactab (Tactile + Tab) is a high-performance, local, and open-source bridge that gives any MCP-compliant AI client (Cursor, Claude Desktop, Antigravity, Windsurf, Cline, etc.) live visual eyes and tactile hands in your Google Chrome browser using the Model Context Protocol (MCP), WebSockets, and a Chrome Extension (Manifest V3).
ā” Highlights
š Works with Any Plan (Free or Paid): Seamlessly connects whether you are on free tiers or paid Pro/Enterprise plans. Zero subscriptions or paid cloud automation platforms (like Browserbase or MultiOn) required.
šļø Visual Multimodal Vision & DOM Control: Enables your AI to capture high-res PNG screenshots for visual layout inspection, alongside full DOM scraping, button clicks, form filling, and navigation.
š Private & 100% Local: Operates strictly over local loopback (
127.0.0.1). Your session cookies, authenticated tabs (AWS, Jira, GitHub), and local dev servers (localhost:3000) never leave your machine.š§© Zero Extension Bloat (One Extension for All Agents): Eliminates the need to install separate, heavy browser extensions for each AI tool. Tactab acts as a single, lightweight gateway that connects Claude, Cursor, Antigravity, or custom agents through standard MCP.
š Universal MCP Standard: Plug-and-play with Claude Desktop, Cursor, Antigravity, Windsurf, or custom AI agents over standard I/O (
stdio).
Related MCP server: real-browser-mcp
šļø Architecture
āāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāā STDIO (JSON-RPC) āāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāā
ā AI Agent / IDE Client ā āāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāŗ ā Tactab MCP Server (Node) ā
ā (Cursor, Claude, Antigravity) ā ā tactab/server ā
āāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāā āāāāāāāāāāāāāāāāā¬āāāāāāāāāāāāāāāā
ā
WebSocket (ws://127.0.0.1:8765)
ā
āāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāā chrome.tabs.sendMessage āāāāāāāāāāāāāāāāāā¼āāāāāāāāāāāāāāāā
ā Webpage DOM (Tab) ā āāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāŗ ā Tactab Extension (MV3) ā
ā (content.js) ā ā (background.js + badge) ā
āāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāā āāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāAI Agent invokes the
automate_websitetool over standard input/output (stdio).Tactab Bridge Server (
bridge.js) listens on127.0.0.1:8765and translates tool calls into JSON WebSocket messages.Tactab Chrome Extension Service Worker (
background.js) receives commands, manages connection status, and routes to active tabs.Content Script (
content.js) executes actions inside the active web page and returns structured results back up the pipeline.
š Project Structure
tactab/
ā
āāā extension/ # Chrome Extension (Manifest V3)
ā āāā manifest.json # Extension manifest declaration
ā āāā background.js # Service worker WebSocket client & badge manager
ā āāā content.js # DOM interaction & scraping execution engine
ā
āāā server/ # Local Node.js MCP Server
ā āāā package.json # Dependencies (@modelcontextprotocol/sdk, ws)
ā āāā bridge.js # MCP Stdio transport + WebSocket bridge
ā
āāā AGENTS.md # Universal Multi-Agent Handoff Guardrail
āāā PLAN.md # Active milestone tracker & Decision Log
āāā mcp_config.example.json # Universal MCP configuration template
āāā LICENSE # MIT License
āāā .gitignore # Git ignore file (excludes plane.md, node_modules)
āāā README.md # Documentation & setup guideš Quick Start Guide
1. Install Server Dependencies
Open your terminal in the server/ directory:
cd server
npm install(On Windows systems where script execution policies restrict npm, run npm.cmd install)
2. Load the Chrome Extension
Open Google Chrome and navigate to
chrome://extensions/.Toggle on Developer mode in the top-right corner.
Click the Load unpacked button in the top-left corner.
Select the
extension/folder in this repository.Tactab ā Browser MCP Bridge will now appear in your extensions list.
When connected to the local MCP server, the badge displays green ON.
When disconnected, it displays red OFF and automatically retries every 5 seconds.
3. Configure Your AI Client
Add Tactab to your AI client's MCP configuration using your absolute path to server/bridge.js:
A. Claude Desktop
Config file location:
Windows:
%APPDATA%\Claude\claude_desktop_config.jsonmacOS:
~/Library/Application Support/Claude/claude_desktop_config.jsonLinux:
~/.config/Claude/claude_desktop_config.json
{
"mcpServers": {
"tactab": {
"command": "node",
"args": [
"/ABSOLUTE/PATH/TO/tactab/server/bridge.js"
]
}
}
}Windows Note: Use escaped backslashes in paths, e.g.:
"C:\\path\\to\\tactab\\server\\bridge.js"
B. Cursor IDE
Open Cursor Settings ā Features ā MCP Servers (or edit .cursor/mcp.json):
{
"mcpServers": {
"tactab": {
"command": "node",
"args": [
"/ABSOLUTE/PATH/TO/tactab/server/bridge.js"
]
}
}
}C. Antigravity / Windsurf / Cline
Add the exact same JSON block into your environment's MCP server configuration file.
š ļø Supported Browser Actions
The automate_website tool supports the following actions on any standard webpage:
Action | Required Payload | Optional Payload | Description |
| (none) | (none) | Captures a high-resolution visual PNG screenshot of the active tab for AI multimodal vision analysis. |
| (none) |
| Extracts page title, URL, meta description, and clean inner text content. |
|
|
| Clicks an element by CSS selector with fallback matching by inner text. |
|
| (none) | Sets input/textarea values and dispatches standard input/change events for modern frameworks (React, Vue, Angular). |
| (none) |
| Scrapes multiple elements matching a selector, returning text, links, classes, and IDs. |
| (none) |
| Scrolls the active tab smoothly. |
|
| (none) | Navigates the active tab to a new URL. |
| (none) |
| Extracts outer HTML of a specific element or the entire page. |
š” Example Prompts to Ask Your AI
Once configured, simply instruct your AI in natural language:
šø "Take a screenshot of my current tab and inspect the layout design."
š "Look at the charts on my active tab and summarize the visual data."
š "Look at my active Chrome tab and summarize what you see."
āļø "Fill in the login form with test credentials and submit."
š "Extract all product titles and links visible on the page."
š "Scroll down 800 pixels and extract the pricing table."
š¤ Multi-Agent Handoff Protocol (The "Triple-Anchor" Architecture)
Whether you are switching between specialized models (e.g. Cursor for rapid coding, Claude for system architecture), managing token budgets, or navigating provider rate limits, this repository includes the Triple-Anchor Architecture for seamless context transfer with zero loss of progress or hallucination.
āāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāā
ā Anchor 1: AGENTS.md (Universal Agent Guardrail) ā
ā -> Tells any newly opened agent: "Read PLAN.md and Git first" ā
āāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāā¬āāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāā
ā
āāāāāāāāāāāāāāāāāāāāāāāāā“āāāāāāāāāāāāāāāāāāāāāāāā
ā¼ ā¼
āāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāā āāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāā
ā Anchor 2: PLAN.md ā ā Anchor 3: Git Status & Diff ā
ā (The Intent & Architecture) ā ā (The Ground Truth of Code) ā
ā ⢠Completed tasks ā ā ⢠Exact lines of code changed ā
ā ⢠Next pending tasks ā ā ⢠Clean syntax state ā
ā ⢠"Decision Log" & Gotchas ā ā ⢠Zero hallucinated changes ā
āāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāā āāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāHow to Use in Any Project
Copy AGENTS.md and PLAN.md into the root of any codebase:
Anchor 1 (
AGENTS.md): Automatically read by modern AI tools (Cursor, Antigravity, Claude, ChatGPT CLI). It instructs the model to inspectPLAN.mdandgit statusbefore modifying any files.Anchor 2 (
PLAN.md): Contains the live checklist and Decision Log. The Decision Log is critical: it prevents the incoming agent from accidentally reverting intentional architecture choices (such as why a heartbeat ping was added).Anchor 3 (Git Milestones): When your current agent approaches its token limit, tell it:
"Commit your progress to Git and update PLAN.md."
Instant Continuation: Open your next AI agent and simply prompt:
"Continue." The new model reads
AGENTS.mdāPLAN.mdāgit status, and picks up immediately with zero lost context.
š§ Troubleshooting
"Chrome extension is disconnected": Make sure Google Chrome is open, the extension is loaded, and you have an active standard website open (not a restricted internal page like
chrome://extensionsorabout:blank).Port Customization: By default, the bridge uses WebSocket port
8765(avoiding conflict with standard HTTP 8080 proxies). To use a different port, set theWS_PORTenvironment variable before launching (e.g.set WS_PORT=8090/export WS_PORT=8090) and update the port inextension/background.js.
ā ļø Legal Disclaimer & Responsible Use
IMPORTANT NOTICE: This software is provided for personal workflow automation, educational research, and authorized testing purposes only.
Compliance with Terms of Service: Users are solely responsible for ensuring that their automated actions, scraping requests, and web interactions comply with all applicable local, national, and international laws, as well as the Terms of Service, Acceptable Use Policies, and
robots.txtguidelines of any websites visited.No Unauthorized Access: This software must not be used to bypass authentication barriers, paywalls, CAPTCHAs, rate limits, or security controls, nor to access or extract proprietary, copyrighted, or sensitive personal data without explicit permission.
Limitation of Liability: The author(s) and contributor(s) of this project assume no liability or responsibility for any misuse, website bans, legal disputes, data loss, damages, or consequences resulting from the installation or execution of this software. By using this project, you agree to assume all associated risks and responsibilities.
š License
Released under the MIT License. Free for open source and commercial use.
Available Tools
1 toolautomate_websiteB
Executes custom browser actions on the currently active webpage via the local Chrome extension bridge.
| Name | Required | Description | Default |
|---|---|---|---|
| action | Yes | The targeted script action. Supported actions: 'take_screenshot' (capture visual PNG screenshot of the tab for multimodal vision), 'scrape_data' (get text content), 'click_button' (click element by selector), 'fill_form' (type into input/textarea), 'get_elements' (query elements), 'scroll' (scroll down/up), 'navigate' (open URL), 'get_html' (extract HTML). | |
| payload | No | Key-value parameters needed for the action (e.g. { selector: 'button.submit' }, { selector: 'input[name=q]', value: 'search text' }, { direction: 'down', amount: 500 }, { url: 'https://example.com' }) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It only discloses the transport (local Chrome extension bridge), omitting that several actions are mutating (click_button, fill_form, navigate), whether they require user-visible permission, what happens on failure, or what the result looks like. For a tool that can drive page state, that is a significant omission.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler. It conveys verb, resource, and mechanism without redundancy, which is exactly the right size for this tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The parameters are fully documented via schema, but with no annotations and no output schema the description should still flag operational requirements (extension installed, tab focused) and the mutating nature of several actions. It leaves an agent to infer those, so it is minimally adequate rather than complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the action enum-like list plus payload examples are fully documented in the schema, so the baseline of 3 applies. The description adds no syntax, format, or default information beyond what the schema already conveys.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: executes browser actions on the currently active webpage. It also names the execution channel (local Chrome extension bridge), which tells an agent what this tool is. It does not enumerate the action types, but the schema covers that, so the gap is minor and there are no siblings to differentiate from.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use or when-not-to-use guidance, no alternatives, and no stated prerequisite such as an active tab or an installed extension. The phrase 'currently active webpage' gestures at context but does not say when an agent should reach for this tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
1 tool update
v1.0.0- First observed
automate_website
TDQS
Scored across 1 tool
Only one tool exists, so there is no possibility of selecting the wrong tool. Its purpose is clearly stated as executing browser actions via a Chrome extension bridge.
The single tool follows a clear verb_noun snake_case pattern (automate_website). There are no mixed conventions or naming conflicts to penalize.
A single tool is borderline thin for a browser automation server. It could be intentional as a generic executor, but the set lacks distinct operations that agents would typically expect.
The surface provides one generic action executor with no explicit operations, parameters, or lifecycle coverage. Agents lack discoverable building blocks for common browser automation tasks, creating likely gaps.
Maintenance
Related MCP Connectors
Live browser debugging for AI assistants ā DOM, console, network via MCP.
AI-powered browser automation ā navigate, click, fill forms, and extract data from any website.
Hosted browser for AI agents: screenshots, post-JS DOM, console, WCAG. No install, no API key.
- openhelmOAuthai.openhelm
Autonomous cloud agent tasks: real browser + your tools, structured evidence-backed results.
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceAn MCP server and Chrome extension that allows users to select browser DOM elements via a keyboard shortcut to provide detailed technical context to AI coding tools. It captures HTML attributes, CSS styles, and React component metadata, enabling agents to analyze and modify web elements directly.598MIT
- AlicenseAqualityCmaintenanceMCP server + Chrome extension that gives AI coding agents control of your real browser with existing sessions, logins, and cookies. Works with Cursor, Claude, Windsurf.1896 npm53MIT
- AlicenseNot gradedqualityDmaintenanceAn extension-based MCP server that enables AI assistants to control your browser, leveraging existing sessions and login states for automation and content analysis. It provides over 20 tools for semantic tab search, interactive element manipulation, and network monitoring directly within your daily Chrome environment.MIT
- AlicenseNot gradedqualityCmaintenanceThe Zero-Setup Local Browser MCP. Enables AI agents to control web browsers via CDP with zero vision tokens and high-speed DOM mapping.17 npmMIT