Agentic Browser
Offers specialized tools for interacting with GitHub, such as navigating repository tabs and verifying action outcomes on GitHub pages.
Provides tools for LangGraph agents to control a Chromium browser for web automation tasks.
Generates tool schemas compatible with OpenAI's Chat Completions API, enabling agents to control a browser using OpenAI models.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Agentic Browsersearch for 'climate change' on Wikipedia"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Agentic Browser
The web, as an LLM wants to see it
Stable refs · compact observations · outcome-verified actions · MCP for any agent host
Not another HTML scraper. A real Chromium browser controlled with semantic tools — so Claude, Cursor, LangGraph, or your own agent can see, click, type, and verify without drowning in 100k+ tokens of markup.
Repo: github.com/applejuice093/Agentic-browser · Branch: dev
Why this exists
Traditional stacks hand the model raw HTML or brittle CSS. Agents need:
Small, structured observations (roles, labels, refs)
Actions that mean success (URL/DOM outcomes, not “no exception”)
A way to plug into any host (MCP + OpenAI/Anthropic tool schemas)
Agentic Browser is built for that loop.
navigate → observe (~1–2k tokens) → click_text / type → outcome_verified?
↘ if page_gate = challenge → stop (don’t hallucinate)Related MCP server: cowork-qa-mcp
Highlights — measured positives
Numbers from live benchmarks on public sites (approx. tokens ≈ chars/4). Reproduce with scripts under examples/.
Token efficiency (LLM feed)
Scenario | Raw HTML → model | Our compact observation / structured feed | Reduction |
Quotes scrape (structured answers) | ~2.8k–6.2k | ~0.45k–1.3k | ~78–84% fewer tokens |
Rockstar GTA VI landing | ~225,000 | ~1,300 | ~99.4% fewer tokens |
GitHub | ~110,000 | ~1,900 | ~98.3% fewer tokens |
Takeaway: Dumping HTML into an LLM is the wrong default. Observation mode is ~98–99% smaller on heavy modern pages.
Task quality
Metric | Traditional HTTP + BS4 | Agentic Browser | Notes |
Quote scrape field completeness | 100% (with site CSS) | 100% | Parity on friendly sites |
Plain-text “LLM scrape” completeness | 0% tags lost | 100% via tools | Text dumps drop structure |
Content signals on GitHub repo | ~70% | ~80% | Slight edge after JS + observe |
Actionable element refs | 0 | 20–90+ per page | Only we can click by ref |
GitHub “open Issues” (after v0.3.2) | N/A | Verified → URL | Outcome-checked, not false-ok |
Network / XHR visibility | 0% | Full request log + GraphQL flags | Agent can wait on APIs |
Product completeness
Area | Coverage |
Milestone roadmap M1–M10 | 100% of planned milestones delivered (plus agent-native v0.3–0.4) |
Automated tests on | 118 passed (full suite) |
MCP tools for hosts | 10 tools (navigate, observe, click, type, wait, find, network, …) |
Observation token budget (default) | ~2,000 (configurable) |
Positives at a glance
~98–99% token cut vs raw HTML on complex landings (Rockstar, GitHub)
~78–84% token cut vs HTML for structured extract tasks
100% structured field parity with best-case CSS scrapers on demo sites
Outcome verification —
okonly when post-conditions hold (e.g./issuesin URL)Scoped grounding — nav-first find so PR/commit text doesn’t steal “Issues”
Page gates —
js_challenge/ captcha / login wall detected and reportedMCP-ready — Claude Desktop, Cursor, any MCP host
OpenAI + Anthropic tool schemas —
tools_as_openai()/tools_as_anthropic()Cookie/CMP dismiss, SPA settle budgets, stale-ref recovery, network intelligence
Privacy defaults: mask
Authorization/ cookies in network logs
Bottlenecks & negatives (no sugarcoating)
We measure these so you don’t ship blind.
Issue | Reality | Impact |
Latency | Real browser is ~6–16× slower than | Bad for high-QPS crawl; fine for agent steps |
Bot walls | Reddit-class JS challenges: both HTTP and browser get ~17% content signals | Not a bypass tool — we surface |
SSR-only read | On some marketing pages, plain HTTP already has the copy | Agent wins on actions, not always on pure read speed |
Full semantic snapshot | Can be larger than HTML if you dump everything | Always use |
Speed vs design goal | Sub-500 ms action latency is not met end-to-end on heavy SPAs | Dominated by page load + settle, not Python |
False success (historical) | Pre-0.3.2 GitHub “Issues” click could report ok without navigating | Fixed with outcome verification + GitHub skill |
Skills coverage | GitHub tab skill is first-class; other domains need packs | Extensible under |
Vision | OCR optional; not in the default observe loop | Canvas-heavy UIs still weaker |
Honest product line:
Best as an LLM action + compact perception layer.
Not a replacement for bulk HTTP scraping, and not a captcha solver.
Architecture
┌──────────────────────────────────────────────────────┐
│ Claude · Cursor · LangGraph · OpenAI tool loop │
└──────────────────────────┬───────────────────────────┘
│ MCP stdio or function tools
▼
┌──────────────────────────────────────────────────────┐
│ agent-browser MCP · tools_as_openai() │
│ AgentSession: observe · click_text · wait · network │
└──────────────────────────┬───────────────────────────┘
▼
┌──────────────────────────────────────────────────────┐
│ Semantic DOM · scoped grounding · outcomes · gates │
│ Network monitor · memory · settle · overlay dismiss │
└──────────────────────────┬───────────────────────────┘
▼
Playwright · ChromiumInstallation
Requirements
Python 3.11+
Git
~Windows / macOS / Linux
From PyPI (recommended)
PyPI package name is agentic-browser (import and CLI stay agent_browser / agent-browser).
pip install agentic-browser
playwright install chromium
# with MCP extras
pip install "agentic-browser[mcp]"
# with vision / OCR extras
pip install "agentic-browser[vision]"From source (library + dev)
git clone https://github.com/applejuice093/Agentic-browser.git
cd Agentic-browser
git checkout main
python -m venv .venv
# Windows
.venv\Scripts\activate
# macOS / Linux
# source .venv/bin/activate
pip install -e ".[dev]"
playwright install chromium
pytest -qWith MCP from source (Claude / Cursor)
pip install -e ".[mcp,dev]"
playwright install chromiumOptional vision from source (local OCR)
pip install -e ".[vision]"
# Install system Tesseract; set TESSERACT_CMD on WindowsExtras summary
Extra | Install | Provides |
default |
| Core browser + agent API |
|
| MCP server for hosts |
|
| pytest, ruff, mypy |
|
| Pillow + pytesseract |
|
| Everything |
Add to your agent (3 ways)
A) MCP — Claude Desktop / Cursor (recommended)
python -m agent_browser.mcp
# or
agent-browser-mcpCursor — .cursor/mcp.json (use your venv Python path):
{
"mcpServers": {
"agent-browser": {
"command": "C:\\A\\PROJECT\\Agentic Browser\\.venv\\Scripts\\python.exe",
"args": ["-m", "agent_browser.mcp"],
"env": {
"AGENT_BROWSER_HEADLESS": "true",
"AGENT_BROWSER_MAX_TOKENS": "2000",
"AGENT_BROWSER_ALLOWED_HOSTS": "github.com,example.com,quotes.toscrape.com"
}
}
}
}Claude Desktop — claude_desktop_config.json:
{
"mcpServers": {
"agent-browser": {
"command": "C:\\A\\PROJECT\\Agentic Browser\\.venv\\Scripts\\python.exe",
"args": ["-m", "agent_browser.mcp"],
"env": { "AGENT_BROWSER_HEADLESS": "true" }
}
}
}Full guide (env vars, system prompt, troubleshooting): docs/mcp.md
B) Python library
import asyncio
from agent_browser import Browser, tools_as_openai
async def main():
async with Browser(headless=True) as browser:
agent = await browser.open_agent("https://github.com/vercel/next.js")
obs = await agent.observe(max_tokens=2000)
if obs.page_gate not in (None, "open", "cookie_wall", "unknown"):
print("Blocked:", obs.page_gate, obs.page_gate_hint)
return
result = await agent.click_text("Issues", scope="nav", intent="issues")
print(result.ok, result.url_after, result.outcome_verified)
asyncio.run(main())OpenAI / Anthropic tool schemas:
from agent_browser import tools_as_openai, tools_as_anthropic
tools = tools_as_openai() # Chat Completions
# tools = tools_as_anthropic() # Messages APIC) CLI
agent-browser version
agent-browser open https://example.com
agent-browser scrape https://example.com -o data/out.jsonCore concepts
Concept | Meaning |
ref | Stable integer id for click/type (from |
Observation | Compact LLM payload: interactive refs, headings, summary, gate, network hints |
ActionResult |
|
page_gate |
|
scope |
|
MCP / tool surface
browser_navigate · browser_observe · browser_click · browser_click_text · browser_type · browser_wait · browser_find · browser_network · browser_resync · browser_prepare
Environment variables
Variable | Default | Purpose |
|
| Headless Chromium |
|
| Observation budget |
|
|
|
|
| SPA settle budget |
| (all) | Comma-separated allowlist |
|
| Action timeout |
| — | OCR binary path (vision extra) |
Project layout
src/agent_browser/
browser.py / page.py / cli.py
agent/ # AgentSession, MCP tools bridge, settle, overlays, grounding, outcomes, skills
mcp/ # FastMCP stdio server
observation/ # Compact observation builder
semantic/ accessibility/ events/ network/ vision/ memory/ planning/
docs/
mcp.md # Add to Claude / Cursor / custom agents
USER_GUIDE.md
agent-native-loop.md
security.md
examples/
agent_loop_demo.py
openai_tool_agent.py
mcp_config_cursor.json
benchmark_hard_site.py
compare_rockstar_vi.py
tests/ # 118+ automated testsDevelopment
git checkout dev
pip install -e ".[dev,mcp]"
playwright install chromium
pytest -q
ruff check src testsDoc | Topic |
Install into your agent (MCP + tools) | |
Full product guide | |
Observe / act / gates / skills | |
Security notes | |
Version history | |
Original design research |
Roadmap status
Track | Status |
M1–M10 foundation | Done |
Agent-native loop (v0.3) | Done |
Overlays / settle / recovery (v0.3.1) | Done |
Grounding + outcomes + GitHub skill (v0.3.2) | Done |
MCP host integration (v0.4.0) | Done |
More domain skills, vision-in-loop, pool/scale | Next |
License
MIT — see LICENSE.
This server cannot be deployed
Maintenance
Related MCP Connectors
Web search, browser automation, scraping, crawling and CAPTCHA solving for AI agents.
Real Chrome for agents: start a browser, read pages as numbered markdown, click, type, hand off.
- openhelmOAuthai.openhelm
Autonomous cloud agent tasks: real browser + your tools, structured evidence-backed results.
Stealth web automation for AI agents. Login, signup, navigate, screenshot.
Related MCP Servers
- AlicenseAqualityDmaintenanceEnables AI agents to control a browser with 30 tools for navigation, interaction, extraction, and tab management, supporting human-like browser automation.362 npmMIT
- AlicenseAqualityDmaintenanceProvides an LLM with a real Chromium browser to perform web tasks, recording every action into a structured trace for later verification of goal completion.522 npmMIT
- AlicenseAqualityCmaintenanceEnables AI agents to control a persistent Chromium browser or attach to an existing Chrome with sandboxed JavaScript and structured tools for web interaction and automation.16MIT
- AlicenseNot gradedqualityCmaintenanceEnables language models to control a real visible browser with a full set of 24 tools for navigation, clicking, form filling, screenshots, and reading page structure, while preserving login states and supporting configurable browser policies.171 npmMIT