Skip to main content
Glama
fangsylar-pixel

browser-search-mcp

Browser Search MCP

Real browser search for AI agents. Browser Search MCP lets Claude Desktop, Cursor, Codex, Ollama, and other MCP clients search the web through Chrome/Edge, read result pages, expand vague user questions, and report whether the results actually match the user's intent.

Search like a real browser. Read like a research agent.

基于真实浏览器的 MCP 搜索引擎服务器 - 让任何支持 MCP 的大模型都能搜索网页内容。 Browser Search MCP - Web search via real browser for any LLM.

GitHub Repo stars License Python PyPI

Built on the same CDP extension bridge architecture as browser-takeover-bridge.

Launch Week

If you are discovering this project from a post, start here:

  • Install in 30 seconds: pip install browser-search-mcp

  • Best default tool: web_research

  • Preview query expansion: web_search_plan

  • Use cases: local agents, private research assistants, authenticated browser search, citation-ready page reading

  • Launch assets: Launch Week Plan and Ready-to-post copy

Why agents use it

Need

Browser Search MCP

No search API key

Uses your browser by default

Real web pages

Reads and cleans top result pages

Vague user questions

Expands natural-language intent into search queries

Quality control

Marks results as strict or partial and reports missing anchors

Private/local workflows

Runs as an MCP server on your machine

Logged-in browser context

Can integrate with browser-takeover bridge

Example: plan before searching

{
  "query": "home projector vs TV which is better",
  "intent": {
    "topic": ["projector", "TV"],
    "task": ["compare", "buying guide", "recommendation"]
  },
  "candidate_queries": [
    "home projector vs TV which is better",
    "projector TV comparison pros cons buying guide",
    "projector TV comparison recommendation"
  ]
}

Example: research-ready result

{
  "ok": true,
  "query": "2026 exam English study plan",
  "quality": "strict",
  "diagnostics": {
    "missing_anchor_groups": []
  },
  "results": [
    {
      "title": "Example result",
      "url": "https://example.com",
      "page": {
        "ok": true,
        "content": "Cleaned page text for the agent..."
      }
    }
  ]
}

Related MCP server: Google Search Tool

Why?

Local LLMs (Ollama, etc.) cant search the web. HTTP-based search tools get blocked by anti-bot measures. This project uses a real browser to search - no API keys, no blocking, no fake results.

Quick Start

pip install browser-search-mcp

# Start the MCP server
browser-search-mcp

# Optional CLI helpers
browser-search-mcp --help
browser-search-mcp status
browser-search-mcp doctor
browser-search-mcp http 9090

Then configure in any MCP client:

{
  "mcpServers": {
    "browser-search": {
      "command": "browser-search-mcp"
    }
  }
}

Bridge Integration (browser-takeover extension)

When installed alongside the browser-takeover-bridge extension, browser-search-mcp automatically detects the extension and routes searches through it instead of launching a headless CDP browser.

Why use the bridge?

  • Authenticated sessions - search while logged into services (e.g., intranet, social media, internal tools)

  • Faster startup - no need to launch a new browser; reuses the extension's existing connection

  • Lower resource usage - share one browser session instead of spawning a separate headless instance

How it works:

LLM/Agent -> MCP Client -> browser-search-mcp -> bridge (extension) -> user's browser -> search engine

The bridge check runs automatically at startup. If the extension is not detected, the server falls back to the standard CDP path (launching its own browser). No configuration needed.

Detection: Run web_search_status to see if the bridge is active: json { "bridge": { "available": true, "search_available": true } }

Quick Demo

Search the web in seconds from any MCP-compatible LLM:

pip install browser-search-mcp
browser-search-mcp

Live search result (Bing, ~5s):

[
  {
    "title": "What is the Model Context Protocol (MCP)?",
    "url": "https://modelcontextprotocol.io/",
    "snippet": "MCP is an open standard for connecting AI applications to external systems."
  },
  {
    "title": "MCP Server Guide",
    "url": "https://example.com/mcp-guide",
    "snippet": "Complete guide to setting up MCP servers for web search."
  },
  {
    "title": "Browser Search MCP",
    "url": "https://github.com/fangsylar-pixel/browser-search",
    "snippet": "Open source MCP server using real browser for web search."
  }
]

No API keys required. No blocking. Just a real browser doing real searches.

Features

Feature

Status

Description

Google, Bing, Baidu, DuckDuckGo

Yes

DOM + JS extraction

Persistent browser session

Yes

Reuses CDP connection

Result caching

Yes

LRU with configurable TTL

Config file

Yes

JSON + env vars

Auto-reconnect

Yes

Transparent reconnection

browser-takeover bridge

Yes

Auto-detected extension bridge

CAPTCHA detection

Yes

Auto fallback on CAPTCHA

Engine fallback

Yes

Automatic on failure/CAPTCHA

Deep mode

Yes

Auto-extracts top 2 result content

Pagination

Yes

Multi-page search support

Time filters

Yes

hour/day/week/month/year

Intent planning

Yes

Platform/topic/task extraction plus query expansion

Engine health check

Yes

Tracks per-engine availability

Cross-engine dedup

Yes

Deduplicate multi-engine results

API providers (Tavily/Brave)

Yes

Faster, API-key based

HTTP API

Yes

FastAPI + OpenAI compatible

Codex plugin

Yes

Auto-install as Codex plugin

Fallback parsers

Yes

Text-based when JS fails

Retry on failure

Yes

Exponential backoff

MCP Tools

Tool

Description

web_search_plan

Analyze intent and candidate queries without launching a browser

web_search

Search a single engine, returns JSON results

web_search_multi

Search multiple engines simultaneously

web_search_read_page

Read structured page content for a URL

web_research

Search and read top result pages for citation-ready research

web_search_status

Check browser, bridge, and cache status

web_search_discover_browsers

Find CDP-enabled browsers

web_search_plan is useful before an expensive search: it returns parsed intent, required coverage anchors, and expanded candidate queries for broad customer requests such as creator topics, product comparisons, tutorials, trends, and recommendations.

web_research is the best default for agents that need grounded answers. It rewrites natural-language questions into search queries, retries across engines, filters off-topic results, and labels result quality as strict or partial based on whether the platform/topic/task anchors were covered. It also returns diagnostics, cleaned page content for the top results, title, final URL, description, detected publish time, and truncation metadata.

Configuration

Config file: ~/.browser-search-mcp/config.json

Browser Mode (default)

{
  "browser": {
    "name": "edge",
    "headless": false,
    "port": 9222
  },
  "cache": {
    "enabled": true,
    "ttl": 300
  }
}

API Mode (faster, needs API key)

{
  "provider": {
    "name": "tavily",
    "tavily_api_key": "tvly-your-key-here"
  }
}

Environment Variables

Variable

Example

Description

BROWSER_SEARCH_HEADLESS

true

Run browser headless

BROWSER_SEARCH_PROVIDER

tavily

Choose provider: browser/tavily/brave

BROWSER_SEARCH_TAVILY_KEY

tvly-xxx

Tavily API key

BROWSER_SEARCH_BRAVE_KEY

...

Brave Search API key

BROWSER_SEARCH_CACHE_TTL

300

Cache TTL in seconds

BROWSER_SEARCH_DEFAULT_ENGINE

bing

Default search engine

BROWSER_SEARCH_BROWSER

edge

Browser executable name

BROWSER_SEARCH_PORT

9333

CDP remote debugging port

BROWSER_SEARCH_USER_DATA_DIR

/tmp/browser-search-profile

Browser profile directory

BROWSER_SEARCH_BROWSER_PATH

/path/to/browser

Explicit browser executable path

How It Works

LLM/Agent -> MCP Client -> browser-search-mcp -> Browser (CDP) -> Search Engine
                                                    | (optional)
                                          browser-takeover extension
  1. MCP server finds or launches a Chrome/Edge browser with remote debugging

  2. Navigates to the search engine

  3. Extracts structured results via JavaScript DOM parsing

  4. Returns title, url, snippet as JSON

  5. Results cached for 5 minutes by default

Project Structure

browser-search-mcp/
  browser_search_mcp/
    config.py    Configuration via JSON file + env vars
    cdp.py       CDP browser control with persistent sessions
    bridge.py    Browser-takeover extension bridge client
    search.py    Search orchestration with caching and retry
    parsers.py   Text-based search result parsers (fallback)
    bridge_provider.py  Bridge-based search provider (optional, auto-detected)
    providers.py API search providers (Tavily, Brave)
    server.py    FastMCP server with 5 search tools
    http_api.py  HTTP API server (FastAPI + OpenAI-compatible endpoint)
    setup_assistant.py  Prerequisites check
  .codex-plugin/ Codex plugin packaging
  .github/      CI and issue templates
  website/      Promotional website (GitHub Pages)
  README.md, CONTRIBUTING.md, LICENSE, SUPPORT.md

Requirements

  • Python 3.11+

  • Chrome or Edge installed

  • Optional: browser-takeover-bridge extension (for authenticated sessions)

License

MIT

Architecture

flowchart TB
    subgraph User[User Environment]
        LLM[LLM / Agent]
        MCP[MCP Client]
    end
    subgraph Browser[Browser Layer]
        BRIDGE[Bridge Extension]
        CDP[Chrome/Edge CDP]
    end
    subgraph Search[Search Layer]
        BSM[browser-search-mcp]
        API_PROV[API Providers]
    end
    subgraph Engines[Search Engines]
        G[Google]
        B[Bing]
        BA[Baidu]
        D[DuckDuckGo]
    end

    LLM --> MCP --> BSM
    BSM -->|Priority 1| BRIDGE --> CDP --> Engines
    BSM -->|Priority 2| API_PROV --> Engines
    BSM -->|Priority 3| CDP --> Engines

    style BSM fill:#10b981,color:#fff
    style BRIDGE fill:#6366f1,color:#fff
    style API_PROV fill:#f59e0b,color:#fff

Support

If this project helps you, optional support is welcome:

Support on Afdian

Bug reports and contributions are welcome. See CONTRIBUTING.md and SUPPORT.md.

Available Tools

5 tools
web_search_discover_browsersA

Scan for browsers with CDP enabled on common ports.

Checks ports 9222, 9223, 9333 for Chrome/Edge instances with remote debugging enabled.

Returns: JSON list of discovered browser instances

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description fully bears the burden of behavioral disclosure. It details the ports checked (9222, 9223, 9333) and targeted browsers (Chrome/Edge with remote debugging), and states the return type (JSON list). This provides sufficient transparency about its read-only scanning behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise (three sentences) and front-loaded with the core purpose. Each sentence adds value: purpose, port details, and return format. No superfluous content.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given zero parameters and an existing output schema, the description covers all necessary aspects: what the tool does, how it does it (ports, browsers), and what it returns. It is complete without gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has zero parameters, and schema description coverage is 100%. The description does not need to add parameter details, and the baseline score of 4 applies as there is no missing information.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool scans for browsers with CDP enabled on specific ports. It uses a specific verb ('Scan') and resource ('browsers with CDP enabled'), and the purpose is distinct from sibling tools like web_search or web_search_read_page, which focus on web content retrieval.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage when needing to discover Chrome/Edge instances for remote debugging. It does not explicitly state when not to use or provide alternatives, but the context (sibling tools) and specificity make the intended use clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

web_search_multiA

Search multiple search engines and combine results. Searches multiple engines in sequence and returns combined results grouped by engine.

ParametersJSON Schema
NameRequiredDescriptionDefault
queryYesThe search query string
enginesNoComma-separated list of engines (google,bing,baidu,duckduckgo)google,bing,duckduckgo
max_results_per_engineNoResults per engine (1-10)
headlessNoRun browser in headless mode
pageNo
time_rangeNo
deep_modeNo
deduplicateNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It discloses that searches are performed 'in sequence' and results are 'grouped by engine,' which is useful. However, it does not mention failure handling, rate limits, or other behavioral traits that could affect agent decisions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely concise: two sentences (21 words) that front-load the core purpose in the first sentence and add detail in the second. Every word earns its place, with no redundancy or filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists but is not described. The description covers the basic workflow of multi-engine search and grouping but omits details like error behavior, pagination across engines, or impact of parameters like 'deep_mode'. For a tool with 8 parameters, it is minimally complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 50%, but the description adds no additional parameter meaning beyond what the schema provides. Parameters like 'page', 'time_range', 'deep_mode', and 'deduplicate' remain unexplained in context of multi-engine search, leaving gaps for the agent.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool searches multiple search engines and combines results. It uses the verb 'search' and resource 'multiple search engines,' and distinguishes from sibling 'web_search' (likely single engine) by specifying multi-engine and grouping.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies when to use this tool (when you need results from multiple engines) but does not explicitly state when not to use it or mention alternatives. It provides enough context for an agent to differentiate it from siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

web_search_read_pageA

Navigate to a URL and return the visible page text. Useful for reading the full content of a search result.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYesThe URL to navigate to
max_lengthNoMaximum characters to return

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description bears the full burden. It only says 'return the visible page text' without disclosing behaviors like JavaScript execution, error handling, or limitations (e.g., paywalls, dynamic content). This is a significant gap for a navigation tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences, front-loaded with the action, and contains no redundant or extraneous information. Every sentence adds value.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has an output schema (not shown but indicated), which reduces the need to describe return values. However, the description lacks details on error scenarios, permission requirements, or behavior with invalid URLs. It is minimally adequate for a relatively simple tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 100% description coverage for both parameters (url, max_length). The description adds no extra semantic value beyond 'Navigate to a URL,' which is already implied by the tool name and schema. Baseline score of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states 'Navigate to a URL and return the visible page text,' specifying the action (navigate, return) and resource (URL, page text). It distinguishes from siblings like 'web_search' which likely executes searches, and other tools with different scopes.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description includes 'Useful for reading the full content of a search result,' giving a clear usage context. However, it does not explicitly state when not to use it or mention alternatives, so it scores slightly below the top tier.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

web_search_statusA

Check the current browser, bridge, and provider status.

Detects available CDP browser instances, checks if the browser-takeover extension bridge is running, and reports the active provider type.

Returns: JSON status information

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description fully shoulders behavioral disclosure. It describes the detection actions and return type accurately. It could be improved by explicitly stating it is non-destructive and read-only, but the current description is reasonably transparent for a status check.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise with three sentences, no fluff, and front-loads the main purpose. Every sentence provides necessary information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has no parameters, an output schema exists, and the description covers the essential status-checking functionality, it is fully complete for its purpose.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has zero parameters, so the description does not need to add parameter details. Baseline 4 is appropriate as it adds no superfluous information.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states it checks browser, bridge, and provider status, listing specific detection targets (CDP instances, bridge extension, provider type). It distinguishes well from sibling tools like web_search which perform searches.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage for checking status before other operations but provides no explicit guidance on when to use versus alternatives. No exclusions or conditions are mentioned.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 2 tool updatesv0.2.0
    • Changedweb_search3 fields changed
      • addedInput schema / properties / deep_mode
        Added value: +{
        +  "default": false,
        +  "type": "boolean"
        +}
      • addedInput schema / properties / page
        Added value: +{
        +  "default": 1,
        +  "type": "integer"
        +}
      • addedInput schema / properties / time_range
        Added value: +{
        +  "default": "",
        +  "type": "string"
        +}
    • Changedweb_search_multi4 fields changed
      • addedInput schema / properties / deduplicate
        Added value: +{
        +  "default": false,
        +  "type": "boolean"
        +}
      • addedInput schema / properties / deep_mode
        Added value: +{
        +  "default": false,
        +  "type": "boolean"
        +}
      • addedInput schema / properties / page
        Added value: +{
        +  "default": 1,
        +  "type": "integer"
        +}
      • addedInput schema / properties / time_range
        Added value: +{
        +  "default": "",
        +  "type": "string"
        +}
  2. 5 tool updatesv0.1.0
    • First observedweb_search
    • First observedweb_search_discover_browsers
    • First observedweb_search_multi
    • First observedweb_search_read_page
    • First observedweb_search_status

TDQS

A3.8/5.0

Scored across 5 tools

Disambiguation4/5

web_search and web_search_multi are related but clearly separated by the multi-engine behavior. web_search_status and web_search_discover_browsers overlap somewhat in browser inspection, but their purposes remain distinct. web_search_read_page is unambiguous.

Naming Consistency4/5

Most tools use the web_search_ prefix with descriptive suffixes like multi, read_page, status, and discover_browsers. The base web_search tool lacks a suffix, but the pattern remains predictable and readable overall.

Tool Count5/5

Five tools is well-scoped for a browser-based search server. Search, multi-engine search, page reading, status, and browser discovery each earn their place without unnecessary overlap or bloat.

Completeness4/5

The core workflow of searching, combining results, and reading pages is well covered. Browser status and discovery support the environment side. Minor gaps like explicit browser launch/close management exist, but they are outside the apparent search-focused purpose.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    D
    maintenance
    Enables AI assistants to perform real-time Google searches with anti-bot protection. Bypasses search engine restrictions using advanced browser automation to extract search results locally without requiring paid API services.
    7
    MIT
  • A
    license
    B
    quality
    D
    maintenance
    Enables LLMs and AI agents to access real-time web data, search websites, and navigate the web without getting blocked. Includes 5,000 free monthly requests and supports web scraping, browser automation, and bypassing geo-restrictions.
    60
    7,869
    1
    MIT