blowsh-mcp
blowsh-mcp
Browshを使ったJS対応ターミナルブラウジングのためのModel Context Protocolサーバー
blowsh-mcpとは?
blowsh-mcpは、完全にJavaScriptに対応したターミナルブラウザであるBrowshのパワーを、あらゆるAIエージェント、IDEエージェント、MCPクライアントに公開するModel Context Protocol (MCP) サーバーです。このプロジェクトにより、AIはJavaScriptを必要とするものを含む最新のWebページを取得・レンダリングし、解析しやすいプレーンテキスト、HTML、Markdownとして結果を受け取ることができます。
記憶のための語呂合わせ: “blowsh” = Browshを搭載したMCPサーバー。
Related MCP server: openmcp
主な機能
fetch_web ツール: 完全なJSレンダリング後に、読みやすいプレーンテキスト、HTML、またはMarkdown抽出のための統合ツール。CSS
selector抽出、max_chars出力上限、wait_msJS安定待機ポーリングに対応。search_web ツール: レンダリングされた検索エンジン(Bingフォールバック付きDuckDuckGo HTML)でページを発見 — URLとスニペット付きのランキング結果。
extract_links ツール: ナビゲーション追従用に、任意のJSレンダリングページからハイパーリンク(テキスト + 絶対URL)を一覧表示。
fetch_web_batch ツール: 1回の呼び出しで最大10URLを取得、URLごとにエラーを分離。
SSRFガード: ループバック、プライベート、リンクローカル、予約済みアドレスへのリクエスト(DNS解決後)を拒否し、サーバー側ブラウザを保護。
AI最適化ツールドキュメント: シームレスなエージェント自動化のために設計された入力、出力、図解付きユースケース。ツールはHTTPステータスコード付きの構造化エラーをスローします(MCP応答の
isError)。堅牢なBrowsh管理: Browshを一度起動して稼働を維持し、RAM/CPU負荷の低いシングルトンを再利用、終了時にグレースフルシャットダウン。
TTL付きインメモリレンダリングキャッシュ: 繰り返しのフェッチは再レンダリングなしですぐに提供。
PaaS、クラウド、ローカルAIツール、IDEエージェント向けに設計。
リンク
Browsh CLI Browser — レンダリングエンジン。
Firefox — Browshのバックエンドとして必要。
Model Context Protocol (MCP) Specification — エージェント/サーバープロトコル。
動作の仕組み
AI/エージェントが
fetch_web(単一URL)、search_web(クエリ)、extract_links(URL)、またはfetch_web_batch(最大10URL)のMCPリクエストを行います。blowsh-mcpはBrowshをHTTPサーバーモードで起動し(最初の使用時)、以後すべての呼び出しで再利用します。
blowsh-mcpは
X-Browsh-Raw-Mode: PLAIN(テキスト用)、DOM(HTML用)を使ってBrowshから生の出力を要求するか、HTMLを取得してMarkdownに変換します。ページ(完全なJS実行後)は、ターミナル用プレーンテキスト、リッチなHTML DOM、またはクリーンなMarkdownとして返されます。AI/エージェントは後続処理に合わせて出力タイプを選択します。
結果はインメモリ(TTL)にキャッシュされるため、繰り返しのフェッチは即座に実行されます。すべてのリクエストはブラウザに到達する前にSSRFチェックを受けます。
クイックスタート(Docker — プリビルドイメージ)
イメージはGitHub Container Registryに公開され、GitHub Actionsによってmainへのプッシュのたびに自動再ビルドされます。ホスト側のFirefox/Browsh/html2markdownは不要です:
docker pull ghcr.io/mokhtarabadi/blowsh-mcp:latest
docker run --rm -i ghcr.io/mokhtarabadi/blowsh-mcp:latest
-iフラグは必須です。MCPサーバーはstdin/stdoutでJSON-RPCを通信します。インタラクティブモードを維持してリクエストをパイプするか、MCPクライアントからこのサーバーを指定してください(下記のAIクライアント設定を参照)。
使用例
Claude、Cursor、またはMCP対応エージェントから:
{
"tool": "search_web",
"params": { "query": "bitcoin price today", "max_results": 5 }
}
// → Ranked results with URLs + snippets → feed top URL to fetch_web
{
"tool": "fetch_web",
"params": { "url": "https://coindesk.com/price/bitcoin/", "type": "plain" }
}
// → Returns readable plain text (live price as text table, etc)
{
"tool": "fetch_web",
"params": { "url": "https://coindesk.com/price/bitcoin/", "type": "markdown", "selector": "main", "wait_ms": 3000 }
}
// → Markdown of <main> only, after JS settles ("# Bitcoin Price\n\n| Time | Price | ...")
{
"tool": "extract_links",
"params": { "url": "https://example.com", "limit": 20 }
}
// → [{"text": "Learn more", "url": "https://iana.org/domains/example"}, ...]
{
"tool": "fetch_web_batch",
"params": { "urls": ["https://a.com", "https://b.com"], "type": "markdown" }
}
// → Per-URL results; a failing page never fails the batchAIは次を受け取ります:
type: plainの場合: 純粋な読み取り可能テキスト(表、リスト、本文コンテンツ。NLP/要約やターミナルコンテキスト取り込みに最適)。type: htmlの場合: すべてのJavaScript実行後の完全なHTMLマークアップ。要素解析、リンクグラフ構築、複雑なスクレイピングなどに使用。type: markdownの場合: クリーンなMarkdownバージョン。LLMコンテキストチャンク、セマンティックパイプライン、AI向きの消費/ワークフローに最適。エラーは構造化されます: MCP応答は、利用可能な場合はHTTPステータスを含む
FetchErrorメッセージとともにisError: trueを設定します。
プロジェクト構成
src/server.ts— ツールを公開するMCPサーバー。src/browshManager.ts— Browshの起動、監視、シャットダウン。src/tools/fetchWeb.ts— fetchWebツールの実装(plain、html、markdown、selector/max_chars/wait_ms)。src/tools/searchWeb.ts— search_web(DuckDuckGo HTML + Bingフォールバックパーサー)。src/tools/extractLinks.ts— extract_links(レンダリングされたDOMからハイパーリンクを抽出)。src/tools/fetchWebBatch.ts— fetch_web_batch(複数URL、URLごとのエラー分離)。src/tools/html2markdownManager.ts— html2markdown CLIのラッパー。src/ssrf.ts— SSRFガード(プライベート/ループバック/予約済みターゲットをブロック)。src/cache.ts— インメモリTTLレンダリングキャッシュ。src/extract.ts— 本文抽出、セレクターヘルパー、切り詰め。src/errors.ts—FetchError+ メッセージフォーマット。README.md— このファイル。Dockerfile— マルチステージコンテナ(TSをビルドし、Firefox、Browsh、html2markdownをバンドル)。.github/workflows/docker-publish.yml— CI/CD:main/v*でghcr.ioにイメージをビルドして公開。.env— 設定の上書き。全オプションは.env.exampleを参照。
インストール
要件:
Node.js >= 20.18
Firefoxがインストールされ、PATHに含まれていること
Browsh CLI がインストールされ、PATHに含まれていること
html2markdown CLI がインストールされ、PATHに含まれていること
Debian/Ubuntuでは、次のコマンドでインストールします:
wget -O /tmp/html2markdown.deb "https://github.com/JohannesKaufmann/html-to-markdown/releases/download/v2.5.2/html2markdown_2.5.2_linux_amd64.deb" sudo apt-get install -y /tmp/html2markdown.deb rm /tmp/html2markdown.debまたは、お使いのOS向けのプリビルドバイナリをリリースページから使用します。
Dockerをお好みですか?ホスト側のインストールはすべて不要です。マルチステージイメージにはFirefox、Browsh、html2markdownがバンドルされています。最速の方法は公開イメージ(
ghcr.io/mokhtarabadi/blowsh-mcp:latest、クイックスタートを参照)を使うことです。自分でビルドするには:docker build -t blowsh-mcp:latest . docker run --rm -i blowsh-mcp:latest
git clone https://github.com/mokhtarabadi/blowsh-mcp.git
cd blowsh-mcp
npm install
npm run buildMCPサーバーの実行
ビルド後、次のコマンドでサーバーを起動します:
node dist/server.jsビルド出力が異なる場合は、dist/server.js を正しいパスに置き換えてください。
必要に応じて設定用の.envファイルを作成します。例:
MCP_TRANSPORT=stdio
BROWSH_FIREFOX_PATH=/usr/bin/firefox-esr
HTML2MARKDOWN_PATH=html2markdown
CACHE_TTL_MS=300000
BROWSH_REQUEST_TIMEOUT_MS=30000
ALLOW_PRIVATE_URLS=false
NODE_ENV=productionBROWSH_FIREFOX_PATHは、ヘッドレス/HTTP動作中にBrowshが使用するFirefox実行ファイルをカスタマイズできます。HTML2MARKDOWN_PATHは、html2markdownバイナリへのカスタムパスを指定できます(デフォルト: PATH内のhtml2markdown)。CACHE_TTL_MS、BROWSH_REQUEST_TIMEOUT_MS、ALLOW_PRIVATE_URLSは、それぞれレンダリングキャッシュ、リクエストごとのタイムアウト、SSRFガードを調整します。BrowshのHTTPポート/ホストは設定できません。
プロジェクトドキュメント
ファイル | 対象 | 目的 |
| エージェント | 運用ルール、ガードレール、タスクライフサイクル |
| 全員 | MCP応答/出力のデザイン言語 |
| 開発者 | システム概要、コンポーネント配線 |
| 開発者 | ツールの入出力スキーマとエラーモデル |
| 開発者 | DateTime標準、SOLIDガイドライン |
| 全員 | バージョン履歴(Keep a Changelog) |
| チーム | カンバンタスクファイル(バックログ → アーカイブ) |
このREADMEはユーザー向けのエントリーポイントです。エージェント向けのルールは AGENTS.md にあり、実装前の必読事項です。
ツールAPI
名前 | パラメータ | AIユースケース/説明 |
fetch_web |
| JSレンダリング後の1ページをテキスト/HTML/Markdownとして取得します。 |
search_web |
| Webを検索し(DuckDuckGo HTML + Bingを同時にレンダリング)、 |
extract_links |
| JSレンダリングされたページ上のすべてのハイパーリンク( |
fetch_web_batch |
| 1回の呼び出しで最大10URLを取得します(キャッシュ対応)。URLごとに |
戻り値
type: plain: ターミナルスタイルのJS実行済み読み取り可能テキスト(またはエラー文字列)。type: html: JS実行後のHTMLマークアップ文字列(またはエラー文字列)。selectorを使用すると、一致した要素のHTMLのみ。type: markdown: 本文または選択した要素のMarkdown変換(またはエラー文字列)。リンク、見出し、リスト、ページ構造はAI向きコンテキストとして保持されます。type: pdf: PDFドキュメントから抽出されたプレーンテキスト(pdftotext経由、上限20 MB)。エラーは構造化されています:
isError: trueを含むMCP応答と、HTTPステータスが判明している場合はそれを含むFetchErrorメッセージ(黙って空文字列にすることはありません)。
環境変数
これらは .env(自動的に読み込まれます)または環境変数で設定します:
変数 | デフォルト | 説明 |
|
| Browsh が使用する Firefox バイナリ(例: |
|
| html2markdown バイナリへのパス。 |
|
| レンダリングごとのリクエストタイムアウト(ms)。 |
|
|
|
|
| ブラウザプロセスを再利用するまでのリクエスト数。 |
|
| ブラウザプロセスが終了されるまでのアイドル時間(ms)(10分)。 |
|
| インメモリのレンダリングキャッシュ TTL(ms)。 |
|
| ループバック/プライベートターゲットに対する SSRF ガードを無効にするには |
|
| トランスポートタイプ(実装済みは |
|
| Node 環境。 |
AI ガイドによるツール選択
まず
search_webを使用: ページを 見つける には、クエリを実行して最適な結果の URL を選び、それから取得します。単一ページには
fetch_webを使用: 要約/分類用にすばやく読み取り可能な出力が必要な場合はplain。要素・リンク・テーブルを解析するにはhtml。LLM 向けのコンテキストチャンクにはmarkdown。selector/max_chars/wait_msを追加して、トークンを効率的に保ち、安定した関連コンテンツを取得します。深いクロールの前に
extract_linksを使用: 完全な DOM を取得する代わりに、ナビゲーションを低コストで辿ります。複数のソースには
fetch_web_batchを使用: N 回のラウンドトリップの代わりに 1 回の呼び出しで済みます。失敗は URL ごとに分離されます。
エラーハンドリング:
ツールは FetchError をスローし、MCP はアクション可能なメッセージとともに isError: true を返します。無効なプロトコル、SSRF ブロック、セレクタ不一致、HTTP ステータスコード、レンダリング失敗は決して黙って無視されません。
MCP プロトコル: AI クライアント設定
AI クライアント(Claude、Cursor など)を設定する前に、以下を実行する必要があります:
依存関係をインストール:
npm installプロジェクトをビルド:
npm run buildビルド出力から MCP サーバーを起動:
node dist/server.js
Claude Desktop または Cursor の設定例:
{
"mcpServers": {
"blowsh": {
"command": "node",
"args": ["dist/server.js"],
"env": {}
}
}
}opencode の設定例(プロジェクト opencode.json):
{
"$schema": "https://opencode.ai/config.json",
"mcp": {
"blowsh": {
"type": "local",
"command": ["docker", "run", "--rm", "-i", "ghcr.io/mokhtarabadi/blowsh-mcp:latest"],
"enabled": true,
"timeout": 120000
}
},
"permission": { "blowsh_*": "allow" }
}Docker 形式ではホスト側のバイナリは不要です。イメージには Firefox、Browsh、html2markdown が同梱されています。保存後、opencode を再起動してください(設定は起動時に一度だけ読み込まれます)。
グレースフルシャットダウン
blowsh-mcp は SIGINT/SIGTERM を捕捉し、Browsh がクリーンに終了することを保証します。孤児のブラウザプロセスは残りません。
セキュリティと考慮事項
サーバーは Browsh をローカルで実行し、HTTP localhost 経由で取得します。
SSRF ガード: デフォルトでは、
fetch_web/search_web/extract_links/fetch_web_batchは、ループバック、プライベート、リンクローカル、予約済み IP 範囲に解決される URL を拒否します(DNS で確認)。無効にするにはALLOW_PRIVATE_URLS=trueを設定しますが、推奨されません。MCP HTTP/ストリーミング可能サーバーが明示的に設定されている場合を除き、公開は行われません。
ファイアウォールなしでポートをオープンな Web に公開しないでください。
シークレットや設定には環境変数を使用してください。
拡張
src/tools/ に新しいツールを追加し、src/server.ts でエクスポートして、文書化してください。
AI クライアントは docstring を自動的に検出します。
トラブルシューティング
fetchPlain が 404 を返すか JS のレンダリングに失敗する場合: Firefox と Browsh がインストールされ、PATH にあることを確認してください。
Firefox が見つからないか起動に失敗する場合は、
.envのBROWSH_FIREFOX_PATHに Firefox インストール先のフルパスを指定してください。Browsh のポート/ホストは固定されています。これらを変更する環境変数や CLI 設定はありません。
最大限のセキュリティを確保するには、コンテナ内で実行してください。
ライセンス
MIT
著者: Mohammad Reza Mokhtarabadi mmokhtarabadi@gmail.com
Available Tools
1 toolfetch_webFetch Web (plain, html, markdown)A
Fetch a web page and return its content as plain text, HTML, or Markdown. Uses a JS-capable browser for dynamic sites.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | The HTTP/HTTPS web URL to fetch | |
| type | Yes | The output type: plain, html, or markdown |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries full burden for behavioral traits. It discloses the use of a JS-capable browser, which is critical for understanding behavior with dynamic sites. It does not mention rate limits or error handling, but the core behavioral trait is well covered.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences, front-loading the main purpose and adding the browser capability as a key differentiator. Every sentence earns its place without redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given 2 simple parameters, no output schema, and no annotations, the description is sufficient. It covers the purpose, output types, and a notable behavior (JS browser). Minor missing details like response size limits or timeout are not critical for a basic fetch tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% with both parameters well described. The description adds 'plain text, HTML, or Markdown' but that is a restatement of the enum values. No additional nuance is provided beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action (fetch a web page) and the resource (web page content), and specifies three output types (plain, HTML, Markdown). It distinguishes the tool by mentioning JS-capable browser for dynamic sites, which sets it apart from simple fetchers.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No explicit guidance on when to use this tool vs alternatives or when not to use it. Given no sibling tools are listed, it is minimally adequate but lacks context like 'use for public pages only' or 'prefer for dynamic content'.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
1 tool update
v1.0.0- First observed
fetch_web
TDQS
Scored across 1 tool
With only one tool, there is no possibility of ambiguity between tools. The tool's purpose is clear and distinct.
The single tool name 'fetch_web' follows a clear verb_noun pattern. With only one tool, there is no inconsistency to evaluate.
A single tool is borderline for a server. While it serves a specific purpose, it feels thin compared to typical MCP servers that offer multiple related operations.
The tool provides core web fetching functionality with output format options. A minor gap might be the lack of custom headers or request methods, but agents can work around this for most use cases.
Maintenance
Related MCP Connectors
- CrawioOAuthcom.crawio
Web pages as Markdown, text or HTML, plus Google Maps places and reviews, for AI agents.
Unblocking and fresh web data for agents: URL to Markdown, YouTube, Maps, Amazon, jobs. Pay per call
Web scraping for AI agents. Converts URLs to clean, LLM-ready Markdown with anti-bot bypass.
Headless browser primitives for AI agents when sites need real JS rendering.
Related MCP Servers
AlicenseAqualityFmaintenanceA Model Context Protocol server that enables AI agents to fetch live web content with JavaScript rendering, proxy rotation, and anti-bot evasion.979 npm58MIT- AlicenseNot gradedqualityDmaintenanceEnables AI agents to automate web tasks such as browsing, clicking, typing, and taking screenshots via the Model Context Protocol.1MIT

Browseagent MCPofficial
AlicenseAqualityDmaintenanceEnables AI agents to control web browsers through the Model Context Protocol, supporting navigation, clicking, typing, and screenshots.127 npm1MIT- AlicenseAqualityDmaintenanceEnables AI agents to fetch any web page as clean markdown or screenshot it, turning URLs into LLM-ready context.27 npmMIT