Skip to main content
Glama
tanamurayuuki

Gemini URL Context & Search MCP Server

Gemini URL Context & Search MCP Server

Google AI Studio の URL context 機能と Google Search を MCP (Model Context Protocol) サーバーとして実装し、Claude Code からWebページのテキスト抽出と検索を可能にします。

📁 プロジェクト構造

詳細なフォルダ・ファイル構成は PROJECT_STRUCTURE.md をご覧ください。

Related MCP server: Browser Automation MCP Server

🎯 機能

  • 📄 URL Content Extract: Webページのテキストと画像URLを全抽出

  • 🔍 Google Search: Webを検索して関連情報を取得

  • 🏗️ 構造化出力: JSON形式でページ情報を整理

  • 🔗 複数URL対応: 複数URLの一括処理

  • ⚡ 高品質アーキテクチャ: ドメイン駆動設計とTDD

📦 インストール

npxで即座に使用(推奨)

# Claude Code で一発セットアップ
claude mcp add gemini-url-context -s user -e GEMINI_API_KEY="your-key" -- npx @yourcompany/gemini-url-context-mcp@latest

手動インストール

npm install -g @yourcompany/gemini-url-context-mcp

🔧 セットアップ

1. APIキー取得

  1. Google AI Studio にアクセス

  2. "Get API key" → "Create API key"

  3. キーをコピー

2. 自動セットアップ(Claude Code)

# セットアップスクリプトを実行
export GEMINI_API_KEY="your-api-key"
./scripts/setup-claude-code.sh

3. 設定ファイル生成(他のクライアント)

# 各クライアント用設定ファイルを生成
node scripts/generate-configs.js

🚀 使用方法

URL Content Extract

Claude Code で話しかけるだけ:
「https://example.com のテキストと画像を全部抽出して」
Claude Code で話しかけるだけ:
「最新のAI技術について検索して」

応用例

「以下のサイトを比較分析して:
- https://site1.com
- https://site2.com」

「Next.js 14の最新情報を検索して、
関連記事の内容も抽出して」

🛠️ 対応クライアント

  • Claude Code (CLI) - ワンライナーセットアップ

  • Cursor - .cursor/mcp.json

  • VS Code - MCP拡張

  • Claude Desktop - 標準設定

  • LM Studio - MCP Server追加

🏗️ アーキテクチャ

ドメイン駆動設計

  • Domain Layer: Url, Page, ModelName 値オブジェクト

  • Use Case Layer: ビジネスロジック分離

  • Adapter Layer: 外部API統合

  • Infrastructure: MCP プロトコル実装

品質保証

  • TDD: テスト駆動開発

  • 型安全: TypeScript厳密モード

  • エラーハンドリング: 分類された例外処理

  • Value Objects: 不変性保証

🔍 API仕様

url_context_extract

{
  "urls": ["https://example.com"],
  "query": "要約して",
  "model": "gemini-2.0-flash-exp",
  "maxCharsPerPage": 8000
}
{
  "query": "検索キーワード",
  "instruction": "処理指示",
  "model": "gemini-2.0-flash-exp"
}

🎯 MCP作成のベストプラクティス

この実装から学べる要素:

🔥 必須要素

  1. npx対応: "bin" でCLIツール化

  2. 複数クライアント対応: 設定ファイル自動生成

  3. ワンライナー セットアップ: ユーザビリティ最優先

  4. エラーハンドリング: 型付きエラーで安全性

🏗️ アーキテクチャ

  1. ドメイン駆動設計: ビジネスロジック分離

  2. Value Object: 型安全と不変性

  3. Adapter Pattern: 外部依存の抽象化

  4. Factory Pattern: 実装切り替え

🧪 品質管理

  1. TDD: テスト先行開発

  2. 統合テスト: 実動作確認

  3. 型安全: TypeScript活用

  4. lint/format: コード品質

📦 配布戦略

  1. npmパッケージ: 即座にインストール可能

  2. 設定自動化: スクリプトで一発セットアップ

  3. ドキュメント: 使用例とトラブルシューティング

  4. 段階的ロールアウト: パイロット→本格展開

🏆 他実装との差別化

項目

この実装

一般的実装

アーキテクチャ

DDD + Clean Architecture

手続き型

テスト

TDD + 統合テスト

テストなし

型安全

Value Object

文字列ベース

エラー処理

型付きドメインエラー

try-catch

ユーザビリティ

ワンライナーセットアップ

手動設定


開発者: あなた
アーキテクチャ: Claude Code AI
品質: エンタープライズ級
使いやすさ: コンシューマー級

🎉 完璧なMCPサーバーの完成です!

Available Tools

2 tools
url_context_extractC

Extract content from URLs using Gemini AI and return structured JSON with pages, answer, and metadata

ParametersJSON Schema
NameRequiredDescriptionDefault
maxCharsPerPageNoMaximum characters per page (optional, defaults to 8000)
modelNoGemini model name to use (optional, defaults to gemini-2.0-flash-exp)
queryNoOptional query to guide content extraction and summary
urlsYesArray of URLs to extract content from

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. It mentions using Gemini AI and returning structured JSON, but lacks details on rate limits, authentication needs, error handling, or what 'extract content' entails (e.g., web scraping, API calls). For a tool with no annotation coverage, this is a significant gap in transparency.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, efficient sentence that front-loads the core purpose. It avoids unnecessary words and directly communicates the tool's function. However, it could be slightly more structured by separating key components (e.g., input, process, output) for clarity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity (AI-powered extraction with 4 parameters), lack of annotations, and no output schema, the description is incomplete. It doesn't explain the output structure (what 'pages, answer, and metadata' contain), error cases, or behavioral traits like rate limits. For a tool with no structured support, the description should provide more context to be fully helpful.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema description coverage is 100%, so the schema already documents all four parameters (urls, query, model, maxCharsPerPage) with descriptions. The tool description adds no additional parameter semantics beyond what's in the schema, such as examples or constraints. With high schema coverage, the baseline score of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: 'Extract content from URLs using Gemini AI and return structured JSON with pages, answer, and metadata.' It specifies the verb (extract), resource (content from URLs), technology used (Gemini AI), and output format (structured JSON). However, it doesn't explicitly differentiate from the sibling tool 'google_search', which likely serves a different purpose (searching vs. extracting).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives. It doesn't mention the sibling tool 'google_search' or any other tools, nor does it specify prerequisites, exclusions, or typical use cases. The agent must infer usage from the purpose alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 2 tool updatesv1.0.0
    • First observedgoogle_search
    • First observedurl_context_extract

TDQS

B3.3/5.0

Scored across 2 tools

Disambiguation5/5

The two tools have clearly distinct purposes: one performs web searches, while the other extracts content from specific URLs. There is no overlap in functionality, making it easy for an agent to select the correct tool based on whether it needs general search results or URL-specific content extraction.

Naming Consistency5/5

Both tool names follow a consistent verb_noun pattern: 'google_search' and 'url_context_extract'. The naming is clear, descriptive, and adheres to a uniform style, making the tools easily identifiable and predictable in their naming convention.

Tool Count2/5

With only two tools, the server feels thin for its purpose of 'Gemini URL Context & Search'. While the tools cover search and extraction, the scope suggests potential for more operations, such as managing search history or handling multiple URLs, making the tool count insufficient for comprehensive coverage.

Completeness3/5

The tools cover basic search and URL content extraction, but there are notable gaps. For example, there's no tool for saving or managing extracted data, refining searches, or handling batch URL processing. This limits the server's ability to support more advanced workflows within its domain.

Related MCP Connectors

Related MCP Servers

  • A
    license
    B
    quality
    D
    maintenance
    Provides web search capabilities using Google Custom Search API, enabling users to perform searches through a Model Context Protocol server.
    2
    219 npm
    68
    MIT
  • F
    license
    Not graded
    quality
    D
    maintenance
    Enables intelligent web scraping through a browser automation tool that can search Google, navigate to webpages, and extract content from various websites including GitHub, Stack Overflow, and documentation sites.
    1
    -
  • A
    license
    A
    quality
    A
    maintenance
    Extract content from URLs, documents, videos, and audio files using intelligent auto-engine selection. Supports web pages, PDFs, Word docs, YouTube transcripts, and more with structured JSON responses.
    2
    172
    MIT