mcp-research
mcp-research
Webリサーチ、学術論文、Twitter/X、YouTube、およびファイル取り込みのためのMCPサーバーです。AIアシスタント向けの8つのツールをすべてMCP stdioプロトコル経由で提供します。機関アクセス用の認証情報保管庫(クレデンシャル・ボルト)、CAPTCHA検出、トークン効率の高い出力を備えています。
ツール
ツール | 説明 |
| 3段階の検索カスケード: Brave API → DuckDuckGo → HTMLスクレイパー |
| URL取得 → クリーンなMarkdown変換(SSRF保護および24時間キャッシュ付き) |
| 複合パイプライン: クエリ書き換え → 検索 → 並列取得 → 要約 → 合成 |
| YouTube動画 → トランスクリプト、要約、要点、チャプター、引用 |
| ファイルからのテキスト抽出: PDF、DOCX、XLSX、PPTX、音声、動画、画像 |
| DOI / ArXiv / PubMedの解決 → メタデータ + 機関アクセス経由の全文 |
| X.com/Twitterからのツイートおよびスレッドの抽出 |
| ロードされた認証プロファイルと依存関係の状態を表示(シークレットは決して公開しません) |
すべてのツールは読み取り専用です。コンテンツを取得および変換するだけで、何も変更しません。
Related MCP server: The Web MCP
インストール
pip install mcp-researchまたは uvx を使用して直接実行します(インストール不要):
uvx mcp-researchオプションの追加機能:
pip install 'mcp-research[twitter]' # yt-dlp for Twitter extraction
pip install 'mcp-research[youtube]' # yt-dlp + faster-whisper for YouTube
pip install 'mcp-research[academic]' # PyPDF2 for academic PDFs
pip install 'mcp-research[ingest]' # PDF, DOCX, XLSX, PPTX, audio support
pip install 'mcp-research[all]' # everythingセットアップの確認:
mcp-research doctorClaude Codeでの使用
Claude CodeのMCP設定(~/.claude/settings.json またはプロジェクトの .mcp.json)に追加します:
{
"mcpServers": {
"research": {
"command": "uvx",
"args": ["mcp-research"],
"env": {
"BRAVE_API_KEY": "BSA...",
"OLLAMA_URL": "http://localhost:11434"
}
}
}
}Claude Desktopでの使用
claude_desktop_config.json に追加します:
{
"mcpServers": {
"research": {
"command": "uvx",
"args": ["mcp-research"],
"env": {
"BRAVE_API_KEY": "BSA..."
}
}
}
}設定
すべての設定は環境変数経由で行います(オプションのボルトを除き、設定ファイルは不要です)。
変数 | デフォルト | 説明 |
| (空) | Brave Search APIキー。未設定の場合はDuckDuckGoにフォールバックします。 |
|
| 要約/合成用のOllamaエンドポイント。無効にするには空に設定します。 |
|
| 要約および合成に使用するモデル。 |
|
| URL取得キャッシュディレクトリ。 |
|
| キャッシュのTTL(時間単位)。 |
|
| 検索ログディレクトリ(NDJSON)。 |
|
| デフォルトの最大検索結果数。 |
|
| 認証情報保管庫ファイルのパス。 |
|
| ファイル変更時にボルトを自動リロードします。 |
|
| セッションのアイドルタイムアウト(秒単位)。 |
ツールの詳細
web_search
web_search(query, max_results=5, summarize=False, auto_fetch_top=False)最大限の信頼性を確保するため、3段階のカスケードを使用してWebを検索します:
Brave Search API — 高速で高品質(
BRAVE_API_KEYが必要)DuckDuckGoライブラリ — APIキー不要、レート制限時に再試行
DuckDuckGo HTMLスクレイパー — 最後の手段としてのフォールバック
オプション:
summarize: Ollamaを使用して結果を要約します(Ollamaの実行が必要)auto_fetch_top: 上位結果の全コンテンツも取得して返します
fetch_url
fetch_url(url, summarize=False, max_chars=15000)URLを取得し、クリーンなMarkdownに変換します:
SSRF保護: ローカルホスト、プライベートIP、非HTTPスキームをブロック
スマート再試行: 429/5xxエラー時の指数バックオフ、ホップごとのリダイレクト検証
24時間キャッシュ: SHA-256キー付き、TTL設定可能
コンテンツサポート: HTML → Markdown、JSON → コードブロック、バイナリ → 拒否
スマート切り詰め: テキストの途中ではなく、見出しや段落の境界で分割
CAPTCHA検出: Cloudflare、hCaptcha、reCAPTCHA、Akamaiの壁をフラグ付け
トークン効率: デフォルト15K文字(約4Kトークン)、
max_charsで調整可能
research
research(query, depth="standard", context="")複合リサーチパイプライン:
クエリ書き換え — Ollamaが質問を検索キーワードに最適化
Web検索 — 関連ページを検索(結果ゼロ時の再試行拡張あり)
並列取得 — 上位Nページを並列で取得
要約 — Ollamaが各ページを要約
合成 — Ollamaが最終的な引用付き回答を作成
深度レベル:
深度 | ページ数 | 合成 |
| 2 | なし |
| 5 | あり |
| 10 | あり |
すべてのステップはOllamaなしでも正常に機能します。その場合でも検索結果とページコンテンツは取得可能です。
youtube_essence
youtube_essence(url, mode="standard")YouTube動画から構造化されたコンテンツを抽出します:
トランスクリプト: 自動字幕またはWhisperによる文字起こし(ローカル、プライベート)
要約: OllamaによるAI要約
要点: 箇条書きのまとめ
チャプター: タイムスタンプ付きセグメント
引用: 注目すべき引用(ディープモード)
モード: quick (TL;DR), standard (+チャプター), deep (+引用)
yt-dlpが必要です。オプション: 音声のみの動画にはfaster-whisper、メディア抽出にはffmpeg。
deep_ingest
deep_ingest(path, include_types="", max_files=200, summarize=False)ディレクトリ内のファイルまたは単一ファイルからテキストを抽出します:
テキストファイル:
.txt,.md,.json,.csv, ソースコードなどPDF: PyPDF2経由(オプションの依存関係)
Office:
.docx,.xlsx,.pptx(オプションの依存関係)音声/動画: Whisper文字起こし(オプション)
画像: OllamaビジョンモデルによるOCR(オプション)
タイプフィルター: text, pdf, audio, video, image, office
academic_lookup
academic_lookup(identifier, fetch_fulltext=True)複数の識別子タイプから学術論文を解決します:
DOI:
10.xxxx/...→ Crossrefメタデータ + 出版社リダイレクトArXiv:
2301.12345→ アブストラクト + PDFPubMed: PMID → E-utilitiesメタデータ → DOIチェーン
URL: 出版社ページの検出
認証情報保管庫経由の全文アクセス:
EZproxy書き換え(プレフィックスおよびサフィックスモード)
Bearerトークン、APIキー、基本認証、クッキーJar
自動出版社検出(IEEE, Springer, Elsevier, ACM, Wiley, Nature, JSTORなど)
twitter_extract
twitter_extract(url, include_thread=False)戦略カスケードを使用してX.com/Twitterからツイートとスレッドを抽出します:
yt-dlp (プライマリ) — 認証済みアクセスのためにクッキーJarと連携
Twitter API v2 — ボルトにBearerトークンが設定されている場合
HTML取得 — クッキーベースの最後の手段
戻り値: テキスト、作成者、タイムスタンプ、指標(いいね、リツイート、返信)、メディアURL。
vault_status
vault_status()ロードされた認証プロファイル、一致パターン、認証タイプを表示します。シークレットは決して公開しません。また、オプションの依存関係の可用性もチェックします。
認証情報保管庫(Credential Vault)
~/.mcp-research/vault.yaml を作成して、保護されたソースの認証を設定します:
version: 1
profiles:
# University EZproxy for IEEE
ieee-university:
match: "*.ieee.org/**"
ezproxy:
base_url: "https://ezproxy.myuniversity.edu/login?url="
mode: prefix
# Springer via API key
springer:
match: "*.springer.com/**"
auth:
type: api_key
header: "X-ApiKey"
value: "${SPRINGER_API_KEY}"
# X.com via browser cookies
twitter:
match: "*.x.com/**"
auth:
type: cookie_jar
path: "${HOME}/.mcp-research/cookies/twitter.txt"${VAR}は環境変数から解決されます。シークレットはプレーンテキストで保存されません。最初に一致したプロファイルが優先されます(順序が重要です)
認証タイプ:
bearer,basic,api_key,cookie_jar,headersEZproxyモード:
prefix(ベースURLを先頭に追加)またはsuffix(ドメイン書き換え)ホットリロード: ボルトファイルの変更は自動的に反映されます
トークン効率
すべてのツールは、AIのコンテキストウィンドウのトークンを無駄にしないよう、デフォルトでコンパクトな出力を生成します:
ツール | デフォルト出力 | オーバーライド |
| 約15K文字(約4Kトークン) |
|
| ソースあたり約500トークン | 生コンテンツより要約を優先 |
| 約10K文字の全文 | 通知付きで切り詰め |
| 15ファイル、300文字の抜粋 |
|
| 3K文字のトランスクリプト抜粋 | 結果オブジェクトに全文トランスクリプト |
安全性と堅牢性
SSRF保護: すべてのホップでローカルホスト、プライベートIP、リンクローカル、非HTTPスキームをブロック
CAPTCHA検出: Cloudflare、hCaptcha、reCAPTCHA、Akamai、DDoS-Guardの壁を識別
入力検証: サイズ制限、URL検証、安全なリダイレクト追跡
eval/execなし: 動的なコード実行は行いません
ボルトのセキュリティ: シークレットは環境変数から解決され、
repr()ですべての認証値を秘匿しますキャッシュの分離: 所有者のみのディレクトリ権限 (0o700)
正常な劣化: オプションの依存関係が欠けていてもクラッシュせず、明確なメッセージで機能が劣化します
CLI
mcp-research serve # Run MCP stdio server (default)
mcp-research search "query" # Search the web
mcp-research fetch https://example.com # Fetch URL to markdown
mcp-research youtube https://youtu.be/... # Extract YouTube video
mcp-research ingest ./docs/ # Extract text from files
mcp-research academic "10.1109/..." # Resolve academic paper
mcp-research tweet https://x.com/.../123 # Extract tweet
mcp-research vault # Show vault profiles
mcp-research doctor # Check dependencies開発
git clone https://github.com/MABAAM/Maibaamcrawler.git
cd Maibaamcrawler
pip install -e ".[all]"
pytest tests/ -v
python -m mcp_research変更履歴
v0.3.0
認証情報保管庫:
~/.mcp-research/vault.yamlでのYAML設定。環境変数の補間、glob URLマッチング、EZproxy書き換え、ホットリロードに対応セッションプーリング: ボルト認証注入、クッキーJarサポート、アイドル時の破棄を備えたドメインごとのセッション
CAPTCHA検出: Cloudflare、hCaptcha、reCAPTCHA、Akamai、DDoS-Guard、一般的なボットの壁を識別
学術検索: DOI/ArXiv/PubMedの解決、Crossrefメタデータ、ボルト経由の機関全文アクセス
Twitter/X抽出: yt-dlp、API v2、スレッドサポート付きのクッキーベースアクセス
トークン効率: AIコンテキストを維持するためのデフォルト出力上限(取得時は約4Kトークン、リサーチソースあたり約500トークン)
Doctorコマンド:
mcp-research doctorで依存関係と設定をすべてチェックWindowsエンコーディング修正: UTF-8 stdout/stderrラッパーによりcp1252クラッシュを防止
v0.2.0
YouTube essence: トランスクリプト抽出、AI要約、要点、チャプター、引用
Deep ingest: PDF, DOCX, XLSX, PPTX, 音声, 動画, 画像テキスト抽出
Ollama統合: クエリ書き換え、要約、合成、ビジョンOCR
検索ログ: すべての操作に対するNDJSONイベントログ
Brave Search: APIキーサポートを備えたプライマリ検索階層
v0.1.0
初回リリース: 3つのツール (web_search, fetch_url, research)、SSRF保護、キャッシュ
ライセンス
MIT
Available Tools
8 toolsacademic_lookupARead-onlyIdempotent
Resolve a DOI, ArXiv ID, or PubMed ID. Fetch paper via institutional access if configured in vault.
Args: identifier: DOI (10.xxxx/...), ArXiv ID (2301.12345), PubMed ID (12345678), or publisher URL. fetch_fulltext: Attempt to fetch the full paper text via vault credentials / EZproxy.
| Name | Required | Description | Default |
|---|---|---|---|
| identifier | Yes | ||
| fetch_fulltext | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description adds behavioral context beyond annotations: it mentions attempting to fetch full text via vault credentials/EZproxy, which is a key side effect. Annotations already declare readOnlyHint=true and idempotentHint=true, so there is no contradiction. The description supplements annotations well.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise and front-loaded with the primary purpose, followed by parameter details. It contains no extraneous text. Slightly more structure (e.g., separating args clearly) could improve scannability, but it is already efficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the presence of an output schema, the description appropriately focuses on input behavior. It covers the main use cases and mentions the vault configuration requirement. Minor gaps exist (e.g., what happens if fetch_fulltext fails), but overall it is sufficiently complete for a well-annotated tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 0% description coverage, so the description must compensate. It explains that 'identifier' can be a DOI, ArXiv ID, PubMed ID, or publisher URL, and that 'fetch_fulltext' defaults to true. This provides necessary semantics that the schema alone lacks.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool resolves specific academic identifiers (DOI, ArXiv ID, PubMed ID) and optionally fetches full text via institutional access. The verb 'Resolve' and listing of identifier types provide a specific purpose that distinguishes it from siblings like web_search and fetch_url.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly states when to use the tool (for resolving academic identifiers and fetching papers with vault access). While it does not provide explicit 'when not to use' guidance, the sibling tools offer natural alternatives, and the context is clear enough for an AI agent to infer appropriate usage.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
deep_ingestARead-onlyIdempotent
Extract text from files in a directory or single file. Supports text, PDF, DOCX, XLSX, PPTX, audio, video, images.
Args: path: Directory or file path to process. include_types: Comma-separated type filter (text,pdf,audio,video,image,office). Empty = all. max_files: Maximum files to process (1-5000). summarize: If true, generate an AI summary of the combined content.
| Name | Required | Description | Default |
|---|---|---|---|
| path | Yes | ||
| max_files | No | ||
| summarize | No | ||
| include_types | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations confirm read-only, idempotent, non-destructive behavior. The description adds value by detailing the extraction process (text from various formats) and the optional AI summarization feature, which annotations do not cover.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise: a one-line overview followed by a clean bullet-style Args section. Each sentence serves a purpose, and the essential information is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (many file types, 4 parameters, optional summarize), the description sufficiently covers purpose, parameters, and behavior. An output schema exists, so return values need not be detailed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, but the description provides detailed parameter docs (path, include_types, max_files, summarize) with defaults and examples (e.g., 'Comma-separated type filter... Empty = all'). This fully compensates for the schema gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states that the tool extracts text from files (directories or single files), listing supported formats (text, PDF, DOCX, etc.). This distinguishes it from sibling tools like fetch_url (URLs), web_search (web queries), and youtube_essence (YouTube).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains the tool's scope (local file processing) and supported types, providing clear context. However, it does not explicitly state when not to use it or mention alternatives beyond implied differences from siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
fetch_urlARead-onlyIdempotent
Fetch a URL, convert to markdown. SSRF-protected and cached.
Args: url: The URL to fetch. summarize: If true and Ollama is available, include a summary. max_chars: Maximum content chars (default ~15K/4K tokens). Set higher for full pages.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | ||
| max_chars | No | ||
| summarize | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Description adds SSRF protection, caching, and conditional summarization beyond annotations' readOnly/idempotent hints. No contradictions. More details on error handling would improve, but current info is solid.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Extremely concise: single opening sentence plus a three-line bullet list. No fluff, every sentence adds value. Perfect structure for quick scanning.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given output schema exists, description doesn't need return details. It covers security (SSRF), caching, and parameter nuances. Missing authentication or error info, but overall adequate for a fetch tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description fully explains each parameter: url is the URL, summarize has Ollama condition, max_chars includes default and advice to increase for full pages. Adds significant value beyond schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states 'Fetch a URL, convert to markdown' with specific verb and resource. It distinguishes from siblings like web_search and academic_lookup by focusing on fetching a single URL rather than searching or academic data.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage: fetch a specific URL for markdown conversion. It doesn't explicitly compare to siblings but provides enough context (e.g., Ollama availability for summarization) to guide appropriate use.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
researchARead-onlyIdempotent
Compound research: search → fetch top pages → summarize → synthesize.
Args: query: The research question. depth: Research depth — "quick" (2 pages), "standard" (5 pages), or "deep" (10 pages). context: Optional context from prior research to inform synthesis.
| Name | Required | Description | Default |
|---|---|---|---|
| depth | No | standard | |
| query | Yes | ||
| context | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already indicate readOnlyHint, openWorldHint, idempotentHint, and destructiveHint. The description adds valuable behavioral context: the multi-step process (search, fetch, summarize, synthesize) and the meaning of depth, which goes beyond the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is very concise with a front-loaded pipeline overview and bullet points for arguments. Every sentence adds value; no wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given that there is an output schema (not shown) and annotations cover safety, the description explains the tool's composite nature, parameter meanings, and pipeline stages. It is complete for an agent to understand and invoke the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description carries full burden. It explains all three parameters: query is the research question, depth with three options, and context as optional prior research. This fully compensates for missing schema descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool does 'Compound research: search → fetch top pages → summarize → synthesize', which is a specific verb+resource and distinguishes it from sibling tools like web_search, fetch_url, or academic_lookup that perform only individual steps.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides the research pipeline and explains the depth parameter with clear options. It implies use for comprehensive research combining multiple steps, but does not explicitly state when not to use or compare directly with siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
twitter_extractARead-onlyIdempotent
Extract tweet or thread from X.com/Twitter. Supports yt-dlp, API, and cookie-based access.
Args: url: Tweet URL (x.com/user/status/id or twitter.com/user/status/id). include_thread: If true, fetch the full conversation thread.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | ||
| include_thread | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already convey read-only, idempotent, and non-destructive behavior. The description adds that it supports multiple access methods, which is useful context beyond annotations, but doesn't detail error handling or rate limits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is very concise: two sentences for purpose and two bullet-point args. No wasted words, front-loaded with main purpose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has only 2 simple params and an output schema (not shown), the description covers the essential behavior and parameter semantics. It's mostly complete, though could mention output format briefly, but output schema covers that.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description fully compensates by explaining the url format (x.com/user/status/id) and the purpose of include_thread (fetch full thread). Both parameters are clearly described.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb 'Extract' and the resource 'tweet or thread from X.com/Twitter', distinguishing it from siblings like fetch_url by being Twitter-specific.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides technical details (yt-dlp, API, cookie-based access) but lacks explicit guidance on when to use this tool versus alternatives like fetch_url. No when-not-to-use or exclusion criteria.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
vault_statusARead-onlyIdempotent
Show credential vault status, loaded profiles, and optional dependency availability. Never exposes secrets.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and idempotentHint=true. Description adds security assurance 'Never exposes secrets', which is valuable beyond annotations. No contradiction.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with purpose, second adds critical security note. Efficient and well-structured.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Zero parameters, good annotations, output schema exists. Description fully covers the tool's behavior and safety. No gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
No parameters, schema coverage 100%. Description adds meaning by specifying what the tool shows (status, profiles, dependencies) beyond the empty schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description uses specific verb 'Show' and resource 'credential vault status', with clear scope including loaded profiles and dependency availability. Distinguishes from siblings by being the only vault-related tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage context is clear: a status tool to check vault state. No explicit alternatives or exclusions, but the purpose implies when to use. Slight lack of when-not guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
web_searchARead-onlyIdempotent
Search the web using a 3-tier cascade (Brave → DuckDuckGo → scraper).
Args: query: Search query string. max_results: Maximum number of results to return (1-20). summarize: If true and Ollama is available, summarize the results. auto_fetch_top: If true, also fetch the full content of the top result.
| Name | Required | Description | Default |
|---|---|---|---|
| query | Yes | ||
| summarize | No | ||
| max_results | No | ||
| auto_fetch_top | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already indicate read-only, idempotent, and non-destructive behavior. The description adds valuable transparency by revealing the 3-tier cascade (Brave, DuckDuckGo, scraper) and optional Ollama summarization, which are beyond what annotations provide.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is succinct, uses a bullet list for arguments, and front-loads the cascade mechanism. Every sentence provides value without redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the presence of an output schema (not shown but indicated), the description adequately covers input parameters and behavior. It is complete for an AI agent to invoke the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must explain parameters. It does so effectively: query as search string, max_results (1-20), summarize (if Ollama available), and auto_fetch_top (fetch top result content). This adds substantial meaning beyond the JSON schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states it searches the web using a 3-tier cascade, differentiating it from siblings like academic_lookup, fetch_url, and research. The verb 'Search' and resource 'web' are specific, and the cascade detail adds precision.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
While the purpose is clear, the description does not provide explicit guidance on when to use this tool versus alternatives like fetch_url or research. It lacks when-not conditions or comparative context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
youtube_essenceARead-onlyIdempotent
Extract essence from a YouTube video: transcript, summary, key points, chapters, quotes.
Args: url: YouTube URL (youtube.com/watch?v=, youtu.be/, youtube.com/shorts/). mode: Extraction depth — "quick" (TL;DR), "standard" (+ chapters), or "deep" (+ quotes).
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | ||
| mode | No | standard |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, openWorldHint, idempotentHint, and destructiveHint. The description adds no behavioral traits beyond these, such as external API dependency or rate limits. Despite annotations covering safety, the description misses contextual details like needing internet access.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is very concise: a single sentence defining purpose followed by a well-structured Args list. Every sentence is meaningful, and the structure is front-loaded with the core action.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's simplicity (2 parameters, no nested objects), the description covers purpose, parameters, and output types. It lacks information on error handling or return format, but the existence of an output schema mitigates this. Overall, it is adequately complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description fully compensates by explaining the 'url' parameter with allowed formats and the 'mode' parameter with three depth levels and their effects. This adds significant meaning beyond the raw schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Extract essence') and the resource ('YouTube video'), followed by a list of outputs (transcript, summary, key points, chapters, quotes). This distinguishes it from siblings like twitter_extract or fetch_url which target different sources or actions.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear usage context via parameter explanations (allowed URL formats and mode options). However, it does not explicitly mention when to use this tool over alternatives or exclude scenarios, though the specificity to YouTube serves as implicit guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections. Dates show when Glama detected each change.
8 tool updates
v0.1.1- Removed
academic_lookup - Removed
deep_ingest - Removed
fetch_url - Removed
research - Removed
twitter_extract - Removed
vault_status - Removed
web_search - Removed
youtube_essence
6 tool updates
v0.3.0- Added
academic_lookup - Added
deep_ingest - Changed
fetch_url1 field changed- changed
Input schema / properties / max_chars / defaultPrevious value: -50000New value: +0
- Added
twitter_extract - Added
vault_status - Added
youtube_essence
3 tool updates
v0.1.0- First observed
fetch_url - First observed
research - First observed
web_search
TDQS
Each tool targets a distinct source or operation: academic references, local files, URLs, compound research, Twitter, vault status, web search, and YouTube. There is no ambiguity between tools.
Tool names use mixed conventions: verb_noun (fetch_url, web_search), noun_noun (vault_status, youtube_essence), platform_verb (twitter_extract), and single word (research). No consistent pattern.
8 tools is an appropriate scope for a research assistant, covering key sources (web, academic, social media, local files) without being overwhelming.
The toolset covers major research workflows: search, fetch, extract, and synthesize. Minor gaps like result organization or citation management are not critical for core functionality.
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Connectors
Web search, fetch, extract, and research for AI agents. Markdown output + AI-synthesized answers.
Web research for agents: quality-scored Google search, webpage extraction, and deep research.
LLM-ready web search + instant answers + URL-to-clean-text fetch for agents and RAG.
Agent-native search engine with live web research optimized for AI agents.
Related MCP Servers
- FlicenseAqualityNot gradedmaintenanceEnables AI assistants to perform comprehensive research by searching Google, mining Reddit discussions, scraping web content with JS rendering, and synthesizing findings with citations into structured context.51653-
- AlicenseAqualityDmaintenanceEnables AI assistants to access real-time web data through search, markdown scraping, and browser automation while bypassing anti-bot protections. It provides tools for web research, e-commerce monitoring, and data extraction from across the globe.47,8695MIT
- AlicenseNot gradedqualityDmaintenanceEnables AI agents to search the web, fetch pages, and synthesize research through three tools: web_search, fetch_page, and research_topic, all in a single pay-per-use API call.2MIT
- AlicenseNot gradedqualityCmaintenanceProvides web search with content extraction, YouTube subtitles, and optional LLM summarization for AI assistants.Apache 2.0
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/MABAAM/Maibaamcrawler'
If you have feedback or need assistance with the MCP directory API, please join our Discord server