google-scholar-labs-ajg-mcp
Google Scholar Labs Search (AJG 2024 MCP 対応版)
ローカル大規模言語モデルと AI エージェント(Agent)向けの Model Context Protocol (MCP) サービスです。ユーザーがログイン済みのローカルブラウザセッションを通じて Google Scholar Labs の学術文献を検索し、AJG 2024 (Academic Journal Guide / ABS) の権威あるジャーナル格付けディレクトリに厳密に基づいて、査読付きジャーナルの選別を行います。
ターミナル Dry-Run オフラインデモ

Related MCP server: Gemini Research MCP Server
主な特徴
厳格な AJG 2024 ジャーナル格付けフィルタリング:検索された文献の出版物(Venue)を公式の AJG 2024 ディレクトリと厳密に照合します(デフォルトは ABS2+:
2、3、4、4*)。カスタムのスター評価しきい値と学問分野別フィルタリング(例:FINANCE、ACCOUNT、STRAT、ECON、ORMANなど)に対応しています。透明な除外記録(Exclusion Transparency):条件を満たさない文献(プレプリントの arXiv/SSRN、AJG に収録されていないジャーナル、設定したしきい値を下回るスター評価、または分野不一致など)はすべて
exclusionsに完全に記録され、具体的な理由が説明されます。中核でないジャーナルを適格な文献として誤判定することを防ぎます。適格結果の全件出力:現在の検索ページで格付け条件を満たすすべての論文を返します。固定の上位 3 件に人為的に切り詰めることはありません。
人間参加型の安全なハンドオフ(Human-in-the-Loop Handoff):Google のログイン認証または CAPTCHA に遭遇した場合は、直ちに安全に一時停止し、
handoff_required: trueを返します。ユーザーがローカルブラウザの画面で手動認証を完了します。回避を強行したり認証情報を窃取したりすることは決してありません。ローカル優先とゼロテレメトリー:完全にローカル環境で動作し、標準の Stdio JSON-RPC 2.0 で通信します。認証情報や検索履歴を第三者サーバーにアップロードすることは一切ありません。
外部依存ゼロのコア解析:中核ジャーナルディレクトリと純粋な標準ライブラリによる XLSX パーサーを内蔵しており、外部の Excel ファイルが存在しない CI やクリーンな環境でも完全な決定論的テストを実行できます。
アーキテクチャとワークフロー
[ AI 智能体 (Codex / Claude / Cursor / Windsurf) ]
│
(Stdio JSON-RPC 2.0)
▼
[ ScholarLabsMCPServer ]
│ │
│ (Dry-Run / Mock) │ (浏览器自动化模式)
▼ ▼
[ 快速 Schema 验证 ] [ CloakBrowser 会话 ]
│ (本地持久化 Profile)
▼
[ Google Scholar Labs ]
│ (HTML DOM 卡片提取)
▼
[ 候选论文卡片 ]
│
▼
[ AJG 2024 匹配引擎 ]
┌──────────┴──────────┐
▼ ▼
[ 合格文献列表 ] [ 剔除记录 ]
└──────────┬──────────┘
▼
[ 结构化 JSON 响应结果 ]インストールと設定
環境要件
Python 3.10 以降
(実際の自動検索を実行する場合に任意)
cloakbrowserライブラリと Chromium ブラウザ環境
1. ソースコードからのインストール
git clone https://github.com/divenire990/Google-scholar-labs-ajg-mcp.git
cd Google-scholar-labs-ajg-mcp
pip install -e .開発用・ビルド用の依存関係をインストール:
pip install -e ".[dev]"
# 或者仅安装打包构建依赖:
pip install -e ".[build]"2. 配布パッケージのビルド(sdist & wheel)
ソース配布パッケージ(.tar.gz)とバイナリ Wheel(.whl)をビルド:
pip install build
python -m buildビルドで生成されたファイルは dist/ ディレクトリに配置されます(.gitignore によって自動的に無視されます)。
3. 環境変数の設定(任意)
.env.example をコピーして .env を作成するか、ターミナルで環境変数を設定します:
# 本地浏览器持久化 Profile 路径(保存 Google 登录态)
export SCHOLAR_LABS_BROWSER_PROFILE="$HOME/.scholar-labs/browser-profile"
# 自定义 AJG2024.xlsx 数据文件路径(未设置时自动使用内置核心期刊或 data/AJG2024.xlsx)
export AJG_DATA_PATH="/path/to/AJG2024.xlsx"MCP クライアント設定
google-scholar-labs-ajg-mcp を AI クライアントの設定に追加します:
Claude Desktop / Claude Code (claude_desktop_config.json)
{
"mcpServers": {
"google-scholar-labs-ajg-mcp": {
"command": "python",
"args": ["-m", "scholar_labs.mcp_server"],
"env": {
"SCHOLAR_LABS_BROWSER_PROFILE": "/path/to/your/browser-profile",
"AJG_DATA_PATH": "/path/to/AJG2024.xlsx"
}
}
}
}Codex / Windsurf / Cursor (mcp.json または .toml)
[mcp_servers.google_scholar_labs_ajg_mcp]
command = "python"
args = ["-m", "scholar_labs.mcp_server"]ツールインターフェースの説明:scholar_labs_search
入力パラメータ
パラメータ名 | 型 | デフォルト値 | 説明 |
|
| (必須) | Google Scholar Labs に送信する学術検索のテーマ、質問、またはキーワード。 |
|
|
| 最低限の AJG スター評価フィルタしきい値( |
|
|
| 任意の学問分野コードのリスト(例: |
|
|
| 最初の解析で抽出する最大候補カード数。 |
|
|
| ブラウザをヘッドレスモードで実行するかどうか。 |
|
|
| カスタムの永続化 Profile ディレクトリパス(環境変数を上書き)。 |
|
|
| Dry-run モード:クエリと AJG マッチングエンジンのみを検証し、ブラウザを起動しません。 |
|
|
| オフライン評価とテスト用の Mock HTML コンテンツ。 |
出力レスポンス例
{
"status": "ok | blocked | no_results | error",
"message": "执行结果摘要",
"query": "dynamic strategic deviation and earnings management",
"min_stars": "2",
"fields_filter": ["FINANCE", "ACCOUNT"],
"total_candidates_found": 8,
"qualified_count": 3,
"exclusion_count": 5,
"qualified_papers": [
{
"title": "Corporate Governance and Financial Reporting Quality",
"authors": "J Smith, A Taylor",
"year": 2022,
"venue": "Journal of Financial Economics",
"scholar_url": "https://doi.org/10.1016/j.jfineco.2022.01.001",
"annotation": "Investigates the causal link between strategic board adjustments and reporting accuracy.",
"citation_signal": "Cited by 142",
"position": 1,
"raw_text": "...",
"ajg_info": {
"official_title": "Journal of Financial Economics",
"ajg_star": "4*",
"field": "FINANCE",
"is_ft50": true,
"is_utd24": true,
"print_issn": "0304-405X"
},
"rank_score": 51.9
}
],
"exclusions": [
{
"title": "Machine Learning in Financial Forecasting",
"venue": "arXiv preprint arXiv:2104.01234",
"reason": "unmatched_venue",
"details": "Venue 'arXiv preprint' not found in AJG 2024 journal index",
"position": 3
}
],
"handoff_required": false,
"handoff_url": null
}オフラインテストと検証
決定論的ユニットテストを実行:
python -m unittest discover -s tests -p "test_*.py"すべてのテストは 2 秒以内に完了し、ネットワークやブラウザへの依存は一切ありません。
プライバシー、セキュリティ、およびコンプライアンスに関する声明
安全なハンドオフと回避ゼロの原則:本ツールは Google CAPTCHA を自動的にクラックしようとは決してせず、ユーザーの Google アカウントのパスワードを収集、エクスポート、または送信することも一切ありません。認証要求に遭遇した場合は直ちに一時停止し、ユーザーに手動での対応を促します。
ローカル資格情報の分離:すべての Cookie とログインセッションは、ユーザー指定のローカル Profile ディレクトリに保存され、リモート同期は一切行われません。
コンプライアンスに関する注意:Google Scholar Labs は Google 傘下の実験的な学術製品です。利用者は Google の利用規約と学術検索規範を自ら遵守する必要があります。
アップストリームの帰属とオープンソースライセンス
本プロジェクトは MIT License の下でオープンソースとして公開されています。詳細は LICENSE ファイルを参照してください。
帰属と謝辞
本プロジェクトは、元の Scholar Labs Search プロジェクトのコンセプトに基づいて発展・拡張された独立した適合版であり、以下の機能を追加しています:
AJG 2024 (ABS) 学術ジャーナルの格付けフィルタリングと加重ソート
構造化された除外分類メカニズム(Exclusion Transparency)
標準の Model Context Protocol (MCP) JSON-RPC プロトコルへの適合
決定論的オフラインテストスイートと安全なハンドオフアーキテクチャ
Available Tools
1 toolscholar_labs_searchA
Search Google Scholar Labs through a logged-in CloakBrowser session and filter results strictly against the AJG (Academic Journal Guide) 2024 rankings. Returns all qualifying papers (default ABS2+: 2, 3, 4, 4*) and detailed exclusion records for unmatchable or sub-threshold candidates. Supports manual handoff if CAPTCHA or Google login is required.
| Name | Required | Description | Default |
|---|---|---|---|
| query | Yes | The search topic, research question, or keyword query for Scholar Labs. | |
| fields | No | Optional list of AJG fields to filter journals (e.g. ['ACCOUNT', 'FINANCE', 'ECON', 'ORMAN', 'STRAT']). | |
| dry_run | No | If true, validates query and matcher setup without launching browser. | |
| headless | No | Run CloakBrowser in headless mode. Set to false if interactive takeover or visual inspection is desired. | |
| min_stars | No | Minimum AJG star rating required for qualification ('1', '2', '3', '4', '4*'). Default is '2' (ABS2+). | 2 |
| mock_html | No | Mock HTML content for non-network / offline testing and verification. | |
| profile_dir | No | Path to persistent browser profile directory (defaults to SCHOLAR_LABS_BROWSER_PROFILE or ~/.scholar-labs/browser-profile). | |
| ajg_data_path | No | Path to AJG2024.xlsx data file (defaults to AJG_DATA_PATH or data/AJG2024.xlsx). | |
| max_candidates | No | Maximum raw candidate cards to extract from the first visible Scholar Labs results page. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden and largely succeeds: it discloses the logged-in-session requirement, the AJG strict-filtering behavior, and the CAPTCHA/manual-handoff scenario. It adds context beyond what structured fields offer, though it stops short of mentioning rate limits or failure modes beyond CAPTCHA. No contradiction with annotations exists since none are present.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, each earning its place: purpose, return behavior, and fallback handling. The primary purpose is front-loaded in sentence one. No filler or redundancy. Slightly more could be trimmed but it is appropriately tight for a tool of this complexity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex browser-automation tool with 9 parameters, no output schema, and no annotations, the description covers the core workflow (search, AJG filtering, return of qualifying/excluded records) and the critical handoff path. It lacks an exact return-format spec, but the high-level return description partially compensates for the missing output schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all nine parameters are already documented in the schema with types, defaults, and descriptions. The tool description adds no additional parameter-level detail beyond restating the ABS2+ default that min_stars already encodes. Baseline 3 applies; the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource ('Search Google Scholar Labs through a logged-in CloakBrowser session') and adds the distinctive filtering behavior ('filter results strictly against the AJG 2024 rankings'). It also specifies the return scope (qualifying papers plus exclusion records). Clear, specific, and unambiguous even without siblings to differentiate.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description states what the tool does and notes the manual-handoff path for CAPTCHA or login, which gives context on when a human may need to step in. However, with no sibling tools listed and no explicit when-to-use vs when-not-to-use statements, the usage guidance is implicit rather than directive. The handoff note is a behavioral fallback, not a usage-exclusion rule.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
1 tool update
v0.1.0- First observed
scholar_labs_search
TDQS
Scored across 1 tool
With only a single tool, there is no possible confusion between competing choices. The tool's purpose is clear and distinct by default.
The name `scholar_labs_search` follows a consistent domain/action pattern. With only one tool, there are no naming conflicts or inconsistencies to evaluate.
A single tool is at the low end of the typical range, but it provides a comprehensive search-and-filter operation for a narrowly scoped server. It is slightly under the usual 3-15 tools yet reasonable for this focused purpose.
The tool covers the full search workflow including AJG filtering, exclusion records, and authentication/CAPTCHA handoff. Within the stated domain of AJG-filtered Google Scholar search, there are no obvious missing operations.
Maintenance
Related MCP Connectors
MCP server for Firecrawl — web search, scraping, and biomedical/arXiv paper search.
Scrape, crawl and search the web for AI agents via MCP.
MCP server for Google search results via SERP API
Free web search for AI agents. No API key required. Hosted MCP in active development.
Related MCP Servers
- AlicenseBqualityDmaintenanceA local MCP server that allows users to search Google Scholar for academic papers by topic, author, and year range without requiring API keys. It utilizes web scraping to provide paginated results for research and academic exploration through natural language.2MIT
- AlicenseAqualityBmaintenanceMCP server for AI-powered research using Gemini. Provides fast grounded web search, deep autonomous research, URL extraction, and session management.627 PyPI9MIT
- AlicenseAqualityDmaintenanceMCP server for the OpenAlex scholarly database, providing AI agents with tools to search and retrieve academic works, authors, and institutions via natural language queries.8MIT
- AlicenseNot gradedqualityCmaintenanceAn MCP server for searching Google Scholar, enabling paper search, author lookup, citation tracking, and BibTeX export for AI assistants and automation workflows.48 PyPI2MIT