google-scholar-labs-ajg-mcp
Google Scholar Labs Search (AJG 2024 MCP 适配版)
面向本地大模型与 AI 智能体(Agent)的 Model Context Protocol (MCP) 服务。通过用户已登录的本地浏览器会话检索 Google Scholar Labs 学术文献,并严格依据 AJG 2024 (Academic Journal Guide / ABS) 权威期刊分级目录进行同行评议期刊筛选。
终端 Dry-Run 离线演示

Related MCP server: Gemini Research MCP Server
核心特性
严格的 AJG 2024 期刊分级筛选:将检索到的文献出版物(Venue)与官方 AJG 2024 目录严格匹配(默认 ABS2+:
2、3、4、4*),支持自定义星级门槛与学科大类筛选(如FINANCE、ACCOUNT、STRAT、ECON、ORMAN等)。透明的剔除记录(Exclusion Transparency):不符合条件的文献(如预印本 arXiv/SSRN、未被 AJG 收录的期刊、星级低于设定门槛或学科不匹配)均在
exclusions中完整记录并说明具体原因,杜绝将非核心刊物误判为合格文献。全量合格结果输出:返回当前检索页所有满足评级条件的论文,不人为截断为固定前 3 篇。
人机协同安全交接(Human-in-the-Loop Handoff):遇到 Google 登录验证或验证码(CAPTCHA)时立即安全暂停,返回
handoff_required: true,由用户在本地浏览器界面中手工完成验证,绝不尝试暴力绕过或窃取凭证。本地优先与零遥测:完全运行于本地环境,通过标准 Stdio JSON-RPC 2.0 通信,不向任何第三方服务器上传凭证或搜索记录。
零外部依赖核心解析:内置核心期刊目录与纯标准库 XLSX 解析器,即使在无外部 Excel 文件的 CI 或纯净环境中也能执行完整的确定性测试。
架构与工作流
[ AI 智能体 (Codex / Claude / Cursor / Windsurf) ]
│
(Stdio JSON-RPC 2.0)
▼
[ ScholarLabsMCPServer ]
│ │
│ (Dry-Run / Mock) │ (浏览器自动化模式)
▼ ▼
[ 快速 Schema 验证 ] [ CloakBrowser 会话 ]
│ (本地持久化 Profile)
▼
[ Google Scholar Labs ]
│ (HTML DOM 卡片提取)
▼
[ 候选论文卡片 ]
│
▼
[ AJG 2024 匹配引擎 ]
┌──────────┴──────────┐
▼ ▼
[ 合格文献列表 ] [ 剔除记录 ]
└──────────┬──────────┘
▼
[ 结构化 JSON 响应结果 ]安装与配置
环境要求
Python 3.10 或更高版本
(执行真实自动化搜索时可选)
cloakbrowser库与 Chromium 浏览器环境
1. 源码安装
git clone https://github.com/divenire990/Google-scholar-labs-ajg-mcp.git
cd Google-scholar-labs-ajg-mcp
pip install -e .安装开发与构建依赖:
pip install -e ".[dev]"
# 或者仅安装打包构建依赖:
pip install -e ".[build]"2. 构建分发包 (sdist & wheel)
构建源码分发包(.tar.gz)与二进制 Wheel(.whl):
pip install build
python -m build构建生成的文件位于 dist/ 目录中(已被 .gitignore 自动忽略)。
3. 环境变量配置(可选)
复制 .env.example 为 .env 或在终端中配置环境变量:
# 本地浏览器持久化 Profile 路径(保存 Google 登录态)
export SCHOLAR_LABS_BROWSER_PROFILE="$HOME/.scholar-labs/browser-profile"
# 自定义 AJG2024.xlsx 数据文件路径(未设置时自动使用内置核心期刊或 data/AJG2024.xlsx)
export AJG_DATA_PATH="/path/to/AJG2024.xlsx"MCP 客户端配置
将 google-scholar-labs-ajg-mcp 添加至您的 AI 客户端配置中:
Claude Desktop / Claude Code (claude_desktop_config.json)
{
"mcpServers": {
"google-scholar-labs-ajg-mcp": {
"command": "python",
"args": ["-m", "scholar_labs.mcp_server"],
"env": {
"SCHOLAR_LABS_BROWSER_PROFILE": "/path/to/your/browser-profile",
"AJG_DATA_PATH": "/path/to/AJG2024.xlsx"
}
}
}
}Codex / Windsurf / Cursor (mcp.json 或 .toml)
[mcp_servers.google_scholar_labs_ajg_mcp]
command = "python"
args = ["-m", "scholar_labs.mcp_server"]工具接口说明:scholar_labs_search
输入参数
参数名 | 类型 | 默认值 | 描述 |
|
| (必填) | 提交给 Google Scholar Labs 的学术检索主题、问题或关键词。 |
|
|
| 最低 AJG 星级筛选门槛( |
|
|
| 可选的学科领域代码列表(例如 |
|
|
| 首次解析提取的最大候选卡片数。 |
|
|
| 是否以无头模式运行浏览器。 |
|
|
| 自定义持久化 Profile 目录路径(覆盖环境变量)。 |
|
|
| Dry-run 模式:仅验证查询与 AJG 匹配引擎,不启动浏览器。 |
|
|
| 用于离线评估与测试的 Mock HTML 内容。 |
输出响应示例
{
"status": "ok | blocked | no_results | error",
"message": "执行结果摘要",
"query": "dynamic strategic deviation and earnings management",
"min_stars": "2",
"fields_filter": ["FINANCE", "ACCOUNT"],
"total_candidates_found": 8,
"qualified_count": 3,
"exclusion_count": 5,
"qualified_papers": [
{
"title": "Corporate Governance and Financial Reporting Quality",
"authors": "J Smith, A Taylor",
"year": 2022,
"venue": "Journal of Financial Economics",
"scholar_url": "https://doi.org/10.1016/j.jfineco.2022.01.001",
"annotation": "Investigates the causal link between strategic board adjustments and reporting accuracy.",
"citation_signal": "Cited by 142",
"position": 1,
"raw_text": "...",
"ajg_info": {
"official_title": "Journal of Financial Economics",
"ajg_star": "4*",
"field": "FINANCE",
"is_ft50": true,
"is_utd24": true,
"print_issn": "0304-405X"
},
"rank_score": 51.9
}
],
"exclusions": [
{
"title": "Machine Learning in Financial Forecasting",
"venue": "arXiv preprint arXiv:2104.01234",
"reason": "unmatched_venue",
"details": "Venue 'arXiv preprint' not found in AJG 2024 journal index",
"position": 3
}
],
"handoff_required": false,
"handoff_url": null
}离线测试与验证
运行确定性单元测试:
python -m unittest discover -s tests -p "test_*.py"所有测试均在 2 秒内完成,无任何网络或浏览器依赖。
隐私、安全与合规声明
安全交接与零绕过原则:本工具绝不尝试自动化破解 Google CAPTCHA 验证码,绝不收集、导出或传输用户 Google 账号密码。遇验证要求时立即暂停并提示用户手动处理。
本地凭证隔离:所有 Cookie 和登录会话均保存在用户指定的本地 Profile 目录中,不进行任何远程同步。
合规提示:Google Scholar Labs 为 Google 旗下实验性学术产品,使用者须自行遵守 Google 服务条款与学术检索规范。
上游归属与开源协议
本项目基于 MIT License 开源。详见 LICENSE 文件。
归属致谢
本项目是在原 Scholar Labs Search 项目概念基础上演进与扩展的独立适配版本,新增了:
AJG 2024 (ABS) 学术期刊分级筛选与加权排序
结构化剔除分类机制(Exclusion Transparency)
标准 Model Context Protocol (MCP) JSON-RPC 协议适配
确定性离线测试套件与安全交接架构
Available Tools
1 toolscholar_labs_searchA
Search Google Scholar Labs through a logged-in CloakBrowser session and filter results strictly against the AJG (Academic Journal Guide) 2024 rankings. Returns all qualifying papers (default ABS2+: 2, 3, 4, 4*) and detailed exclusion records for unmatchable or sub-threshold candidates. Supports manual handoff if CAPTCHA or Google login is required.
| Name | Required | Description | Default |
|---|---|---|---|
| query | Yes | The search topic, research question, or keyword query for Scholar Labs. | |
| fields | No | Optional list of AJG fields to filter journals (e.g. ['ACCOUNT', 'FINANCE', 'ECON', 'ORMAN', 'STRAT']). | |
| dry_run | No | If true, validates query and matcher setup without launching browser. | |
| headless | No | Run CloakBrowser in headless mode. Set to false if interactive takeover or visual inspection is desired. | |
| min_stars | No | Minimum AJG star rating required for qualification ('1', '2', '3', '4', '4*'). Default is '2' (ABS2+). | 2 |
| mock_html | No | Mock HTML content for non-network / offline testing and verification. | |
| profile_dir | No | Path to persistent browser profile directory (defaults to SCHOLAR_LABS_BROWSER_PROFILE or ~/.scholar-labs/browser-profile). | |
| ajg_data_path | No | Path to AJG2024.xlsx data file (defaults to AJG_DATA_PATH or data/AJG2024.xlsx). | |
| max_candidates | No | Maximum raw candidate cards to extract from the first visible Scholar Labs results page. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden and largely succeeds: it discloses the logged-in-session requirement, the AJG strict-filtering behavior, and the CAPTCHA/manual-handoff scenario. It adds context beyond what structured fields offer, though it stops short of mentioning rate limits or failure modes beyond CAPTCHA. No contradiction with annotations exists since none are present.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, each earning its place: purpose, return behavior, and fallback handling. The primary purpose is front-loaded in sentence one. No filler or redundancy. Slightly more could be trimmed but it is appropriately tight for a tool of this complexity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex browser-automation tool with 9 parameters, no output schema, and no annotations, the description covers the core workflow (search, AJG filtering, return of qualifying/excluded records) and the critical handoff path. It lacks an exact return-format spec, but the high-level return description partially compensates for the missing output schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all nine parameters are already documented in the schema with types, defaults, and descriptions. The tool description adds no additional parameter-level detail beyond restating the ABS2+ default that min_stars already encodes. Baseline 3 applies; the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource ('Search Google Scholar Labs through a logged-in CloakBrowser session') and adds the distinctive filtering behavior ('filter results strictly against the AJG 2024 rankings'). It also specifies the return scope (qualifying papers plus exclusion records). Clear, specific, and unambiguous even without siblings to differentiate.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description states what the tool does and notes the manual-handoff path for CAPTCHA or login, which gives context on when a human may need to step in. However, with no sibling tools listed and no explicit when-to-use vs when-not-to-use statements, the usage guidance is implicit rather than directive. The handoff note is a behavioral fallback, not a usage-exclusion rule.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
1 tool update
v0.1.0- First observed
scholar_labs_search
TDQS
Scored across 1 tool
With only a single tool, there is no possible confusion between competing choices. The tool's purpose is clear and distinct by default.
The name `scholar_labs_search` follows a consistent domain/action pattern. With only one tool, there are no naming conflicts or inconsistencies to evaluate.
A single tool is at the low end of the typical range, but it provides a comprehensive search-and-filter operation for a narrowly scoped server. It is slightly under the usual 3-15 tools yet reasonable for this focused purpose.
The tool covers the full search workflow including AJG filtering, exclusion records, and authentication/CAPTCHA handoff. Within the stated domain of AJG-filtered Google Scholar search, there are no obvious missing operations.
Maintenance
Related MCP Connectors
MCP server for Firecrawl — web search, scraping, and biomedical/arXiv paper search.
Scrape, crawl and search the web for AI agents via MCP.
MCP server for Google search results via SERP API
Free web search for AI agents. No API key required. Hosted MCP in active development.
Related MCP Servers
- AlicenseBqualityDmaintenanceA local MCP server that allows users to search Google Scholar for academic papers by topic, author, and year range without requiring API keys. It utilizes web scraping to provide paginated results for research and academic exploration through natural language.2MIT
- AlicenseAqualityBmaintenanceMCP server for AI-powered research using Gemini. Provides fast grounded web search, deep autonomous research, URL extraction, and session management.627 PyPI9MIT
- AlicenseAqualityDmaintenanceMCP server for the OpenAlex scholarly database, providing AI agents with tools to search and retrieve academic works, authors, and institutions via natural language queries.8MIT
- AlicenseNot gradedqualityCmaintenanceAn MCP server for searching Google Scholar, enabling paper search, author lookup, citation tracking, and BibTeX export for AI assistants and automation workflows.48 PyPI2MIT