google-scholar-labs-ajg-mcp
Google Scholar Labs Search (AJG 2024 MCP 어댑터 버전)
로컬 대형 언어 모델과 AI 에이전트(Agent)를 위한 Model Context Protocol (MCP) 서비스입니다. 사용자가 이미 로그인한 로컬 브라우저 세션을 통해 Google Scholar Labs 학술 문헌을 검색하고, AJG 2024 (Academic Journal Guide / ABS) 권위 있는 저널 등급 디렉터리에 따라 동료 심사 저널을 엄격하게 필터링합니다.
터미널 Dry-Run 오프라인 데모

Related MCP server: Gemini Research MCP Server
핵심 기능
엄격한 AJG 2024 저널 등급 필터링: 검색된 문헌의 게재지(Venue)를 공식 AJG 2024 디렉터리와 엄격히 매칭하며(기본 ABS2+:
2,3,4,4*), 사용자 정의 별점 임계값 및 학문 분야 필터링(예:FINANCE,ACCOUNT,STRAT,ECON,ORMAN등)을 지원합니다.투명한 제외 기록(Exclusion Transparency): 조건에 맞지 않는 문헌(예: 사전 인쇄본 arXiv/SSRN, AJG에 등재되지 않은 저널, 별점이 설정된 임계값보다 낮거나 학문 분야가 일치하지 않는 경우)은 모두
exclusions에 완전히 기록되고 구체적인 사유가 설명되므로, 비핵심 학술지를 적격 문헌으로 오판하는 것을 방지합니다.전체 적격 결과 출력: 현재 검색 페이지에서 등급 조건을 충족하는 모든 논문을 반환하며, 임의로 고정 상위 3편으로 잘라내지 않습니다.
인간-기계 협업 안전 인계(Human-in-the-Loop Handoff): Google 로그인 인증이나 CAPTCHA를 만나면 즉시 안전하게 일시 중지하고
handoff_required: true를 반환하여, 사용자가 로컬 브라우저 인터페이스에서 수동으로 인증을 완료하도록 합니다. 절대 무차별 우회나 자격 증명 탈취를 시도하지 않습니다.로컬 우선 및 제로 원격 측정: 완전히 로컬 환경에서 실행되며 표준 Stdio JSON-RPC 2.0 통신을 통해 자격 증명이나 검색 기록을 어떤 제3자 서버에도 업로드하지 않습니다.
제로 외부 의존성 핵심 파싱: 핵심 저널 디렉터리와 순수 표준 라이브러리 XLSX 파서가 내장되어 있어, 외부 Excel 파일이 없는 CI 또는 깨끗한 환경에서도 완전한 결정적 테스트를 실행할 수 있습니다.
아키텍처 및 워크플로
[ AI 智能体 (Codex / Claude / Cursor / Windsurf) ]
│
(Stdio JSON-RPC 2.0)
▼
[ ScholarLabsMCPServer ]
│ │
│ (Dry-Run / Mock) │ (浏览器自动化模式)
▼ ▼
[ 快速 Schema 验证 ] [ CloakBrowser 会话 ]
│ (本地持久化 Profile)
▼
[ Google Scholar Labs ]
│ (HTML DOM 卡片提取)
▼
[ 候选论文卡片 ]
│
▼
[ AJG 2024 匹配引擎 ]
┌──────────┴──────────┐
▼ ▼
[ 合格文献列表 ] [ 剔除记录 ]
└──────────┬──────────┘
▼
[ 结构化 JSON 响应结果 ]설치 및 구성
환경 요구 사항
Python 3.10 이상
(실제 자동화 검색 시 선택 사항)
cloakbrowser라이브러리 및 Chromium 브라우저 환경
1. 소스 설치
git clone https://github.com/divenire990/Google-scholar-labs-ajg-mcp.git
cd Google-scholar-labs-ajg-mcp
pip install -e .개발 및 빌드 의존성 설치:
pip install -e ".[dev]"
# 或者仅安装打包构建依赖:
pip install -e ".[build]"2. 배포 패키지 빌드 (sdist & wheel)
소스 배포 패키지(.tar.gz)와 바이너리 Wheel(.whl) 빌드:
pip install build
python -m build빌드로 생성된 파일은 dist/ 디렉터리에 있습니다 (.gitignore에 의해 자동으로 무시됨).
3. 환경 변수 구성 (선택 사항)
.env.example를 .env로 복사하거나 터미널에서 환경 변수를 구성하세요:
# 本地浏览器持久化 Profile 路径(保存 Google 登录态)
export SCHOLAR_LABS_BROWSER_PROFILE="$HOME/.scholar-labs/browser-profile"
# 自定义 AJG2024.xlsx 数据文件路径(未设置时自动使用内置核心期刊或 data/AJG2024.xlsx)
export AJG_DATA_PATH="/path/to/AJG2024.xlsx"MCP 클라이언트 구성
google-scholar-labs-ajg-mcp를 AI 클라이언트 구성에 추가하세요:
Claude Desktop / Claude Code (claude_desktop_config.json)
{
"mcpServers": {
"google-scholar-labs-ajg-mcp": {
"command": "python",
"args": ["-m", "scholar_labs.mcp_server"],
"env": {
"SCHOLAR_LABS_BROWSER_PROFILE": "/path/to/your/browser-profile",
"AJG_DATA_PATH": "/path/to/AJG2024.xlsx"
}
}
}
}Codex / Windsurf / Cursor (mcp.json 또는 .toml)
[mcp_servers.google_scholar_labs_ajg_mcp]
command = "python"
args = ["-m", "scholar_labs.mcp_server"]도구 인터페이스 설명: scholar_labs_search
입력 매개변수
매개변수명 | 유형 | 기본값 | 설명 |
|
| (필수) | Google Scholar Labs에 제출할 학술 검색 주제, 질문 또는 키워드. |
|
|
| 최소 AJG 별점 필터링 임계값 ( |
|
|
| 선택적 학문 분야 코드 목록 (예: |
|
|
| 첫 번째 파싱에서 추출할 최대 후보 카드 수. |
|
|
| 브라우저를 헤드리스 모드로 실행할지 여부. |
|
|
| 사용자 정의 영구 Profile 디렉터리 경로 (환경 변수 재정의). |
|
|
| Dry-run 모드: 쿼리와 AJG 매칭 엔진만 검증하고 브라우저를 시작하지 않습니다. |
|
|
| 오프라인 평가 및 테스트에 사용할 Mock HTML 콘텐츠. |
출력 응답 예시
{
"status": "ok | blocked | no_results | error",
"message": "执行结果摘要",
"query": "dynamic strategic deviation and earnings management",
"min_stars": "2",
"fields_filter": ["FINANCE", "ACCOUNT"],
"total_candidates_found": 8,
"qualified_count": 3,
"exclusion_count": 5,
"qualified_papers": [
{
"title": "Corporate Governance and Financial Reporting Quality",
"authors": "J Smith, A Taylor",
"year": 2022,
"venue": "Journal of Financial Economics",
"scholar_url": "https://doi.org/10.1016/j.jfineco.2022.01.001",
"annotation": "Investigates the causal link between strategic board adjustments and reporting accuracy.",
"citation_signal": "Cited by 142",
"position": 1,
"raw_text": "...",
"ajg_info": {
"official_title": "Journal of Financial Economics",
"ajg_star": "4*",
"field": "FINANCE",
"is_ft50": true,
"is_utd24": true,
"print_issn": "0304-405X"
},
"rank_score": 51.9
}
],
"exclusions": [
{
"title": "Machine Learning in Financial Forecasting",
"venue": "arXiv preprint arXiv:2104.01234",
"reason": "unmatched_venue",
"details": "Venue 'arXiv preprint' not found in AJG 2024 journal index",
"position": 3
}
],
"handoff_required": false,
"handoff_url": null
}오프라인 테스트 및 검증
결정적 단위 테스트 실행:
python -m unittest discover -s tests -p "test_*.py"모든 테스트는 2초 이내에 완료되며 네트워크 또는 브라우저 의존성이 없습니다.
개인정보 보호, 보안 및 규정 준수 고지
안전 인계 및 제로 우회 원칙: 이 도구는 Google CAPTCHA를 자동으로 해독하려 시도하지 않으며, 사용자의 Google 계정 비밀번호를 수집, 내보내기 또는 전송하지 않습니다. 인증 요청이 발생하면 즉시 일시 중지하고 사용자에게 수동 처리를 안내합니다.
로컬 자격 증명 격리: 모든 쿠키와 로그인 세션은 사용자가 지정한 로컬 Profile 디렉터리에 저장되며 원격 동기화를 수행하지 않습니다.
규정 준수 안내: Google Scholar Labs는 Google의 실험적 학술 제품이므로, 사용자는 Google 서비스 약관 및 학술 검색 규정을 준수해야 합니다.
업스트림 귀속 및 오픈소스 라이선스
이 프로젝트는 MIT License로 오픈소스입니다. 자세한 내용은 LICENSE 파일을 참조하세요.
귀속 및 감사
이 프로젝트는 원래 Scholar Labs Search 프로젝트의 개념을 기반으로 발전 및 확장된 독립 어댑터 버전으로, 다음을 새로 추가했습니다:
AJG 2024 (ABS) 학술 저널 등급 필터링 및 가중치 정렬
구조화된 제외 분류 메커니즘 (Exclusion Transparency)
표준 Model Context Protocol (MCP) JSON-RPC 프로토콜 어댑테이션
결정적 오프라인 테스트 스위트 및 안전 인계 아키텍처
Available Tools
1 toolscholar_labs_searchA
Search Google Scholar Labs through a logged-in CloakBrowser session and filter results strictly against the AJG (Academic Journal Guide) 2024 rankings. Returns all qualifying papers (default ABS2+: 2, 3, 4, 4*) and detailed exclusion records for unmatchable or sub-threshold candidates. Supports manual handoff if CAPTCHA or Google login is required.
| Name | Required | Description | Default |
|---|---|---|---|
| query | Yes | The search topic, research question, or keyword query for Scholar Labs. | |
| fields | No | Optional list of AJG fields to filter journals (e.g. ['ACCOUNT', 'FINANCE', 'ECON', 'ORMAN', 'STRAT']). | |
| dry_run | No | If true, validates query and matcher setup without launching browser. | |
| headless | No | Run CloakBrowser in headless mode. Set to false if interactive takeover or visual inspection is desired. | |
| min_stars | No | Minimum AJG star rating required for qualification ('1', '2', '3', '4', '4*'). Default is '2' (ABS2+). | 2 |
| mock_html | No | Mock HTML content for non-network / offline testing and verification. | |
| profile_dir | No | Path to persistent browser profile directory (defaults to SCHOLAR_LABS_BROWSER_PROFILE or ~/.scholar-labs/browser-profile). | |
| ajg_data_path | No | Path to AJG2024.xlsx data file (defaults to AJG_DATA_PATH or data/AJG2024.xlsx). | |
| max_candidates | No | Maximum raw candidate cards to extract from the first visible Scholar Labs results page. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden and largely succeeds: it discloses the logged-in-session requirement, the AJG strict-filtering behavior, and the CAPTCHA/manual-handoff scenario. It adds context beyond what structured fields offer, though it stops short of mentioning rate limits or failure modes beyond CAPTCHA. No contradiction with annotations exists since none are present.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, each earning its place: purpose, return behavior, and fallback handling. The primary purpose is front-loaded in sentence one. No filler or redundancy. Slightly more could be trimmed but it is appropriately tight for a tool of this complexity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex browser-automation tool with 9 parameters, no output schema, and no annotations, the description covers the core workflow (search, AJG filtering, return of qualifying/excluded records) and the critical handoff path. It lacks an exact return-format spec, but the high-level return description partially compensates for the missing output schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all nine parameters are already documented in the schema with types, defaults, and descriptions. The tool description adds no additional parameter-level detail beyond restating the ABS2+ default that min_stars already encodes. Baseline 3 applies; the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource ('Search Google Scholar Labs through a logged-in CloakBrowser session') and adds the distinctive filtering behavior ('filter results strictly against the AJG 2024 rankings'). It also specifies the return scope (qualifying papers plus exclusion records). Clear, specific, and unambiguous even without siblings to differentiate.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description states what the tool does and notes the manual-handoff path for CAPTCHA or login, which gives context on when a human may need to step in. However, with no sibling tools listed and no explicit when-to-use vs when-not-to-use statements, the usage guidance is implicit rather than directive. The handoff note is a behavioral fallback, not a usage-exclusion rule.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
1 tool update
v0.1.0- First observed
scholar_labs_search
TDQS
Scored across 1 tool
With only a single tool, there is no possible confusion between competing choices. The tool's purpose is clear and distinct by default.
The name `scholar_labs_search` follows a consistent domain/action pattern. With only one tool, there are no naming conflicts or inconsistencies to evaluate.
A single tool is at the low end of the typical range, but it provides a comprehensive search-and-filter operation for a narrowly scoped server. It is slightly under the usual 3-15 tools yet reasonable for this focused purpose.
The tool covers the full search workflow including AJG filtering, exclusion records, and authentication/CAPTCHA handoff. Within the stated domain of AJG-filtered Google Scholar search, there are no obvious missing operations.
Maintenance
Related MCP Connectors
MCP server for Firecrawl — web search, scraping, and biomedical/arXiv paper search.
Scrape, crawl and search the web for AI agents via MCP.
MCP server for Google search results via SERP API
Free web search for AI agents. No API key required. Hosted MCP in active development.
Related MCP Servers
- AlicenseBqualityDmaintenanceA local MCP server that allows users to search Google Scholar for academic papers by topic, author, and year range without requiring API keys. It utilizes web scraping to provide paginated results for research and academic exploration through natural language.2MIT
- AlicenseAqualityBmaintenanceMCP server for AI-powered research using Gemini. Provides fast grounded web search, deep autonomous research, URL extraction, and session management.627 PyPI9MIT
- AlicenseAqualityDmaintenanceMCP server for the OpenAlex scholarly database, providing AI agents with tools to search and retrieve academic works, authors, and institutions via natural language queries.8MIT
- AlicenseNot gradedqualityCmaintenanceAn MCP server for searching Google Scholar, enabling paper search, author lookup, citation tracking, and BibTeX export for AI assistants and automation workflows.48 PyPI2MIT