Skip to main content
Glama
afrise

Academic Paper Search MCP Server

by afrise

학술 논문 검색 MCP 서버

대장간 배지

다양한 소스에서 학술 논문 정보를 검색하고 불러올 수 있는 MCP(Model Context Protocol) 서버입니다.

서버는 LLM에 다음을 제공합니다.

  • 실시간 학술 논문 검색 기능

  • 논문 메타데이터 및 초록에 대한 액세스

  • 사용 가능한 경우 전체 텍스트 콘텐츠를 검색하는 기능

  • MCP 사양에 따른 구조화된 데이터 응답

MCP 사양은 주로 Anthropic의 Claude Desktop 클라이언트와의 통합을 위해 설계되었지만, 도구/함수 호출 기능(예: OpenAI의 API)을 지원하는 다른 AI 모델 및 클라이언트와의 잠재적인 호환성을 허용합니다.

참고 : 이 소프트웨어는 현재 개발 중입니다. 기능 및 특징은 변경될 수 있습니다.

특징

이 서버는 다음 도구를 제공합니다.

  • search_papers : 여러 출처에서 학술 논문 검색

    • 매개변수:

      • query (str): 검색 쿼리 텍스트

      • limit (int, 선택 사항): 반환할 최대 결과 수(기본값: 10)

    • 반환: 문서 세부 정보가 포함된 형식화된 문자열

  • fetch_paper_details : 특정 논문에 대한 자세한 정보를 검색합니다.

    • 매개변수:

      • paper_id (str): 논문 식별자(DOI 또는 Semantic Scholar ID)

      • source (str, 선택 사항): 데이터 소스("crossref" 또는 "semantic_scholar", 기본값: "crossref")

    • 반환: 다음을 포함한 포괄적인 논문 메타데이터가 포함된 형식화된 문자열:

      • 제목, 저자, 연도, DOI

      • 장소, 오픈 액세스 상태, PDF URL(Semantic Scholar 전용)

      • 초록 및 TL;DR 요약(가능한 경우)

  • search_by_topic : 날짜 범위 필터를 선택하여 주제별로 논문 검색

    • 매개변수:

      • topic (str): 검색어 텍스트(최대 300자)

      • year_start (int, 선택 사항): 날짜 범위의 시작 연도

      • year_end (int, 선택 사항): 날짜 범위의 종료 연도

      • limit (int, 선택 사항): 반환할 최대 결과 수(기본값: 10)

    • 반환: 다음을 포함한 검색 결과가 포함된 형식화된 문자열:

      • 논문 제목, 저자 및 연도

      • 초록 및 TL;DR 요약(가능한 경우)

      • 장소 및 오픈 액세스 정보

Related MCP server: Research MCP

설정

Smithery를 통해 설치

Smithery를 통해 Claude Desktop용 학술 논문 검색 서버를 자동으로 설치하려면:

지엑스피1

이 방법은 아직 거의 테스트되지 않았습니다 . 서버에 문제가 있는 듯합니다. 스미서리 문제가 해결될 때까지는 독립 실행형 지침을 따르세요.

uv를 통해 설치(수동 설치):

  1. 종속성 설치:

uv add "mcp[cli]" httpx
  1. 환경이나 .env 파일에 필요한 API 키를 설정하세요.

#  These are not actually implemented
SEMANTIC_SCHOLAR_API_KEY=your_key_here 
CROSSREF_API_KEY=your_key_here  # Optional but recommended
  1. 서버를 실행합니다:

uv run server.py

Claude Desktop과 함께 사용

  1. Claude Desktop 구성( claude_desktop_config.json )에 서버를 추가합니다.

{
  "mcpServers": {
    "academic-search": {
      "command": "uv",
      "args": ["run ", "/path/to/server/server.py"],
      "env": {
        "SEMANTIC_SCHOLAR_API_KEY": "your_key_here",
        "CROSSREF_API_KEY": "your_key_here"
      }
    }
  }
}
  1. Claude Desktop을 다시 시작하세요

개발

이 서버는 다음을 사용하여 구축되었습니다.

  • 파이썬 MCP SDK

  • 간소화된 서버 구현을 위한 FastMCP

  • API 요청에 대한 httpx

API 소스

  • 의미론적 학술 API

  • CrossRef API

특허

이 프로젝트는 GNU Affero General Public License v3.0(AGPL-3.0)에 따라 라이선스가 부여됩니다. 이 라이선스는 다음을 보장합니다.

  • 본 소프트웨어는 자유롭게 사용, 수정, 배포할 수 있습니다.

  • 모든 수정 사항은 동일한 라이선스에 따라 오픈 소스로 공개되어야 합니다.

  • 이 소프트웨어를 사용하여 네트워크 서비스를 제공하는 모든 사람은 소스 코드를 공개해야 합니다.

  • 상업적 사용은 허용되지만 소프트웨어와 파생물은 무료이며 오픈 소스로 유지되어야 합니다.

전체 라이센스 텍스트는 LICENSE 파일을 참조하세요.

기여하다

참여를 환영합니다! 다음과 같은 방법으로 도움을 주세요.

  1. 저장소를 포크하세요

  2. 기능 브랜치를 생성합니다( git checkout -b feature/amazing-feature )

  3. 변경 사항을 커밋하세요( git commit -m 'Add amazing feature' )

  4. 브랜치에 푸시( git push origin feature/amazing-feature )

  5. 풀 리퀘스트 열기

참고사항:

  • 기존 코드 스타일과 규칙을 따르세요

  • 새로운 기능에 대한 테스트를 추가합니다.

  • 필요에 따라 문서를 업데이트하세요

  • 변경 사항이 AGPL-3.0 라이센스 조건을 준수하는지 확인하세요.

이 프로젝트에 기여함으로써 귀하는 귀하의 기여가 AGPL-3.0 라이선스에 따라 라이선스되는 데 동의하게 됩니다.

Available Tools

3 tools
fetch_paper_detailsB

Get detailed information about a specific paper.

Args:
    paper_id: Paper identifier (DOI for Crossref, paper ID for Semantic Scholar)
    source: Source database ("semantic_scholar" or "crossref")
ParametersJSON Schema
NameRequiredDescriptionDefault
paper_idYes
sourceNosemantic_scholar

TDQS

B3.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. It states the tool 'Get[s] detailed information,' which implies a read-only operation, but it doesn't disclose any behavioral traits such as authentication needs, rate limits, error handling, or what 'detailed information' includes. This leaves significant gaps in understanding how the tool behaves beyond its basic purpose.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is appropriately sized and front-loaded, starting with a clear purpose statement followed by a concise 'Args' section that lists parameters with brief explanations. Every sentence earns its place by providing essential information without unnecessary details, making it efficient and easy to parse.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's moderate complexity (2 parameters, no annotations, no output schema), the description is partially complete. It covers the purpose and parameters well, but it lacks information on behavioral aspects like what 'detailed information' entails, potential errors, or usage constraints. Without an output schema, the description should ideally hint at the return structure, but it doesn't, leaving some context gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The description adds meaningful semantics beyond the input schema, which has 0% description coverage. It explains that 'paper_id' is a 'Paper identifier (DOI for Crossref, paper ID for Semantic Scholar)' and 'source' is a 'Source database' with options 'semantic_scholar' or 'crossref'. This clarifies the purpose and format of the parameters, compensating well for the lack of schema descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: 'Get detailed information about a specific paper.' This specifies the verb ('Get') and resource ('paper'), making it easy to understand what the tool does. However, it doesn't explicitly differentiate from sibling tools like 'search_by_topic' or 'search_papers', which likely return lists rather than details for a specific paper.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage by specifying that it's for a 'specific paper' and lists the required 'paper_id' and optional 'source' parameters. This suggests it should be used when you have a known paper identifier, but it doesn't explicitly state when to use this tool versus the sibling search tools or provide any exclusions or alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_by_topicB

Search for papers by topic with optional date range.

Note: Query length is limited to 300 characters. Longer queries will be automatically truncated.

Args:
    topic (str): Search query (max 300 chars)
    year_start (int, optional): Start year for date range
    year_end (int, optional): End year for date range  
    limit (int, optional): Maximum number of results to return (default 10)
    
Returns:
    str: Formatted search results or error message
ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
topicYes
year_endNo
year_startNo

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It adds useful context: the query length limit (300 characters with truncation) and the return type (formatted search results or error message). However, it lacks details on permissions, rate limits, error conditions beyond truncation, or pagination behavior, which are important for a search tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is appropriately sized and well-structured. It starts with a clear purpose statement, adds a critical behavioral note (query length limit), and then lists parameters and returns in a formatted way. Every sentence adds value without redundancy, making it easy for an agent to parse quickly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's moderate complexity (4 parameters, no output schema, no annotations), the description is partially complete. It covers parameters and basic behavior but lacks output details (e.g., result format beyond 'formatted'), error handling specifics, and differentiation from siblings. For a search tool, this leaves gaps in guiding the agent effectively.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The description adds significant meaning beyond the input schema, which has 0% description coverage. It explains each parameter's purpose: 'topic' as the search query with a character limit, 'year_start' and 'year_end' for date range, and 'limit' for maximum results with a default. This compensates well for the schema's lack of descriptions, though it could note that year parameters are optional integers.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: 'Search for papers by topic with optional date range.' It specifies the verb ('search'), resource ('papers'), and scope ('by topic with optional date range'), making the intent unambiguous. However, it doesn't explicitly differentiate from sibling tools like 'search_papers' or 'fetch_paper_details,' which would be needed for a perfect score.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives like 'search_papers' or 'fetch_paper_details.' It mentions optional parameters like date range and limit, but doesn't explain scenarios where this tool is preferred over siblings or any prerequisites for usage. This leaves the agent without context for tool selection.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_papersC

Search for papers across multiple sources.

args: 
    query: the search query
    limit: the maximum number of results to return (default 10)
ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
queryYes

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It mentions searching 'across multiple sources' but does not cover critical aspects such as authentication needs, rate limits, pagination, or what the response format looks like. This leaves significant gaps in understanding the tool's behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is brief and front-loaded with the main purpose, but the 'args' section is somewhat redundant as it repeats parameter names without adding new insights. It could be more structured to avoid duplication and enhance clarity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity of a search tool with no annotations and no output schema, the description is incomplete. It lacks details on result format, error handling, source specifics, and behavioral traits, making it inadequate for full contextual understanding.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The description adds meaningful context for both parameters: 'query' is explained as 'the search query,' and 'limit' as 'the maximum number of results to return (default 10).' Since schema description coverage is 0%, this compensates well by clarifying parameter purposes beyond the bare schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states the tool 'Search for papers across multiple sources,' which provides a clear verb ('Search') and resource ('papers'). However, it does not differentiate from sibling tools like 'search_by_topic' or specify what 'multiple sources' entails, making it somewhat vague in distinguishing its unique scope.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives like 'search_by_topic' or 'fetch_paper_details.' The description lacks context on scenarios, prerequisites, or exclusions, leaving usage decisions unclear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 3 tool updatesv1.0.0
    • First observedfetch_paper_details
    • First observedsearch_by_topic
    • First observedsearch_papers

TDQS

C2.8/5.0

Scored across 3 tools

Disambiguation2/5

The tools 'search_by_topic' and 'search_papers' have significant overlap in purpose—both search for papers, with only minor differences in parameters. This creates ambiguity, as an agent might struggle to choose between them. The 'fetch_paper_details' tool is distinct, but the two search tools are not clearly differentiated.

Naming Consistency3/5

The naming is mixed: 'fetch_paper_details' uses a verb_noun pattern, while 'search_by_topic' and 'search_papers' use verb_preposition_noun and verb_noun styles, respectively. This inconsistency makes the set less predictable, though the names are still readable and descriptive.

Tool Count2/5

With only 3 tools, the server feels thin for an academic paper search domain. It lacks essential operations like filtering by author, journal, or citation count, and there's no update or delete functionality, which limits its utility for comprehensive paper management.

Completeness2/5

The tool surface is incomplete for academic paper search. It covers basic fetch and search operations but misses key features such as author-based searches, citation tracking, paper categorization, or integration with reference managers. This will likely cause agent failures in complex workflows.

Maintenance

ActivityInactive
ResponsivenessUnresponsive

Related MCP Connectors

Related MCP Servers

  • F
    license
    Not graded
    quality
    D
    maintenance
    Enables AI assistants to search across multiple academic databases (PubMed, arXiv, bioRxiv, medRxiv, Semantic Scholar) through a unified interface. Supports advanced filtering, metadata retrieval, PDF downloads, and comprehensive research workflows with citation analysis.
    5
    -
  • A
    license
    Not graded
    quality
    C
    maintenance
    Enables LLMs to search, analyze, and summarize academic research papers in real-time from arXiv, Semantic Scholar, and PubMed. Provides automatic deduplication, citation analysis, and BibTeX generation across multiple research databases.
    46 npm
    MIT
  • A
    license
    Not graded
    quality
    D
    maintenance
    Enables searching and downloading academic papers from multiple sources including arXiv, PubMed, bioRxiv, Google Scholar, and Semantic Scholar. Provides standardized tools compatible with OpenAI Deep Research and ChatGPT connectors.
    15
    MIT
  • A
    license
    A
    quality
    C
    maintenance
    Enables retrieval of academic paper metadata, PDFs, full text, citations, and references by title via Semantic Scholar, arXiv, and other sources.
    6
    1
    MIT