Skip to main content
Glama
tfscharff

DOI Citation Verifier

by tfscharff

🚀 빠른 설치

npx -y github:tfscharff/doi-mcp

또는 Claude Desktop 설정에 추가하세요:

{
  "mcpServers": {
    "doi-mcp": {
      "command": "npx",
      "args": ["-y", "github:tfscharff/doi-mcp"]
    }
  }
}

Related MCP server: CiteStamp MCP server

해결하는 문제

거대 언어 모델은 때때로 존재하지 않는 논문을 인용하거나, 실제 제목을 잘못된 저자에게 귀속시키거나, 출판 세부 정보를 혼동하는 등 학술 인용을 "환각"할 수 있습니다. 이 MCP 서버는 다음과 같은 방법으로 이 문제를 해결합니다:

  1. 9개 데이터베이스 검증: CrossRef, OpenAlex, PubMed, zbMATH, ERIC, HAL, INSPIRE-HEP, Semantic Scholar, DBLP 전반에 걸쳐 인용을 확인합니다.

  2. 병렬 검색: 모든 데이터베이스를 동시에 쿼리하여 빠른 결과(~1초)를 제공합니다.

  3. 포괄적인 범위: STEM, 인문학, 사회과학, 교육학을 포함한 모든 분야의 6억 개 이상의 출판물을 다룹니다.

  4. DOI 기반 인용: 검증된 모든 인용에는 유효하고 클릭 가능한 DOI가 포함됩니다.

기능

  • 9개 데이터베이스 검색: CrossRef, OpenAlex, PubMed, zbMATH, ERIC, HAL, INSPIRE-HEP, Semantic Scholar, DBLP

  • 인용 검증: 특정 세부 정보가 포함된 논문이 모든 데이터베이스에 실제로 존재하는지 확인

  • 검증된 논문 찾기: 주제에 대한 실제 논문을 검색하고 검증된 인용만 확보

  • 병렬 처리: 모든 데이터베이스 쿼리가 동시에 실행되어 속도 극대화

  • 성능 최적화: 스마트 캐싱 및 조기 종료 전략으로 25-35% 더 빠른 검증

  • 소스 선택: 모든 데이터베이스 검색 또는 특정 소스 타겟팅

  • 인용 형식 지정: DOI가 포함된 적절한 형식의 인용 반환

  • 제로 구성: API 키 없이 모든 데이터베이스를 즉시 사용 가능

작동 방식

AI 어시스턴트가 연구나 인용에 대해 질문을 받으면:

  1. 이 MCP가 없을 때: 어시스턴트는 존재하지 않는 논문을 참조하며 "Nature지에 실린 Smith et al. (2023)에 따르면..."과 같이 인용할 수 있습니다.

  2. 이 MCP가 있을 때: 어시스턴트는 먼저 verifyCitation을 사용하여 9개 데이터베이스를 병렬로 검색하고 다음을 반환합니다:

    • 전체 DOI와 함께 검증된 일치 항목 → 인용 가능

    • 일치 항목 없음 → 인용 불가; 대신 실제 논문을 검색해야 함

도구

verifyCitation

주요 환각 방지 도구 - 인용이 언급되기 전에 여러 데이터베이스에 걸쳐 존재하는지 검증합니다.

입력:

  • title (문자열, 선택 사항): 논문 제목 (부분 일치 허용)

  • authors (배열, 선택 사항): 저자 이름 (성만으로도 충분)

  • year (숫자, 선택 사항): 출판 연도

  • doi (문자열, 선택 사항): DOI를 알고 있는 경우

  • journal (문자열, 선택 사항): 저널 이름

반환 JSON:

  • verified: true/false

  • verified=true인 경우: DOI, 제목, 저자, 연도, 저널, URL, 소스 데이터베이스

  • verified=false인 경우: 일치하는 출판물을 찾을 수 없다는 경고 메시지

  • 투명성을 위한 일치 품질 지표

성공적인 검증 예시:

{
  "verified": true,
  "doi": "10.1038/s41586-023-06004-9",
  "title": "Accurate structure prediction of biomolecular interactions...",
  "authors": ["John Jumper", "Richard Evans", "..."],
  "year": 2023,
  "journal": "Nature",
  "url": "https://doi.org/10.1038/s41586-023-06004-9",
  "source": "crossref",
  "message": "✓ Citation verified"
}

findVerifiedPapers

주제에 대한 실제 논문을 검색하고 여러 데이터베이스에서 DOI가 포함된 검증된 인용만 반환합니다.

입력:

  • query (문자열): 검색 쿼리 (주제, 키워드, 저자 이름)

  • source (문자열, 선택 사항): 검색할 데이터베이스 - "all" (기본값), "crossref", "openalex", "pubmed", "zbmath", "eric", "hal", "inspirehep", "semanticscholar", 또는 "dblp"

  • limit (숫자, 선택 사항): 소스당 결과 수 (1-20, 기본값: 5)

  • yearFrom (숫자, 선택 사항): 최소 출판 연도

  • yearTo (숫자, 선택 사항): 최대 출판 연도

반환: 소스를 포함한 완전한 인용 정보와 함께 지정된 데이터베이스에서 검증된 논문 배열

예시:

// Search all 9 databases
findVerifiedPapers({ query: "CRISPR gene editing", limit: 5 })

// Search only PubMed for biomedical papers
findVerifiedPapers({ query: "cancer immunotherapy", source: "pubmed", limit: 10 })

// Search zbMATH for mathematics papers
findVerifiedPapers({ query: "algebraic topology", source: "zbmath" })

// Search DBLP for computer science papers
findVerifiedPapers({ query: "neural networks", source: "dblp", yearFrom: 2020 })

// Search ERIC for education research
findVerifiedPapers({ query: "active learning pedagogy", source: "eric" })

// Search HAL for French/European humanities research
findVerifiedPapers({ query: "phenomenology Husserl", source: "hal" })

// Search INSPIRE-HEP for high-energy physics papers
findVerifiedPapers({ query: "Higgs boson", source: "inspirehep" })

설치

Claude Desktop 설정 파일에 추가하세요:

Windows: %APPDATA%\Claude\claude_desktop_config.json macOS: ~/Library/Application Support/Claude/claude_desktop_config.json Linux: ~/.config/Claude/claude_desktop_config.json

{
  "mcpServers": {
    "doi-mcp": {
      "command": "npx",
      "args": ["-y", "github:tfscharff/doi-mcp"]
    }
  }
}

Claude Desktop을 다시 시작하면 서버를 사용할 수 있습니다.

대안: 전역 설치

npm install -g github:tfscharff/doi-mcp

그런 다음 이 설정을 사용하세요:

{
  "mcpServers": {
    "doi-mcp": {
      "command": "doi-mcp"
    }
  }
}

대안: 로컬 복제

git clone https://github.com/tfscharff/doi-mcp.git
cd doi-mcp
npm install
npm run build

로컬 설치를 위한 설정:

{
  "mcpServers": {
    "doi-mcp": {
      "command": "node",
      "args": ["/absolute/path/to/doi-mcp/dist/index.js"]
    }
  }
}

문제 해결

서버가 연결되지 않는 경우

  1. Node.js가 설치되어 있는지 확인: node --version (v18+ 필요)

  2. Claude Desktop 로그 확인:

    • Windows: %APPDATA%\Claude\logs\

    • macOS: ~/Library/Logs/Claude/

    • Linux: ~/.config/Claude/logs/

npx 명령 실패

npm cache clean --force

로컬 테스트

npx @modelcontextprotocol/inspector node dist/index.js

개발

# Install dependencies
npm install

# Build
npm run build

# Development with watch mode
npm run dev

사용 예시

이 MCP 이전 (인용 환각):

User: "Tell me about recent AlphaFold research"
Assistant: "According to Johnson et al. (2024) in Science, AlphaFold3 achieved..."
           ❌ This paper doesn't exist

이 MCP 이후 (검증된 인용만):

User: "Tell me about recent AlphaFold research"
Assistant: [Uses findVerifiedPapers tool]
           "According to Jumper et al. (2023) in Nature (DOI: 10.1038/s41586-023-06004-9), 
            AlphaFold3 achieved..."
           ✓ Real paper with valid DOI verified across databases

검증을 통해 가짜 인용을 잡아냄:

User: "Can you verify this citation: Smith et al. (2024), 'Quantum AI', Nature"
Assistant: [Uses verifyCitation tool - searches all 9 databases in parallel]
           "⚠ I cannot verify this citation - no matching publication found in
            any of the 9 databases. This citation may be incorrect."

데이터베이스 범위

모든 데이터베이스는 속도 극대화를 위해 병렬로 쿼리됩니다(총 ~1초):

일반 데이터베이스

  • CrossRef: 모든 분야에 걸친 1억 5천만 개 이상의 학술 출판물

  • OpenAlex: 모든 분야에 걸친 2억 5천만 개 이상의 학술 저작물

  • Semantic Scholar: AI 기반 검색을 제공하는 2억 개 이상의 논문

전문 데이터베이스

  • PubMed: 3천 5백만 개 이상의 생물의학 및 생명과학 출판물

  • zbMATH: 4백만 개 이상의 수학 출판물

  • DBLP: 포괄적인 컴퓨터 과학 서지 (저널 및 컨퍼런스)

  • ERIC: 170만 개 이상의 교육 연구 출판물

  • HAL: 440만 개 이상의 프랑스/유럽 학술 문서 (250만 개 영어)

  • INSPIRE-HEP: 170만 개 이상의 고에너지 물리학 출판물

총 범위

STEM, 컴퓨터 과학, 생물의학, 수학 및 교육 연구 분야의 전문적인 깊이를 갖춘 모든 학문 분야의 6억 개 이상의 출판물.

라이선스

MIT

기여

기여를 환영합니다! 이슈를 제출하거나 풀 리퀘스트를 보내주세요.

관련 자료

API 문서

리소스

Available Tools

3 tools
batchVerifyCitationsA
Read-onlyIdempotent

Verify multiple citations in a single call. More efficient than calling verifyCitation multiple times. Returns verification status for each citation.

ParametersJSON Schema
NameRequiredDescriptionDefault
citationsYesArray of citations to verify

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description adds context beyond annotations by specifying that it 'Returns verification status for each citation,' which clarifies the output behavior. Annotations already indicate it's read-only, idempotent, and non-destructive, so the description doesn't need to repeat those traits, but it usefully describes the return format.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise and front-loaded, consisting of two sentences that efficiently convey the tool's purpose, efficiency benefit, and return value without any wasted words. Every sentence adds value, making it easy for an agent to parse quickly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's moderate complexity, the description is complete enough: it covers purpose, usage guidelines, and output behavior. With annotations handling safety traits and no output schema, the description fills gaps by explaining the return format. However, it could briefly mention error handling or limits for full completeness.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The description mentions 'citations' as the input but doesn't add semantic details beyond what the schema provides. With 100% schema description coverage, the schema fully documents the 'citations' array and its nested properties, so the baseline score of 3 is appropriate as the description doesn't compensate with extra parameter insights.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose with a specific verb ('Verify multiple citations') and resource ('citations'), distinguishing it from sibling tools like 'verifyCitation' by emphasizing batch processing efficiency. It explicitly mentions the return value ('verification status for each citation'), which adds clarity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides explicit guidance on when to use this tool versus alternatives: it states 'More efficient than calling verifyCitation multiple times,' directly comparing it to a sibling tool. This helps the agent choose this tool for batch operations over single-citation verification.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

findVerifiedPapersA
Read-onlyIdempotent

Search multiple academic databases (CrossRef, OpenAlex, PubMed, zbMATH, ERIC, HAL, INSPIRE-HEP, Semantic Scholar, DBLP) for papers and return only verified, real citations with DOIs.

ParametersJSON Schema
NameRequiredDescriptionDefault
queryYesSearch query (topic, keywords, author names)
limitNoNumber of results per source
yearFromNoMinimum publication year
yearToNoMaximum publication year
sourceNoWhich source to searchall

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate read-only, idempotent, and non-destructive behavior. The description adds valuable context beyond this: it specifies the multiple databases searched (CrossRef, OpenAlex, etc.) and the verification requirement (only papers with DOIs are returned). This helps the agent understand the tool's scope and output quality, though it doesn't mention rate limits or authentication needs.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, dense sentence that efficiently conveys the tool's purpose, scope, and key behavior. It lists all databases upfront and specifies the verification requirement without unnecessary words. Every element earns its place, making it highly concise and well-structured.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (searching multiple databases with verification), annotations cover safety (read-only, non-destructive), and schema fully documents parameters, the description provides good contextual completeness. It explains the multi-source approach and DOI verification, though without an output schema, it doesn't detail the return format (e.g., what fields are included).

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, providing full parameter documentation. The description doesn't add any parameter-specific details beyond what's in the schema (e.g., it doesn't explain query syntax or source differences). With high schema coverage, the baseline score of 3 is appropriate as the description doesn't compensate but doesn't need to.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('search multiple academic databases'), the resource ('papers'), and a key distinguishing feature ('return only verified, real citations with DOIs'). It differentiates from siblings by focusing on multi-source search with verification, unlike batchVerifyCitations and verifyCitation which likely handle verification of existing citations rather than searching.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage context by specifying it searches 'multiple academic databases' and returns 'verified, real citations with DOIs', suggesting it's for finding reliable academic sources. However, it doesn't explicitly state when to use this tool versus its siblings (batchVerifyCitations, verifyCitation), which likely handle different verification scenarios.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

verifyCitationA
Read-onlyIdempotent

CRITICAL: Use this to verify ANY academic citation before mentioning it. Checks multiple databases (CrossRef, OpenAlex, PubMed, zbMATH, ERIC, HAL, INSPIRE-HEP, Semantic Scholar, DBLP) if a paper exists. Returns null if not found.

ParametersJSON Schema
NameRequiredDescriptionDefault
titleNoPaper title (partial matches accepted)
authorsNoAuthor names (last names sufficient)
yearNoPublication year
doiNoDOI if known
journalNoJournal name

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description adds valuable behavioral context beyond annotations: it lists the specific databases checked (CrossRef, OpenAlex, etc.) and states that it 'returns null if not found,' which clarifies the output behavior. Annotations already indicate it's read-only, idempotent, and non-destructive, so the description doesn't need to repeat those traits, but it enhances understanding with operational details.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is highly concise and well-structured: it starts with a critical warning, states the purpose and usage in a single sentence, lists databases efficiently, and ends with return behavior. Every sentence adds essential information without redundancy, making it front-loaded and easy to parse.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (verifying citations across multiple databases) and the absence of an output schema, the description is mostly complete: it explains the purpose, usage, databases checked, and return behavior. However, it lacks details on error handling, rate limits, or authentication needs, which could be useful for full transparency. The annotations cover safety aspects, so it's adequate but not exhaustive.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 100% schema description coverage, the input schema fully documents all 5 parameters (title, authors, year, doi, journal), including details like 'partial matches accepted' for title and 'last names sufficient' for authors. The description adds no additional parameter information, so it meets the baseline of 3 by not duplicating schema content.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose with a specific verb ('verify') and resource ('academic citation'), explicitly distinguishes it from siblings by specifying it's for verifying citations before mentioning them (unlike batchVerifyCitations or findVerifiedPapers), and provides critical context about checking multiple databases. The 'CRITICAL' prefix emphasizes its importance.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly states when to use this tool ('before mentioning [a citation]') and provides clear alternatives by naming sibling tools (batchVerifyCitations, findVerifiedPapers), though it doesn't detail when to use those instead. The 'CRITICAL' label implies it should be used for any citation verification, making the guidance comprehensive.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 3 tool updatesv1.0.0
    • First observedbatchVerifyCitations
    • First observedfindVerifiedPapers
    • First observedverifyCitation

TDQS

A4.3/5.0

Scored across 3 tools

Disambiguation5/5

Each tool has a clearly distinct purpose: batchVerifyCitations handles multiple citations efficiently, verifyCitation checks individual citations, and findVerifiedPapers searches databases for verified papers. There is no overlap in functionality, making tool selection straightforward for an agent.

Naming Consistency4/5

The naming follows a consistent verb_noun pattern (batchVerifyCitations, findVerifiedPapers, verifyCitation), with all tools using camelCase. However, verifyCitation lacks a noun suffix like 'Citation' in its verb part, which is a minor deviation from perfect consistency.

Tool Count4/5

With 3 tools, the count is reasonable for a DOI citation verification server, covering core operations (verify single, verify batch, search verified). It might be slightly thin, as additional tools for managing results or databases could enhance completeness, but it's well-scoped for the basic purpose.

Completeness4/5

The tool set covers key verification tasks: single and batch verification, plus searching for verified papers. Minor gaps exist, such as tools for updating or deleting verification data, but the core workflow of verifying and finding citations is adequately supported without dead ends.

Maintenance

ActivityStale
ResponsivenessUnresponsive

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    B
    maintenance
    Ground citations before your agent emits them by checking references against public scholarly registries and flagging hallucinated or retracted ones.
    MIT
  • F
    license
    Not graded
    quality
    D
    maintenance
    Fabrication-free, DOI-backed citations for AI content and agents, using openAlex public-domain data with resolvable DOIs. Includes an API and planned MCP server for agent-native citation retrieval.
    -