DOI Citation Verifier
🚀 クイックインストール
npx -y github:tfscharff/doi-mcpまたは、Claude Desktopの設定に追加してください:
{
"mcpServers": {
"doi-mcp": {
"command": "npx",
"args": ["-y", "github:tfscharff/doi-mcp"]
}
}
}Related MCP server: CiteStamp MCP server
このツールが解決する問題
大規模言語モデルは、存在しない論文を引用したり、実在するタイトルを誤った著者に帰属させたり、出版情報を混同したりする「ハルシネーション(幻覚)」を起こすことがあります。このMCPサーバーは、以下の方法でその問題を解決します:
9つのデータベースによる検証: CrossRef、OpenAlex、PubMed、zbMATH、ERIC、HAL、INSPIRE-HEP、Semantic Scholar、DBLPを横断して引用をチェックします
並列検索: すべてのデータベースを同時にクエリし、高速な結果(約1秒)を提供します
包括的なカバレッジ: STEM、人文科学、社会科学、教育学を含む全分野で6億件以上の出版物をカバー
DOI付きの引用: 検証されたすべての引用には、有効でクリック可能なDOIが含まれます
機能
9つのデータベース検索: CrossRef、OpenAlex、PubMed、zbMATH、ERIC、HAL、INSPIRE-HEP、Semantic Scholar、DBLP
引用の検証: 特定の詳細情報を持つ論文が、すべてのデータベース上で実際に存在するかを確認
検証済み論文の検索: トピックに基づいて実在する論文を検索し、検証済みの引用のみを取得
並列処理: すべてのデータベースクエリを同時に実行し、速度を最大化
パフォーマンス最適化: スマートキャッシュと早期終了戦略により、検証速度を25〜35%向上
ソース選択: すべてのデータベースを検索、または特定のソースをターゲットに設定可能
引用フォーマット: DOI付きの適切にフォーマットされた引用を返却
設定不要: すべてのデータベースがすぐに使用可能で、APIキーは不要
仕組み
AIアシスタントが研究や引用について尋ねられた場合:
このMCPがない場合: アシスタントは「Nature誌のSmithら(2023)によると…」のように、存在しない論文を引用する可能性があります
このMCPがある場合: アシスタントはまず
verifyCitationを使用し、9つのデータベースを並列検索して以下を返します:完全なDOIと一致 → 引用可能
一致なし → 引用不可。代わりに実在する論文を検索する必要がある
ツール
verifyCitation
主要なハルシネーション防止ツール - 引用を言及する前に、複数のデータベースにわたって存在するかを検証します。
入力:
title(string, オプション): 論文タイトル(部分一致可)authors(array, オプション): 著者名(姓のみでも可)year(number, オプション): 出版年doi(string, オプション): DOI(既知の場合)journal(string, オプション): ジャーナル名
返却されるJSON:
verified: true/falseverified=trueの場合: DOI、タイトル、著者、年、ジャーナル、URL、ソースデータベース
verified=falseの場合: 一致する出版物が見つからなかった旨の警告メッセージ
透明性のためのマッチ品質指標
検証成功の例:
{
"verified": true,
"doi": "10.1038/s41586-023-06004-9",
"title": "Accurate structure prediction of biomolecular interactions...",
"authors": ["John Jumper", "Richard Evans", "..."],
"year": 2023,
"journal": "Nature",
"url": "https://doi.org/10.1038/s41586-023-06004-9",
"source": "crossref",
"message": "✓ Citation verified"
}findVerifiedPapers
トピックに基づいて実在する論文を検索し、複数のデータベースからDOI付きの検証済み引用のみを返します。
入力:
query(string): 検索クエリ(トピック、キーワード、著者名)source(string, オプション): 検索対象データベース - "all" (デフォルト), "crossref", "openalex", "pubmed", "zbmath", "eric", "hal", "inspirehep", "semanticscholar", "dblp"limit(number, オプション): ソースごとの結果数 (1-20, デフォルト: 5)yearFrom(number, オプション): 出版年の下限yearTo(number, オプション): 出版年の上限
返却: 指定されたデータベースからの、ソースを含む完全な引用情報を持つ検証済み論文の配列
例:
// Search all 9 databases
findVerifiedPapers({ query: "CRISPR gene editing", limit: 5 })
// Search only PubMed for biomedical papers
findVerifiedPapers({ query: "cancer immunotherapy", source: "pubmed", limit: 10 })
// Search zbMATH for mathematics papers
findVerifiedPapers({ query: "algebraic topology", source: "zbmath" })
// Search DBLP for computer science papers
findVerifiedPapers({ query: "neural networks", source: "dblp", yearFrom: 2020 })
// Search ERIC for education research
findVerifiedPapers({ query: "active learning pedagogy", source: "eric" })
// Search HAL for French/European humanities research
findVerifiedPapers({ query: "phenomenology Husserl", source: "hal" })
// Search INSPIRE-HEP for high-energy physics papers
findVerifiedPapers({ query: "Higgs boson", source: "inspirehep" })インストール
Claude Desktopの設定ファイルに追加してください:
Windows: %APPDATA%\Claude\claude_desktop_config.json
macOS: ~/Library/Application Support/Claude/claude_desktop_config.json
Linux: ~/.config/Claude/claude_desktop_config.json
{
"mcpServers": {
"doi-mcp": {
"command": "npx",
"args": ["-y", "github:tfscharff/doi-mcp"]
}
}
}Claude Desktopを再起動すると、サーバーが利用可能になります。
代替案: グローバルインストール
npm install -g github:tfscharff/doi-mcpその後、以下の設定を使用してください:
{
"mcpServers": {
"doi-mcp": {
"command": "doi-mcp"
}
}
}代替案: ローカルでクローン
git clone https://github.com/tfscharff/doi-mcp.git
cd doi-mcp
npm install
npm run buildローカルインストール用の設定:
{
"mcpServers": {
"doi-mcp": {
"command": "node",
"args": ["/absolute/path/to/doi-mcp/dist/index.js"]
}
}
}トラブルシューティング
サーバーが接続できない場合
Node.jsがインストールされているか確認:
node --version(v18以上が必要)Claude Desktopのログを確認:
Windows:
%APPDATA%\Claude\logs\macOS:
~/Library/Logs/Claude/Linux:
~/.config/Claude/logs/
npxコマンドが失敗する場合
npm cache clean --forceローカルでのテスト
npx @modelcontextprotocol/inspector node dist/index.js開発
# Install dependencies
npm install
# Build
npm run build
# Development with watch mode
npm run dev使用例
このMCP導入前(引用のハルシネーション):
User: "Tell me about recent AlphaFold research"
Assistant: "According to Johnson et al. (2024) in Science, AlphaFold3 achieved..."
❌ This paper doesn't existこのMCP導入後(検証済みの引用のみ):
User: "Tell me about recent AlphaFold research"
Assistant: [Uses findVerifiedPapers tool]
"According to Jumper et al. (2023) in Nature (DOI: 10.1038/s41586-023-06004-9),
AlphaFold3 achieved..."
✓ Real paper with valid DOI verified across databases検証により偽の引用を検知:
User: "Can you verify this citation: Smith et al. (2024), 'Quantum AI', Nature"
Assistant: [Uses verifyCitation tool - searches all 9 databases in parallel]
"⚠ I cannot verify this citation - no matching publication found in
any of the 9 databases. This citation may be incorrect."データベースのカバレッジ
すべてのデータベースは並列でクエリされ、速度を最大化します(合計約1秒):
一般データベース
CrossRef: 全分野にわたる1億5000万件以上の学術出版物
OpenAlex: 全分野にわたる2億5000万件以上の学術著作物
Semantic Scholar: AI駆動検索を備えた2億件以上の論文
専門データベース
PubMed: 3500万件以上の生物医学および生命科学の出版物
zbMATH: 400万件以上の数学出版物
DBLP: 包括的なコンピュータサイエンス文献目録(ジャーナルおよび会議録)
ERIC: 170万件以上の教育研究出版物
HAL: 440万件以上のフランス/欧州の学術文書(英語250万件)
INSPIRE-HEP: 170万件以上の高エネルギー物理学出版物
合計カバレッジ
6億件以上の出版物を全学術分野でカバーし、STEM、コンピュータサイエンス、生物医学、数学、教育研究に特に強みを持っています。
ライセンス
MIT
貢献
貢献を歓迎します!Issueやプルリクエストを自由にお送りください。
関連情報
APIドキュメント
リソース
Available Tools
3 toolsbatchVerifyCitationsARead-onlyIdempotent
Verify multiple citations in a single call. More efficient than calling verifyCitation multiple times. Returns verification status for each citation.
| Name | Required | Description | Default |
|---|---|---|---|
| citations | Yes | Array of citations to verify |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description adds context beyond annotations by specifying that it 'Returns verification status for each citation,' which clarifies the output behavior. Annotations already indicate it's read-only, idempotent, and non-destructive, so the description doesn't need to repeat those traits, but it usefully describes the return format.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise and front-loaded, consisting of two sentences that efficiently convey the tool's purpose, efficiency benefit, and return value without any wasted words. Every sentence adds value, making it easy for an agent to parse quickly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's moderate complexity, the description is complete enough: it covers purpose, usage guidelines, and output behavior. With annotations handling safety traits and no output schema, the description fills gaps by explaining the return format. However, it could briefly mention error handling or limits for full completeness.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description mentions 'citations' as the input but doesn't add semantic details beyond what the schema provides. With 100% schema description coverage, the schema fully documents the 'citations' array and its nested properties, so the baseline score of 3 is appropriate as the description doesn't compensate with extra parameter insights.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose with a specific verb ('Verify multiple citations') and resource ('citations'), distinguishing it from sibling tools like 'verifyCitation' by emphasizing batch processing efficiency. It explicitly mentions the return value ('verification status for each citation'), which adds clarity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides explicit guidance on when to use this tool versus alternatives: it states 'More efficient than calling verifyCitation multiple times,' directly comparing it to a sibling tool. This helps the agent choose this tool for batch operations over single-citation verification.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
findVerifiedPapersARead-onlyIdempotent
Search multiple academic databases (CrossRef, OpenAlex, PubMed, zbMATH, ERIC, HAL, INSPIRE-HEP, Semantic Scholar, DBLP) for papers and return only verified, real citations with DOIs.
| Name | Required | Description | Default |
|---|---|---|---|
| query | Yes | Search query (topic, keywords, author names) | |
| limit | No | Number of results per source | |
| yearFrom | No | Minimum publication year | |
| yearTo | No | Maximum publication year | |
| source | No | Which source to search | all |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already indicate read-only, idempotent, and non-destructive behavior. The description adds valuable context beyond this: it specifies the multiple databases searched (CrossRef, OpenAlex, etc.) and the verification requirement (only papers with DOIs are returned). This helps the agent understand the tool's scope and output quality, though it doesn't mention rate limits or authentication needs.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, dense sentence that efficiently conveys the tool's purpose, scope, and key behavior. It lists all databases upfront and specifies the verification requirement without unnecessary words. Every element earns its place, making it highly concise and well-structured.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (searching multiple databases with verification), annotations cover safety (read-only, non-destructive), and schema fully documents parameters, the description provides good contextual completeness. It explains the multi-source approach and DOI verification, though without an output schema, it doesn't detail the return format (e.g., what fields are included).
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, providing full parameter documentation. The description doesn't add any parameter-specific details beyond what's in the schema (e.g., it doesn't explain query syntax or source differences). With high schema coverage, the baseline score of 3 is appropriate as the description doesn't compensate but doesn't need to.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('search multiple academic databases'), the resource ('papers'), and a key distinguishing feature ('return only verified, real citations with DOIs'). It differentiates from siblings by focusing on multi-source search with verification, unlike batchVerifyCitations and verifyCitation which likely handle verification of existing citations rather than searching.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage context by specifying it searches 'multiple academic databases' and returns 'verified, real citations with DOIs', suggesting it's for finding reliable academic sources. However, it doesn't explicitly state when to use this tool versus its siblings (batchVerifyCitations, verifyCitation), which likely handle different verification scenarios.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
verifyCitationARead-onlyIdempotent
CRITICAL: Use this to verify ANY academic citation before mentioning it. Checks multiple databases (CrossRef, OpenAlex, PubMed, zbMATH, ERIC, HAL, INSPIRE-HEP, Semantic Scholar, DBLP) if a paper exists. Returns null if not found.
| Name | Required | Description | Default |
|---|---|---|---|
| title | No | Paper title (partial matches accepted) | |
| authors | No | Author names (last names sufficient) | |
| year | No | Publication year | |
| doi | No | DOI if known | |
| journal | No | Journal name |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description adds valuable behavioral context beyond annotations: it lists the specific databases checked (CrossRef, OpenAlex, etc.) and states that it 'returns null if not found,' which clarifies the output behavior. Annotations already indicate it's read-only, idempotent, and non-destructive, so the description doesn't need to repeat those traits, but it enhances understanding with operational details.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is highly concise and well-structured: it starts with a critical warning, states the purpose and usage in a single sentence, lists databases efficiently, and ends with return behavior. Every sentence adds essential information without redundancy, making it front-loaded and easy to parse.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (verifying citations across multiple databases) and the absence of an output schema, the description is mostly complete: it explains the purpose, usage, databases checked, and return behavior. However, it lacks details on error handling, rate limits, or authentication needs, which could be useful for full transparency. The annotations cover safety aspects, so it's adequate but not exhaustive.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 100% schema description coverage, the input schema fully documents all 5 parameters (title, authors, year, doi, journal), including details like 'partial matches accepted' for title and 'last names sufficient' for authors. The description adds no additional parameter information, so it meets the baseline of 3 by not duplicating schema content.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose with a specific verb ('verify') and resource ('academic citation'), explicitly distinguishes it from siblings by specifying it's for verifying citations before mentioning them (unlike batchVerifyCitations or findVerifiedPapers), and provides critical context about checking multiple databases. The 'CRITICAL' prefix emphasizes its importance.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly states when to use this tool ('before mentioning [a citation]') and provides clear alternatives by naming sibling tools (batchVerifyCitations, findVerifiedPapers), though it doesn't detail when to use those instead. The 'CRITICAL' label implies it should be used for any citation verification, making the guidance comprehensive.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
3 tool updates
v1.0.0- First observed
batchVerifyCitations - First observed
findVerifiedPapers - First observed
verifyCitation
TDQS
Scored across 3 tools
Each tool has a clearly distinct purpose: batchVerifyCitations handles multiple citations efficiently, verifyCitation checks individual citations, and findVerifiedPapers searches databases for verified papers. There is no overlap in functionality, making tool selection straightforward for an agent.
The naming follows a consistent verb_noun pattern (batchVerifyCitations, findVerifiedPapers, verifyCitation), with all tools using camelCase. However, verifyCitation lacks a noun suffix like 'Citation' in its verb part, which is a minor deviation from perfect consistency.
With 3 tools, the count is reasonable for a DOI citation verification server, covering core operations (verify single, verify batch, search verified). It might be slightly thin, as additional tools for managing results or databases could enhance completeness, but it's well-scoped for the basic purpose.
The tool set covers key verification tasks: single and batch verification, plus searching for verified papers. Minor gaps exist, such as tools for updating or deleting verification data, but the core workflow of verifying and finding citations is adequately supported without dead ends.
Maintenance
Related MCP Connectors
AI research grounded in 300M scientific works — every citation a verifiable DOI.
Checks AI-written references against Crossref, PubMed and OpenAlex. Formats citations, PRISMA.
Catch AI-fabricated citations (real DOI + fake title). Retraction, open-access, 10,000+ CSL styles.
Real-time fact-check, citation verification, and source-freshness for AI agents.
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceVerifies citations in reference lists by checking DOIs against public registries to catch AI-hallucinated or mismatched citations.MIT

CiteStamp MCP serverofficial
AlicenseNot gradedqualityBmaintenanceGround citations before your agent emits them by checking references against public scholarly registries and flagging hallucinated or retracted ones.MIT- FlicenseNot gradedqualityDmaintenanceFabrication-free, DOI-backed citations for AI content and agents, using openAlex public-domain data with resolvable DOIs. Includes an API and planned MCP server for agent-native citation retrieval.-

SciWeave MCP Serverofficial
FlicenseNot gradedqualityCmaintenanceGround Claude, ChatGPT, Cursor, and Windsurf in 300M scientific works — every citation a verifiable DOI.-