bibtex-mcp
Fetches multiple BibTeX citations from Google Scholar for a given paper query, ranking archival records over preprints.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@bibtex-mcpget bibtex for Attention Is All You Need"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
bibtex-mcp
CLI and MCP server for fetching multiple BibTeX options from Google Scholar's visible citation export flow.
What Was Reverse Engineered
Google Scholar search results include a per-result Scholar id in the result container. The visible Cite button loads:
https://scholar.google.com/scholar?q=info:{scholarId}:scholar.google.com/&output=cite&scirp={rank}&hl=enThat citation dialog HTML contains a signed scholar.bib link on scholar.googleusercontent.com. This package follows that same two-step flow:
Search Scholar for the paper title.
Read the result ids from the Scholar result HTML.
Follow
All versionslinks for preprint/unknown rows so archival versions hidden behind an arXiv result can be found.Fetch the matching citation dialog HTML.
Extract and fetch Scholar's BibTeX link.
Rank options so archival conference/journal/publisher records come before preprints when Scholar exposes both.
The tool does not bypass CAPTCHA, paywalls, rate limits, or access controls.
Related MCP server: scholar-toolkit-mcp
Install
npm install
npm run buildCLI
npm run cli -- "Attention Is All You Need"
npm run cli -- --exact --json "Attention Is All You Need"
npm run cli -- --limit 20 --json "Attention Is All You Need"After a global install or npm link:
scholar-bibtex "Attention Is All You Need"MCP
Build first, then configure your MCP client to run:
node /absolute/path/to/bibtex-mcp/dist/mcp.jsInstall With add-mcp
For a local checkout:
npm install
npm run build
npx add-mcp "node $(pwd)/dist/mcp.js" --name scholar-bibtex -a codex -yFor a published npm package:
npx add-mcp bibtex-mcp --name scholar-bibtex -a codex -yEquivalently, use the explicit stdio command form:
npx add-mcp "npx -y bibtex-mcp" --name scholar-bibtex -a codex -yYou can also install directly from the public GitHub repo:
npx add-mcp "npx -y github:aryankeluskar/bibtex-mcp" --name scholar-bibtex -a codex -yChange -a codex to another supported agent such as claude-code, cursor, vscode, or opencode.
The server exposes one tool:
google_scholar_bibtexInput:
{
"query": "Attention Is All You Need",
"maxResults": 10,
"exactTitle": false
}Agent Skill
This repo includes an installable skill for agents that use the skills CLI:
npx skills add aryankeluskar/bibtex-mcp --skill google-scholar-bibtex-mcp --full-depthFrom a local checkout, list or install it with:
npx skills add . --list --full-depth
npx skills add . --skill google-scholar-bibtex-mcp --full-depth -yThe skill tells agents how to install this MCP with add-mcp, call google_scholar_bibtex, and select archival records over preprints.
Stress Test
Run the deterministic stress test without hitting Google Scholar:
npm run stressThe stress test uses fixture mode to validate the CLI, stdio MCP concurrency, archival ranking, and add-mcp Codex project config generation.
Accuracy Notes
The BibTeX is sourced from Google Scholar's citation export, which is better than generated citations but still not infallible. Scholar can contain duplicate records, incomplete metadata, or venue variants.
By default the tool returns multiple options and ranks likely archival records above preprints. It also follows Scholar All versions clusters for non-archival rows so the conference or journal version can outrank a top-level preprint result. The ranking uses signals from the Scholar row, such as venue text and publisher/proceedings domains. For camera-ready bibliographies, pick the highest-ranked archival option whose title and venue match the paper PDF, proceedings page, DOI/Crossref, or publisher page.
Available Tools
1 toolgoogle_scholar_bibtexGoogle Scholar BibTeXA
Search Google Scholar for a paper title and return multiple BibTeX options from Scholar's citation export flow, ranked to prefer archival conference/journal records over preprints. Does not bypass CAPTCHA or access controls.
| Name | Required | Description | Default |
|---|---|---|---|
| query | Yes | Paper title or Scholar search query. | |
| exactTitle | No | Prefer exact title matches before archival ranking. | |
| maxResults | No | Number of Scholar options to return. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Without annotations, the description carries full burden. It discloses the ranking logic and the limitation about CAPTCHA/access controls, but could mention side effects (none) or rate limits. Still, it provides substantial behavioral context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two concise sentences, front-loaded with main function and ranking, then limitation. No wasted words, earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 3-parameter tool with no output schema, the description is adequate but incomplete. It lacks specifics on output format (e.g., list of BibTeX strings) and error handling (invalid queries, empty results).
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3. The description adds context about ranking that complements the 'exactTitle' and 'maxResults' parameters, enhancing understanding beyond individual schema descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb 'Search', the resource 'Google Scholar', and the output ('multiple BibTeX options'). It specifies ranking preference and limitations, distinguishing it from generic search tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No sibling tools are provided, so differentiation is not required. The description implies usage when a BibTeX citation from Google Scholar is needed, but lacks explicit guidance on when not to use or prerequisites.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
1 tool update
v0.1.0- First observed
google_scholar_bibtex
TDQS
Scored across 1 tool
With only one tool, there is no possible ambiguity between tools. The single tool serves a distinct purpose.
The single tool name 'google_scholar_bibtex' follows a clear pattern (source_action) and is self-explanatory, making it consistent by default.
The server has a narrow purpose of fetching BibTeX from Google Scholar, and a single tool is slightly minimal but appropriate for its focused scope.
The tool covers the core functionality implied by the server name—searching Google Scholar and returning BibTeX citations. There are no obvious gaps for its intended use.
Maintenance
Related MCP Connectors
Federated search of books and papers, BibTeX/RIS citations, open-access retrieval and reading.
ArXiv preprints + Google Scholar papers, with citation counts in one query.
Crossref MCP — wraps the Crossref REST API (academic papers, free, no auth)
MCP server for Altmetric APIs - track research attention across news, policy, social media, and more
Related MCP Servers
- AlicenseBqualityDmaintenanceRetrieve citation data effortlessly from CiteAs and Google Scholar. Get BibTeX-formatted citations for your resources with just a few commands. Enhance your research workflow by integrating citation retrieval directly into your applications.214MIT
- AlicenseAqualityAmaintenanceComprehensive MCP server for academic research workflows, enabling paper searching across multiple sources, manuscript processing with citation placeholders, search caching, and citation export.11MIT
- AlicenseBqualityDmaintenanceAutomatically enriches BibTeX citations with missing metadata (DOI, venue, etc.) from arXiv, DBLP, and CrossRef via an MCP server for AI assistants.2MIT
- AlicenseNot gradedqualityDmaintenanceEnables AI agents to access academic citations, fetch metadata, and generate BibTeX entries directly from Google Scholar via MCP tools.1MIT