Skip to main content
Glama

bibtex-mcp

CLI and MCP server for fetching multiple BibTeX options from Google Scholar's visible citation export flow.

What Was Reverse Engineered

Google Scholar search results include a per-result Scholar id in the result container. The visible Cite button loads:

https://scholar.google.com/scholar?q=info:{scholarId}:scholar.google.com/&output=cite&scirp={rank}&hl=en

That citation dialog HTML contains a signed scholar.bib link on scholar.googleusercontent.com. This package follows that same two-step flow:

  1. Search Scholar for the paper title.

  2. Read the result ids from the Scholar result HTML.

  3. Follow All versions links for preprint/unknown rows so archival versions hidden behind an arXiv result can be found.

  4. Fetch the matching citation dialog HTML.

  5. Extract and fetch Scholar's BibTeX link.

  6. Rank options so archival conference/journal/publisher records come before preprints when Scholar exposes both.

The tool does not bypass CAPTCHA, paywalls, rate limits, or access controls.

Related MCP server: scholar-toolkit-mcp

Install

npm install
npm run build

CLI

npm run cli -- "Attention Is All You Need"
npm run cli -- --exact --json "Attention Is All You Need"
npm run cli -- --limit 20 --json "Attention Is All You Need"

After a global install or npm link:

scholar-bibtex "Attention Is All You Need"

MCP

Build first, then configure your MCP client to run:

node /absolute/path/to/bibtex-mcp/dist/mcp.js

Install With add-mcp

For a local checkout:

npm install
npm run build
npx add-mcp "node $(pwd)/dist/mcp.js" --name scholar-bibtex -a codex -y

For a published npm package:

npx add-mcp bibtex-mcp --name scholar-bibtex -a codex -y

Equivalently, use the explicit stdio command form:

npx add-mcp "npx -y bibtex-mcp" --name scholar-bibtex -a codex -y

You can also install directly from the public GitHub repo:

npx add-mcp "npx -y github:aryankeluskar/bibtex-mcp" --name scholar-bibtex -a codex -y

Change -a codex to another supported agent such as claude-code, cursor, vscode, or opencode.

The server exposes one tool:

google_scholar_bibtex

Input:

{
  "query": "Attention Is All You Need",
  "maxResults": 10,
  "exactTitle": false
}

Agent Skill

This repo includes an installable skill for agents that use the skills CLI:

npx skills add aryankeluskar/bibtex-mcp --skill google-scholar-bibtex-mcp --full-depth

From a local checkout, list or install it with:

npx skills add . --list --full-depth
npx skills add . --skill google-scholar-bibtex-mcp --full-depth -y

The skill tells agents how to install this MCP with add-mcp, call google_scholar_bibtex, and select archival records over preprints.

Stress Test

Run the deterministic stress test without hitting Google Scholar:

npm run stress

The stress test uses fixture mode to validate the CLI, stdio MCP concurrency, archival ranking, and add-mcp Codex project config generation.

Accuracy Notes

The BibTeX is sourced from Google Scholar's citation export, which is better than generated citations but still not infallible. Scholar can contain duplicate records, incomplete metadata, or venue variants.

By default the tool returns multiple options and ranks likely archival records above preprints. It also follows Scholar All versions clusters for non-archival rows so the conference or journal version can outrank a top-level preprint result. The ranking uses signals from the Scholar row, such as venue text and publisher/proceedings domains. For camera-ready bibliographies, pick the highest-ranked archival option whose title and venue match the paper PDF, proceedings page, DOI/Crossref, or publisher page.

Available Tools

1 tool
google_scholar_bibtexGoogle Scholar BibTeXA

Search Google Scholar for a paper title and return multiple BibTeX options from Scholar's citation export flow, ranked to prefer archival conference/journal records over preprints. Does not bypass CAPTCHA or access controls.

ParametersJSON Schema
NameRequiredDescriptionDefault
queryYesPaper title or Scholar search query.
exactTitleNoPrefer exact title matches before archival ranking.
maxResultsNoNumber of Scholar options to return.

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Without annotations, the description carries full burden. It discloses the ranking logic and the limitation about CAPTCHA/access controls, but could mention side effects (none) or rate limits. Still, it provides substantial behavioral context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two concise sentences, front-loaded with main function and ranking, then limitation. No wasted words, earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 3-parameter tool with no output schema, the description is adequate but incomplete. It lacks specifics on output format (e.g., list of BibTeX strings) and error handling (invalid queries, empty results).

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so baseline is 3. The description adds context about ranking that complements the 'exactTitle' and 'maxResults' parameters, enhancing understanding beyond individual schema descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb 'Search', the resource 'Google Scholar', and the output ('multiple BibTeX options'). It specifies ranking preference and limitations, distinguishing it from generic search tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No sibling tools are provided, so differentiation is not required. The description implies usage when a BibTeX citation from Google Scholar is needed, but lacks explicit guidance on when not to use or prerequisites.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 1 tool updatev0.1.0
    • First observedgoogle_scholar_bibtex

TDQS

A4.3/5.0

Scored across 1 tool

Disambiguation5/5

With only one tool, there is no possible ambiguity between tools. The single tool serves a distinct purpose.

Naming Consistency5/5

The single tool name 'google_scholar_bibtex' follows a clear pattern (source_action) and is self-explanatory, making it consistent by default.

Tool Count4/5

The server has a narrow purpose of fetching BibTeX from Google Scholar, and a single tool is slightly minimal but appropriate for its focused scope.

Completeness5/5

The tool covers the core functionality implied by the server name—searching Google Scholar and returning BibTeX citations. There are no obvious gaps for its intended use.

Maintenance

ActivityStale
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    B
    quality
    D
    maintenance
    Retrieve citation data effortlessly from CiteAs and Google Scholar. Get BibTeX-formatted citations for your resources with just a few commands. Enhance your research workflow by integrating citation retrieval directly into your applications.
    2
    14
    MIT
  • A
    license
    A
    quality
    A
    maintenance
    Comprehensive MCP server for academic research workflows, enabling paper searching across multiple sources, manuscript processing with citation placeholders, search caching, and citation export.
    11
    MIT
  • A
    license
    B
    quality
    D
    maintenance
    Automatically enriches BibTeX citations with missing metadata (DOI, venue, etc.) from arXiv, DBLP, and CrossRef via an MCP server for AI assistants.
    2
    MIT