starbuck-mcp
Resolves and checks arXiv identifiers and preprint records, verifying citation metadata and detecting whether a cited preprint has a published version.
Resolves DOI identifiers via doi.org and checks them against Crossref/DataCite records for existence, metadata agreement, retractions, corrections, and expressions of concern.
Resolves and checks PubMed PMIDs and records, verifying citation metadata, retraction or concern status, and retrieving abstracts for claim-support checks.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@starbuck-mcpcheck the references in paper.qmd for retractions"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Starbuck
Reference-integrity checks for manuscripts. Starbuck reads a Markdown or Quarto manuscript with its bibliography and checks each cited reference against public scholarly records. It writes a Quarto report that renders to HTML or Word.
"I will have no man in my boat," said Starbuck, "who is not afraid of a whale." Herman Melville, Moby-Dick, chapter 26
What it checks
Level | Question | Sources |
1. Exists | Does the DOI, PMID, arXiv ID or ISBN resolve? Without an identifier, can the work be found? | Crossref, DataCite, doi.org, PubMed, arXiv, Open Library, Crossref search |
2. Matches | Do the cited title, first author and year agree with the record? | The record from level 1 |
3. Stands | Is the work retracted, under an expression of concern or corrected? Is it a notice itself? Was a cited preprint published? | Crossref (with Retraction Watch data), PubMed, OpenAlex, arXiv |
4. Supports (optional) | Does the source support the sentence that cites it? | Europe PMC open-access full text or the abstract (PubMed, Europe PMC, Crossref, OpenAlex); a language model judges |
Each reference gets one result:
Result | Meaning |
Fail | The identifier does not exist, belongs to a different work, two identifiers disagree, or the work was retracted. |
Check | A person should look: a field differs, an expression of concern, a possible match only, an unreachable web page. |
Note | Information that needs no change in most cases: a correction, a published version of a preprint, a DOI found by search. |
Pass | Nothing found. |
Not checked | A service did not answer. Run the check again later. |
The report also lists citations without a bibliography entry, entries that
the text does not cite, and sentences that state a finding without a citation.
A sentence is listed when it has at least 50 characters, cites nothing, and
has a number with a unit or a percentage, an effect measure (OR, HR, RR, CI)
or a claim word such as found, showed, increased, reduced, associated with,
risk or prevalence. Headings, tables, figures, captions, code, the reference
list and sections headed Methods or Results are left out: the authors' own
methods and data need no citation. The rules are simple and English only, so
the list is for a person to read; it never makes a reference Fail or Check.
The JSON keeps the first 50 sentences under uncited_sentences (the key
uncited lists the bibliography entries that are not cited).
Level 4 adds a verdict for each citing sentence: Supported, Partly supported, Not supported or Cannot assess. Level 4 never makes a reference Fail: Not supported is Check and Partly supported is Note, so a person decides. See Level 4.
Related MCP server: CiteStamp MCP server
Install
Starbuck needs Python 3.11 or later and uv. For HTML and Word reports, install Quarto. If Quarto is not installed, Starbuck uses Pandoc when it is available.
Install Starbuck:
uv tool install starbuckSet a contact address. The services give faster, more reliable access to requests that identify their sender:
export STARBUCK_EMAIL=you@example.orgCheck the installation:
starbuck --version
Use
Check a manuscript. The bibliography comes from the bibliography: field in the
YAML header:
starbuck check paper.qmdStarbuck writes the report to a folder named _starbuck next to the manuscript.
Quarto ignores folders that start with an underscore, so the report does not
become part of a Quarto project.
More options:
starbuck check paper.qmd --to html,docx # HTML and Word
starbuck check paper.md --bib refs.bib # bibliography not in the YAML header
starbuck check paper.qmd --all-entries # also check entries that are not cited
starbuck check paper.qmd --out reports/ # another report folder
starbuck ids 10.1183/09031936.00080312 arXiv:1706.03762
starbuck check paper.qmd --claims # also level 4 (see below)To check a Word manuscript, convert it to Markdown first:
quarto pandoc paper.docx -o paper.mdInput
Starbuck reads two citation styles:
Citation keys: Pandoc citations such as
[@key],[see @key, p. 3; @other]and@keyin running text. The bibliography can be BibTeX or BibLaTeX (.bib), CSL JSON (.json) or CSL YAML (.yaml), orreferences:in the YAML header.Numbered references:
[1],[2, 5],[3-4]in the text, with a numbered list under a heading such as References or Sources. AI research agents often write this format.
Output
File | Contents |
| The report. Edit it or render it again with Quarto. |
| The report as one self-contained web page. |
| The report for Word, with |
| Every result, finding and record, for other tools. |
| Level 4 only: passages, source texts and the model's answers. |
The report starts with a summary, then the references that need attention, each with the reason, the cited and the recorded metadata side by side, the sentence that cites it, and a suggested action. A section at the end states the rules and the services used.
Exit codes
Code | Meaning |
0 | No reference failed. |
1 | At least one reference failed, or a cited key has no entry. |
2 | The check could not run (for example, no citations found). |
3 | No reference failed, but some checks did not run. |
Level 4: claim support
For each sentence that cites a source, Starbuck takes the best text of that source: a local text file you give, open-access full text from Europe PMC, or otherwise the abstract. It ranks the passages of that text against the sentence (BM25) and gives the best five to a language model. The model answers Supported, Partly supported, Not supported or Cannot assess, and quotes the passage it relies on. Starbuck then checks that the quotation appears word for word in the source text. A quotation that is not there is rejected, and the verdict becomes Cannot assess. The report states for each verdict whether it rests on the full text or on the abstract only.
The report also gives two scores over the judged citations (Cannot assess is left out). Citation recall: the share of citing sentences with at least one citation judged Supported or Partly supported. Citation precision: the share of citations judged Supported or Partly supported. A citation judged Not supported in a sentence that another of its sources supports gets a Note (overcitation): it adds no support. The idea comes from ScholarQABench (OpenScholar).
Set the model. Any OpenAI-compatible endpoint works (a hosted provider, or a local server such as Ollama):
export STARBUCK_JUDGE_URL=https://api.example.org/v1 export STARBUCK_JUDGE_MODEL=model-name export STARBUCK_JUDGE_KEY=... # if the endpoint needs a keyRun the check with
--claims:starbuck check paper.qmd --claims starbuck check paper.qmd --claims --text smith2020=smith2020.txt # full text you have
Without --claims, or without the model settings, level 4 does not run and the
report says so. Level 4 sends the citing sentences and the source passages to
the model's provider. The file <name>-claims.json in _starbuck keeps the
passages, the source texts and the model's answers.
A language model can be wrong. Read the source before you change the text.
Settings
Variable | Purpose |
| Contact address for the services. |
| Higher request rate for PubMed. |
| OpenAlex API key. |
| Path to Quarto when it is not on the PATH. |
| Level 4 model (OpenAI-compatible endpoint). |
Use from an agent (MCP)
starbuck-mcp is an MCP server with four tools:
check_manuscript: checks a manuscript and writes the report.check_references: checks references given directly (an identifier, citation details, or both). An agent can use it to check the sources of a text it wrote.prepare_claims: level 4, step 1. Checks the references and returns each citing sentence with the best passages of its source. A local full text can be given per citation key.record_claims: level 4, step 2. Takes the agent's verdicts, checks the quoted passages and adds level 4 to the report.
Inside an agent, the agent's own model judges, so no STARBUCK_JUDGE_*
settings are needed.
In Sub-Sub, turn on reference checks in
subsub init or with the switch in the web view; Sub-Sub then runs Starbuck as
its verify server.
Example configuration for an MCP client:
{
"mcpServers": {
"starbuck": {
"command": "uvx",
"args": ["--from", "starbuck", "starbuck-mcp"],
"env": { "STARBUCK_EMAIL": "you@example.org" }
}
}
}Privacy
Starbuck sends identifiers and reference details (title, authors, year, or the
reference as written) to the services listed above, and it opens cited web
pages. Levels 1 to 3 do not send the manuscript text. Level 4 (only with
--claims) sends the citing sentences and source passages to the model you
set. Starbuck collects no usage data.
Demo
examples/demo.qmd cites real works, each set up to show one kind of result:
a correct reference, a wrong year, a DOI that belongs to another paper, a DOI
that does not exist, a retracted paper, a reference without a DOI, an arXiv
preprint and a citation without an entry.
starbuck check examples/demo.qmd --to html,docxDevelopment
uv sync
uv run pytest -qThe tests use canned service responses and do not need the network.
License
MIT © Tiago Jacinto
Available Tools
4 toolscheck_manuscriptAIdempotent
Check every cited reference of a Markdown or Quarto manuscript and write a report. Reads [@citekey] citations with the bibliography named in the YAML header (or given here), or numbered [1] citations with a numbered list under a References/Sources heading. Writes -references.qmd and .json to report_dir (default: _starbuck next to the manuscript) and renders formats (default ["html"]; "docx" for Word). Also returns the count and the first 10 sentences that state a finding without a citation.
| Name | Required | Description | Default |
|---|---|---|---|
| path | Yes | ||
| formats | No | ||
| report_dir | No | ||
| bibliography | No | ||
| include_uncited | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Despite sparse annotations, the description discloses concrete side effects: it writes '<name>-references.qmd' and '.json' into report_dir (default '_starbuck' next to the manuscript) and renders formats (default html, docx for Word). This matches and enriches the non-readOnly/openWorld annotation profile. It stops short of stating overwrite semantics or whether rendering failures are fatal.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action and effect, followed by input syntax and then output/return details. Sentences are dense and each carries information, though the second sentence packs two citation dialects together in a way that reads slightly run-on.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a five-parameter, no-output-schema tool the description covers input dialects, output artifacts and their location, rendering formats, and even the return payload (count plus first 10 uncited-finding sentences). An agent has enough to invoke it correctly without inspecting the schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description must carry the load, and it does for four of five parameters: path (implicit), formats (default ["html"], "docx" for Word), report_dir (default '_starbuck next to the manuscript'), and bibliography (from YAML header or supplied here). Only 'include_uncited' is left undocumented, which is the one gap keeping this from a 5.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource pair ('Check every cited reference of a Markdown or Quarto manuscript and write a report') and details the two supported citation syntaxes, so an agent knows exactly what the tool consumes and produces. It does not, however, distinguish itself from the close sibling 'check_references', leaving the agent to guess which to pick.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implicitly scopes usage by enumerating supported inputs (Markdown/Quarto, [@citekey] with YAML bibliography, or numbered [1] citations with a References heading) and overrides (bibliography 'given here'), which helps the agent decide applicability. But it never states when to use this vs check_references, prepare_claims, or record_claims, and gives no exclusions or prerequisites.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
check_referencesARead-only
Check references given directly: an identifier, citation details, or both. With citation details, level 2 compares them with the record; with an identifier alone, only levels 1 and 3 apply. Use this for the sources of a report you wrote.
| Name | Required | Description | Default |
|---|---|---|---|
| references | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish this as a safe, read-only, open-world operation, so the bar is lower. The description adds a real behavioral nuance — the number of comparison levels depends on whether citation details are supplied — but the 'levels' are never defined, so the agent cannot act on that distinction. No return shape or failure behavior is described.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences with the accepted-input forms front-loaded and the usage cue last. The unexplained 'level 2'/'levels 1 and 3' phrasing adds words whose meaning is not resolved, but there is no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only checking tool with no output schema and no annotations beyond read-only/open-world, the definition covers inputs and usage but omits what a check actually returns (pass/fail? per-level verdicts?) and leaves 'level 1/2/3' undefined. An agent can call it, but cannot anticipate the result.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Reported schema coverage is 0% at the top level, so the description must carry weight, and it does tell the agent that a lone identifier and full citation details are both valid inputs. However, it never explains the relationship between 'raw' and the individual fields, nor how 'identifier' formats (DOI/pmid/arXiv/isbn) are consumed.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('check') and resource ('references') and clarifies the accepted inputs (identifier, citation details, or both). It does not name the sibling check_manuscript, though the closing sentence ('sources of a report you wrote') hints at the distinction. The unexplained 'level 1/2/3' jargon keeps it from being fully self-explanatory.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives a concrete triggering condition ('Use this for the sources of a report you wrote') and specifies which comparison applies under each input mode (citation details vs. identifier alone). It stops short of naming the alternative tool or stating when not to use it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
prepare_claimsAIdempotent
Level 4, step 1. Checks the references (levels 1 to 3) and returns, for each citing sentence, the best passages of the cited source (open-access full text, a local text file from texts {citekey: path}, or the abstract). Judge each claim from its passages, then call record_claims. Results come in pages: use offset and limit; total gives the number of claims. Writes -references.json and -claims.json to report_dir (default _starbuck).
| Name | Required | Description | Default |
|---|---|---|---|
| path | Yes | ||
| limit | No | ||
| texts | No | ||
| offset | No | ||
| report_dir | No | ||
| bibliography | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Goes beyond annotations by disclosing concrete side effects: it writes <name>-references.json and <name>-claims.json to report_dir (default _starbuck), and that results are paged via offset/limit with a total count. readOnlyHint=false and idempotentHint=true are consistent with this file-writing, paginated read. It does not cover auth requirements or whether existing report files are overwritten.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four sentences that are dense with useful workflow, paging, and output-file detail, and the purpose leads. However, 'Level 4, step 1' and 'levels 1 to 3' consume space with pipeline jargon an agent may not be able to resolve, diluting the front-loaded clarity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 6-parameter, file-writing, paginated tool with no output schema and 0% schema coverage, the description supplies workflow, return shape, paging, and artifact names — but the required 'path' and 'bibliography' inputs remain unexplained, leaving a gap the agent cannot fill elsewhere.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across 6 parameters, so the description must carry the load — and it only explains texts ({citekey: path}), offset/limit, report_dir, and the default. The required 'path' parameter and 'bibliography' are never described, leaving the most important input and one other undocumented in both places.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific action (check the references at levels 1–3) and a concrete return (best passages for each citing sentence), and names the follow-up sibling record_claims. The 'Level 4, step 1' framing is opaque without pipeline context, but the core verb+resource is discernible and clearly separated from record_claims.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly sequences the workflow: 'Judge each claim from its passages, then call record_claims', which tells the agent what to do with the output and which tool comes next. It stops short of stating when NOT to use this tool or what alternative exists if the references are not already checked.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
record_claimsAIdempotent
Level 4, step 2. Checks each quoted passage against the source text (a quote that is not there turns the verdict into cannot_assess), adds the verdicts to the report and renders it. Give the verdicts of all claims from prepare_claims in one call.
| Name | Required | Description | Default |
|---|---|---|---|
| path | Yes | ||
| formats | No | ||
| verdicts | Yes | ||
| judged_by | No | the agent's model | |
| report_dir | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare the safety profile (non-destructive, idempotent, open-world, mutating), and the description adds value beyond them by disclosing a specific business rule: a quote that is not found in the source downgrades the verdict to cannot_assess. It also reveals the side effect of appending verdicts to the report and rendering it, which is consistent with readOnlyHint=false.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three dense sentences that front-load position and behavior with little waste. The 'Level 4, step 2' prefix is somewhat opaque but functions as an ordering cue, and every remaining clause (quote check, fallback, render, call guidance) earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
A mutating, file-writing tool with 5 parameters, no output schema, and zero schema descriptions needs more from the description than it provides, particularly on path/formats/report_dir. The workflow and fallback rule are covered, but the input surface is left largely unexplained.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 5 parameters and 0% top-level schema description coverage, the description must explain path, formats, judged_by and report_dir, but it only gestures at 'verdicts' and the source text. The agent is left guessing what path points to (source? report?), what formats does, and the role of judged_by.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States specific actions (checks each quote against the source, records verdicts, renders the report) on a clear resource (claims), and names the upstream sibling prepare_claims. It is distinguishable from check_manuscript and check_references, though the cryptic 'Level 4, step 2' opener only carries meaning if the agent already knows the pipeline.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit call directive: 'Give the verdicts of all claims from prepare_claims in one call,' which tells the agent the prerequisite step and that it should be a single batched invocation. It does not explicitly contrast with the other siblings (check_manuscript, check_references), so there is no when-not guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
4 tool updates
v0.2.0- First observed
check_manuscript - First observed
check_references - First observed
prepare_claims - First observed
record_claims
TDQS
Scored across 4 tools
check_manuscript and check_references both check references, but their inputs differ clearly (manuscript vs. direct reference details), and the descriptions distinguish them. prepare_claims and record_claims are explicitly sequential steps, leaving little ambiguity.
All four tools use a consistent snake_case verb_noun pattern: check_manuscript, check_references, prepare_claims, record_claims. No mixed conventions or vague verbs.
Four tools is well-scoped for a specialized citation- and claim-checking workflow. Each tool maps to a distinct stage or input mode, and none feels redundant.
The surface covers manuscript checking, direct reference checking, claim preparation, and claim recording, which is strong lifecycle coverage for this domain. Minor gaps exist around report retrieval or configuration management, but agents can work around them.
Maintenance
Related MCP Connectors
Checks AI-written references against Crossref, PubMed and OpenAlex. Formats citations, PRISMA.
Real-time fact-check, citation verification, and source-freshness for AI agents.
Cited, versioned knowledge for agents: retrieve sourced passages and propose owner-approved fixes.
Verifies legal citations vs primary sources: existence, quote match, proposition support.
Related MCP Servers
- FlicenseNot gradedqualityDmaintenanceEnables AI agents to fact-check claims, verify citations, and check source freshness using Wikipedia, Wikidata, Crossref, and Wayback Machine.1-

CiteStamp MCP serverofficial
AlicenseNot gradedqualityBmaintenanceGround citations before your agent emits them by checking references against public scholarly registries and flagging hallucinated or retracted ones.MIT- AlicenseNot gradedqualityBmaintenanceEnables agents to verify their own output mid-task by checking every claim against provided sources, returning supported, partial, unsupported, or contradicted verdicts with exact citations.MIT
- AlicenseAqualityBmaintenanceEnables agents to verify whether citations exist and match canonical records, and whether URLs resolve and contain expected content, with evidence-backed confirmed, contradicted, or unknown verdicts.2MIT