Phở Research Tools
Server Details
Checks AI-written references against Crossref, PubMed and OpenAlex. Formats citations, PRISMA.
- Status
- Healthy
- Last Tested
- Transport
- Streamable HTTP · MCP 2025-11-25
- URL
TDQS
Scored across 5 tools
Primary purposes are distinct: searching, verifying, formatting, retraction screening, and PRISMA diagramming. However, three tools (verify_references, check_retractions, format_citations) all touch reference verification, and the descriptions must explicitly note that verify_references 'already includes' the retraction check and that format_citations verifies by default to prevent misselection.
Four of five tools follow a clean verb_noun snake_case pattern (check_retractions, format_citations, search_papers, verify_references). prisma_flow_diagram breaks the pattern as a noun phrase, but it is still readable and clearly snake_case.
Five tools is well-scoped for a research/reference workflow. Each tool covers a distinct, non-redundant step (search, verify, screen retractions, format, diagram) with no filler.
The surface covers a full research-reference lifecycle: discovery, verification, retraction screening, formatting, and PRISMA reporting. Minor gaps exist for systematic-review workflows such as deduplication of search results and title/abstract screening, but core operations are present.
Available Tools
5 toolscheck_retractionsCheck retractionsARead-onlyIdempotentInspect
Check DOIs for retractions, expressions of concern and corrections, using Crossref (which includes Retraction Watch data). Status per DOI: retracted, concern, corrected, other_notice, no_notice (nothing on record), not_in_crossref, invalid_doi, rate_limited or lookup_failed (could not be checked; say so, do not report it as clean). Use it before citing a paper whose status you are unsure of, or to screen a whole reference list. verify_references already includes this check.
| Name | Required | Description | Default |
|---|---|---|---|
| dois | Yes | DOIs, bare (10.1000/xyz) or as https://doi.org/ links. |
Output Schema
| Name | Required | Description |
|---|---|---|
| quota | Yes | |
| counts | Yes | |
| results | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations cover read-only/idempotent/open-world, but the description goes further by enumerating every possible status and, crucially, warning that failure modes (rate_limited, lookup_failed) must not be reported as clean. That is genuine behavioral guidance an agent cannot get from structured fields.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with purpose, then the status vocabulary, then usage guidance; every element is relevant. The status list is somewhat long, but it encodes non-obvious failure handling rather than padding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists so return structure needn't be re-explained, annotations carry the safety profile, and the description supplies source, status semantics, failure handling, and usage context. Nothing needed to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There is only one parameter and schema coverage is 100%, with the DOI format (bare vs https://doi.org/ link) already documented in the schema. The description adds no extra syntax, batching, or ordering semantics beyond what the schema provides, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (check) plus the resource and scope (DOIs for retractions, expressions of concern and corrections) and names the underlying source (Crossref/Retraction Watch). It is clearly distinguishable from siblings such as verify_references and search_papers.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly gives two use cases ('before citing a paper whose status you are unsure of', 'screen a whole reference list') and routes the agent away from redundant use by noting 'verify_references already includes this check.' That is an explicit when-to-use plus an alternative.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
format_citationsFormat citationsARead-onlyIdempotentInspect
Format references in a citation style (APA 7, MLA 9, Chicago author-date, Harvard, Vancouver, IEEE, AMA) or export them as BibTeX, RIS or CSL-JSON. By default every reference is first verified against Crossref, PubMed and OpenAlex, and verified references are formatted from the database record, so wrong years, pages or author names in the input are corrected. References that could not be verified are formatted exactly as written and flagged in entries[].warnings; tell the user about them. Returns the bibliography and, for text and html output, the matching in-text citation for each reference.
| Name | Required | Description | Default |
|---|---|---|---|
| style | No | Citation style: apa (APA 7), mla (MLA 9), chicago (Chicago author-date), harvard (Harvard), vancouver (Vancouver), ieee (IEEE), ama (AMA). Ignored for bibtex/ris/csl-json output. | apa |
| output | No | text (plain bibliography), html, bibtex, ris or csl-json. | text |
| verify | No | Verify each reference before formatting (recommended). Set false only for references already verified in this conversation; they are then formatted exactly as given. | |
| references | Yes | References in any form: a pasted list in any style, BibTeX, RIS, CSL-JSON, or DOIs one per line. |
Output Schema
| Name | Required | Description |
|---|---|---|
| quota | No | |
| style | Yes | |
| output | Yes | |
| content | Yes | |
| entries | Yes | |
| warnings | Yes | |
| style_title | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Goes well beyond the readOnly/idempotent/openWorld annotations: it discloses the Crossref/PubMed/OpenAlex verification pipeline, the correction behavior (wrong years, pages, author names are fixed from the database record), the exact fallback for unverified entries, and the warnings flag location. Auth, side effects and trust implications are all covered.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three front-loaded sentences covering output, default verification behavior, and return values, with no filler. Slightly dense given the style list is duplicated from the enum, but every sentence carries information an agent needs.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, the description need not enumerate return fields, yet it still calls out the bibliography, in-text citations, and `entries[].warnings`. Verification semantics, fallback behavior and the required `references` input are all sufficiently covered for a 4-parameter tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is already 100%, so the baseline is 3; the description adds real meaning by explaining what `verify` actually triggers and by reinforcing that style is irrelevant for bibtex/ris/csl-json output. The `references` parameter is only lightly described here beyond the schema's 'any form' note.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (format) and resource (references) plus the supported styles and export formats, making the tool's scope concrete. It does not explicitly distinguish itself from the sibling verify_references, even though verification is a core part of what it does, so differentiation is left to inference.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear operational context: verification runs by default, `verify: false` is only for references already verified in this conversation, and unverified entries should be surfaced to the user. It stops short of stating when to pick this over verify_references or check_retractions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
prisma_flow_diagramPRISMA 2020 flow diagramARead-onlyIdempotentInspect
Draw a PRISMA 2020 flow diagram for a new systematic review from the record counts, as a PNG image plus an SVG and download links. Give the counts you have; boxes that follow arithmetically from others are filled in and marked as derived. The tool checks that the numbers add up (records identified minus removed equals screened, and so on) and returns a warning naming each box that does not, with the value it should have. Show those warnings to the user; the diagram always shows the numbers exactly as given. Add otherMethods counts to draw the second column (websites, organisations, citation searching).
| Name | Required | Description | Default |
|---|---|---|---|
| counts | Yes | ||
| footnotes | No | Print the template footnotes and source attribution. | |
| other_methods_column | No | Force the other-methods column on or off. Default: drawn when otherMethods counts are given. |
Output Schema
| Name | Required | Description |
|---|---|---|
| svg | Yes | |
| width | Yes | |
| counts | Yes | |
| height | Yes | |
| svg_url | No | |
| warnings | Yes | |
| image_url | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations cover only the safety profile (readOnly, idempotent, closed-world); the description goes well beyond that with real behavioral disclosure: derived boxes are auto-filled and marked, numbers are validated arithmetically with per-box warnings naming the expected value, and the diagram always renders the raw numbers even when inconsistent. It also discloses the returned artifacts (PNG, SVG, download links) and the recommendation to surface warnings.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the action and output, then layers the derivation rule, validation behavior, and the other-methods column in five dense sentences. Every sentence carries operational information; there is no filler or restatement of the name.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with deeply nested count objects and an output schema, the description covers exactly what the schema cannot: how missing values are derived, how inconsistencies are reported, and how the optional other-methods column is activated. Return-value detail is delegated correctly to the output schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 67%, so the schema carries much of the load, but the description adds genuine semantics the schema does not: 'otherMethods' counts trigger the second column (websites, organisations, citation searching), and other_methods_column / footnotes behavior is hinted at. It clarifies the derived-box convention that governs how partial `counts` input is interpreted.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and artifact ('Draw a PRISMA 2020 flow diagram ... as a PNG image plus an SVG and download links') and the input it consumes ('from the record counts'). It is trivially distinguishable from all four siblings (check_retractions, format_citations, search_papers, verify_references), none of which produce a diagram.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear situational context: it is for 'a new systematic review' and tells the agent to 'Give the counts you have' rather than requiring all counts up front. It also states what to do with the output warnings ('Show those warnings to the user'). No alternative tool is named or excluded, but no sibling overlaps this function, so the absence is minor.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_papersSearch papersARead-onlyInspect
Search PubMed and OpenAlex for real papers. Every result carries a DOI or PubMed ID that resolves, so nothing returned can be fabricated. Use it to find sources for a claim or a topic instead of recalling references from memory. Filters: publication year range, open-access only, study types, and the number of results. Retracted papers are left out by default. Preprints are included and marked preprint: true; tell the user a preprint is not peer reviewed. Results are ranked by the databases' relevance, alternating between the two sources.
| Name | Required | Description | Default |
|---|---|---|---|
| query | Yes | What to look for, in English keywords (e.g. "sepsis early antibiotics mortality"). | |
| year_to | No | Latest publication year, inclusive. | |
| year_from | No | Earliest publication year, inclusive. | |
| max_results | No | Number of results, 1–50. | |
| study_types | No | Study designs. Known ids use database filters: rct, clinical-trial, systematic-review, meta-analysis, review, guideline, cohort, case-control, case-report, observational. Any other text is added as keywords. | |
| open_access_only | No | Only papers with a free, legal full text. | |
| include_retracted | No | Keep retracted papers (flagged) instead of leaving them out. |
Output Schema
| Name | Required | Description |
|---|---|---|
| query | Yes | |
| status | Yes | |
| results | Yes | |
| excluded | Yes | |
| warnings | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only cover readOnly/openWorld, but the description adds substantive behavior: retracted papers excluded by default, preprints included and flagged with a user-facing caveat, and results ranked by alternating source relevance. These are exactly the traits an agent needs before quoting output to a user.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The differentiator (lookup vs. recall, verifiable identifiers) is front-loaded, then filters, then caveats, in descending order of importance. No sentence is redundant and none repeats structured field data verbatim.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return-shape explanation is unnecessary; the description still supplies the output-relevant facts an agent must convey (retraction exclusion, preprint flag, relevance ordering). Nothing needed to invoke or report results is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds meaning beyond it: it explains that retracted papers are dropped *by default* (clarifying the include_retracted toggle's intent) and that study types act as filters. The remaining filter mentions (year range, max results) largely restate the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Opens with a specific verb+resource pair ('Search PubMed and OpenAlex for real papers'), naming both underlying databases and the core guarantee that results carry resolvable DOIs/PMIDs. This clearly separates it from siblings like verify_references or check_retractions, which validate rather than discover.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit use case ('find sources for a claim or a topic instead of recalling references from memory'), which is a strong positive routing cue. It stops short of naming when to prefer a sibling (e.g., verify_references for a citation you already have), so it is clear context without exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
verify_referencesVerify referencesARead-onlyIdempotentInspect
Check every reference in a list against Crossref, PubMed and OpenAlex and report, for each one, whether the paper really exists and whether the details are right. Call this before presenting any reference list to the user, including references you wrote yourself. Input: a pasted reference list in any common style (APA, Vancouver, numbered, mixed), BibTeX, RIS, CSL-JSON, or one DOI per line. Each reference gets exactly one status: verified, verified_with_corrections (real paper, some details wrong; corrected metadata included), likely_fabricated (no such paper exists), not_found (could not be confirmed either way), retracted, or ambiguous (several papers fit). Never present a likely_fabricated, not_found or ambiguous reference as real, and always tell the user about a retracted one. At most 50 references per call.
| Name | Required | Description | Default |
|---|---|---|---|
| references | Yes | The reference list exactly as written: any citation style, BibTeX, RIS, CSL-JSON, or DOIs one per line. |
Output Schema
| Name | Required | Description |
|---|---|---|
| quota | Yes | |
| summary | Yes | |
| references | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only cover the read-only/idempotent/open-world profile; the description adds substantial behavioral detail: the full status taxonomy (verified, verified_with_corrections, likely_fabricated, not_found, retracted, ambiguous), what corrections mean, and a concrete throughput limit of 50 references per call. No contradiction with readOnlyHint=true.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with purpose, then trigger, then input formats, then status semantics and limits. Every sentence carries information, though the status enumeration is dense and the passage runs long for a single-parameter tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, return formatting need not be explained, and the description still supplies the status vocabulary, correction behavior, and per-call limit. Nothing an agent needs to invoke or interpret this tool correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% for the single parameter, and the description's input paragraph largely restates it, though it does add style examples (APA, Vancouver, numbered, mixed) and the 50-item cap that the schema does not express. Baseline 3 is appropriate since the schema carries the parameter meaning.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (check/verify) and resource (every reference in a list) plus the external services used (Crossref, PubMed, OpenAlex). An agent can distinguish it from check_retractions (retraction-only) or format_citations (style formatting) without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a clear, explicit trigger — 'call this before presenting any reference list to the user, including references you wrote yourself' — plus a prohibition ('never present a likely_fabricated... reference as real'). It does not name sibling alternatives such as check_retractions or format_citations, so it stops short of full routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
5 tool updates
- First observed
check_retractions - First observed
format_citations - First observed
prisma_flow_diagram - First observed
search_papers - First observed
verify_references
Related MCP Connectors
Catch AI-fabricated citations (real DOI + fake title). Retraction, open-access, 10,000+ CSL styles.
PRISMA 2020 systematic reviews as MCP tools - verified citations, IMRAD papers
Search arXiv/Semantic Scholar/OpenAlex + medical evidence (PubMed/Europe PMC) + LaTeX/PDF tools.
Free mechanical checks for AI text: unnamed counts, dangling references, bad arithmetic, misquotes.
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceVerifies citations in reference lists by checking DOIs against public registries to catch AI-hallucinated or mismatched citations.MIT
- AlicenseAqualityBmaintenanceEnables verifying citations in reference lists and BibTeX by detecting fabricated or mismatched references and retractions, and returning corrected BibTeX.3MIT

CiteStamp MCP serverofficial
AlicenseNot gradedqualityBmaintenanceGround citations before your agent emits them by checking references against public scholarly registries and flagging hallucinated or retracted ones.MIT- AlicenseNot gradedqualityCmaintenancereference and citation validation, verification, enrichment, replacement and improvement1MIT
Glama MCP Gateway
Add one secure layer between your agents and this server.