Skip to main content
Glama

omniseek_paper_enrich

Retrieve open-access PDF, retraction status, and citation count for a paper using DOI or arXiv ID. Verify integrity and full-text access before citing.

Instructions

Use WHEN you need ONE paper's open-access full-text PDF, retraction / integrity status, or citation count — signals omniseek_search / field_skeleton do NOT give cleanly. Keyless, mechanical: YOU decide when + on which papers.

Fully-qualified MCP name: mcp__omniseek__omniseek_paper_enrich (server name is omniseek; there is no omniseek-eye server).

Pass DOIs and/or arXiv ids (e.g. "2306.08543", "10.1145/3292500.3330701"; use a node's doi from omniseek_field_skeleton, or metadata.paper_id/metadata.doi from an openalex omniseek_search result — NOT its source_id, the OpenAlex W-id, which is not a DOI/arXiv id). Enrich only the handful you care about, not a whole map. For each id: • is_oa / pdf_url — the open-access full text (arXiv always OA; real DOIs via Unpaywall). Feed pdf_url to omniseek_read (or read it yourself) to get the WHOLE paper, not just the abstract — then YOU synthesize. (This thin PDF primitive is why we did NOT add a synthesis engine.) For FIGURES / architecture diagrams / result plots: download the PDF and Read its pages with your own VISION — they render in context with captions, so no figure-extraction channel is needed. • integrity.retracted + integrity.notices (retraction / expression_of_concern / correction / …) from Crossref's Retraction Watch feed — check before trusting a high-stakes citation. (retracted=None means "not checked" / backend unreachable; notices=[] means clean. arXiv ids are checked too: an author withdrawal marker plus the journal DOI, when present, run through the same Crossref retraction path.) • citation_count — this paper's citation count (DOI: Crossref is-referenced-by-count; arXiv: S2 citationCount). The single-paper count's home, so you need NOT repurpose omniseek_field_skeleton to read one node's count. (None when the backend was unreachable.)

Returns: {"results": [{id, kind, doi, is_oa, pdf_url, oa_url, citation_count, integrity:{retracted, notices}}, ...]} (or {id, error} for an unrecognized id).

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
idsYes

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.2.0

TDQS

A5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description fully discloses key behaviors: keyless mechanical operation, no synthesis engine, OA semantics per ID type, retracted=None meaning 'not checked' versus notices=[] meaning clean, citation count sources, backend-unreachable behavior, and a clear result/error return shape. This is highly transparent beyond what structured fields provide.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long but densely informative, front-loaded with the core purpose and organized with labeled bullets for each returned signal. Every sentence carries operational or decision-making value, and the return format is clearly laid out.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Despite no output schema, the description defines the exact result structure, error cases, edge-case semantics for retraction status, and downstream usage. For a single-parameter enrichment tool with no annotations, this is complete enough for an agent to select and invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% and there is only one parameter, but the description thoroughly defines it: pass DOIs and/or arXiv ids, gives concrete examples, maps ID sources to specific fields in sibling tool outputs, and explicitly warns against using an OpenAlex W-id. This fully compensates for the bare schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific use case: retrieving ONE paper's open-access full-text PDF, retraction/integrity status, or citation count. It explicitly contrasts these signals with those provided by omniseek_search and field_skeleton, making the tool's scope and distinction from siblings clear.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives explicit when-to-use guidance ('Use WHEN you need ONE paper's...'), warns against enriching a whole map, and explains which IDs to pass versus which to avoid (not source_id/W-id). It also names sibling signals that do not cleanly provide this data and tells the agent to feed pdf_url to omniseek_read.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.