Skip to main content
Glama
michalhron

Scopus Plus MCP

by michalhron

citation_network

Build a direct-citation network from a set of papers. Fetch references, keep only internal citations, and compute main paths (SPC weights, local/global paths, key routes).

Instructions

Direct-citation network within a set of papers, in one call: fetches every paper's reference list, keeps only the references to other papers in the set, and runs main-path analysis (SPC weights, local and global main path, key routes). Give ids (Scopus IDs/EIDs; with source='openalex', DOIs or OpenAlex IDs) or a query. Completeness: each list is compared with an independent reference count (Crossref, else OpenAlex, else Semantic Scholar; the source is reported per paper); short = fewer than 90% of the comparison count and at least 5 references missing (SCOPUS_COMPLETENESS_RATIO, SCOPUS_COMPLETENESS_MIN_MISSING). Papers whose references could not be loaded or parsed are listed, retried once after rate limits, and the main path is marked provisional while any are missing. Likely duplicate records are listed. Writes JSON, Pajek .net (arcs from cited to citing, SPC weights; Pajek, VOSviewer, Gephi) and an edge CSV. Cost: one reference request per paper (cached). Sets over about 50 papers may return a job ID: poll job_status, then job_result.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
idsNoThe papers (up to 1000). Use this or query.
pageNoWith inline='full': which page of nodes (1-based).
queryNoSearch query defining the set, instead of ids.
scopeNoRestrict to journals, by ISSN: a list of ISSNs, or a basket name ('basket_of_eight' / 'ais8': the AIS Senior Scholars' Basket of Eight). ISSNs are used rather than journal names, which Scopus spells inconsistently.
inlineNoWhat the reply carries besides the file paths. 'summary': counts, flags and paths. 'edges' (default): also every edge as a compact line (up to 2,000). 'nodes': one compact line per paper (ID, author year, venue, references retrieved/reported, comparison count and source, completeness, error) plus the edges: node-level data for callers that cannot read the server's files, about 25k characters for 150 papers. 'full': the corpus as JSON, paged by page/page_size nodes.edges
sourceNoData source. 'scopus' (default) needs subscriber entitlement for search, citations and references. 'openalex' needs none: IDs may be DOIs, OpenAlex work IDs (W...), or Scopus IDs (resolved to a DOI via Scopus metadata), and results carry OpenAlex IDs. Never mix sources within one analysis.scopus
weightNoTraversal weight for the main paths and key routes: 'spc' (search path count, source-to-sink paths), 'splc' (search path link count: paths starting at any paper), 'spnp' (search path node pair: paths between any two papers). Liu & Lu 2012.spc
page_sizeNoWith inline='full': nodes per page (default 50).
key_routesNoNumber of top-SPC key edges to extend into key-route main paths (Liu & Lu 2012; default 10, 0 = none). Key edges that extend into the same route are merged.
robustnessNoAlso compute the global main path under all three weights and report the papers they share: a path that survives a change of weight is a finding, one that does not is partly an artefact of the weight.
corpus_fileNoA corpus file from import_records, instead of ids or query.
max_resultsNoWith query: how many papers to include (default 300).
edge_contextsNoAlso gather citation contexts for every edge (Semantic Scholar; one request per citing paper, cached), give each a draft transmission label, and write a coding sheet for two coders. Slow without a Semantic Scholar key; runs as a job past the sync budget.
construct_termsNoWith edge_contexts: the construct for the draft labels, e.g. ['organizing vision'].
key_route_searchNoHow key edges are extended: 'local' follows the heaviest adjoining edge; 'global' takes the heaviest whole path to and from the key edge.local
check_retractionsNoFlag retracted, withdrawn or concern-flagged papers (Crossref / Retraction Watch; default true).
max_context_edgesNoWith edge_contexts: most edges to examine, heaviest SPC first (default 300).
check_completenessNoCompare each reference list with an independent count (default true).

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.1.0

TDQS

A4.2/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and discharges it: it discloses completeness thresholds (90%, 5 missing), retry-after-rate-limit behavior, provisional main-path marking, duplicate listing, output formats, one cached request per paper, and job-ID escalation. Exceptionally rich behavioral disclosure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with purpose, then method, then caveats, then outputs, then cost. It is long, but the length is justified by 18 parameters and substantial hidden behavior; each sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a very complex tool with no output schema and no annotations, the description covers compute method, completeness semantics, error states, job escalation, output artifacts, and cost — everything an agent needs to call and interpret it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and the schema descriptions are themselves highly detailed, so the baseline is 3; the description adds only marginal extra meaning (ID formats by source), largely restating what the schema already documents.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource — 'Direct-citation network within a set of papers' — and details the three-step mechanics (fetch references, filter to set, run main-path analysis). This is clearly distinguishable from coupling, co-citation, and lineage siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives operational routing ('give ids ... or a query', 'never mix sources') and the >50-paper job polling path with job_status/job_result, but never states when to prefer this tool over bibliographic_coupling, co_citation, or citation_lineage. Usage context is implied, not contrasted with alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.