Skip to main content
Glama
michalhron

Scopus Plus MCP

by michalhron

rpys

Count cited references by year, subtract the five-year median, and identify peak years with the most-cited works to reveal a paper set's historical roots; writes CSV and PNG.

Instructions

Reference Publication Year Spectroscopy (Marx et al. 2014): counts every cited reference of a set of papers by the year the cited work appeared, subtracts the five-year median, and reports the peak years with the works cited most from each: the set's historical roots. Give ids or a query (e.g. the confirmed citers from resolve_citers). Writes a CSV of the spectrogram and a PNG. Cost: one reference request per paper (cached; shared with citation_network). With source='openalex' the cited works are fetched in batches of 50.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
idsNoThe papers (Scopus IDs/EIDs; with source='openalex', DOIs or OpenAlex IDs). Use this or query.
queryNoSearch query defining the set, instead of ids.
scopeNoRestrict to journals, by ISSN: a list of ISSNs, or a basket name ('basket_of_eight' / 'ais8': the AIS Senior Scholars' Basket of Eight). ISSNs are used rather than journal names, which Scopus spells inconsistently.
sourceNoData source. 'scopus' (default) needs subscriber entitlement for search, citations and references. 'openalex' needs none: IDs may be DOIs, OpenAlex work IDs (W...), or Scopus IDs (resolved to a DOI via Scopus metadata), and results carry OpenAlex IDs. Never mix sources within one analysis.scopus
to_yearNoLatest cited year (default: this year).
from_yearNoEarliest cited year to count (default 1900).
top_peaksNoPeaks to report (default 10).
corpus_fileNoA corpus file from import_records, instead of ids or query.
max_resultsNoWith query: how many papers to include (default 300).

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.1.0

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does well: it discloses side effects (writes a CSV and PNG), the cost model (one reference request per paper), caching behavior, cache sharing with citation_network, and OpenAlex batching of 50. It omits permission/error behavior, but the resource and mutation profile is unusually well covered.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is a single dense paragraph, but it is front-loaded with the method and purpose before cost/side-effect details, and every sentence carries information. Slightly run-on, but no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 9-parameter analytical tool with no output schema and no annotations, the description covers method, inputs, outputs (CSV/PNG), cost, and source handling adequately. It could say more about failure modes or what the returned job/report object contains, but nothing essential to calling it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so baseline is 3; the description still adds value by clarifying what ids should be (a set of papers, e.g. confirmed citers) and by noting the OpenAlex batch-of-50 fetch behavior tied to source. Most other parameter meaning is already in the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific technique (Reference Publication Year Spectroscopy), states the exact algorithm (count cited references by year, subtract five-year median, report peak years with top cited works), and frames the output as the set's historical roots. This is clearly distinguishable from siblings like citation_lineage or bibliographic_coupling.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It tells the agent what to feed it (ids, a query, or a corpus) and gives a concrete sourcing example ('the confirmed citers from resolve_citers'). However, it never explicitly states when to prefer rpys over alternatives such as historiograph or citation_network, so routing is left partly to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.