Skip to main content
Glama
michalhron

Scopus Plus MCP

by michalhron

co_citation

Build a co-citation graph from seed papers to map a field's intellectual base; edge weights count papers citing both seeds, output GraphML and CSV edge lists.

Instructions

Build a co-citation graph for a set of seed papers. Two seeds are co-cited when a later paper cites both; edge weight = count of co-citing papers, cosine = Salton index. Maps the intellectual base of a field. max_citing_per_seed bounds the API quota used per seed. Output: GraphML + CSV edge list written to disk.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
sourceNoData source. 'scopus' (default) needs subscriber entitlement for search, citations and references. 'openalex' needs none: IDs may be DOIs, OpenAlex work IDs (W...), or Scopus IDs (resolved to a DOI via Scopus metadata), and results carry OpenAlex IDs. Never mix sources within one analysis.scopus
seed_idsYesSeed papers: Scopus IDs (bare numeric or SCOPUS_ID: prefixed). With source='openalex', DOIs and OpenAlex work IDs also work.
min_sharedNoMinimum co-citing papers for an edge to be emitted (default 2).
max_citing_per_seedNoCap on citing papers fetched per seed (default 500). Limits quota usage.

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.1.0

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does reasonably well: it discloses the quota-bounding role of max_citing_per_seed and that output is persisted to disk as GraphML + CSV rather than returned inline. It omits runtime/job behavior (whether this is an async job like its siblings job_status/job_result imply) and error/retry semantics, which is a notable gap for a network-fetching tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Five tightly packed sentences, front-loaded with the core method definition, then the weight/cosine semantics, then the practical bounds and output artifacts. No filler and every sentence carries information an agent needs.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with no output schema, the description usefully names the returned artifacts (GraphML + CSV edge list on disk) and the source-entitlement split is covered by the schema. It does not clarify whether invocation is blocking or returns a job handle, which matters given the sibling job_status/job_result tools.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents source, seed_ids, min_shared and max_citing_per_seed in detail. The description's only parameter remark ('max_citing_per_seed bounds the API quota used per seed') largely repeats the schema's own 'Limits quota usage' note, adding little beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Build a co-citation graph for a set of seed papers') and defines the method precisely (co-cited when a later paper cites both), which implicitly separates it from related graph tools like bibliographic_coupling. However, it never names a sibling or explicitly distinguishes this technique from the adjacent citation_network/citation_lineage tools, so differentiation is left to inference.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The clause 'Maps the intellectual base of a field' implies the analytic context in which this tool is appropriate, but there is no explicit when-to-use, when-not-to-use, or named alternative (e.g., bibliographic_coupling). Usage is only implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.