Skip to main content
Glama

DataCite Works Statistics

datacite.works.stats
Read-onlyIdempotent

Get aggregated statistics across DataCite research outputs. Returns total DOI count, breakdown by resource type (datasets 70M+, software, preprints, journal articles, dissertations), registration year distribution, top contributing organizations (arXiv, CERN/Zenodo, Figshare, Mendeley), and top repositories. Optionally scope to a keyword query or specific resource type/year to get focused analytics. Useful for understanding the research landscape in a topic area — e.g. how many climate datasets exist, which repositories host the most neuroscience data.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
queryNoOptional keyword query to scope the statistics (e.g. "machine learning" returns stats only for ML-related DOIs). If omitted, stats cover all 70M+ DataCite works.
resource_typeNoScope statistics to a single resource type (e.g. "dataset" to see dataset-specific provider breakdowns).
publication_yearNoScope statistics to a specific publication year.

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault
errorNoPresent only when the call failed. Includes error code, message, request_id, and any provider-specific extras.
resultNoTool response payload. Shape varies per tool — consult the tool description and inputSchema. May be an object, array, string, or number depending on the upstream provider response.

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observed

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, covering safety. The description adds value by detailing what data is returned (breakdowns, distributions, top lists) without contradicting any annotations. It does not mention pagination or rate limits, but with annotations covering the read-only, non-destructive nature, this is sufficient.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single well-structured paragraph that front-loads the primary purpose, then enumerates return dimensions, then explains optional scoping and a use case. Every sentence contributes information without redundancy, making it efficient and easy to parse.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, the description does not need to explain return format details. It covers the key use case, scoping options, and what statistics are included, making it complete for an agent to decide when and how to invoke this read-only stats tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% for all three optional parameters (query, resource_type, publication_year), each fully documented in the schema. The description reinforces their purpose ('Optionally scope to a keyword query or specific resource type/year') but does not add substantial new meaning beyond the schema, so the baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's function: 'Get aggregated statistics across DataCite research outputs.' It enumerates specific return dimensions (DOI count, resource type breakdown, registration year distribution, top organizations, top repositories), making the tool's purpose concrete and distinguishable from sibling tools like datacite.doi.lookup or datacite.doi.search, which focus on individual records.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear usage context, noting it is 'Useful for understanding the research landscape in a topic area' and gives a concrete example. It explains that statistics can be scoped by keyword, resource type, or year. However, it does not explicitly contrast this tool with alternatives or state when not to use it, though the stats-focused nature is implicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.