Skip to main content
Glama

cern-opendata-mcp-server

Search CERN Open Data Records

cern_opendata_search_records
Read-onlyIdempotent

Search the CERN Open Data Portal's datasets, software, environments, documentation and supplementary records with exact-vocabulary filters and an optional full-text query. The filters are experiment, record type, physics category, keyword, collision energy and type, file format, data-taking year, event count, availability, collection, and LHCb magnet polarity, stripping stream and stripping version. Returns compact hits with recids plus live facet counts. Each facet ignores its own filter, so its counts show the alternatives under the other filters. Filter values are exact upstream; common spellings are normalized, and cern_opendata_list_reference lists the vocabulary. Paging reaches the first 10,000 matches.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
pageNoPage number, from 1. page × limit may not exceed 10,000.
sortNobestmatch (relevance), mostrecent (by date_published, newest first), title (A-Z) or title_desc (Z-A). Omitted: bestmatch with a query, mostrecent (newest first) without.
typeNoRecord types (OR): a primary (Dataset, Documentation, Environment, Software, Supplementaries, News) or Primary::Secondary such as Dataset::Collision or Dataset::Simulated. Array or comma-separated string, up to 7. Omit for every served type; Glossary is not served.
limitNoHits per page, 1-50.
queryNoFull-text query, sent verbatim as an OpenSearch query_string (AND between terms; title matches weigh double). Field forms such as title:"…", recid:(1 OR 2) or run_period:("Run2012B") work; cern_opendata_list_reference with topic query_syntax lists them. Omit to browse by filters alone.
year_toNoLatest data-taking year, inclusive. Alone, it means up to this year.
categoryNoPhysics categories of simulated datasets (OR): a primary such as Exotica, Higgs Physics, Standard Model Physics or Supersymmetry, or Primary::Secondary such as Higgs Physics::Standard Model. Array or comma-separated string, up to 20; case and the separator are normalized, and a known value holding a comma is kept whole. Heavy-Ion Physics also matches its leading-space spelling. cern_opendata_list_reference topic categories lists every value.
keywordsNoRecord keywords (OR), exact and case-sensitive (Education and education differ), sent as given. Array or comma-separated string, up to 10; pass a keyword holding a comma in an array.
file_typeNoFile formats and data tiers (OR), such as nanoaod, miniaod, aod, root, DAOD_PHYSLITE or csv. Array or comma-separated string, up to 20; case is normalized for known values.
year_fromNoEarliest data-taking year, inclusive. Alone, it means this year onward; for one year, set year_from and year_to to it.
collectionNoPortal collections (OR), exact and case-sensitive, such as CMS-Validated-Runs; copy spellings from a record's collections field. Array or comma-separated string, up to 10.
experimentNoExperiments (OR): ALICE, ATLAS, CMS, DELPHI, JADE, LHCb, OPERA, PHENIX, TOTEM. Array or comma-separated string, up to 9; case is normalized.
max_eventsNoMaximum number of events, inclusive.
min_eventsNoMinimum number of events, inclusive.
availabilityNoRecord availability (OR): online, partial, ondemand (on tape, requested before download) or requested. Array or comma-separated string, up to 4.
collision_typeNoCollision types (OR): pp, PbPb, pPb, e+e-, Interfill. PbPb also matches the Pb-Pb spelling. Array or comma-separated string, up to 6.
magnet_polarityNoLHCb magnet polarities (OR): MagDown or MagUp, set on LHCb collision datasets. Array or comma-separated string, up to 2; case is normalized.
collision_energyNoCollision energies (OR), such as 7TeV, 8TeV, 13TeV or 5.02TeV. Array or comma-separated string, up to 15; "13TeV, 13.6TeV" is one upstream value and is kept whole.
stripping_streamNoLHCb stripping streams (OR), such as BHADRON, CHARM, DIMUON, EW or LEPTONIC; they also match LHCb stripping documentation. Array or comma-separated string, up to 11; case is normalized. cern_opendata_list_reference topic lhcb lists them.
stripping_versionNoLHCb stripping versions (OR), such as stripping21 or stripping21r1p2. Array or comma-separated string, up to 12; case is normalized.

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault
capNoThe page size (limit) applied to this call.
hitsNoMatching records on this page.
pageNoThe page returned.
errorNoPresent when the call failed. Absent on success.
shownNoNumber of items returned on this page.
facetsNoLive facet counts. Each facet ignores its own filter, so its counts show the alternatives under the other filters. Terms facets list the first 10 values alphabetically (file_type up to 100). other_count is 0 for year and number_events.
noticeNoGuidance for the next call: how to page on, why nothing matched and what to change, or a caveat about the results.
has_moreNoTrue when the next page holds matches and lies within the first 10,000. False on the last reachable page even when more matches exist; truncated and notice say so.
truncatedNoTrue when more results remain past this page; notice says how to reach them.
totalCountNoTotal matches for the query and filters.
applied_filtersNoThe filters as the server applied them.

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observed

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With readOnly/idempotent/openWorld annotations already covering the safety profile, the description adds genuinely useful behavior: live facet counts where 'each facet ignores its own filter', the 10,000-match paging ceiling, and spelling normalization vs exact-match semantics for specific filters. It stops short of describing ranking or performance characteristics, but the facet behavior is a real value-add.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with purpose and then progressively more detailed behavior; nearly every sentence adds something an agent needs (facet semantics, paging limit, exact vs normalized filters, reference tool pointer). The one long enumeration of filter names largely duplicates the schema, which is minor waste but not padding.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 20-parameter, filter-heavy search tool with an output schema and full annotation coverage, the description covers purpose, filter behavior, pagination ceiling, and result shape ('compact hits with recids plus live facet counts') without redundantly explaining return fields. Nothing needed to call it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and each parameter already documents enums, OR semantics, maxItems, and normalization, so the schema does the heavy lifting. The description's filter enumeration and normalization notes mostly restate schema content rather than adding syntax or format guidance beyond it.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Starts with a specific verb+resource ('Search the CERN Open Data Portal's datasets, software, environments, documentation and supplementary records') and enumerates the filterable dimensions. It also signals its relationship to the sibling cern_opendata_list_reference for vocabulary lookup, letting an agent place it without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explains that filters can be used alone ('Omit to browse by filters alone'), that the full-text query is optional, and routes vocabulary questions to cern_opendata_list_reference. It does not, however, explicitly contrast with get_records or get_analysis_env, so sibling selection for non-search cases is left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.