cern-opendata-mcp-server
Server Details
Search CERN Open Data, fetch records, files, analysis environments, CMS good-run lists, HLT paths.
- Status
- Healthy
- Last Tested
- Transport
- Streamable HTTP · MCP 2025-11-25
- URL
- Repository
- cyanheads/cern-opendata-mcp-server
- GitHub Stars
- 0
- Server Listing
- cern-opendata-mcp-server
TDQS
Scored across 7 tools
Each tool targets a clearly distinct resource or action: search (records vs trigger paths), retrieval (records vs analysis environment), file listing, vocabulary reference, validated runs, and trigger paths. Overlaps between get_records and get_analysis_env are resolved by the latter's explicit assembly purpose, so an agent can reliably choose the right tool.
All tool names use the same cern_opendata_ prefix followed by snake_case verb_noun patterns (get_records, list_files, search_records, etc.). The minor variation between get_analysis_env and get_records is still consistent with the overall verb_noun convention.
Seven tools is well-scoped for a specialized metadata server covering search, retrieval, file listing, vocabulary, validated runs, trigger paths, and analysis environment assembly. No tool appears redundant and each earns its place.
The surface covers discovery, metadata retrieval, file listing, vocabulary decoding, validated-run lists, trigger-path lookup, and analysis-environment assembly, which is strong lifecycle coverage for a read-only portal. Minor gaps remain, such as no direct tool to request on-demand files or to fetch actual file contents, but these are likely outside the intended metadata scope.
Available Tools
7 toolscern_opendata_get_analysis_envGet CERN Open Data Analysis EnvironmentARead-onlyIdempotentInspect
Assemble what is needed to analyse a record: its container images, CMSSW release and global tag; the condition-data, VM and validated-run records for its run periods; example software that declares it works with the record; and quoted sections of the guides the record links (the first two are fetched). Container images, software and guide code are licensed separately from the CC0 data. Use cern_opendata_list_files for the record's files.
| Name | Required | Description | Default |
|---|---|---|---|
| recid | Yes | Record id: up to 12 digits (6004), optionally after an experiment prefix (cms-93956); recid:6004 or a portal record URL also work. cern_opendata_search_records and cern_opendata_get_records return it. |
Output Schema
| Name | Required | Description |
|---|---|---|
| type | No | Record type as the portal classifies it. |
| error | No | Present when the call failed. Absent on success. |
| recid | No | The record id. |
| title | No | Record title. |
| guides | No | Every portal guide linked in the usage section. |
| notice | No | What could not be assembled and how to get it: an empty environment, linked records that could not be read, guides not fetched, not found or cut, a guide anchor no heading carries, or more linked records than shown. |
| software | No | The record's own software environment (system_details). |
| experiment | No | Experiments the record belongs to. |
| run_period | No | Run periods the record covers, such as Run2012B. |
| other_links | No | The other usage links (getting-started pages and the like). |
| license_note | No | How the environment is licensed. |
| example_software | No | Software records that declare they work with this record. |
| environment_records | No | Condition, VM and validation records for the run periods, and others linking here. |
| separately_licensed | No | Always true: images, software and guide code are licensed apart from the CC0 data. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly=true, idempotent=true and openWorld=true, so the safety profile is covered. The description adds genuinely non-obvious behavior: only the first two sub-collections are fetched, and container images/software/guide code carry licenses distinct from the CC0 data. Return details are left to the output schema, which is acceptable.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the purpose before a colon-delimited inventory of components, so structure is strong and each clause carries information. The licensing sentence is slightly tangential and the run-on list is dense, but nothing is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an aggregate read tool with a rich input schema and an existing output schema, the description need not explain return values, and it correctly leaves them to the schema while describing scope. The only minor gap is the lack of guidance on the sibling aggregate calls, but the essentials for invoking it correctly are present.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Single parameter with 100% schema description coverage; the schema already documents the recid formats, prefixes, and where to obtain it. The description adds nothing about recid semantics, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Assemble) and a precise resource (analysis environment for a record) and then enumerates the exact artifacts returned: container images, CMSSW release/global tag, condition-data/VM/validated-run records, example software, quoted guide sections. This clearly distinguishes it from siblings, and it even routes the file-listing case elsewhere.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Names an explicit alternative ('Use cern_opendata_list_files for the record's files') with the condition that selects it, which is real routing guidance. However, it does not clarify when to prefer get_records or get_validated_runs (broadly overlapping siblings) over this aggregate call, so it stops short of full when-to-use/when-not coverage.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
cern_opendata_get_recordsGet CERN Open Data RecordsARead-onlyIdempotentInspect
Fetch full metadata for 1-20 records in one call, by recid, DOI, CMS dataset path (/Primary/Era/TIER) or documentation slug. Returns the description, run periods, collision and distribution details, related records, a software-environment summary, the license and a ready citation, plus what a record states of its variable dictionary, physics category, pile-up, keywords, and LHCb magnet polarity and stripping. Documentation and news pages include their markdown body in slices of at most 30,000 characters; body_offset with that one id reads on from where a slice stops. Each response holds to 64,000 bytes: records past it are left out whole and listed under deferred, to pass back as ids. File lists are not included; use cern_opendata_list_files. Identifiers that do not resolve come back under missing with guidance; they do not fail the call.
| Name | Required | Description | Default |
|---|---|---|---|
| ids | Yes | Identifiers to resolve, 1-20: an array, or one comma-separated string. Forms may be mixed; duplicates collapse. | |
| body_offset | No | With exactly one documentation or news id: where its body slice starts, in characters (UTF-16 code units), usually the body_next_offset an earlier call returned. 0 (the default) is the body's start and is accepted with any ids; above 0 needs exactly one id. |
Output Schema
| Name | Required | Description |
|---|---|---|
| error | No | Present when the call failed. Absent on success. |
| notice | No | Caveats about the returned records: a documentation body with more to read, or records deferred by the response budget, each with the call that continues it. |
| missing | No | Identifiers that resolved to no record; empty when every id resolved. |
| records | No | Resolved records, in the order of their first matching input. |
| deferred | No | Inputs whose records were left out to hold the response to 64,000 bytes, in input order; pass them as ids to fetch those records. Empty when every resolved record is returned. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only cover safety/idempotency, yet the description discloses real operational behavior: a 64,000-byte response cap with overflow records returned whole under 'deferred' to pass back as ids, unresolved identifiers surfaced under 'missing' rather than failing the call, and 30,000-character body slicing with body_offset continuation. This is exactly the context an agent needs to drive the tool in a loop.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose and identifier forms are front-loaded, and every subsequent sentence carries operational payload (slicing, size cap, missing handling). It is dense and readable, though the very long second sentence could be split for scanability.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return-field documentation is not required, yet the description still explains response structure, pagination, truncation and error semantics. For a read-only lookup tool with full schema coverage, nothing an agent needs to invoke it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both parameters are already fully documented, including the identifier forms and the body_offset constraint. The description reinforces the same details but adds no syntax or edge-case meaning beyond the schema, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (fetch) plus resource (full metadata for records) and enumerates the four identifier forms accepted. It explicitly differentiates itself from the sibling cern_opendata_list_files, so an agent can route correctly without opening schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit exclusion and alternative ('File lists are not included; use cern_opendata_list_files') and describes the batching limit (1-20 ids). It does not contrast with cern_opendata_search_records, so the when-to-search-vs-when-to-fetch boundary is left implicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
cern_opendata_get_validated_runsGet CMS Validated RunsARead-onlyIdempotentInspect
Get a CMS validated-run (good-run) list, which certifies the luminosity sections that are good for physics in each run. Select it by a CMS collision dataset recid, a validated-run list recid, or a run period such as Run2012B (give exactly one of recid and run_period). Choose the full validation or the muons-only variant, and narrow to a run range; a dataset recid defaults the range to the first and last run the dataset lists. Returns the runs with their luminosity-section ranges and the list file download URL. CMS only.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Runs to return, 1-2000. | |
| recid | No | A CMS collision dataset recid (its linked list is used) or a validated-run list recid (used as named): up to 12 digits, optionally after an experiment prefix (cms-93956); recid:N or a portal record URL also work. Give this or run_period. | |
| run_max | No | Highest run to return, inclusive. Defaults with run_min for a dataset recid. | |
| run_min | No | Lowest run to return, inclusive. With a dataset recid and neither bound set, run_min and run_max default to the first and last run the dataset lists; run_bounds echoes the range applied. | |
| variant | No | full (every detector certified) or muons_only (certified for muon physics). Omitted: a list recid is used as named; a dataset recid or run_period selects full. | |
| run_period | No | A CMS run period such as Run2012B, matched case-insensitively; 2012B also matches Run2012B. Give this or recid. cern_opendata_list_reference with topic run_periods lists them. |
Output Schema
| Name | Required | Description |
|---|---|---|
| cap | No | The page size (limit) applied to this call. |
| list | No | The selected list, when exactly one matched. |
| runs | No | Certified runs, ascending, inside run_bounds and cut at limit; empty when no single list was selected. |
| error | No | Present when the call failed. Absent on success. |
| shown | No | Number of items returned on this page. |
| notice | No | Guidance for the next call: how to page on, why nothing matched and what to change, or a caveat about the results. |
| dataset | No | The dataset, when recid named a dataset rather than a list. |
| summary | No | The whole selected list, before run_bounds. |
| truncated | No | True when more results remain past this page; notice says how to reach them. |
| run_bounds | No | The run range runs was filtered to; absent when no bound applies. |
| totalCount | No | Runs in the selected list inside run_bounds. |
| matched_lists | No | Every list the selector matched in the requested variant. Several means none was read: call again with recid set to one of them. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover safety (readOnly, idempotent, openWorld), so the bar is lower; the description adds useful behavioral context beyond them: the CMS-only scope, the implicit run-range defaulting for a dataset recid, and the variant defaulting rules. It stops short of noting rate limits or how large a returned list can be, so it is not exhaustive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
It is a dense block but front-loads the core purpose and then the selection rules in a logical progression; every clause carries information. It is slightly long for the amount of guidance and could be split into sentences for faster scanning.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present the description needn't explain return values, but it still summarizes them (runs with luminosity-section ranges and list file download URL). Combined with annotations and full schema coverage, an agent has everything needed to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3, but the description adds interpretive value: it clarifies the dual meaning of recid (collision dataset vs named list) and how variant and run bounds default differently per selection mode.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb and resource ('Get a CMS validated-run (good-run) list') and immediately states what the artifact certifies (luminosity sections good for physics), which no sibling tool does. It is clearly distinguishable from the record/file/reference lookup siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It states the mutually exclusive selection rules explicitly ('give exactly one of recid and run_period'), explains the variant choice, and routes the agent to a sibling ('cern_opendata_list_reference with topic run_periods lists them') for discovering run periods. That is concrete when/how guidance, not inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
cern_opendata_list_filesList CERN Open Data Record FilesARead-onlyIdempotentInspect
List one record's files: its file indexes (groups of up to ~1,500 files) with their XRootD URI-list URLs, and per file the XRootD URI, HTTPS download URL, size, adler32 checksum and availability. Without index, returns the record's indexes and its regular files; with index, reads only that index and pages through its files. Files marked on demand sit on tape and must be requested on the record's portal page before download.
| Name | Required | Description | Default |
|---|---|---|---|
| index | No | A file index key from the indexes list, matched exactly (a .txt ending is read as .json). Omit to list the record's indexes and regular files. | |
| limit | No | Files per page, 1-500. | |
| recid | Yes | Record id: up to 12 digits (6004), optionally after an experiment prefix (atlas-160006); recid:6004 or a portal record URL also work. cern_opendata_search_records and cern_opendata_get_records return it. | |
| cursor | No | next_cursor from the previous page, unchanged, with the same recid and index. Omit for the first page. |
Output Schema
| Name | Required | Description |
|---|---|---|
| cap | No | The page size (limit) applied to this call. |
| error | No | Present when the call failed. Absent on success. |
| files | No | This page of files: regular files in record scope, the index's files in index scope. |
| recid | No | The record id. |
| scope | No | record: the record's indexes and regular files; index: one index's files. |
| shown | No | Number of items returned on this page. |
| title | No | Record title. |
| notice | No | Guidance for the next call: how to page on, why nothing matched and what to change, or a caveat about the results. |
| indexes | No | Every file index in record scope; only the selected one in index scope. |
| children | No | Set only for an umbrella record holding no files itself: the child recids whose files make it up. Empty otherwise. |
| has_more | No | True when more files remain past this page. |
| truncated | No | True when more results remain past this page; notice says how to reach them. |
| portal_url | No | The record page, where on-demand (tape) files are requested. |
| totalCount | No | Files in scope: regular files in record scope, index files in index scope. |
| next_cursor | No | Pass as cursor, with the same recid and index, for the next page. |
| availability | No | Record-level availability: online, partial, ondemand or requested. |
| availability_details | No | File counts by availability state for the whole record. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and openWorldHint, so the safety profile is covered. The description adds valuable behavior beyond annotations: the paging mechanism via cursor with same recid/index, the index file limit (~1,500 files per index), and the tape/on-demand requirement to request files on the portal before download. This is meaningful operational context not present in the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two dense sentences that front-load the primary purpose and then describe the mode split and the on-demand tape caveat. There is almost no waste, though the enumeration of every returned field is lengthy and partly redundant with the output schema, slightly reducing conciseness.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is relatively complex (4 params, paging, index modes) and the description plus schema plus output schema cover most needs. It explains the index/no-index behavior, paging cursor usage, and the on-demand availability caveat. Minor gaps remain, such as error behavior or how the ~1,500-file index grouping affects paging, but these are not critical for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, with detailed per-parameter descriptions (index exact match and .txt-to-.json handling, limit 1-500, recid pattern, cursor semantics). The description mentions the index mode but does not add syntax or format details beyond the schema. A baseline 3 is appropriate because the schema fully documents parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource ('List one record's files') and immediately enumerates the return contents (indexes with XRootD URI-list URLs, per-file XRootD URI, HTTPS URL, size, adler32, availability). This clearly distinguishes it from siblings like get_records and search_records, which operate at the record level rather than the file level.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explains the two modes: 'Without index, returns the record's indexes and its regular files; with index, reads only that index and pages through its files.' This is useful conditional guidance tied to the index parameter. However, it does not explicitly name alternatives or state when not to use this tool (e.g., to find a record, use search_records).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
cern_opendata_list_referenceList Reference VocabularyARead-onlyIdempotentInspect
Decode the vocabulary the other cern_opendata tools accept: experiments, record types, collision energies and types, file formats and data tiers, availability states, physics categories, LHCb magnet polarities and stripping streams and versions, identifier forms, query syntax, licensing, and the CMS run periods that have validated-run lists. Static and offline; omit topic for every table.
| Name | Required | Description | Default |
|---|---|---|---|
| topic | No | One table to return: experiments, record_types, collision_energies, collision_types, file_types, availability, categories, lhcb, identifiers, query_syntax, licensing or run_periods. Omit for every table. |
Output Schema
| Name | Required | Description |
|---|---|---|
| error | No | Present when the call failed. Absent on success. |
| topics | No | The requested table, or every table when topic is omitted. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint/idempotentHint/openWorldHint=false, and the description usefully reinforces this with 'Static and offline', which confirms results are deterministic and safe to repeat. It adds real behavioral context (no network, cached reference data) beyond the safety flags, though nothing about limits or output shape.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core purpose, and the trailing 'omit topic for every table' clause is a useful default up front. The long topic enumeration is somewhat redundant with the enum but does convey the breadth of the vocabulary in one pass.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, return-value explanation is not required, and the description covers scope, default behavior, and offline/static nature. An agent has enough to call it correctly; only a pointer to a concrete alternative tool is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the single enum param is fully documented there, so baseline is 3. The description enumerates the same topic values in prose (plus extras like LHCb magnet polarities that map under 'lhcb') and restates the omit-for-all default, adding little beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (decode/list) and resource (the controlled vocabulary) and explains exactly what it contains for the cern_opendata tool family. It is clearly distinguishable from siblings like get_records or search_records, which consume this vocabulary rather than define it.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explains where this fits: it decodes the terms the other cern_opendata tools accept, so an agent knows to consult it before constructing queries. It also states the no-argument default ('omit topic for every table'). It stops short of naming a specific alternative tool to use instead, so not a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
cern_opendata_search_recordsSearch CERN Open Data RecordsARead-onlyIdempotentInspect
Search the CERN Open Data Portal's datasets, software, environments, documentation and supplementary records with exact-vocabulary filters and an optional full-text query. The filters are experiment, record type, physics category, keyword, collision energy and type, file format, data-taking year, event count, availability, collection, and LHCb magnet polarity, stripping stream and stripping version. Returns compact hits with recids plus live facet counts. Each facet ignores its own filter, so its counts show the alternatives under the other filters. Filter values are exact upstream; common spellings are normalized, and cern_opendata_list_reference lists the vocabulary. Paging reaches the first 10,000 matches.
| Name | Required | Description | Default |
|---|---|---|---|
| page | No | Page number, from 1. page × limit may not exceed 10,000. | |
| sort | No | bestmatch (relevance), mostrecent (by date_published, newest first), title (A-Z) or title_desc (Z-A). Omitted: bestmatch with a query, mostrecent (newest first) without. | |
| type | No | Record types (OR): a primary (Dataset, Documentation, Environment, Software, Supplementaries, News) or Primary::Secondary such as Dataset::Collision or Dataset::Simulated. Array or comma-separated string, up to 7. Omit for every served type; Glossary is not served. | |
| limit | No | Hits per page, 1-50. | |
| query | No | Full-text query, sent verbatim as an OpenSearch query_string (AND between terms; title matches weigh double). Field forms such as title:"…", recid:(1 OR 2) or run_period:("Run2012B") work; cern_opendata_list_reference with topic query_syntax lists them. Omit to browse by filters alone. | |
| year_to | No | Latest data-taking year, inclusive. Alone, it means up to this year. | |
| category | No | Physics categories of simulated datasets (OR): a primary such as Exotica, Higgs Physics, Standard Model Physics or Supersymmetry, or Primary::Secondary such as Higgs Physics::Standard Model. Array or comma-separated string, up to 20; case and the separator are normalized, and a known value holding a comma is kept whole. Heavy-Ion Physics also matches its leading-space spelling. cern_opendata_list_reference topic categories lists every value. | |
| keywords | No | Record keywords (OR), exact and case-sensitive (Education and education differ), sent as given. Array or comma-separated string, up to 10; pass a keyword holding a comma in an array. | |
| file_type | No | File formats and data tiers (OR), such as nanoaod, miniaod, aod, root, DAOD_PHYSLITE or csv. Array or comma-separated string, up to 20; case is normalized for known values. | |
| year_from | No | Earliest data-taking year, inclusive. Alone, it means this year onward; for one year, set year_from and year_to to it. | |
| collection | No | Portal collections (OR), exact and case-sensitive, such as CMS-Validated-Runs; copy spellings from a record's collections field. Array or comma-separated string, up to 10. | |
| experiment | No | Experiments (OR): ALICE, ATLAS, CMS, DELPHI, JADE, LHCb, OPERA, PHENIX, TOTEM. Array or comma-separated string, up to 9; case is normalized. | |
| max_events | No | Maximum number of events, inclusive. | |
| min_events | No | Minimum number of events, inclusive. | |
| availability | No | Record availability (OR): online, partial, ondemand (on tape, requested before download) or requested. Array or comma-separated string, up to 4. | |
| collision_type | No | Collision types (OR): pp, PbPb, pPb, e+e-, Interfill. PbPb also matches the Pb-Pb spelling. Array or comma-separated string, up to 6. | |
| magnet_polarity | No | LHCb magnet polarities (OR): MagDown or MagUp, set on LHCb collision datasets. Array or comma-separated string, up to 2; case is normalized. | |
| collision_energy | No | Collision energies (OR), such as 7TeV, 8TeV, 13TeV or 5.02TeV. Array or comma-separated string, up to 15; "13TeV, 13.6TeV" is one upstream value and is kept whole. | |
| stripping_stream | No | LHCb stripping streams (OR), such as BHADRON, CHARM, DIMUON, EW or LEPTONIC; they also match LHCb stripping documentation. Array or comma-separated string, up to 11; case is normalized. cern_opendata_list_reference topic lhcb lists them. | |
| stripping_version | No | LHCb stripping versions (OR), such as stripping21 or stripping21r1p2. Array or comma-separated string, up to 12; case is normalized. |
Output Schema
| Name | Required | Description |
|---|---|---|
| cap | No | The page size (limit) applied to this call. |
| hits | No | Matching records on this page. |
| page | No | The page returned. |
| error | No | Present when the call failed. Absent on success. |
| shown | No | Number of items returned on this page. |
| facets | No | Live facet counts. Each facet ignores its own filter, so its counts show the alternatives under the other filters. Terms facets list the first 10 values alphabetically (file_type up to 100). other_count is 0 for year and number_events. |
| notice | No | Guidance for the next call: how to page on, why nothing matched and what to change, or a caveat about the results. |
| has_more | No | True when the next page holds matches and lies within the first 10,000. False on the last reachable page even when more matches exist; truncated and notice say so. |
| truncated | No | True when more results remain past this page; notice says how to reach them. |
| totalCount | No | Total matches for the query and filters. |
| applied_filters | No | The filters as the server applied them. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With readOnly/idempotent/openWorld annotations already covering the safety profile, the description adds genuinely useful behavior: live facet counts where 'each facet ignores its own filter', the 10,000-match paging ceiling, and spelling normalization vs exact-match semantics for specific filters. It stops short of describing ranking or performance characteristics, but the facet behavior is a real value-add.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with purpose and then progressively more detailed behavior; nearly every sentence adds something an agent needs (facet semantics, paging limit, exact vs normalized filters, reference tool pointer). The one long enumeration of filter names largely duplicates the schema, which is minor waste but not padding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 20-parameter, filter-heavy search tool with an output schema and full annotation coverage, the description covers purpose, filter behavior, pagination ceiling, and result shape ('compact hits with recids plus live facet counts') without redundantly explaining return fields. Nothing needed to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and each parameter already documents enums, OR semantics, maxItems, and normalization, so the schema does the heavy lifting. The description's filter enumeration and normalization notes mostly restate schema content rather than adding syntax or format guidance beyond it.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Starts with a specific verb+resource ('Search the CERN Open Data Portal's datasets, software, environments, documentation and supplementary records') and enumerates the filterable dimensions. It also signals its relationship to the sibling cern_opendata_list_reference for vocabulary lookup, letting an agent place it without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explains that filters can be used alone ('Omit to browse by filters alone'), that the full-text query is optional, and routes vocabulary questions to cern_opendata_list_reference. It does not, however, explicitly contrast with get_records or get_analysis_env, so sibling selection for non-search cases is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
cern_opendata_search_trigger_pathsSearch CMS Trigger PathsARead-onlyIdempotentInspect
Look up CMS High-Level Trigger paths by exact name (HLT_IsoMu24) or prefix pattern (HLT_IsoMu*), optionally for one data-taking year. Paths outside the HLT_ family, such as AlCa_EcalPi0, DST_ and DQM_ paths, output modules and HLTriggerFinalPath, are found by their own names, in the case given. Each match is a per-year path record parsed into the primary datasets its title names, the first and last run the path was seen online, per-version run ranges, the L1 seed, and links to the HLT menu records. Covers CMS open data from 2011-2016. Prescale tables are not published.
| Name | Required | Description | Default |
|---|---|---|---|
| page | No | Page number, from 1. page × limit may not exceed 10,000. | |
| path | Yes | Trigger path name (HLT_IsoMu24, AlCa_EcalPi0) or prefix with one trailing wildcard (HLT_IsoMu*): a letter or digit, then letters, digits and underscores. Trimmed. A path starting HLT_ in any case is searched as given, with the prefix case fixed, and matches path names in any case. Any other path is matched against record path names in the case given, the whole name or, with a trailing *, its start (AlCa_EcalPi0, AlCa_*), and is also searched with HLT_ prepended (IsoMu24 finds HLT_IsoMu24, 60Jet10 finds HLT_60Jet10). A trailing version suffix (_v3 or _v*) is dropped, since records list versions as V<n>. | |
| year | No | Data-taking year, such as 2012. Trigger records cover 2011-2016. | |
| limit | No | Records per page, 1-50. |
Output Schema
| Name | Required | Description |
|---|---|---|
| cap | No | The page size (limit) applied to this call. |
| page | No | The page returned. |
| error | No | Present when the call failed. Absent on success. |
| shown | No | Number of items returned on this page. |
| notice | No | Guidance for the next call: how to page on, why nothing matched and what to change, or a caveat about the results. |
| has_more | No | True when the next page holds matches and lies within the first 10,000. False on the last reachable page even when more matches exist; truncated and notice say so. |
| triggers | No | Matching trigger path records on this page. |
| truncated | No | True when more results remain past this page; notice says how to reach them. |
| totalCount | No | Trigger path records matching the path (and year). |
| effectiveQuery | No | The query sent, after normalization (version suffix dropped): an HLT_ path as given; any other path as title clauses matching it as a record path name, OR HLT_ plus the path. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare read-only, idempotent, and open-world behavior, and the description adds significant matching and coverage details beyond them: case behavior, wildcard handling, dropped version suffixes, per-year records, run ranges, L1 seeds, and the 2011–2016 CMS open-data scope. It also discloses a real limitation: 'Prescale tables are not published.' No behavior contradicts the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but front-loaded, starting with the core lookup action and matching modes before covering exceptions, return record contents, and source limitations. Every sentence contributes domain-relevant information for a specialized search tool, with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's domain complexity, the description is complete: it explains what is searched, how names and prefixes match, what a match record contains, the available year range, and a known data limitation. An output schema exists, so return-value structure need not be repeated in prose.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all four parameters in detail, including path pattern rules, year range, and pagination limits. The description adds useful domain context about trigger-path naming but does not expand parameter syntax or semantics beyond what the schema provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'Look up CMS High-Level Trigger paths by exact name ... or prefix pattern.' It clearly distinguishes this tool from generic record or file searches by focusing on CMS HLT path records and naming the HLT_ family. An agent can identify the tool's narrow purpose without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context for when the tool applies: exact trigger-path names or prefix patterns, optional year filtering, and how non-HLT_ paths such as AlCa_EcalPi0 or DQM_ are handled. It does not explicitly compare this tool against siblings like cern_opendata_search_records, so it stops short of full alternative routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
7 tool updates
- First observed
cern_opendata_get_analysis_env - First observed
cern_opendata_get_records - First observed
cern_opendata_get_validated_runs - First observed
cern_opendata_list_files - First observed
cern_opendata_list_reference - First observed
cern_opendata_search_records - First observed
cern_opendata_search_trigger_paths
Related MCP Connectors
Search INSPIRE-HEP papers, authors, experiments, HEPData records; get citation metrics and BibTeX.
811Search NOAA CDO stations and datasets, fetch historical weather observations.
Search and query government open-data portals (Socrata SODA API).
Search NOAA climate stations and datasets, fetch historical weather observations.
Related MCP Servers
- AlicenseNot gradedqualityBmaintenanceEnables AI agents to search CERN's open research data repository and retrieve records, files, and curated communities by DOI or keyword, covering datasets, software, papers, and presentations. Works keyless over HTTP or as a local stdio server, with no account required for initial calls.169 npmMIT

hepdata-mcpofficial
AlicenseAqualityBmaintenanceEnables discovery of HEPData records, tables, and data access with read-only operations and export links.9GPL 2.0- AlicenseAqualityAmaintenanceSearches and fetches research datasets across Zenodo, DataCite (Dryad/Figshare/Dataverse/OSF), NCBI omics archives (GEO/SRA/BioProject), and the literature (PubMed/OpenAIRE) through one normalized model — deduplicating by DOI, expanding organism queries with NCBI Taxonomy synonyms, and bridging papers to the datasets they produced. Resolves citations and open-access full text, and downloads files.6310 PyPI4MIT
- AlicenseNot gradedqualityAmaintenanceEnables searching high-energy-physics literature, author profiles, experiments, and HEPData measurement records, fetching full paper dossiers, computing citation summaries and h-indices, and exporting BibTeX or LaTeX entries through chainable query identifiers. It runs as a stdio process or a local Streamable HTTP server, exposing eight tools and one literature resource.1Apache 2.0
Glama MCP Gateway
Add one secure layer between your agents and this server.