cern-opendata-mcp-server
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@cern-opendata-mcp-serverFind CMS datasets about Higgs boson from 2012"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Public Hosted Server: https://cern-opendata.caseyjhand.com/mcp
Overview
Particle-physics data from the CERN Open Data Portal: collision, simulated and derived datasets, analysis software, environments and documentation from ALICE, ATLAS, CMS, LHCb and other experiments. Search with exact-vocabulary filters and live facet counts, open records with license and citation, list data files, assemble a record's analysis environment, and look up CMS good-run lists and trigger paths. Runs as a stdio process, a local Streamable HTTP server, or the public hosted endpoint above.
Tools
Tool | Description |
| Search datasets, software, environments, documentation and supplementary records with exact-vocabulary filters and live facet counts |
| Fetch full metadata for 1–20 records by recid, DOI, CMS dataset path or documentation slug, with license and citation |
| Page through a record's file indexes and files: XRootD URIs, HTTPS URLs, sizes, checksums, tape availability |
| Assemble a record's analysis environment: container images, CMSSW release, global tag, linked environment and software records, guide sections |
| Get a CMS validated-run (good-run) list for a dataset, a list or a run period, with luminosity-section ranges |
| Look up CMS High-Level Trigger paths by name or prefix, parsed into run ranges, versions and L1 seeds |
| Decode the vocabulary the other tools accept: experiments, record types, energies, formats, physics categories, LHCb stripping, identifiers, query syntax, licensing, run periods |
Resources
Resource | Description |
| One record's metadata, license and citation, in the |
Tool-only clients get the same data from cern_opendata_get_records.
Related MCP server: data-bs-mcp
Capability reference
cern_opendata_search_records tool
Optional
query(an OpenSearchquery_string, up to 500 characters) plus OR-list filterstype,experiment,category(physics category,Higgs Physics::Standard Model),keywords,collision_energy,collision_type,file_type,availability,collection, and the LHCbmagnet_polarity,stripping_streamandstripping_version, each an array or a comma-separated string;year_from/year_toandmin_events/max_eventsbound the data-taking year and event countsort(bestmatch,mostrecentfor newestdate_publishedfirst,title,title_desc),limit1–50 (default 10) andpagefrom 1;page × limitpast 10,000 fails aspage_window_exceededCompact
hitswith recids, plus thirteen livefacetsthat each ignore their own filter (typeandcategorywith their secondary values);applied_filtersechoes what ran, and values outside the verified vocabulary appear underunrecognized_values
cern_opendata_get_records tool
ids: 1–20 recids, DOIs, CMS dataset paths (/Primary/Era/TIER) or documentation slugs, mixed in one array or comma-separated stringEach record carries a
licensewith itsbasis(record,cern_terms_default,not_stated) and, when it has a DOI, a readycitationDocumentation and news bodies come in slices of up to 30,000 characters:
body_offsetwith that one id reads on from thebody_next_offsetthe last slice returnedEach response stays within 64,000 bytes: records past the budget are left out whole and listed in
deferred, to pass back asids(the first record always comes back whole)Records also return what they state of a
variablesdictionary (name, type, unit, description), a physicscategory, pile-up (pileup_html, with the pile-up datasets underlinks),keywords, and the LHCbmagnet_polarityandstrippingstream and versionUnresolved identifiers land in
missingwithinterpreted_asand guidance instead of failing the call; file lists come fromcern_opendata_list_files
cern_opendata_list_files tool
recidrequired; withoutindex, returns the record's file indexes and regular files, and with an index key, that index's files, read without the rest of the recordlimit1–500 (default 50), continued withnext_cursor; each file carriesxrootd_uri,size_in_bytes,checksumandavailability, anhttps_urlwhen one can be built from its key or EOS path, andkeyorfilenameas the portal states them, and each index auri_list_urllisting every XRootD URI in itFiles marked
on demandsit on tape and must be requested on the record's portal page first; an umbrella record with no files of its own returns its child recids underchildren
cern_opendata_get_analysis_env tool
recidrequired;softwarecarries the record's own container images, CMSSW release, global tag and environment recidenvironment_records(condition, VM, validation) for the record's run periods andexample_softwarethat declares it works with the record, up to 50 between them;guidesquotes the linked section of the first two portal guides, each capped at 12,000 characters, and a cut names thecern_opendata_get_recordsbody_offsetthat reads on from itAlways
separately_licensed: true; linked records or guides that can't be read leave anoticeinstead of failing the call
cern_opendata_get_validated_runs tool
Exactly one of
recid(a CMS collision dataset or a validated-run list) orrun_period(Run2012B;2012Balso matches);variantfullormuons_only;run_min/run_max;limit1–2000 (default 200)A dataset
recidbounds the runs to the dataset's first and last listed run, echoed inrun_bounds; when several lists match,matched_listsnames them and no runs are readEach run carries
lumi_sectionsandlumi_ranges;list.https_urldownloads the whole list file. CMS only: other records fail asno_validated_runs
cern_opendata_search_trigger_paths tool
path: an exact name (HLT_IsoMu24,AlCa_EcalPi0) or a prefix with one trailing*(HLT_IsoMu*); a name without theHLT_prefix is matched against record path names in the case given and also searched withHLT_added, and a_v<n>version suffix is dropped; optionalyear,limit1–50 (default 10) andpageEach per-year record is parsed into the primary
datasetsits title names,first_seen,last_seen, per-version run ranges with theirl1_seed, and HLT menu record links;parsed: falsemarks a record to read from itsabstract_htmlCMS open data from 2011–2016 only
cern_opendata_list_reference tool
Optional
topic:experiments,record_types,collision_energies,collision_types,file_types,availability,categories,lhcb,identifiers,query_syntax,licensingorrun_periods; omit it for every tableStatic and offline;
categories,lhcbandrun_periodsare dated snapshots, while the search facets andcern_opendata_get_validated_runsread the live portal
cern-opendata://record/{recid} resource
One record by
recid(6004, or a prefixed recid such asatlas-160006; leading zeros ignored) asapplication/json, in thecern_opendata_get_recordsrecord shape: metadata (variable dictionary, physics category, pile-up and LHCb run conditions included), license and citation, without file listsrecidcomes fromcern_opendata_search_records; reads carry a 15-minute public cache hint
Features
Built on @cyanheads/mcp-ts-core: stdio and Streamable HTTP transports, pluggable auth (none / jwt / oauth), swappable storage (in-memory, filesystem, Supabase, Cloudflare KV/R2/D1), structured logging with optional OpenTelemetry tracing.
CERN Open Data-specific:
Keyless and read-only; it never stages tape files or writes anything
One shared pacer at 50 requests a minute, under the portal's published 60 per client IP, and one 50-second deadline per call across queue wait and retries
Filter values canonicalized against the portal's verified vocabulary (
13 tev→13TeV,lhcb→LHCb,Pb-Pb→PbPb,dataset/collision→Dataset::Collision); unknown values are sent as given and flaggedTape-resident (
ondemand) records are included in every search and lookup, where the portal otherwise drops them silentlyFile manifests, and file indexes read on their own, are cached for 15 minutes, so paging through a record's or an index's files costs one read
Agent-friendly output:
Provenance on every response:
portal_urlon each hit and record, alicensewith itsbasis, a DOIcitation, andapplied_filtersoreffectiveQueryechoing what ranGraceful partial results: unresolved ids land under
missingwith guidance, and unreadable linked records or guides incern_opendata_get_analysis_envsurface as anoticerather than a failureDiscriminated outputs:
kind,license.basis,interpreted_as,scope,variant,run_bounds.sourceandparsedlet callers branch on data, not string parsingPortal text kept as data: titles, descriptions, guide sections and file names are fenced or escaped in
content[]and relayed as received (HTML in_htmlfields) instructuredContent
Data and licensing
Portal metadata and datasets are CC0 under the CERN Open Data Terms of Use. Software, container images, documentation and guide code are licensed separately, per record (software is commonly GPL). cern_opendata_get_records reports each record's license and its basis: record when the record states one, cern_terms_default for a dataset that states none (CC0), and not_stated otherwise.
CERN asks reusers to cite each dataset's DOI in applications and publications; cern_opendata_get_records returns a ready citation for every record with a DOI.
This server is an independent project and is not affiliated with or endorsed by CERN.
Known limitations
60 requests a minute per client IP. The portal publishes this limit; the server paces itself to 50 a minute, and a call that cannot start within its deadline fails as
rate_limitedwithretryAfter. A hosted deployment shares one budget across every user behind its egress IP.cern_opendata_get_analysis_envandcern_opendata_get_validated_runscost 2–4 requests each.10,000-result window. Search and trigger-path paging reach only the first 10,000 matches; narrow deeper result sets with filters.
Facet lists are partial. Terms facets return the first 10 values alphabetically (
file_typeup to 100), with the rest counted inother_count.Tape-resident files. Files with availability
on demandmust be requested on the record's portal page before download; staging them is a write and out of scope. A record whose availability isondemandlists none of its files through the API, socern_opendata_list_filesreports only the count and size its metadata states.CMS-only run lists and triggers. Trigger records cover 2011–2016, and no prescale tables are published, so trigger detail is limited to what each record's abstract states. Muons-only run lists do not exist for Commissioning2010, Run2010B or the 2011 ReReco list.
Glossary entries are not served. The portal's glossary links answer 404, so search excludes them.
Getting started
Public Hosted Instance
A public instance is available at https://cern-opendata.caseyjhand.com/mcp — no installation required. Point any MCP client at it via Streamable HTTP:
{
"mcpServers": {
"cern-opendata-mcp-server": {
"type": "streamable-http",
"url": "https://cern-opendata.caseyjhand.com/mcp"
}
}
}Every caller of the hosted instance shares one portal request budget of 50 requests a minute. For sustained use, run your own instance.
Self-Hosted / Local
Add the following to your MCP client configuration file.
{
"mcpServers": {
"cern-opendata-mcp-server": {
"type": "stdio",
"command": "bunx",
"args": ["@cyanheads/cern-opendata-mcp-server@latest"],
"env": {
"MCP_TRANSPORT_TYPE": "stdio",
"MCP_LOG_LEVEL": "info"
}
}
}
}Or with npx (no Bun required):
{
"mcpServers": {
"cern-opendata-mcp-server": {
"type": "stdio",
"command": "npx",
"args": ["-y", "@cyanheads/cern-opendata-mcp-server@latest"],
"env": {
"MCP_TRANSPORT_TYPE": "stdio",
"MCP_LOG_LEVEL": "info"
}
}
}
}Or with Docker:
{
"mcpServers": {
"cern-opendata-mcp-server": {
"type": "stdio",
"command": "docker",
"args": ["run", "-i", "--rm", "-e", "MCP_TRANSPORT_TYPE=stdio", "ghcr.io/cyanheads/cern-opendata-mcp-server:latest"]
}
}
}For Streamable HTTP, set the transport and start the server:
MCP_TRANSPORT_TYPE=http MCP_HTTP_PORT=3010 bun run start:http
# Server listens at http://localhost:3010/mcpPrerequisites
Bun v1.4.0 or higher (or Node.js v24+).
No API key or account: the CERN Open Data Portal is public.
Installation
Clone the repository:
git clone https://github.com/cyanheads/cern-opendata-mcp-server.gitNavigate into the directory:
cd cern-opendata-mcp-serverInstall dependencies:
bun installConfigure environment:
cp .env.example .env
# optional: adjust transport, logging, or telemetry settingsConfiguration
The server has no settings of its own; these framework variables apply.
Variable | Description | Default |
| Transport: |
|
| HTTP server port. |
|
| HTTP server host. |
|
| HTTP session mode: |
|
| Authentication: |
|
| Log level ( |
|
| Directory for log files (Node.js only). |
|
| Enable OpenTelemetry. |
|
See .env.example for the common framework overrides.
Running the server
Local development
Build and run the production version:
# One-time build bun run rebuild # Run the built server bun run start:http # or bun run start:stdioRun checks and tests:
bun run devcheck # Lints, formats, type-checks, and more bun run test # Runs the test suite
Project structure
Directory | Purpose |
|
|
| Tool definitions ( |
| The |
| Record output schema shared by |
| Portal client (pacing, retries, deadline, caches), normalization, vocabulary tables, text rendering, trigger parsing. |
| Unit and tool tests, mirroring the |
| Tool surface, verified portal behavior, and design decisions. |
Development guide
See CLAUDE.md for development guidelines and architectural rules. The short version:
Handlers throw, framework catches — no
try/catchin tool logicUse
ctx.logfor logging andctx.enrichfor notices and paging contextRegister new tools and resources in the barrels in
src/mcp-server/*/definitions/index.tsWrap external API calls: validate raw → normalize to domain type → return output schema; never fabricate missing fields
Contributing
Issues are welcome. Run checks and tests before submitting:
bun run devcheck
bun run testLicense
This project is licensed under the Apache 2.0 License. See the LICENSE file for details.
This server cannot be deployed
Maintenance
Related MCP Connectors
Query SEC EDGAR filings, XBRL financials, and company data through MCP. STDIO & Streamable HTTP.
Query FDA data on drugs, food, devices, and recalls via openFDA. STDIO or Streamable HTTP.
Governed data discovery, exact queries, decisions, simulations, and runtime utilities over MCP.
Search and retrieve published Alkemata articles, pages, and guidance through a read-only MCP server.
Related MCP Servers
- AlicenseNot gradedqualityAmaintenanceAccess the OpenAlex academic research catalog - 270M+ publications through MCP. Supports STDIO and Streamable HTTP.1,299 npm14Apache 2.0
- FlicenseAqualityBmaintenanceMCP server for querying Huwise/Opendatasoft data portals. Enables dataset search, metadata retrieval, record filtering with ODSQL, and data export.53-
- AlicenseNot gradedqualityAmaintenanceEnables querying UNHCR refugee, IDP, stateless populations, asylum decisions, returns, and resettlement statistics via MCP, including staging large datasets for SQL queries. Supports stdio or Streamable HTTP transports.212 npm1Apache 2.0
- AlicenseNot gradedqualityBmaintenanceEnables searching DataCite datasets and software, fetching DOI metadata, tracing relations, and formatting citations via MCP, supporting stdio or Streamable HTTP.1Apache 2.0