@cyanheads/cern-opendata-mcp-server
MCP serverSearchSearch CERN Open Data, fetch records, files, analysis environments, CMS good-run lists, HLT paths.
Available today. Use it from your connected AI after setup.
No other account needed.
Add ahel to your AI once: Claude, ChatGPT, Cursor, Claude Code or Codex. Then ask it to use this.
Then ask your AI: use the cern opendata search records tool from @cyanheads/cern-opendata-mcp-server
Install @cyanheads/cern-opendata-mcp-server
The server’s own address, for the clients that take one directly. Or connect ahel once and every client you use reads it from one address, with the account kept on ahel rather than in each client’s config.
Claude Code
claude mcp add --transport http --scope user cyanheads-cern-opendata-mcp-serv 'https://cern-opendata.caseyjhand.com/mcp'Run it once in your project, then open /mcp to approve any sign-in the server asks for.
Claude Desktop
https://cern-opendata.caseyjhand.com/mcpAdd a custom connector in Settings, paste this address, and approve the sign-in.
Cursor
cursor://anysphere.cursor-deeplink/mcp/install?name=cyanheads-cern-opendata-mcp-serv&config=eyJ1cmwiOiJodHRwczovL2Nlcm4tb3BlbmRhdGEuY2FzZXlqaGFuZC5jb20vbWNwIn0=Open the link and Cursor adds the server at that address.
ChatGPT
https://cern-opendata.caseyjhand.com/mcpIn Settings, enable Developer mode, create an MCP app, and paste this address. Your plan and workspace must allow custom apps.
Codex
codex mcp add cyanheads-cern-opendata-mcp-serv --url 'https://cern-opendata.caseyjhand.com/mcp'Run it once, then sign in with codex mcp login cyanheads-cern-opendata-mcp-serv if the server asks for an account.
From the project's README
As published by cyanheads/cern-opendata-mcp-server in README.md.
Public Hosted Server: https://cern-opendata.caseyjhand.com/mcp
Overview
Particle-physics data from the CERN Open Data Portal: collision, simulated and derived datasets, analysis software, environments and documentation from ALICE, ATLAS, CMS, LHCb and other experiments. Search with exact-vocabulary filters and live facet counts, open records with license and citation, list data files, assemble a record's analysis environment, and look up CMS good-run lists and trigger paths. Runs as a stdio process, a local Streamable HTTP server, or the public hosted endpoint above.
Tools
| Tool | Description |
|---|---|
cern_opendata_search_records | Search datasets, software, environments, documentation and supplementary records with exact-vocabulary filters and live facet counts |
cern_opendata_get_records | Fetch full metadata for 1–20 records by recid, DOI, CMS dataset path or documentation slug, with license and citation |
cern_opendata_list_files | Page through a record's file indexes and files: XRootD URIs, HTTPS URLs, sizes, checksums, tape availability |
cern_opendata_get_analysis_env | Assemble a record's analysis environment: container images, CMSSW release, global tag, linked environment and software records, guide sections |
cern_opendata_get_validated_runs | Get a CMS validated-run (good-run) list for a dataset, a list or a run period, with luminosity-section ranges |
cern_opendata_search_trigger_paths | Look up CMS High-Level Trigger paths by name or prefix, parsed into run ranges, versions and L1 seeds |
cern_opendata_list_reference | Decode the vocabulary the other tools accept: experiments, record types, energies, formats, physics categories, LHCb stripping, identifiers, query syntax, licensing, run periods |
Resources
| Resource | Description |
|---|---|
cern-opendata://record/{recid} | One record's metadata, license and citation, in the cern_opendata_get_records record shape |
Tool-only clients get the same data from cern_opendata_get_records.
Capability reference
cern_opendata_search_records tool
- Optional
query(an OpenSearchquery_string, up to 500 characters) plus OR-list filterstype,experiment,category(physics category,Higgs Physics::Standard Model),keywords,collision_energy,collision_type,file_type,availability,collection, and the LHCbmagnet_polarity,stripping_streamandstripping_version, each an array or a comma-separated string;year_from/year_toandmin_events/max_eventsbound the data-taking year and event count sort(bestmatch,mostrecentfor newestdate_publishedfirst,title,title_desc),limit1–50 (default 10) andpagefrom 1;page × limitpast 10,000 fails aspage_window_exceeded- Compact
hitswith recids, plus thirteen livefacetsthat each ignore their own filter (typeandcategorywith their secondary values);applied_filtersechoes what ran, and values outside the verified vocabulary appear underunrecognized_values
cern_opendata_get_records tool
ids: 1–20 recids, DOIs, CMS dataset paths (/Primary/Era/TIER) or documentation slugs, mixed in one array or comma-separated string- Each record carries a
licensewith itsbasis(record,cern_terms_default,not_stated) and, when it has a DOI, a readycitation - Documentation and news bodies come in slices of up to 30,000 characters:
body_offsetwith that one id reads on from thebody_next_offsetthe last slice returned - Each response stays within 64,000 bytes: records past the budget are left out whole and listed in
deferred, to pass back asids(the first record always comes back whole) - Records also return what they state of a
variablesdictionary (name, type, unit, description), a physicscategory, pile-up (pileup_html, with the pile-up datasets underlinks),keywords, and the LHCbmagnet_polarityandstrippingstream and version - Unresolved identifiers land in
missingwithinterpreted_asand guidance instead of failing the call; file lists come fromcern_opendata_list_files
cern_opendata_list_files tool
recidrequired; withoutindex, returns the record's file indexes and regular files, and with an index key, that index's files, read without the rest of the recordlimit1–500 (default 50), continued withnext_cursor; each file carriesxrootd_uri,size_in_bytes,checksumandavailability, anhttps_urlwhen one can be built from its key or EOS path, andkeyorfilenameas the portal states them, and each index auri_list_urllisting every XRootD URI in it- Files marked
on demandsit on tape and must be requested on the record's portal page first; an umbrella record with no files of its own returns its child recids underchildren
cern_opendata_get_analysis_env tool
recidrequired;softwarecarries the record's own container images, CMSSW release, global tag and environment recidenvironment_records(condition, VM, validation) for the record's run periods andexample_softwarethat declares it works with the record, up to 50 between them;guidesquotes the linked section of the first two portal guides, each capped at 12,000 characters, and a cut names thecern_opendata_get_recordsbody_offsetthat reads on from it- Always
separately_licensed: true; linked records or guides that can't be read leave anoticeinstead of failing the call
cern_opendata_get_validated_runs tool
- Exactly one of
recid(a CMS collision dataset or a validated-run list) orrun_period(Run2012B;2012Balso matches);variantfullormuons_only;run_min/run_max;limit1–2000 (default 200) - A dataset
recidbounds the runs to the dataset's first and last listed run, echoed inrun_bounds; when several lists match,matched_listsnames them and no runs are read - Each run carries
lumi_sectionsandlumi_ranges;list.https_urldownloads the whole list file. CMS only: other records fail asno_validated_runs
cern_opendata_search_trigger_paths tool
path: an exact name (HLT_IsoMu24,AlCa_EcalPi0) or a prefix with one trailing*(HLT_IsoMu*); a name without theHLT_prefix is matched against record path names in the case given and also searched withHLT_added, and a_v<n>version suffix is dropped; optionalyear,limit1–50 (default 10) andpage- Each per-year record is parsed into the primary
datasetsits title names,first_seen,last_seen, per-version run ranges with theirl1_seed, and HLT menu record links;parsed: falsemarks a record to read from itsabstract_html - CMS open data from 2011–2016 only
cern_opendata_list_reference tool
- Optional
topic:experiments,record_types,collision_energies,collision_types,file_types,availability,categories,lhcb,identifiers,query_syntax,licensingorrun_periods; omit it for every table - Static and offline;
categories,lhcbandrun_periodsare dated snapshots, while the search facets andcern_opendata_get_validated_runsread the live portal
cern-opendata://record/{recid} resource
- One record by
recid(6004, or a prefixed recid such asatlas-160006; leading zeros ignored) asapplication/json, in thecern_opendata_get_recordsrecord shape: metadata (variable dictionary, physics category, pile-up and LHCb run conditions included), license and citation, without file lists recidcomes fromcern_opendata_search_records; reads carry a 15-minute public cache hint
Features
Built on @cyanheads/mcp-ts-core: stdio and Streamable HTTP transports, pluggable auth (none / jwt / oauth), swappable storage (in-memory, filesystem, Supabase, Cloudflare KV/R2/D1), structured logging with optional OpenTelemetry tracing.
CERN Open Data-specific:
- Keyless and read-only; it never stages tape files or writes anything
- One shared pacer at 50 requests a minute, under the portal's published 60 per client IP, and one 50-second deadline per call across queue wait and retries
- Filter values canonicalized against the portal's verified vocabulary (
13 tev→13TeV,lhcb→LHCb,Pb-Pb→PbPb,dataset/collision→Dataset::Collision); unknown values are sent as given and flagged - Tape-resident (
ondemand) records are included in every search and lookup, where the portal otherwise drops them silently - File manifests, and file indexes read on their own, are cached for 15 minutes, so paging through a record's or an index's files costs one read
Agent-friendly output:
- Provenance on every response:
portal_urlon each hit and record, alicensewith itsbasis, a DOIcitation, andapplied_filtersoreffectiveQueryechoing what ran - Graceful partial results: unresolved ids land under
missingwith guidance, and unreadable linked records or guides incern_opendata_get_analysis_envsurface as anoticerather than a failure - Discriminated outputs:
kind,license.basis,interpreted_as,scope,variant,run_bounds.sourceandparsedlet callers branch on data, not string parsing - Portal text kept as data: titles, descriptions, guide sections and file names are fenced or escaped in
content[]and relayed as received (HTML in_htmlfields) instructuredContent
Data and licensing
Portal metadata and datasets are CC0 under the CERN Open Data Terms of Use. Software, container images, documentation and guide code are licensed separately, per record (software is commonly GPL). cern_opendata_get_records reports each record's license and its basis: record when the record states one, cern_terms_default for a dataset that states none (CC0), and not_stated otherwise.
CERN asks reusers to cite each dataset's DOI in applications and publications; cern_opendata_get_records returns a ready citation for every record with a DOI.
This server is an independent project and is not affiliated with or endorsed by CERN.
Known limitations
- 60 requests a minute per client IP. The portal publishes this limit; the server paces itself to 50 a minute, and a call that cannot start within its deadline fails as
rate_limitedwithretryAfter. A hosted deployment shares one budget across every user behind its egress IP.cern_opendata_get_analysis_envandcern_opendata_get_validated_runscost 2–4 requests each. - 10,000-result window. Search and trigger-path paging reach only the first 10,000 matches; narrow deeper result sets with filters.
- Facet lists are partial. Terms facets return the first 10 values alphabetically (
file_typeup to 100), with the rest counted inother_count. - Tape-resident files. Files with availability
on demandmust be requested on the record's portal page before download; staging them is a write and out of scope. A record whose availability isondemandlists none of its files through the API, socern_opendata_list_filesreports only the count and size its metadata states. - CMS-only run lists and triggers. Trigger records cover 2011–2016, and no prescale tables are published, so trigger detail is limited to what each record's abstract states. Muons-only run lists do not exist for Commissioning2010, Run2010B or the 2011 ReReco list.
- Glossary entries are not served. The portal's glossary links answer 404, so search excludes them.
Getting started
Public Hosted Instance
A public instance is available at https://cern-opendata.caseyjhand.com/mcp — no installation required. Point any MCP client at it via Streamable HTTP:
{
"mcpServers": {
"cern-opendata-mcp-server": {
"type": "streamable-http",
"url": "https://cern-opendata.caseyjhand.com/mcp"
}
}
}
Every caller of the hosted instance shares one portal request budget of 50 requests a minute. For sustained use, run your own instance.
Self-Hosted / Local
Add the following to your MCP client configuration file.
{
"mcpServers": {
"cern-opendata-mcp-server": {
"type": "stdio",
"command": "bunx",
"args": ["@cyanheads/cern-opendata-mcp-server@latest"],
"env": {
"MCP_TRANSPORT_TYPE": "stdio",
"MCP_LOG_LEVEL": "info"
}
}
}
}
Or with npx (no Bun required):
{
"mcpServers": {
"cern-opendata-mcp-server": {
"type": "stdio",
"command": "npx",
"args": ["-y", "@cyanheads/cern-opendata-mcp-server@latest"],
"env": {
"MCP_TRANSPORT_TYPE": "stdio",
"MCP_LOG_LEVEL": "info"
}
}
}
}
Or with Docker:
{
"mcpServers": {
"cern-opendata-mcp-server": {
"type": "stdio",
"command": "docker",
"args": ["run", "-i", "--rm", "-e", "MCP_TRANSPORT_TYPE=stdio", "ghcr.io/cyanheads/cern-opendata-mcp-server:latest"]
}
}
}
For Streamable HTTP, set the transport and start the server:
MCP_TRANSPORT_TYPE=http MCP_HTTP_PORT=3010 bun run start:http
# Server listens at http://localhost:3010/mcp
Prerequisites
- Bun v1.4.0 or higher (or Node.js v24+).
- No API key or account: the CERN Open Data Portal is public.
Installation
- Clone the repository:
git clone https://github.com/cyanheads/cern-opendata-mcp-server.git
- Navigate into the directory:
cd cern-opendata-mcp-server
- Install dependencies:
bun install
- Configure environment:
cp .env.example .env
# optional: adjust transport, logging, or telemetry settings
Configuration
The server has no settings of its own; these framework variables apply.
| Variable | Description | Default |
|---|---|---|
MCP_TRANSPORT_TYPE | Transport: stdio or http. | stdio |
MCP_HTTP_PORT | HTTP server port. | 3010 |
MCP_HTTP_HOST | HTTP server host. | 127.0.0.1 |
MCP_SESSION_MODE | HTTP session mode: stateless, stateful, or auto. | stateless |
MCP_AUTH_MODE | Authentication: none, jwt, or oauth. | none |
MCP_LOG_LEVEL | Log level (debug, info, warning, error, etc.). | info |
LOGS_DIR | Directory for log files (Node.js only). | <app-root>/logs |
OTEL_ENABLED | Enable OpenTelemetry. | false |
See .env.example for the common framework overrides.
Running the server
Local development
-
Build and run the production version:
# One-time build bun run rebuild # Run the built server bun run start:http # or bun run start:stdio -
Run checks and tests:
bun run devcheck # Lints, formats, type-checks, and more bun run test # Runs the test suite
Project structure
| Directory | Purpose |
|---|---|
src/index.ts | createApp() entry point: registers the tools and resource, sets the server instructions, starts the portal client. |
src/mcp-server/tools | Tool definitions (*.tool.ts), shared input helpers and list enrichment. |
src/mcp-server/resources | The cern-opendata://record/{recid} resource definition. |
src/mcp-server/record-schema.ts | Record output schema shared by cern_opendata_get_records and the resource. |
src/services/cern-opendata | Portal client (pacing, retries, deadline, caches), normalization, vocabulary tables, text rendering, trigger parsing. |
tests/ | Unit and tool tests, mirroring the src/ structure, run against fixture portal responses. |
docs/design.md | Tool surface, verified portal behavior, and design decisions. |
Development guide
See CLAUDE.md for development guidelines and architectural rules. The short version:
- Handlers throw, framework catches — no
try/catchin tool logic - Use
ctx.logfor logging andctx.enrichfor notices and paging context - Register new tools and resources in the barrels in
src/mcp-server/*/definitions/index.ts - Wrap external API calls: validate raw → normalize to domain type → return output schema; never fabricate missing fields
Contributing
Issues are welcome. Run checks and tests before submitting:
bun run devcheck
bun run test
License
This project is licensed under the Apache 2.0 License. See the LICENSE file for details.
Tools it offers (7)
What this server listed when ahel dialed its public endpoint in Oct 2026, with no key and no account of yours. The names are the server’s own.
cern_opendata_search_recordscern_opendata_get_recordscern_opendata_list_filescern_opendata_get_analysis_envcern_opendata_get_validated_runscern_opendata_search_trigger_pathscern_opendata_list_reference
Signals
- GitHub stars
- 1
- Last commit
- Oct 2026
Advanced
- Delivery
- cern-opendata-mcp-server MCP server → your ahel connector (mcp.ahel.ai) → your AI.
- Item type
- mcp-server
- Key
io-github-cyanheads-cern-opendata-mcp-server- Source
- github.com/cyanheads/cern-opendata-mcp-server
- Hosted endpoint
https://cern-opendata.caseyjhand.com/mcp
github.com/cyanheads/cern-opendata-mcp-server