Skip to main content
Glama
oceanum-io

Oceanum MCP

Official
by oceanum-io

Oceanum MCP

An MCP server package that provides AI assistants with access to the Oceanum platform for ocean/environmental data and cloud storage.

Servers

This package contains multiple MCP servers, selectable at runtime:

Server

Description

datamesh

Search, query, and manage ocean/environmental datasets

storage

List, read, write, and delete files in Oceanum cloud storage

combined

All tools from both servers under a single endpoint (default)

Related MCP server: Code Ocean MCP Server

Prerequisites

Get an API token from oceanum.io. Set it as the DATAMESH_TOKEN environment variable.

Installation

pip install oceanum-mcp

Or run directly with uvx:

uvx oceanum-mcp              # combined server (default)
uvx oceanum-mcp datamesh     # datamesh only
uvx oceanum-mcp storage      # storage only
uvx oceanum-mcp --list       # show available servers

Configuration

Claude Desktop

Add to your claude_desktop_config.json:

macOS: ~/Library/Application Support/Claude/claude_desktop_config.json Windows: %APPDATA%\Claude\claude_desktop_config.json

Combined server (all tools):

{
  "mcpServers": {
    "oceanum": {
      "command": "uvx",
      "args": ["oceanum-mcp"],
      "env": {
        "DATAMESH_TOKEN": "your-token-here"
      }
    }
  }
}

Individual server (datamesh only):

{
  "mcpServers": {
    "oceanum-datamesh": {
      "command": "uvx",
      "args": ["oceanum-mcp", "datamesh"],
      "env": {
        "DATAMESH_TOKEN": "your-token-here"
      }
    }
  }
}

Claude Code

# Combined server
claude mcp add --transport stdio oceanum -- uvx oceanum-mcp

# Individual server
claude mcp add --transport stdio oceanum-datamesh -- uvx oceanum-mcp datamesh

Set the token in your environment:

export DATAMESH_TOKEN=your-token-here

VS Code / Cline / Continue

Use stdio transport with the same command:

{
  "command": "uvx",
  "args": ["oceanum-mcp"],
  "env": {
    "DATAMESH_TOKEN": "your-token-here"
  }
}

Environment Variables

Variable

Required

Description

DATAMESH_TOKEN

Yes

Oceanum API token (shared by all servers)

DATAMESH_SERVICE

No

Custom datamesh service URL (default: https://datamesh.oceanum.io)

STORAGE_SERVICE

No

Custom storage service URL (default: https://storage.oceanum.io)

OCEANUM_DOMAIN

No

Override the base domain for all services (default: oceanum.io)

OCEANUM_MCP_READ_ONLY

No

Set to 1/true to disable write tools (update_metadata, storage write_file/delete_file)

OCEANUM_MCP_MAX_INLINE_BYTES

No

Max staged result size returned inline by query_data (default 50,000,000)

OCEANUM_MCP_MAX_INLINE_ROWS

No

Max rows/records previewed inline before truncation (default 100)

OCEANUM_MCP_EXPORT_DIR

No

If set, export_query may only write inside this directory

OCEANUM_MCP_STAGE_TIMEOUT

No

Seconds a Datamesh staging/query request may wait on the gateway before the tool returns a timeout error (default 120)

OCEANUM_MCP_AUTH

No

Auth scheme for --transport http: auto (default), datamesh, auth0, or none

OCEANUM_MCP_AUTH0_DOMAIN

No

Auth0 tenant domain for auth0 mode (default: auth.oceanum.io)

OCEANUM_MCP_AUTH0_AUDIENCE

No

Auth0 API audience for auth0 mode (default: https://api.oceanum.io)

OCEANUM_MCP_PUBLIC_URL

No

Externally visible base URL (e.g. https://mcp.oceanum.io); enables OAuth discovery metadata for claude.ai connectors

OCEANUM_MCP_CORS_ORIGINS

No

Comma-separated browser origins allowed by CORS on http (default: https://*.oceanum.io,https://*.oceanum.tech,https://vscode.dev; empty string disables)

DATAMESH_TOKEN is required for the stdio transport (and for --transport http with OCEANUM_MCP_AUTH=none); in authenticated http mode each request carries its own credential and no server-side token is needed.

Hosted mode (HTTP transport)

Run any server as a shared, multi-tenant HTTP service:

oceanum-mcp datamesh --transport http --host 0.0.0.0 --port 8000

This serves the MCP streamable-HTTP endpoint at http://<host>:<port>/<server> (/datamesh here; override with --path). Each server owning its own path lets several MCP servers share one domain behind an ingress — e.g. https://mcp.oceanum.io/datamesh and https://mcp.oceanum.io/storage. Every request must present a bearer credential, and all Datamesh/Storage calls are made as that request's user — connections are cached per credential and never shared between tokens.

Auth schemes (OCEANUM_MCP_AUTH):

  • auto (default) — accepts either credential, detected per request: Datamesh tokens arrive in the X-DATAMESH-TOKEN header (the same header the gateway itself uses), Auth0 JWTs in Authorization: Bearer. A Datamesh token sent as the bearer also works (JWT-shape detection routes it), but the dedicated header is the canonical form. Each request runs as the identity its own credential resolves to.

  • datamesh — only Datamesh tokens are accepted. The server validates them against the gateway. Add to Claude Code with:

    claude mcp add --transport http oceanum-datamesh https://mcp.oceanum.io/datamesh \
      --header "X-DATAMESH-TOKEN: <datamesh-token>"
  • auth0 — clients send an Auth0-issued JWT, validated against the tenant JWKS and forwarded to the Datamesh gateway as-is.

  • none — no authentication; every request uses the server's own DATAMESH_TOKEN. Only for trusted-network deployments.

Notes:

  • export_query works over HTTP but changes shape: instead of writing to the server's local filesystem (meaningless for remote clients), it returns a time-limited, self-authenticating gateway download_url that the caller fetches out-of-band. The path argument is ignored on hosted servers. The URL needs no credential, so it is a capability — treat it as a secret. Hosted exports are capped at 10 GB to bound what one link can pull.

  • Combine with OCEANUM_MCP_READ_ONLY=1 to run a read-only public service.

  • Pass --stateless when running behind a load balancer or on autoscaled platforms (Cloud Run, etc.): sessions are otherwise held in instance memory, and consecutive requests routed to different instances would fail.

  • Breaking change for --transport sse (deprecated): sse is a network transport and now behaves like http — authenticated by default and no export_query. Set OCEANUM_MCP_AUTH=none to restore the old unauthenticated behavior.

Serving under an external ASGI server

To run under uvicorn/gunicorn (multiple workers, serverless platforms), use the packaged app factory — it applies the same auth and tool policy as the CLI; serving mcp.http_app() directly would bypass both:

uvicorn --factory oceanum_mcp.app:create_http_app --host 0.0.0.0 --port 8000

Requests on a network transport never fall back to the server's DATAMESH_TOKEN: an unauthenticated request fails unless OCEANUM_MCP_AUTH=none was set explicitly.

Browser clients (CORS)

Browser-based MCP clients send a CORS preflight (OPTIONS) before every cross-origin request. The http app answers preflights for an allowlist of origins, ahead of authentication (a preflight never carries a credential); all other requests still require one. Configure the allowlist with OCEANUM_MCP_CORS_ORIGINS, comma-separated:

OCEANUM_MCP_CORS_ORIGINS="https://*.oceanum.io,https://vscode.dev,http://localhost:6274"
  • Unset: https://*.oceanum.io, https://*.oceanum.tech and https://vscode.dev.

  • Empty string: CORS disabled — no CORS headers, preflights are not answered.

  • Each entry is an exact origin (scheme://host[:port], no path). A leading *. matches exactly one subdomain label: https://*.oceanum.io allows https://app.oceanum.io but not https://oceanum.io, https://a.b.oceanum.io or https://oceanum.io.evil.com. Bare *, single-label wildcards (https://*.io) and malformed entries are rejected at startup; a scheme-default port (:443 on https) is dropped, as browsers omit it. Multi-label public suffixes (https://*.co.uk, https://*.github.io) cannot be detected — never configure them: they would admit every site under that suffix.

  • Credentials travel in Authorization / X-DATAMESH-TOKEN headers, never cookies, so Access-Control-Allow-Credentials is never sent. Mcp-Session-Id and WWW-Authenticate are exposed to browser scripts.

claude.ai custom connectors

Set OCEANUM_MCP_PUBLIC_URL to the server's public base URL and the auto and auth0 modes additionally serve OAuth 2.0 Protected Resource Metadata (RFC 9728) naming the Auth0 tenant, plus a WWW-Authenticate challenge on 401s — which is how claude.ai discovers where to send users to authorize. The Auth0 tenant must allow client registration for Claude (Dynamic Client Registration, or a pre-registered application) and issue JWT access tokens for the configured audience. Header/bearer clients are unaffected either way.

Datamesh Tools

The intended workflow is: search_catalog → get_datasource_info → stage_query (dry run: learn the result size without downloading) → query_data for small results inline, or export_query to write large results to a file that analysis code reads directly.

search_catalog

Search the Datamesh catalog with optional text search, time range, and bounding box filters. Returns a JSON object with count and results; if count equals limit, more matches may exist.

By default each result is a compact summary: id, name, a description cut to 200 characters, tstart/tend, bounds, variable names when the catalog provides them (first 20, with variables_total), and short tags (with tags_total). detail="full" returns each datasource's complete catalog record instead. The total output is bounded (~30k characters for summaries, ~100k for full records): matches beyond the bound are dropped, counted in omitted, and a note asks to refine the search.

Parameter

Type

Description

search

string

Text search for name, description, or tags

time_start

string

ISO 8601 start time

time_end

string

ISO 8601 end time

bbox

list[float]

Bounding box [xmin, ymin, xmax, ymax] in WGS84

limit

int

Max results to return (default 20)

detail

string

summary (default) or full

get_datasource_info

Get the metadata for a datasource: coverage, schema dims and coordinates, variables, and attributes.

The default (detail="summary") is a bounded view with the same fields and shape as the full record. Each variable and coordinate keeps its dims, shape, dtype and units/long_name/standard_name; other attributes are capped at 8 per variable (with attrs_omitted), global attributes and info at 25 entries, and long values are clipped to 200 characters. If the view still exceeds ~60k characters, each variable and coordinate keeps only dims, shape, dtype and units/long_name/standard_name; every variable is listed either way. schema.data_vars/schema.attrs are omitted as duplicates of variables/attributes. detail="full" returns the complete record.

Parameter

Type

Description

datasource_id

string

Datasource ID

detail

string

summary (default) or full

stage_query

Dry-run a query on the Datamesh gateway: reports the result size, container type, and domain length without downloading any data, echoes the canonical query, and recommends the next step (inline query vs export vs narrowing). Accepts the same query parameters as query_data.

query_data

Query a datasource with filters and return small results inline as coordinate-attributed JSON records with explicit truncated/lazy flags. The query is staged first: gridded results above the inline limit are summarized lazily (structure only); tabular results above the limit are refused with the staged size and alternatives. Library warnings (e.g. the 2,000,000-row cap on tabular queries) are included in the response.

Parameter

Type

Description

datasource_id

string

Datasource to query

variables

list[string]

Variables to select

time_start

string

ISO 8601 start of a time range (open-ended if omitted)

time_end

string

ISO 8601 end of a time range (open-ended if omitted)

times

list[string]

Discrete times (series selection); excludes time_start/time_end

time_resolution

string

Server-side temporal downsampling (pandas frequency, e.g. 1h, 1D, 30D; not MS/ME/QS/YS or finer than 1 minute)

time_resample

string

Resampling method for time_resolution: mean, nearest, linear

bbox

list[float]

Bounding box [xmin, ymin, xmax, ymax]

geofilter_feature

object

GeoJSON Feature (Point, MultiPoint, or Polygon) for selection

geofilter_interp

string

Interpolation for feature selection: nearest or linear

geofilter_resolution

float

Max spatial resolution for downsampling, in CRS units

level_min

float

Minimum vertical level

level_max

float

Maximum vertical level

levels

list[float]

Discrete vertical levels (series selection)

level_interp

string

Interpolation for level series: nearest or linear

coord_filters

list[object]

Coordinate selections: [{"coord": "name", "values": [...]}]

crs

string/int

CRS for filter coordinates and returned data

aggregate_operations

list[string]

Aggregation ops: mean, min, max, std, sum

aggregate_spatial

bool

Aggregate over spatial dims (default true)

aggregate_temporal

bool

Aggregate over temporal dims (default true)

limit

int

Keep the last N records (see below)

limit keeps the last N records (Datamesh semantics: the last N steps along time/ensemble), with the same meaning in stage_query, query_data, and export_query. Combined with time_resolution or aggregate_operations it is applied by this server after resampling/aggregation instead of being sent to Datamesh (which would apply it to the native records first), so stage_query reports the unlimited size as an upper bound. Hosted export_query cannot apply it to a gateway download and refuses that combination: drop limit or narrow the time range.

export_query

Export the full result of a query — the data-handle path for results too large to return inline. Behaviour depends on transport:

  • Local (stdio): writes the result to the local file path (required) and returns that path. Gridded datasets stream lazily to NetCDF; tabular results write Parquet or CSV.

  • Hosted (http/sse): returns a time-limited, self-authenticating gateway download_url (choose format; path is ignored). Fetch it out-of-band. Results above 10 GB (for CSV, the estimated text size, about 3x the staged size) are refused (narrow the query); above 2 GB the link is returned with a large-download warning.

Accepts the same query parameters as query_data plus:

Parameter

Type

Description

path

string

Destination file path (parent directories are created)

format

string

netcdf (datasets), parquet or csv (tabular); sensible default

overwrite

bool

Overwrite an existing file (default false)

load_datasource

Summarize an entire datasource. Gridded datasources are opened lazily (no data download); tabular datasources are downloaded only if under the inline size limit.

Parameter

Type

Description

datasource_id

string

Datasource to load

update_metadata

Update metadata on an existing datasource. Only provided fields are changed. Disabled when the server runs with OCEANUM_MCP_READ_ONLY set.

Parameter

Type

Description

datasource_id

string

Datasource to update

name

string

New name

description

string

New description

tags

list[string]

New tags

labels

list[string]

New labels

info

object

Additional metadata object

details

string

URL for datasource details

Storage Tools

list_files

List files and directories in Oceanum cloud storage.

Parameter

Type

Description

path

string

Directory path to list (default: "/")

recursive

bool

List subdirectories recursively

file_exists

Check if a file or directory exists in storage.

Parameter

Type

Description

path

string

Path to check

read_file

Read the contents of a text file from storage.

Parameter

Type

Description

path

string

Path to the file

write_file

Write text content to a file in storage.

Parameter

Type

Description

path

string

Destination path

content

string

Text content to write

delete_file

Delete a file or directory from storage.

Parameter

Type

Description

path

string

Path to delete

recursive

bool

Delete directory contents recursively

file_info

Get metadata about a file or directory.

Parameter

Type

Description

path

string

Path to inspect

Example Workflows

Discover wave data in the Pacific:

  1. search_catalog(search="wave", bbox=[120, -50, 180, 10])

  2. get_datasource_info(datasource_id="some-wave-dataset")

  3. stage_query(datasource_id="some-wave-dataset", variables=["Hs", "Tp"], time_start="2024-01-01", time_end="2024-01-31") to check the result size

  4. query_data(...) with the same parameters if small, or export_query(..., path="waves.nc") if large

Shrink a 40-year hourly time series to something inline-sized:

  1. stage_query(datasource_id="hindcast", variables=["Hs"], time_start="1984-01-01", time_end="2024-01-01") — too large

  2. query_data(..., time_resolution="30D", time_resample="mean") — roughly monthly means, small enough to return inline (calendar aliases such as "1MS" are refused until a Datamesh engine fix ships: the engine currently reads "1MS" as 1 millisecond)

Browse and read files in cloud storage:

  1. list_files(path="/") to see top-level contents

  2. list_files(path="/my-project", recursive=True) to drill down

  3. read_file(path="/my-project/config.json") to read a file

Get a quick summary of a dataset:

  1. get_datasource_info(datasource_id="my-dataset") to see variables and time range

  2. query_data(datasource_id="my-dataset", limit=10) to preview the data

Related MCP Connectors

Related MCP Servers

  • F
    license
    Not graded
    quality
    D
    maintenance
    Access oceanographic and environmental data from 63+ ERDDAP servers worldwide through natural language queries. Search datasets, retrieve metadata, and download scientific data for climate research, marine biology, and coastal management.
    18
    -
  • A
    license
    B
    quality
    B
    maintenance
    Provides tools to search and execute Code Ocean capsules and pipelines while managing platform data assets. It enables users to interact with Code Ocean's computational resources and scientific workflows directly through natural language interfaces.
    26
    5
    MIT
  • A
    license
    Not graded
    quality
    D
    maintenance
    Enables AI agents to interact with the OSDU® data platform through Model Context Protocol, providing tools for search, storage, schema management, and dataset operations.
    1
    Apache 2.0