Skip to main content
Glama

Part of the Swiss Public Data MCP Portfolio — a collection of open-source MCP servers connecting AI agents to Swiss public and open data. This is a private project. It is not affiliated with, endorsed by, or operated on behalf of any employer or public authority.

lindas-mcp

License: MIT Python 3.10+ MCP Data: LINDAS

MCP server for LINDAS — the linked-data knowledge graph of the Swiss administration.

🇩🇪 Deutsche Version


What LINDAS is

LINDAS (Linked Data Service) is the Swiss Confederation's SPARQL knowledge graph, run by the Federal Archives. Instead of tables, it publishes data as RDF triples: around 2000 statistical data cubes (cube.link) from federal offices, plus the geo-linked data that powers visualize.admin.ch.

Mnemonic: «I14Y is the library catalogue, LINDAS is the library itself.» i14y-mcp tells you a dataset exists. LINDAS holds the data and lets you query across all of it at once.

This server wraps LINDAS in guarded tools rather than exposing raw SPARQL, because the store rewards precise queries and times out on broad ones.


Related MCP server: Zurich Open Data MCP Server

🎯 Anchor Demo Query

«Which forest-fire danger level currently applies, who publishes it, and under which licence?»

search_cubes(query="waldbrand")
  → «Waldbrandgefahr» — BAFU, published

get_cube_structure(cube_uri=...)
  → dimensions: Warnregion (key), Gefahrenstufe (measure)
  → licence: fedlex.data.admin.ch/eli/cc/1984/... (a Fedlex URI!)

query_cube_observations(cube_uri=...)
  → Warnregion: "Dorneck / Thierstein (SO)", Gefahrenstufe: "grosse Gefahr"

The codes come back as labels — «grosse Gefahr», not 4. And the licence is a Fedlex URI you can resolve with fedlex-mcp.

Demo

Demo: Claude using search_cubes, get_cube_structure and query_cube_observations


The two-phase access pattern

LINDAS cubes are self-describing but coded, and the server does two steps to read them:

  1. Structure — the cube's SHACL shape: its dimensions (filterable axes), its measures (the numbers), and which dimensions carry code lists.

  2. Data — the observations, with coded values resolved to human labels using that structure.

Mnemonic: «LINDAS speaks in postcodes, not place names.» An observation says region 1805; the server turns that into «Alpennordhang» for you.

Both steps happen inside query_cube_observations. It fetches the structure itself, so a caller does not have to call get_cube_structure first to get readable rows — doing so only repeats the cube-metadata and dimension queries. get_cube_structure is the tool for finding out what a cube contains (dimension names, key vs. measure, licence) and for writing a run_sparql query against it. It reports only whether a dimension has a code list, not the entries, so it is not a decoding aid either.


Architecture

                 ┌──────────────────────────────┐
                 │      MCP Host (Claude)       │
                 └───────────────┬──────────────┘
                                 │ stdio | streamable-http
                 ┌───────────────▼──────────────┐
                 │          lindas-mcp          │
                 │  ┌────────────────────────┐  │
                 │  │ server.py  (7 tools)   │  │  talks only to cube.py
                 │  ├────────────────────────┤  │
                 │  │ lindas/cube.py         │  │  ← vocabulary guardrail,
                 │  │                        │  │    two-phase access,
                 │  │                        │  │    code→label resolution
                 │  ├────────────────────────┤  │
                 │  │ lindas/queries.py      │  │  SPARQL templates,
                 │  │                        │  │    all anchored on a class
                 │  ├────────────────────────┤  │
                 │  │ lindas/client.py       │  │  raw SPARQL over HTTP,
                 │  │                        │  │    knows nothing of cubes
                 │  └────────────────────────┘  │
                 └───────────────┬──────────────┘
                                 │ HTTPS, no auth
                 ┌───────────────▼──────────────┐
                 │  lindas.admin.ch/query       │
                 │  SPARQL 1.1 · ~2000 cubes    │
                 └──────────────────────────────┘

The lindas/ package is deliberately layered so it can be lifted into other LINDAS-backed servers unchanged. client.py knows only HTTP and SPARQL; cube.py knows the cube.link vocabulary; the tools know only cube.py. Raw SPARQL never reaches the agent except through the guarded run_sparql escape hatch.

Architecture decision

Architecture A (live SPARQL only), with a strict vocabulary guardrail.

Verified live on 2026-07-21:

  • The endpoint is stable, needs no authentication, and returns a clean HTTP 400 with a diagnostic on malformed queries.

  • Blind scans (SELECT *, COUNT(*) over the whole store) time out at 60–90 s; the same question anchored on ?x a cube:Cube answers in ~2 s.

Consequences, baked into the tools:

  • Every query template is anchored on a known class. No unbounded scans.

  • Two-phase access is enforced; the agent never sees raw codes.

  • run_sparql is capped at 500 rows and 30 s and marked as advanced.

  • The client timeout sits at 45 s, in front of the store's own 60–90 s abort.

Full probe report: docs/probe-lindas.md.


Tools

Tool

Purpose

search_cubes

Find cubes by topic. Entry point. Deduplicates versions.

get_cube_structure

Phase 1: dimensions, measures, licence.

query_cube_observations

Phase 2: data points with codes resolved to labels.

list_publishers

Federal bodies publishing cubes, with counts.

resolve_municipality

Name ↔ URI ↔ BFS number — the portfolio join key.

run_sparql

Advanced escape hatch. Capped, guarded.

api_status

Reachability check with cube count.

All tools are annotated readOnlyHint: true.

Reading a search_cubes result

returned is a count of what came back, not a statement about what exists. Two fields say how complete the answer is:

Field

Meaning

truncated

false — every match is in cubes. true — there is more, or there may be more and the server could not rule it out.

total_matched

The exact total when one is available on the same unit as returned, null when no comparable number exists.

Check truncated before concluding anything. On true, do not report the result as complete and do not answer "there are N cubes about X" from returned. Widen in this order: raise limit (up to 100); if it is still truncated at 100, narrow with creator_uri from list_publishers and ask one federal body at a time; use latest_only=False only when you actually want the version history, because it widens the result rather than escaping the cut.

total_matched is null more often than you might expect, and on purpose. With latest_only=true (the default) the count LINDAS can answer cheaply counts cube versions, while the tool returns version-collapsed cubes. Measured on 2026-09-20 with German labels: «wald» matches 127 published versions that collapse to 35 logical cubes, «energie» 33 to 13. A total_matched of 127 next to a returned of 20 would claim 107 missing cubes where at most 15 exist to find — so the field says null, which is true, instead of a number that is not. truncated stays reliable there and is the field to act on.

Where the server has seen every matching row — which is the common case, because it fetches one row beyond what it needs — total_matched is exact in both branches, and truncated is derived from it rather than from the fetch limit.


Installation

uvx lindas-mcp

Claude Desktop

{
  "mcpServers": {
    "lindas": {
      "command": "uvx",
      "args": ["lindas-mcp"]
    }
  }
}

Remote deployment

LINDAS_MCP_TRANSPORT=streamable-http PORT=8000 lindas-mcp

LINDAS_MCP_TRANSPORT accepts stdio (default), streamable-http or sse. The transport decides the path: streamable-http serves /mcp, sse serves /sse. Anything else falls through to stdio, which opens no port at all — in a container that surfaces only as a failing health check. Both HTTP transports bind to HOST, default 127.0.0.1; set HOST=0.0.0.0 explicitly to expose it (only behind a reverse proxy). LOG_LEVEL tunes the JSON stderr logs.

Hosting it as a remote connector

The connector URL is https://<host>/mcp. The path is not configurable — it comes from the transport, and streamable-http is the one a current client expects. sse and its /sse path are the specification's superseded transport; nothing was removed and they still serve, but a new connector should not be pointed at them.

A hosted deployment — Railway, Fly, a container behind any reverse proxy — needs three variables, and the third is the one a deployment inherits wrongly because nothing fails loudly without it:

Variable

Hosted value

If unset

LINDAS_MCP_TRANSPORT

streamable-http

Falls through to stdio: the process starts, opens no port, and surfaces only as a failing health check. The container image already sets it.

HOST

0.0.0.0

Binds loopback only, so the published port reaches nothing. The image sets it deliberately (SEC-016).

LINDAS_MCP_ALLOWED_HOSTS

the public hostname

Host validation is switched off entirely — see below.

LINDAS_MCP_ALLOWED_HOSTS is a comma-separated list of hostnames, without scheme and without port: lindas-mcp.example.ch,alias.example.ch, not https://lindas-mcp.example.ch:443. The value is matched literally against the incoming Host header, and behind TLS on port 443 that header carries no port. Loopback forms are added automatically, so the container health check keeps working.

It fails in two opposite directions:

  • Unset on a non-loopback bind: the protection is off altogether. No hostname is derivable in that situation — the server is reached under a service or public DNS name this process does not know, and a guessed list would answer every real request with 421. So no allow-list is installed at all and the Host header is never checked, which is the SDK's own default. The only sign is a startup warning, dns_rebinding_protection_off.

  • Set to the wrong name: every real request gets HTTP 421. The match is exact and port-precise — mcp.example.ch does not cover Host: mcp.example.ch:8443. If the proxy forwards a non-default port, name both forms.

ALLOWED_ORIGINS is a separate question and concerns browser clients only. It is the CORS origin list, comma-separated, and unset means no browser client is permitted at all — that is the default. A client that is not a browser sends no Origin and is unaffected. * is still accepted and logs a warning. One detail worth knowing if you do serve browsers: the origins derived from LINDAS_MCP_ALLOWED_HOSTS are the http:// ones, so an https:// browser origin has to be named in ALLOWED_ORIGINS yourself.

Docker

docker compose up --build          # binds 0.0.0.0 inside the container, publishes :8000

The image runs as a non-root user, read-only, with resource limits and a TCP health check (see Dockerfile and compose.yaml).


Join keys

LINDAS is a connector layer, and two of its identifiers make it composable with the rest of the portfolio:

Key

Where

Joins to

BFS commune number

resolve_municipality → bfs_number

swiss-statistics-mcp, zurich-opendata-mcp

Fedlex URI

cube licence field

fedlex-mcp

The Fedlex link is the quiet surprise: many cubes declare their licence as a legal-basis URI (fedlex.data.admin.ch/eli/cc/...), so you can go from a data point straight to the law that governs it.


Known limitations

Verified live on 2026-07-21.

  1. Broad SPARQL times out. The store aborts unanchored scans at 60–90 s. The guarded tools avoid this; run_sparql warns about it and caps runtime.

  2. Observations are coded. Dimension values are URIs, not labels. The server resolves them via each dimension's code list, but resolution costs one extra query per coded dimension. Set resolve_labels=False to skip it.

  3. No server-side observation filtering by arbitrary value. LINDAS has no cheap way to filter observations by a dimension value inside a cube, so query_cube_observations reads the first N observations. Analytical slicing belongs in run_sparql.

  4. Licences vary per cube and are declared as dcterms:license, frequently a Fedlex URI rather than a plain name. Always surface the licence field.

  5. Version handling is heuristic. search_cubes deduplicates by stripping the version suffix from the cube URI and keeping the highest schema:version among published cubes. Unusual URI shapes may not collapse cleanly; use latest_only=False to inspect every version. The same heuristic is why total_matched can be null: it is a URI-shape guess in Python, and no cheap SPARQL count expresses it, so restating it in a query would put the same guess in a second place where it can drift.


MCP Protocol Version

This server speaks two protocol eras over the same endpoint. The client's first request on a connection decides which one applies; a later claim from the other era is refused.

Era

Revision

Who reaches it

initialize handshake

2024-11-05 … 2025-11-25

What today's clients speak. The server answers with the revision asked for, or with the 2025-11-25 ceiling when the request asks for something newer.

Per-request envelope

2026-07-28

A request carrying the 2026-07-28 _meta envelope opens a modern connection.

Both revisions are pinned in tests/test_protocol_version.py and asserted against the installed SDK, so a Dependabot bump of mcp cannot move either one silently. The handshake ceiling is measured against a live initialize through the assembled ASGI stack, not read off a constant name.

Note that the SDK's LATEST_PROTOCOL_VERSION is an alias for the modern era, not for the handshake era — pinning against it alone would leave the era that current clients actually negotiate free to drift.

What 2026-07-28 changes here

The modern era has no initialize handshake, and therefore no handshake result in which a client learns who it is talking to. Three consequences are served explicitly rather than left at the SDK's defaults:

Surface

Behaviour

serverInfo

Stamped into _meta on every response and into server/discover. Name, title, version, description and website URL are resolved from the installed distribution's metadata, never written by hand. MCPServer defaults version to "" and substitutes nothing, so an unset version is a required field that says nothing.

instructions

Returned by server/discover — the only orientation channel a modern client has. It names the entry point, says that query_cube_observations resolves labels on its own, and says what get_cube_structure is actually for.

Log delivery

logging/setLevel is gone (SEP-2577); a client opts in per request via the reserved _meta key io.modelcontextprotocol/logLevel. Without it the server sends nothing; with debug it sends one notifications/message per tool call on that request's stream.

tools/list, server/discover

Carry ttlMs 300000 and cacheScope public (SEP-2549).

All of it is measured through the assembled stack in tests/test_spec_2026_07_28.py — both eras, both transports, and both branches of the log opt-in. The tool contract itself is guarded independently by tool-definitions.lock.json (SEC-022), and the modern era is asserted to serve exactly the tools listed there, so the two eras cannot drift into two different servers behind one address.

Update policy. When the gate fails, do not edit the constant blindly: read the spec changelog between the two revisions, verify the server still behaves, then move the constant, this section, README.de.md and CHANGELOG.md together. SDK upgrades are a reviewed change for the same reason: any protocol-affecting bump is called out in CHANGELOG.md.


Testing

PYTHONPATH=src pytest tests/ -m "not live"   # offline, used in CI
PYTHONPATH=src pytest tests/ -m "live"       # hits the real endpoint
python -m ruff check src tests

The live tests earn their place: the observationSet indirection (a cube's observations hang off cube:observationSet, never directly off the cube) is a structural assumption that a mock cannot validate. It is covered by a live test.


Contributing

See CONTRIBUTING.md for the ground rules (read-only, one egress host, anchored queries) and the local dev loop. Further reading: EXAMPLES.md for use cases by audience with the tool-selection table, docs/roadmap.md for the project phase, and PUBLISHING.md for the PyPI / MCP Registry release process.


Security

See SECURITY.md for the security posture and how to report a vulnerability.


License

MIT License — see LICENSE. The LINDAS data remains subject to the licence each publisher declares on the cube.


Author

Hayal Oezkan · github.com/malkreide


Licence: MIT. The cube data remains subject to the licence each publisher declares.


MCP Registry

Ownership marker used by the MCP Registry to link this PyPI package to the GitHub namespace:

mcp-name: io.github.malkreide/lindas-mcp

Available Tools

7 tools
api_statusA
Read-onlyIdempotent

Check whether the LINDAS SPARQL endpoint is reachable.

Returns an evaluable status even on failure, so an agent can tell "no data matched" apart from "the endpoint is down".

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
noteYes
sourceNo
endpointYes
reachableYes
cube_countNo
provenanceNo
retrieved_atYes
last_successful_callNo

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and non-destructive behavior. The description adds crucial failure behavior: even on failure, the status is evaluable, allowing differentiation between data absence and endpoint outage. This is valuable beyond the annotations and contains no contradiction.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences, front-loaded with the core purpose and followed by a concise behavioral note. Every sentence adds value; no fluff or repetition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's simplicity (no parameters), the presence of an output schema, and rich annotations, the description fully covers purpose, failure behavior, and the diagnostic use case. No missing critical information.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With zero parameters, the schema is trivially covered. The baseline for no parameters is 4, and the description need not elaborate on parameter meanings. Nothing additional is required.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb+resource: 'Check whether the LINDAS SPARQL endpoint is reachable.' This unambiguously identifies the tool's purpose and differentiates it from siblings that query data or manage cubes.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explains when the tool is valuable: 'Returns an evaluable status even on failure, so an agent can tell "no data matched" apart from "the endpoint is down".' This provides clear context for use, though it does not explicitly name alternative tools or exclusion criteria.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_cube_structureA
Read-onlyIdempotent

Read a cube's dimensions and measures — what the cube contains.

Tells you which dimensions you can filter on (KeyDimension), which values are measured (MeasureDimension), and which dimensions carry code lists. It also returns the licence, which is frequently a Fedlex URI you can resolve with fedlex-mcp.

Call this to understand a cube before reading it, and to write a run_sparql query against it. It is NOT a prerequisite for readable data: query_cube_observations fetches the structure itself and resolves labels on its own. Note that this result reports only whether a dimension has a code list (has_codelist), not the list's entries — so calling it first does not help you decode raw codes either.

Args: cube_uri: A cube URI from search_cubes. language: Language for dimension names and description.

ParametersJSON Schema
NameRequiredDescriptionDefault
cube_uriYes
languageNode

Output Schema

ParametersJSON Schema
NameRequiredDescription
nameNo
sourceNo
statusNo
licenceNoOften a Fedlex URI — joins to fedlex-mcp.
versionNo
cube_uriYes
dimensionsYes
provenanceNo
descriptionNo
creator_nameNo
retrieved_atYes

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and non-destructive behavior. The description adds meaningful behavioral context beyond that: it returns only whether a dimension has a code list, not the entries, and warns that calling it first does not help decode raw codes. This is useful and does not contradict the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is clearly structured: summary, return contents, usage guidance, caveat, and args. It is slightly repetitive ('what the cube contains' / 'Tells you which...'), but every section adds necessary information and the key functional caveat is highlighted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description is complete for this tool's complexity. It covers what the tool returns, when it is useful, when it is not needed, and important limitations like has_codelist not containing entries. The output schema handles return-format details, and annotations cover the read-only safety profile, so nothing critical is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It does: cube_uri is explained as 'A cube URI from search_cubes', and language is explained as 'Language for dimension names and description.' This adds real meaning beyond the raw schema field names, though it leaves enum values and defaults to the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Read a cube's dimensions and measures'. It then details what is returned (KeyDimension, MeasureDimension, code lists, licence), which clearly distinguishes it from siblings like query_cube_observations and run_sparql.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly states when to call this tool: 'Call this to understand a cube before reading it, and to write a run_sparql query against it.' It also explicitly names the alternative and exclusion: query_cube_observations fetches the structure itself)Skip; and it says the tool does not help decode raw codes.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_publishersA
Read-onlyIdempotent

List the federal bodies that publish cubes, with cube counts.

Returns creator URIs you can pass to search_cubes to restrict a search to one authority.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
sourceNo
returnedYes
provenanceNo
publishersYes
retrieved_atYes

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate read-only and non-destructive behavior, and the description adds valuable context about the return value (creator URIs and cube counts) without contradicting the annotations. It could mention more about open-world semantics, but that is already hinted by annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is only two sentences, front-loaded with the core purpose, followed by a practical usage note. Every sentence earns its place with no wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter tool with a rich output schema and annotations, the description fully covers the purpose, what is returned, and how the output integrates with a sibling tool. It is complete for an agent to select and invoke correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, so there is nothing to document. The baseline score of 4 applies, and the description correctly focuses on the output and usage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states a specific verb ('List') and resource ('federal bodies that publish cubes') and adds that it includes cube counts. It distinguishes itself from siblings by explaining the output can be used with search_cubes.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides explicit guidance on how to use the returned creator URIs with search_cubes, implying a clear use case. It does not explicitly state when not to use it or alternatives, but the context is sufficient.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

query_cube_observationsA
Read-onlyIdempotent

Read the actual data points of a cube, with codes resolved to labels.

This is phase 2. Values are returned keyed by human-readable dimension names, and coded dimension values (e.g. region "1805") are replaced by their labels (e.g. "Alpennordhang") unless you turn that off.

For large cubes this reads only the first limit observations. LINDAS has no cheap way to filter observations server-side by arbitrary dimension value, so heavy analytical slicing belongs in run_sparql.

Args: cube_uri: A cube URI from search_cubes. language: Language for labels. limit: Maximum observations to return (1-500). resolve_labels: Replace coded values with human labels.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
cube_uriYes
languageNode
resolve_labelsNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
sourceNo
licenceNo
cube_uriYes
returnedYes
cube_nameNo
provenanceNo
observationsYes
retrieved_atYes
labels_resolvedYes

TDQS

A4.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate read-only, idempotent, and non-destructive behavior. The description adds valuable context about the limit on observations, label resolution toggle, and the inability to filter server-side, going beyond annotation basics.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-structured with a brief overview, a contextual note about limitations, and an Args list. Every sentence provides essential information without redundancy, making it appropriately concise yet informative.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the presence of an output schema and strong annotations, the description fully covers the tool's purpose, limitations, and usage context. It mentions the 'phase 2' pipeline position and provides the needed alternative for heavy queries, making it complete for invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Although the schema lists parameters, the description gives each parameter semantic meaning: cube_uri from search_cubes, language for labels, limit as observation cap, and resolve_labels toggling label replacement. This clarifies how each parameter affects behavior.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool 'Read the actual data points of a cube' with a specific resource and action. It adds details about label resolution and keying by dimension names, distinguishing it from sibling tools like run_sparql.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly says to use this for reading observations and to use run_sparql for heavy analytical slicing, noting the lack of cheap server-side filtering. This provides clear when-to-use and when-not-to-use guidance with a named alternative.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

resolve_municipalityA
Read-onlyIdempotent

Resolve a Swiss municipality to its LINDAS URI and BFS number.

The BFS commune number is the join key across the whole portfolio: the same number identifies the municipality in swiss-statistics-mcp, zurich-opendata-mcp and any cube that references a place. In LINDAS the URI is literally ld.admin.ch/municipality/.

Args: name_or_bfs: A municipality name ("Zürich") or a BFS number ("261"). language: Language for the name.

ParametersJSON Schema
NameRequiredDescriptionDefault
languageNode
name_or_bfsYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
queryYes
sourceNo
returnedYes
match_typeNo'none' when the name/BFS number resolved to nothing.
provenanceNo
suggestionNoActionable next step when match_type is 'none'.
retrieved_atYes
municipalitiesYes

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already establish the operation as read-only and idempotent. The description adds valuable behavioral context beyond annotations, including the exact URI format ('ld.admin.ch/municipality/<BFS>') and the BFS number's role as a universal identifier. No contradiction exists.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise and well-structured: a clear first sentence states the purpose, a short paragraph provides contextual significance, and an Args list documents parameters. Every sentence earns its place with no redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple resolution tool with read-only annotations and an output schema, the description covers the essential purpose, parameter semantics, and contextual significance. The presence of an output schema means return values need not be detailed in the description.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema has no parameter descriptions (0% coverage), so the description's Args section is essential. It fully explains name_or_bfs with examples ("Zürich" or "261") and clarifies that language specifies the language for the name, compensating completely for the schema's lack of detail.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's function: 'Resolve a Swiss municipality to its LINDAS URI and BFS number.' This provides a specific verb and outcome, and it is distinct from sibling tools like search_cubes or query_cube_observations, which focus on data cubes rather than municipality resolution.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explains the BFS number as the join key across the portfolio, strongly implying when to use this tool for cross-tool consistency. It does not explicitly name alternatives or exclusion scenarios, but the context is clear enough to guide an agent.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

run_sparqlA
Read-onlyIdempotent

Run a raw SPARQL SELECT query. Advanced escape hatch — use sparingly.

Prefer the structured tools. This exists for analytical queries the guarded tools cannot express (cross-cube joins, aggregations, custom filters).

Guardrails, learned from probing: LINDAS times out on unanchored scans, so always anchor on a known class such as ?x a <https://cube.link/Cube>. A bare SELECT * WHERE { ?s ?p ?o } will time out. This tool caps the result at 500 rows and the runtime at 30 seconds.

Args: query: A complete SPARQL SELECT query, including its own PREFIX lines.

ParametersJSON Schema
NameRequiredDescriptionDefault
queryYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
noteYes
rowsYes
sourceNo
row_countYes
provenanceNo
retrieved_atYes

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the annotations (readOnlyHint, idempotentHint), the description discloses result caps (500 rows), runtime limit (30 seconds), and timeout behavior on unanchored scans. This is valuable behavioral context that annotations do not provide, and it does not contradict the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is somewhat lengthy but well-structured: it starts with purpose, then usage guidance, guardrails, and parameter details. Each sentence adds value, though the guardrails paragraph could be slightly more compact without losing important information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the existence of an output schema, the description fully addresses purpose, usage rules, parameter requirements, and operational pitfalls (timeouts, caps). No significant gaps remain for a raw query tool of this complexity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema only defines 'query' as a string, but the description adds critical requirements: 'a complete SPARQL SELECT query, including its own PREFIX lines.' This clarifies what the parameter must contain, though it stops short of providing a full example or specifying SPARQL dialect details.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states 'Run a raw SPARQL SELECT query' and distinguishes this from sibling structured tools by positioning it as an 'advanced escape hatch' for analytical queries they cannot express, such as cross-cube joins and aggregations.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly instructs to 'Prefer the structured tools' and specifies when this tool is appropriate: analytical queries the guarded tools cannot express (cross-cube joins, aggregations, custom filters). It also provides concrete implementation guidance, such as always anchoring on a known class to avoid timeouts.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_cubesA
Read-onlyIdempotent

Find statistical data cubes in LINDAS by topic.

The entry point. Returns cube URIs you then pass to get_cube_structure. By default only the newest published version of each cube is returned; set latest_only=False to see every version.

Check truncated before you conclude anything from the result. returned alone cannot tell a complete answer from a first page.

  • truncated: false — cubes holds every cube that matched. You may say so.

  • truncated: true — more cubes matched than you got back, or that could not be ruled out cheaply. Do NOT report the result as the complete set, and do NOT answer "there are N cubes about X" from returned.

On truncated: true, widen in this order:

  1. Raise limit (up to 100) and search again — usually enough.

  2. Still truncated at 100? Narrow instead of paging: add creator_uri from list_publishers to ask one federal body at a time.

  3. Set latest_only=False if you specifically need historical versions. It widens rather than narrows — every version becomes its own hit — so use it to inspect a cube's history, not to escape truncation.

total_matched carries the exact number when one is available and null otherwise. null is not an error and not zero: with latest_only=true the count the store can give cheaply counts cube versions, while this tool returns version-collapsed cubes (measured: 127 versions collapse to 35 cubes for "wald"), so no comparable number exists. truncated is still reliable there — prefer it over guessing from total_matched.

Args: query: Topic term, e.g. "Wald", "Abfluss", "Energie". Matched against cube names and descriptions in the chosen language. language: Language for names and descriptions. creator_uri: Restrict to one publishing body (from list_publishers). limit: Maximum cubes to return (1-100). latest_only: Collapse versions to the newest per cube.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
queryYes
languageNode
creator_uriNo
latest_onlyNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
cubesYes
queryYes
sourceNo
languageYes
returnedYes
truncatedYesTrue when more cubes match than are returned, or when that could not be ruled out. False means `cubes` holds every match. Errs towards True: a wrong True costs one more query, a wrong False silently hides results.
match_typeNo'none' when nothing matched — distinguishes a real miss from an error.
provenanceNo
suggestionNoActionable next step when match_type is 'none' (e.g. which tool to try).
latest_onlyYes
retrieved_atYes
total_matchedNoTotal cubes matching the query, on the same unit as `returned`. None means the store could not be asked cheaply for a comparable number — with latest_only=true the count LINDAS gives counts cube VERSIONS, not the version-collapsed cubes returned here (measured: 127 versions vs 35 cubes for 'wald'), so a number would be misleading. Read `truncated` instead.

TDQS

A5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate `readOnlyHint`, `openWorldHint`, and `idempotentHint`, but the description adds meaningful behavioral detail beyond them: the semantics of `truncated`, why `total_matched` may be `null`, and the measured version-collapse behavior (127 versions to 35 cubes). These are exactly the kind of non-obvious behaviors an agent must know.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long but every section earns its place: purpose is front-loaded, the `truncated` warning is highlighted and prominent, the troubleshooting list is numbered, and parameter explanations are compact. There is no filler or repetition beyond what is necessary for safe use.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a search/discovery tool with 5 parameters and an output schema, the description covers input semantics, output fields (`cubes`, `truncated`, `total_matched`), critical caveats, and next-step routing to `get_cube_structure`. The presence of an output schema means return-value structure need not be restated, and nothing needed for correct invocation is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must carry the meaning, and it does. Each parameter is explained with purpose and constraints: `query` is matched against names and descriptions, `creator_uri` filters by publisher from `list_publishers`, `limit` has range 1-100, and `latest_only` collapses versions. This fully compensates for the bare schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'Find statistical data cubes in LINDAS by topic.' It also identifies the tool as 'The entry point' and explains that it returns cube URIs that feed into `get_cube_structure`, clearly separating it from sibling tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly says when to use it (entry point for finding cubes) and gives concrete guidance for alternatives: passing URIs to `get_cube_structure`, using `list_publishers` to narrow by `creator_uri`, and avoiding `latest_only=False` as a truncation workaround. This is strong when-to-use guidance with exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 1 tool updatev0.3.0
    • Changedsearch_cubes3 fields changed
      • addedOutput schema / properties / total_matched
        Added value: +{
        +  "anyOf": [
        +    {
        +      "type": "integer"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null,
        +  "description": "Total cubes matching the query, on the same unit as `returned`. None means the store could not be asked cheaply for a comparable number — with latest_only=true the count LINDAS gives counts cube VERSIONS, not the version-collapsed cubes returned here (measured: 127 versions vs 35 cubes for 'wald'), so a number would be misleading. Read `truncated` instead.",
        +  "title": "Total Matched"
        +}
      • addedOutput schema / properties / truncated
        Added value: +{
        +  "description": "True when more cubes match than are returned, or when that could not be ruled out. False means `cubes` holds every match. Errs towards True: a wrong True costs one more query, a wrong False silently hides results.",
        +  "title": "Truncated",
        +  "type": "boolean"
        +}
      • changedOutput schema / required
        Previous value: -[
        -  "retrieved_at",
        -  "query",
        -  "language",
        -  "latest_only",
        -  "returned",
        -  "cubes"
        -]New value: +[
        +  "retrieved_at",
        +  "query",
        +  "language",
        +  "latest_only",
        +  "returned",
        +  "truncated",
        +  "cubes"
        +]
  2. 7 tool updatesv0.2.0
    • First observedapi_status
    • First observedget_cube_structure
    • First observedlist_publishers
    • First observedquery_cube_observations
    • First observedresolve_municipality
    • First observedrun_sparql
    • First observedsearch_cubes

TDQS

A4.6/5.0

Scored across 7 tools

Disambiguation5/5

Each tool owns a clearly distinct step in the workflow: discover cubes, list publishers, inspect structure, read observations, resolve municipalities, check status, or run raw SPARQL. Even the structurally related tools are easy to tell apart because their responsibilities are explicitly separated.

Naming Consistency4/5

Six of seven tool names follow a clear verb_noun snake_case pattern like search_cubes, list_publishers, and get_cube_structure. api_status is the only outlier, being a noun phrase rather than a verb-led name, but the overall style remains predictable.

Tool Count5/5

Seven tools is well-scoped for a read-only statistical data cube server: discovery, structure inspection, data reading, publisher filtering, municipality resolution, status checking, and a raw SPARQL escape hatch. Each tool earns its place without redundancy.

Completeness4/5

The set covers the core cube workflow well: search, publisher filtering, structure inspection, observation reading, and key resolution. Minor gaps like first-class code-list entry enumeration or direct version-history browsing are not offered as dedicated tools, though run_sparql and search flags provide workarounds.

Maintenance

ActivityActive
ResponsivenessUnresponsive

Related MCP Connectors

Related MCP Servers

  • A
    license
    A
    quality
    D
    maintenance
    An MCP server for discovering, downloading, querying, and analyzing datasets from Ontario's open data portals, allowing natural language questions and high-performance analytics via DuckDB.
    23
    1
    MIT
  • A
    license
    B
    quality
    A
    maintenance
    An MCP server providing AI-powered access to Open Data from the City of Zurich, enabling queries to 900+ datasets, real-time environmental and mobility data, geodata, parliamentary proceedings, and more.
    26
    166 PyPI
    8
    MIT
  • A
    license
    A
    quality
    C
    maintenance
    MCP server exposing SPARQL query functionalities for LLMs, enabling query execution, validation, and graph exploration across SPARQL endpoints.
    7
    MIT
  • F
    license
    A
    quality
    C
    maintenance
    Enables LLMs to query structured statistical data from the Swiss Federal Archives' Linked Data platform (LINDAS) by translating natural language questions into SPARQL queries against RDF data cubes.
    8
    -