okf-ons
This server provides read-only, metadata-only access to the frozen OKF-ONS corpus, enabling discovery, comparison, and non-executing selection planning for ONS statistics without any network calls or observation data.
okf.descriptor: Return the pinned metadata-only corpus descriptor, scope boundaries, entrypoints, and caveats.
okf.search: Rank compact candidates from a non-empty intent or identifier query, showing nearby alternatives.
okf.get_record: Hydrate one exact frozen metadata record by stable ID, native ID, route, or declared alias.
okf.compare: Produce evidence-backed field contrasts between two to five exact records, without asserting statistical equivalence.
okf.prepare_mcp_plan: Bind a record to its inspection/query tools, audience, purpose, and expiry—preserving unresolved dimensions and never authorising or executing queries.
okf_eval.submit_answer: Validate and return a deterministic, non-persisted evaluation envelope for AI answer submission, separating visible evidence from hidden reasoning.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@okf-onsFind datasets related to unemployment in the UK"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
ONS Open Knowledge Format
okf-ons is a metadata-only discovery layer for public statistics exposed
through Office for National Statistics services. It is designed to help a
person or an agent find the exact dataset, see easily confused alternatives,
understand the statistical-quality evidence that is available, and prepare a
non-executing candidate plan that a downstream service must live-validate and
authorise before retrieval.
The project deliberately separates four jobs:
The static OKF bundle discovers and explains datasets, versions, dimensions, provenance, standards evidence, and alternatives.
OKF Explorer narrows and compares a large static corpus without an LLM or a hosted search service.
The repository's local MCP broker gives AI clients bounded access to the same frozen metadata and prepares a non-executing selection plan.
A downstream live-data MCP server executes an ONS or Nomis query only after the required dataset-specific choices are complete.
The local broker is not the downstream live-data server: it makes no network calls and returns no observation values. No observations, API keys, or private data are stored in the bundle.
OKF 0.2 core and Explorer extensions
The checked-in bundle targets
Open Knowledge Format 0.2.
Its generated portable core starts at
index.md, which declares
okf_version: "0.2" and progressively links to typed concepts for the
catalogue, frozen snapshot, four source lanes, provider state, governance,
standards and MCP selection.
Every concept uses generated and sources. No verified event is claimed:
the automated checks prove structural conformance and frozen-input integrity,
not human review or statistical truth. No stale_after value is invented
while the governed freshness policy remains undefined. Lifecycle status is
therefore explicit and live execution continues to require provider
revalidation.
The large-corpus descriptor, JSON/YAML-LD, static search, facets,
relationships, federation, integrity catalogues and provider datapacks remain
additive Explorer extensions. New concepts also retain the v0.1 timestamp
and # Citations fallbacks so older consumers can continue best-effort
discovery while v0.2 consumers prefer generated and sources.
Related MCP server: ONS + Nomis MCP server
Access and documentation
These are the canonical entry points for the Pages deployment from main.
The immutable application release v0.2.0 predates the OKF specification 0.2
migration and remains available as historical release evidence.
Live and machine-readable access
Subordinate dataset, resource, search and selection-option shards are discovered through the data and search manifests above; their generated paths are not independent stable entry points.
Repositories, release and provenance
Immutable v0.2.0 downloads: bundle ZIP and SHA-256 digest
Source register, OKF publication metadata, release snapshot manifest, current enrichment snapshot manifest, and release changelog
Documentation map
Product and demonstration: demo guide, product contract, ONS hackathon brief, and accessibility statement
Repository operation: repository guide, security policy, release changelog, and CI publication workflow; the AI-client change fragment is retained as pre-release history
Metadata and assurance: metadata model, scope and denominator, and standards register, plus the timed metadata-enrichment campaign and the OKF Explorer campaign handoff, with the machine-readable stopping audit
MCP and AI access: MCP selection contract, MCP client rollout, and agent-access and evaluation proposal
Evaluation: evaluation method, cross-client trial analysis, and the machine-readable study, client registry, tasks, expected results, personas and journeys, issue register, and gold queries
Research and communications: research register, research evidence manifest, architecture manifest, governed-AI research prompt, governed-AI architecture report, and publication draft
Dated AI-client research records: Claude Desktop access trace, Antigravity briefing, and Antigravity postmortem
Portable research outputs: Claude trace DOCX, M365 access briefing DOCX, M365 hosting research DOCX, governed-AI presentation, enterprise-AI presentation, and the discovery-bridge visual
OKF Explorer can load the bundle today through the direct link above. A separate Explorer registry and presentation update is recommended so the bundle is suggested without pasting its URL and the new authority, rights and governance fields are shown as first-class UI rather than only in raw JSON.
SDMX support
The bundle registers SDMX 3.1 (ISO 17369), maps seven canonical OKF fields to
SDMX concepts, and preserves the Nomis lane's SDMX agency, identifier, version,
dimension order, component roles and code-list references. Nomis selections
also retain bounded FREQ code-label options and strict available TIME extents
where that evidence is usable. FREQ is statistical/reference frequency, not
publication cadence, and TIME-period revision annotations are not dataset
revision status. All other required dimensions still need live validation and
selection; completed live execution uses MCP-Geo's nomis_query.
The bundle is not itself an SDMX message. It is serialized as JSON-LD using
DCAT 3, SKOS, PROV-O and RDF Data Cube terms, with no SDMX namespace in its
context. The generated data/standards/sdmx.json reports the exact mapping and
Nomis evidence counts without asserting upstream conformance or certification.
Monday demonstrator
The pre-hackathon demonstrator is frozen so the Monday presentation can use a reproducible, reviewable snapshot rather than changing live acquisitions. It publishes:
a source and coverage ledger that makes incompleteness explicit;
a frozen 5,097-record metadata snapshot from ONS Data API (337), Nomis (1,617), ONS Open Geography (3,035), and ONS Explore Local Statistics (108);
qualified source-producer, service-operator, bundle-publisher and semantic authority roles, with explicit non-endorsement;
deterministic search and first-class filter postings;
evidence-backed alternative/contrast relationships;
exact Census table-code reconciliation across ONS Data API and Nomis;
statistical-quality and standards evidence without false assurance claims;
an MCP selection-plan contract; and
a reproducible retrieval and metadata-quality evaluation report.
The immutable snapshot ID monday-2026-07-17-r2 extends the Friday 17 July
freeze with the pinned ELS metadata projection. The r2 suffix prevents the
expanded corpus from reusing the original v0.1 snapshot identity; “monday” is
the demonstrator name, not a claim that 17 July was a Monday.
The Explore Local Statistics lane is a deterministic, allowlisted projection
of the pinned application submodule. It carries 108 ONS-curated indicators from
24 attributed producers, explains 12 unpublished manifest entries, preserves
46 valid historical aliases, and removes observation values, status arrays,
value domains, binaries and geometry. It is a development fixture rather than
a claim that the internal ELS API is a stable public execution contract. Its
verified Git commitAsOf is kept separate from acquisition time; the frozen
fixture does not invent a retrieval timestamp that was not recorded.
The bundle also publishes a provider datapack that makes the ELS
snapshot/live distinction visible without weakening reproducibility. The
governed snapshot remains pinned to 795eaf2 from 17 July 2026. A separately
reviewed upstream reference at d5f0ac9 records a known example difference:
average house price ends in April 2026 in the snapshot and May 2026 in the
reviewed reference. That reference was last checked on 23 July 2026, is
explicitly external and is not a live validation. Its comparison is a
non-exhaustive reviewed example, not a claim that all 108 indicators were
compared. The descriptor hashes the provider manifest and the manifest hashes
the pack, so the dated review cannot be replaced under the same snapshot ID
without failing Explorer's integrity checks. See the
provider datapack contract.
This experimental bundle is independently published by the OKF ONS project. Source attribution does not imply endorsement by ONS or another producer.
Build
Python 3.11 or later is required. A recursive clone is required because normal GitHub source archives do not contain the pinned ELS submodule. The release bundle ZIP linked above is the self-contained publication artifact.
okf.publication.json records the repository's
publication-method v1 contract: frozen source families, authored and generated
boundaries, dependency planes, reviewed command declarations, documentation
lockstep and Pages target. Command strings in the contract are untrusted
declarations and must be checked against this guide before use. Validate the
local paths, cross-references and acyclic plane graph with:
python scripts/check_publication_contract.pyChanges to controlled source, generator, application or workflow paths must
update the relevant documentation and CHANGELOG.md in the same change.
Unknown paths fail closed and dependency updates have no blanket exemption.
The Pages workflow builds the ignored bundle/ directory once, then validates,
assembles and uploads those same workspace bytes. A clean pre-build --check
is not possible because no generated bundle baseline is tracked. Introducing a
separate immutable baseline or release artefact is the recorded future option;
the workflow does not perform a second identical build merely to compare the
result with itself.
To reproduce v0.2.0 from its checked-in frozen snapshot:
git clone --recurse-submodules --branch v0.2.0 \
https://github.com/chris-page-gov/okf-ons.git
cd okf-ons
python scripts/project_els_snapshot.py \
--submodule-dir vendor/explore-local-statistics-app \
--output output/ons-explore-local-statistics.json
cmp output/ons-explore-local-statistics.json \
source/demo-snapshot/ons-explore-local-statistics.json
python scripts/build_bundle.py \
--snapshot-dir source/demo-snapshot \
--output bundle
python scripts/build_bundle.py \
--snapshot-dir source/demo-snapshot \
--output bundle \
--check
python scripts/check_okf_v02.py bundleThe metadata-enrichment campaign also includes immutable successor snapshot
metadata-enrichment-2026-07-21-r6. Relative to the v0.2.0 snapshot, r3
refreshes the bounded ONS Data API catalogue, r4 adds bounded metadata-only
Nomis compact overviews, and r5 adds only the reviewed dimension projection
from each frozen ONS latest-version URL. r6 follows the exact FREQ and TIME
codelist references in the frozen r5 Nomis cohort and retains only code-label
metadata and explicit period revision-status annotations. The ONS Data API,
Explore Local Statistics and Open Geography source envelopes in r6 are
byte-identical to r5. Rebuild and profile r6 without network access:
python scripts/build_bundle.py \
--snapshot-dir source/metadata-enrichment-2026-07-21-r6 \
--output bundle
python scripts/build_bundle.py \
--snapshot-dir source/metadata-enrichment-2026-07-21-r6 \
--output bundle \
--check
python scripts/profile_metadata_gaps.py \
--bundle bundle \
--output evaluation/metadata-completeness/current.jsonThe live Pages and Explorer URLs above still serve governed release v0.2.0. The r6 campaign snapshot is checked in as governed campaign evidence but remains an undeployed release candidate until the release metadata and Pages-selected snapshot are deliberately switched.
The measured batch history, metric definitions, timings and machine-readable profiles are linked from the metadata-enrichment campaign, including the stopping audit.
For an existing clone, run git submodule update --init --recursive before the
projector. To acquire a future snapshot, keep the raw cache outside the
repository, use --mode refresh, and choose a new immutable identity—never
overwrite an existing snapshot. To refresh only the ONS catalogue while
carrying the other validated source envelopes forward:
python scripts/acquire_snapshot.py \
--cache-dir /path/to/raw-cache \
--output-dir source \
--snapshot-id NEW_UNIQUE_SNAPSHOT_ID \
--base-snapshot source/metadata-enrichment-2026-07-21-r4 \
--source ons-data-api \
--mode refresh \
--page-size 1000 \
--maximum-pages 1 \
--require-completeTo reproduce or refresh the bounded Nomis compact overviews, use the validated pre-enrichment r3 snapshot as the base: replacement envelopes are digest-bound to that exact cohort and enrichment snapshots are not chained as acquisition bases. First acquire into an external replacement envelope, then compose that envelope over r3. The acquisition script independently derives its cohort from the frozen Nomis source and publishes neither raw responses nor cache paths:
python scripts/acquire_nomis_overviews.py \
--snapshot-dir source/metadata-enrichment-2026-07-21-r3 \
--cache-dir /path/to/raw-cache \
--output /path/to/nomis-overviews.json \
--limit 1617 \
--mode refresh \
--request-interval 0.2
python scripts/acquire_snapshot.py \
--cache-dir /path/to/raw-cache \
--output-dir source \
--snapshot-id NEW_UNIQUE_SNAPSHOT_ID \
--base-snapshot source/metadata-enrichment-2026-07-21-r3 \
--replacement-acquisition /path/to/nomis-overviews.json \
--require-completeTo reproduce or refresh the bounded ONS latest-version dimension projection,
use r4 as the validated pre-enrichment base. The acquisition follows only the
337 exact links.latest_version.href values frozen in r4, stores only the
allowlisted projection in an external cache, and does not fetch observations,
dimension options, codelists or downloads:
python scripts/acquire_ons_version_metadata.py \
--snapshot-dir source/metadata-enrichment-2026-07-21-r4 \
--cache-dir /path/to/external-cache \
--output /path/to/ons-version-metadata.json \
--limit 337 \
--mode refresh \
--request-interval 0.5
python scripts/acquire_snapshot.py \
--cache-dir /path/to/external-cache \
--output-dir source \
--snapshot-id NEW_UNIQUE_SNAPSHOT_ID \
--base-snapshot source/metadata-enrichment-2026-07-21-r4 \
--replacement-acquisition /path/to/ons-version-metadata.json \
--require-completeTo reproduce or refresh the bounded Nomis FREQ/TIME codelist projection, use r5 as the exact base. The acquisition follows 3,234 frozen metadata-only codelist references sequentially, retains explicit failed outcomes in the denominator, and stores only its allowlisted projection in a versioned external cache:
python scripts/acquire_nomis_codelists.py \
--snapshot-dir source/metadata-enrichment-2026-07-21-r5 \
--cache-dir /path/to/external-cache \
--output /path/to/nomis-codelists.json \
--limit 1617 \
--mode refresh \
--request-interval 0.2
python scripts/acquire_snapshot.py \
--cache-dir /path/to/external-cache \
--output-dir source \
--snapshot-id NEW_UNIQUE_SNAPSHOT_ID \
--base-snapshot source/metadata-enrichment-2026-07-21-r5 \
--replacement-acquisition /path/to/nomis-codelists.json \
--require-completeFor a full registered-source refresh, first produce the pinned local ELS projection, then acquire all HTTP lanes and include it:
python scripts/project_els_snapshot.py \
--submodule-dir vendor/explore-local-statistics-app \
--output /path/to/els-projection.json
python scripts/acquire_snapshot.py \
--cache-dir /path/to/raw-cache \
--output-dir /path/to/public-snapshots \
--snapshot-id NEW_UNIQUE_SNAPSHOT_ID \
--mode refresh \
--projected-acquisition /path/to/els-projection.json \
--require-completeAcquisition is resumable and external. bundle/ is deterministic from a
validated frozen snapshot.
The sendable ONS hackathon brief explains the work package, demo route and questions for ONS. See also the standards register, evaluation method, and scope and denominator.
AI access research and evaluation
The repository preserves and SHA-256-pins 18 July 2026 trials from Claude
Desktop Cowork, Google Antigravity CLI and Microsoft 365 Copilot Researcher in
research/. Only technically reviewed, sanitized public
DOCX derivatives are published. The original Office packages are withheld
outside Git and identified only by their pinned source hashes. Normalized
observations keep public derivatives, session-reported identity, later
verification and our interpretations distinct.
The case study now seeds a provider-neutral, fixture-safe harness:
python3 scripts/ai_client_harness.py validate
python3 scripts/ai_client_harness.py validate-research
python3 scripts/ai_client_harness.py clients --initial-only
python3 scripts/ai_client_harness.py probe-clients
python3 scripts/ai_client_harness.py tasksThe harness defines OKF, open-web and raw-API arms; eight development smoke tasks; six ONS-specific personas and eight journeys; a ten-item observed-issue register; a common structured answer; answer-bound independent assessment; deterministic component scoring; enforced-versus-self-reported separation; and separate readiness outcomes so a blocked client is not scored as a wrong answer. CI performs no live, authenticated or paid model calls.
A dependency-free, metadata-only MCP broker now exposes deterministic
descriptor, search, exact-record, comparison, read-only MCP-plan and
answer-submission tools. Antigravity CLI (agy) can load it from the
repository-scoped .agents/mcp_config.json. That
configuration was added after the preserved Antigravity trial; no new paid or
live model run is claimed. The broker uses the frozen corpus, makes no network
call, stores no credentials and returns no observation values.
AI-system connection map
The MCP rollout guide is the canonical setup and status guide for every AI surface in the evaluation registry. It distinguishes four access and deployment layers:
Static bundle access: any permitted HTTP client can read the Pages descriptor and JSON, but native fetch limits differ by host.
Local metadata MCP: command-line and desktop clients that support local stdio MCP can start
scripts/okf_ons_mcp.py. The committed AGY workspace configuration is the only repository-scoped client configuration; the guide gives the separate Codex, Claude, Gemini, VS Code/Copilot and Inspector routes without treating installation as a successful trial.Planned remote metadata MCP: ChatGPT custom apps need a supported remote endpoint or tunnel route; Claude Research, remote-session Cowork and Microsoft 365 Copilot Researcher need an authenticated remote endpoint. Desktop-local Cowork may use local MCP, subject to device policy. This repository deploys neither a remote endpoint nor a tunnel. The M365 federated-connector route is documented, but it has not been created or enabled for a tenant.
Downstream live MCP: once selection is complete, a separate ONS or Nomis integration may execute the plan. The repository broker deliberately never does so.
The machine-readable
client-profiles.json is the
complete 16-profile client registry; study.json
pins the evaluation design. The public
tasks.json,
personas-and-journeys.json
and issue-register.json keep the
tasks, users, journeys, observed failures and remediations in lockstep with
that design.
See the cross-client trial analysis for the evidence, issue/remediation matrix, personas, efficiency protocol and prioritized work packages.
What “all” means
The repository may describe itself as covering all public ONS metadata only
when every in-scope source in source/source-register.json has a measured
denominator and the generated coverage ledger reports
unexplained_omissions = 0.
Observation cells, secure microdata, and duplicate binary payloads are intentionally out of scope. See scope and denominator.
Licensing
Code is MIT licensed. Source metadata remains subject to its stated upstream
licence, often the Open Government Licence v3.0. The bundle makes no blanket
licence claim: ELS record rights remain not-evaluated pending source-by-source
review, and generated records preserve the available source and rights evidence.
Available Tools
4 toolsokf.compareCompare exact ONS metadata recordsARead-onlyIdempotent
Return evidence-backed field contrasts for two to five exact records. Never asserts statistical equivalence.
| Name | Required | Description | Default |
|---|---|---|---|
| record_ids | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and idempotentHint=true. The description adds the behavioral caveat 'Never asserts statistical equivalence,' which goes beyond annotations. However, it does not disclose error handling or behavior with missing records.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two concise sentences with the primary action front-loaded. Every sentence adds value, with no redundant information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the simplicity of the tool and rich annotations, the description is adequate but could be more complete by explaining the output format or providing an example. Parameter explanation is minimal.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description must compensate. It adds the context of 'two to five exact records' reinforcing minItems/maxItems constraints but does not explain the format or meaning of the record_ids strings beyond what the schema provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool compares exact ONS metadata records by returning field contrasts. It specifies the range (2-5 records) and differentiates from siblings like okf.get_record which retrieves a single record.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for comparing two to five exact records but does not explicitly state when to use this tool versus alternatives like okf.get_record. It provides no when-not-to-use guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
okf_eval.submit_answerNormalise an AI evaluation answerARead-onlyIdempotent
Validate and return a deterministic, non-persisted evaluation envelope. Accepts visible answers, evidence and caveats, never hidden reasoning.
| Name | Required | Description | Default |
|---|---|---|---|
| answer | Yes | ||
| task_id | Yes | ||
| evidence | Yes | ||
| mcp_plan | No | ||
| contrasts | No | ||
| caveat_ids | Yes | ||
| confidence | Yes | ||
| alternatives | Yes | ||
| substitution | No | ||
| chosen_record_id | Yes | ||
| considered_record_ids | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description adds context beyond annotations: it specifies the tool is non-persisted, deterministic, and never accepts hidden reasoning. Annotations already mark it as readOnlyHint=true, idempotentHint=true, destructiveHint=false, so the description reinforces and elaborates on these traits, providing useful behavioral clarity.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence that effectively front-loads the core purpose (validate, deterministic, non-persisted) and lists accepted inputs. Every phrase adds value with no redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the complexity (11 parameters, 8 required, no output schema, nested objects), the description is insufficient. It does not explain the structure of the return envelope, the meaning of required parameters, or how inputs like task_id, chosen_record_id, and considered_record_ids should be used. The absence of output schema makes this gap more critical.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description carries full responsibility for explaining parameters, but it only mentions 'visible answers, evidence and caveats,' corresponding to answer, evidence, caveat_ids. Eleven parameters exist, including key ones like task_id, chosen_record_id, considered_record_ids, alternatives, mcp_plan, contrasts, substitution, confidence, all unexplained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's function: 'Validate and return a deterministic, non-persisted evaluation envelope.' It specifies what it accepts (visible answers, evidence, caveats) and excludes (hidden reasoning), distinguishing it from siblings like okf.get_record, okf.compare, and okf.prepare_mcp_plan, which serve different purposes.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when or when not to use this tool versus alternatives. No explicit context for usage, prerequisites, or exclusions are mentioned, leaving the agent to infer when submit_answer is appropriate given the sibling tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
okf.get_recordHydrate one exact ONS metadata recordARead-onlyIdempotent
Resolve an exact stable ID, native ID, route or declared evaluation alias and return the complete frozen metadata record.
| Name | Required | Description | Default |
|---|---|---|---|
| identifier | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already provide readOnlyHint and idempotentHint. Description adds 'frozen' to indicate immutability but does not disclose error behavior (e.g., what happens if identifier not found). Minimal additional transparency.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Single sentence, no fluff, front-loaded with action and resource. All words carry meaning.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema; description does not explain return structure beyond 'complete frozen metadata record'. Lacks details on pagination, error responses, or data format. Adequate but not complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema has no description for 'identifier' (0% coverage). Description compensates by listing acceptable identifier types: 'exact stable ID, native ID, route or declared evaluation alias', which adds meaningful context beyond the raw schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description uses specific verb 'Resolve' and lists exact resource types (stable ID, native ID, route, alias) and output ('complete frozen metadata record'). Clearly distinguishes from sibling tools like okf.compare or okf_eval.submit_answer.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No explicit guidance on when to use this tool versus alternatives. While the purpose implies single-record retrieval, it does not state when not to use it or mention alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
okf.prepare_mcp_planPrepare a non-executing MCP selection planBRead-onlyIdempotent
Bind the exact frozen identity and snapshot digest to its inspection/query tools, audience, purpose and caller-declared expiry. Preserve unresolved dimensions and reject credential fields. Never authorises or executes a query.
| Name | Required | Description | Default |
|---|---|---|---|
| purpose | No | ||
| audience | No | ||
| arguments | No | ||
| record_id | Yes | ||
| expires_at | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint. The description adds behavioral context by stating it preserves unresolved dimensions, rejects credential fields, and never authorizes or executes, which goes beyond the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is brief (two sentences) but the first sentence is dense with jargon, making it less clear. It lacks a clear breakdown of what the tool does in simpler terms.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the complexity (5 parameters, no output schema, nested objects), the description lacks details on the return value or how the plan is used afterward. It does not cover what the output contains, leaving the agent with insufficient context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description refers to 'audience, purpose and caller-declared expiry,' linking to three parameters, but does not explain 'arguments,' 'record_id,' or their semantics. With 0% schema coverage, the description should provide more parameter detail.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states it binds a frozen identity and snapshot digest to tools, audience, purpose, and expiry, and emphasizes it never executes. The title 'Prepare a non-executing MCP selection plan' reinforces this. However, jargon like 'frozen identity' and 'snapshot digest' may not be universally understood, slightly reducing clarity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies it should be used for preparing a plan without execution ('Never authorises or executes a query') and mentions preserving unresolved dimensions and rejecting credentials. It does not explicitly contrast with siblings like okf.get_record or okf_eval.submit_answer, but the distinct roles are inferable.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
4 tool updates
v0.2.0- First observed
okf_eval.submit_answer - First observed
okf.compare - First observed
okf.get_record - First observed
okf.prepare_mcp_plan
TDQS
Scored across 4 tools
Each tool targets a distinct operation: record retrieval, comparison, plan preparation, and answer submission. No functional overlap exists.
Names mix dot notation (okf.get_record) and underscore (okf_eval.submit_answer), with inconsistent prefixes (okf vs okf_eval) and verb forms (compare vs prepare_mcp_plan).
4 tools is on the lower side but may be appropriate for a narrow domain. However, the descriptions suggest a more complex system that could benefit from additional tools.
Missing fundamental operations such as creating, updating, listing, or searching records. The tool set feels incomplete for the implied workflow.
Maintenance
Related MCP Connectors
UK Office for National Statistics dataset catalogue + Beta JSON API
Search, sample and query open reproducible datasets published as immutable Parquet with schemas.
Search and query 1,500+ OECD statistical datasets via SDMX. Keyless.
Curated gateway to snapshot-versioned Canadian public data services with source provenance.
Related MCP Servers
- AlicenseNot gradedqualityCmaintenanceEnables querying UK Office for National Statistics datasets and their editions through natural language, with no authentication required.1 npmMIT
- AlicenseAqualityBmaintenanceProvides access to UK economic, social, and labour-market statistics via the ONS beta and Nomis APIs, enabling discovery and retrieval of time series, datasets, and census data.18MIT
- AlicenseAqualityBmaintenanceEnables users to discover and recommend Korea public datasets by describing their project idea, searching a static catalog of 96,056 datasets without API calls.715MIT
- AlicenseNot gradedqualityCmaintenanceProvides access to DataCite DOIs for research datasets, enabling searching and retrieval of dataset metadata.2 npmMIT