Skip to main content
Glama

Met research MCP

Public source repository: https://github.com/evenwestvang/met-mcp. The distribution/package name remains met-research-mcp and the Python module is met_mcp.

A read-only, snapshot-backed MCP server for collection research. It exposes six tools without bundling a catalogue, embeddings, images, or model files:

  • collection_capabilities reports the loaded snapshot, fields, rank methods, coverage, unavailable features, and limits.

  • collection_facets lists exact recorded values, optionally within a selection.

  • collection_select applies bounded typed predicates and returns a snapshot-bound selection receipt.

  • collection_rank ranks that exact eligible selection with an explicitly supported anchor, axis, and method.

  • collection_objects hydrates bounded batches of source-scoped object metadata.

  • collection_evidence resolves source and vector evidence IDs.

Selection is deliberately separate from ranking: hard filters and missing/conflict policies determine eligibility first, then ranking orders only those candidates. Metadata-only methods remain available without vectors. Visual ranking scores only objects with compatible stored vectors and reports the rest as unscored; it does not infer style, authorship, culture, provenance, or historical relationships.

The optional paired text path adds topic_text / visual / siglip2_text_cosine_v1 to collection_rank only when a compatible fixed encoder identity and sidecar are configured. It adds no seventh tool and never falls back to another ranker. Clients must inspect capabilities before use.

This is an independent research tool. It is not affiliated with, endorsed by, or an official service of The Metropolitan Museum of Art.

MCP tool API

These are MCP tools invoked with tools/call; they are not six HTTP REST routes. Connect an MCP client to the Streamable HTTP endpoint (normally /mcp), initialize the session, and call collection_capabilities before relying on fields, coverage, or ranking availability. Every successful response includes contract, datasetMode, snapshot, and queryVersion. Tool failures have isError: true and JSON text containing code, message, and supportedAlternatives.

Use the case-sensitive public argument names below. The pinned MCP SDK currently ignores unknown top-level arguments, unlike the stricter checked-in input schemas; do not rely on this to detect typos. Nested predicates, policies, and anchors reject unknown properties.

Tools and responses

collection_capabilities

  • Arguments: none ({}).

  • Response: tools; predicate fields, operators, and limits; ranking.supported, ranking.unavailable, and optional ranking.textEncoder details; unavailable relations; limits for pages, batches, and cache; coverage; and sourceCoverage. Coverage values and vector denominators describe the loaded snapshot, not a universal or hosted dataset.

collection_facets

  • Arguments: field (required facet-field string); selection (optional selection receipt, default null, at most 8,192 characters); prefix (optional string, default null, at most 100 characters); pageSize (integer, default 25, range 1–100); cursor (optional cursor, default null, at most 8,192 characters).

  • Response: field, values (value and count within the selection, or the whole snapshot when no selection is supplied), returnedCount, truncated, and nextCursor. Values are exact recorded strings; a prefix is Unicode/case/diacritic normalized and must match from the start. Repeated values within one record count once.

collection_select

  • Arguments: predicate (required predicate object); policies (optional object; omitted or null uses the defaults; unknown and conflict each default to "exclude" and accept "exclude" or "include"); pageSize (integer, default 25, range 1–100); cursor (optional, default null, at most 8,192 characters).

  • Response: the opaque selection receipt, normalizedPredicate, applied policies, counts (universeCount, eligibleCount, unknownCount, conflictingCount, returnedCount, truncated), ordered objectIds, and nextCursor.

collection_rank

  • Arguments: selection (required receipt, at most 8,192 characters); anchor (required object described below); axis (required enum); method (required enum); pageSize (integer, default 25, range 1–100); cursor (optional, default null, at most 8,192 characters); maxCandidates (integer, default 2,000, range 1–5,000); includeAnchor (boolean, default false); and includeUnscored (boolean, default false).

  • anchor is exactly {"type":"object_id","objectId":<positive integer>} or {"type":"topic_text","topicText":"<1–500 character string>"}. A configured text encoder also rejects blank/non-UTF-8 text and input over its advertised 64-token budget including EOS; it never truncates.

  • axis is parseable as topic, technique_material, context, visual, or style. method is parseable as lexical_bm25_v1, metadata_rule_v1, siglip2_cosine_v1, or siglip2_text_cosine_v1. Parseability does not mean that every combination is supported.

  • Response: rankerVersion, axis, anchor, counts (eligibleCount, excludedAnchorCount, scorableCount, unscoredCount, returnedCount, truncated), scored results, nextCursor, diagnostic unscoredObjectIds (always present, empty unless requested), and provenance. Each result contains objectId, score, scoreComponents, whyRelated, and evidenceIds.

collection_objects

  • Arguments: objectIds, a required array of 1–100 unique positive integers.

  • Response: objects in requested-ID order for IDs that exist, plus missingObjectIds.

Object response fields

Description

objectId, title, artistDisplayName, objectDate

Source object ID and compatible display summaries.

objectURL, primaryImageSmall, primaryImage

Recorded object and image URLs; no image fetch is performed.

isPublicDomain, rightsAndReproduction

Merged rights summaries, which may be unknown or conflicted.

source, fetchedAt, evidenceIds, qualityFlags

Source summary, optional dated API observation, resolvable evidence IDs, and data-quality flags.

artistMentions

Structured name, role, and attribution observations.

originalFields

Preserved compatible fields from the source CSV, or null without a CSV record.

rightsEvidence, imageEvidence

Separate CSV/API rights observations and optional dated API image observations, each with its evidence ID.

Many summary and nested fields may be null; conflicted merged title, artist, date, or rights values are not silently chosen.

collection_evidence

  • Arguments: evidenceIds, a required array of 1–100 unique strings, each 1–300 characters. IDs returned by other calls are source-scoped, such as object:<objectId>:csv, object:<objectId>:api, and vector:<objectId>:siglip2.

  • Response: evidence plus missingEvidenceIds. An evidence item includes evidenceId, category, objectId, source, retrievedAt, snapshotDate, statements, and derivation. Source-record statements preserve field/value observations; vector evidence describes derivation and is not image-rights evidence.

Query fields and predicates

The following are all currently exposed fields. Faceting is deliberately narrower than filtering.

Kind

Fields

Operators

String or string-list

department, object_name, classification, title, metadata_text, culture, medium, tags, artist_name, artist_role, attribution, geography_city, geography_state, geography_country, geography_region, geography_type, data_quality

eq, one_of, contains_normalized, is_missing

Boolean

is_public_domain, has_known_image

eq, is_missing

Date interval

object_date

overlaps, contained_within, is_missing

Facetable

department, object_name, classification, culture, medium, tags, artist_role, geography_country, data_quality

exact recorded values, plus optional normalized prefix

Field names refer to recorded catalogue data: department, object_name, classification, title, culture, and medium are their corresponding source summaries; tags is the recorded tag list. artist_name, artist_role, and attribution address the respective parts of structured artist mentions. The five geography_* fields address city, state, country, region, and geography type. object_date uses the parsed inclusive begin/end interval, while is_public_domain and has_known_image are recorded booleans. data_quality addresses quality flags. metadata_text is a combined search field built from the recorded title, department, object name, classification, culture, medium, object date, tags, artist attributions, and geography values.

A leaf is {"field": ..., "op": ..., "value": ...}. eq takes one string or boolean as appropriate; one_of takes 1–50 nonempty strings. Strings are at most 500 characters. Both operators compare recorded values exactly, including case and punctuation, and list fields match if any member is exact. contains_normalized does a substring match after Unicode NFKD normalization, case folding, removal of combining marks, and whitespace collapsing. It is not an exact facet lookup.

Date operators take {"begin": integer, "end": integer}, inclusive, with -10000 <= begin <= end <= 10000. overlaps accepts touching intervals; contained_within requires the recorded interval to be wholly inside the query. is_missing needs no value. It tests recorded absence, not general data validity: a nonempty but invalid date interval is not missing. A conflicted field remains conflict even for is_missing; inspect its evidence separately.

Predicates are recursively exactly one of a leaf, {"all":[...]}, {"any":[...]}, or {"not":{...}}. all and any have 1–20 children; a tree is limited to depth 5 and 50 leaves. Evaluation uses true, false, unknown, and a separate conflict state. An ordinary comparison against a missing value—and a date comparison against an invalid recorded interval—is unknown. Three-valued negation does not turn absence into a match: NOT unknown remains unknown; similarly, conflict remains conflict. Use is_missing to select recorded absence.

After the complete tree is evaluated, true records are eligible. The unknown and conflict policies independently include or exclude their states; both exclude by default. An all with any false child is false, while an any with any true child is true; otherwise conflict takes precedence over unknown. The response counts are evaluation-stage diagnostics and are not promised to partition the universe.

Ranking, coverage, and continuation

Selection determines the entire eligible set before ranking. The server errors if that exact set exceeds maxCandidates; it never ranks only a prefix. The check uses the default 2,000 or explicit maximum 5,000 before an object anchor is optionally removed. With the default includeAnchor: false, an object anchor that is eligible is omitted and counted in excludedAnchorCount.

These are the ordinary supported triples:

Anchor

Axis

Method

topic_text

topic

lexical_bm25_v1

object_id

topic

lexical_bm25_v1

object_id

technique_material

metadata_rule_v1

object_id

context

metadata_rule_v1

object_id

visual

siglip2_cosine_v1

Only when collection_capabilities.ranking.supported advertises it, one additional triple is available: topic_text / visual / siglip2_text_cosine_v1. Its ranking.textEncoder capability supplies the fixed identity and query limits. This README makes no claim that any hosted or local deployment currently enables it. style is a parseable axis but has no supported triple; visual similarity is not a style substitute.

Lexical results without a metadata match, rule-ranked records without usable rule components, and visual candidates without compatible stored vectors are unscored. A topic-text lexical anchor must contain at least one searchable term. A visual object anchor without its own stored vector is an error. Text-to-image and visual scores cover only candidates with stored vectors, whose snapshot denominator is reported in capabilities. includeUnscored: true exposes the first pageSize unscored IDs as a diagnostic sample. They are not mixed into scored results, and the rank cursor does not advance this sample: subsequent pages repeat it. Inspect scorableCount, unscoredCount, and provenance rather than treating missing scores as zero. Scores are discovery signals, not factual or calibrated relevance claims.

A selection receipt embeds the normalized predicate, policies, query version, and snapshot. It is immutable, snapshot-bound, and recoverable even if the process-local two-entry selection cache evicts it. The receipt represents the whole eligible set, not just the IDs on the current selection page.

  • Select pages are ordered by ascending objectId. Their cursor binds the selection receipt, hence the predicate and policies.

  • Facet pages sort by normalized value, then original value. Their cursor binds the field, selection (or no selection), and the exact supplied prefix.

  • Rank pages sort by descending score, then ascending objectId. Their cursor binds the selection, anchor, axis, method, and includeAnchor; optional text ranking also binds the encoder identity. pageSize, maxCandidates, and includeUnscored are not cursor bindings, but their limits still apply.

All cursors bind the snapshot and operation. Keep each tool's nextCursor separate; a select cursor cannot continue a rank. To continue, repeat the same arguments with cursor set to the previous nextCursor; stop when it is null (omitting it starts at page one again). A mismatched, malformed, incompatible, or stale receipt/cursor errors and never silently rebases. Each page is at most 100; rank pagination still covers only the eligible set admitted under the candidate cap, not a way around that cap.

Illustrative request flow

These are illustrative MCP tool arguments, not executed results. Replace every angle-bracket placeholder with the value returned by the named earlier call; in particular, do not guess object IDs or invent selection, cursor, or evidence receipts for an arbitrary snapshot. A quoted integer placeholder must be replaced by a JSON number, not a numeric string. Empty results are valid; do not send empty object/evidence batches or assume index [0] exists.

  1. Discover the loaded contract with collection_capabilities:

    {}
  2. Discover exact values with collection_facets:

    {
      "field": "department",
      "pageSize": 10
    }
  3. Pass an exact returned facet value to collection_select:

    {
      "predicate": {
        "all": [
          {
            "field": "department",
            "op": "eq",
            "value": "<exact value from collection_facets.values[].value>"
          },
          {
            "not": {
              "field": "medium",
              "op": "contains_normalized",
              "value": "silver"
            }
          }
        ]
      },
      "policies": {
        "unknown": "exclude",
        "conflict": "exclude"
      },
      "pageSize": 25
    }
  4. Check collection_select.counts.eligibleCount first. If it exceeds 2,000, narrow the predicate—using additional facet values or a date interval—and obtain a new selection, or explicitly allow up to 5,000 candidates. A smaller pageSize does not narrow eligibility. Rank the selection with an ordinarily available text/topic triple using collection_rank:

    {
      "selection": "<collection_select.selection>",
      "anchor": {
        "type": "topic_text",
        "topicText": "garden landscape"
      },
      "axis": "topic",
      "method": "lexical_bm25_v1",
      "pageSize": 10,
      "maxCandidates": 2000,
      "includeUnscored": true
    }
  5. Hydrate IDs actually returned in collection_rank.results with collection_objects:

    {
      "objectIds": [
        "<integer objectId from collection_rank.results[0].objectId>"
      ]
    }
  6. Resolve IDs actually returned in rank or object evidenceIds with collection_evidence:

    {
      "evidenceIds": [
        "<evidenceId from collection_rank.results[0].evidenceIds or collection_objects.objects[0].evidenceIds>"
      ]
    }

To continue step 4, use its returned cursor with the same request, for example:

{
  "selection": "<collection_select.selection>",
  "anchor": {"type": "topic_text", "topicText": "garden landscape"},
  "axis": "topic",
  "method": "lexical_bm25_v1",
  "pageSize": 10,
  "maxCandidates": 2000,
  "includeUnscored": true,
  "cursor": "<non-null collection_rank.nextCursor from the previous rank page>"
}

Only if capabilities advertise the exact optional text/visual triple, a new rank request can instead be:

{
  "selection": "<collection_select.selection>",
  "anchor": {"type": "topic_text", "topicText": "blue ocean and boats"},
  "axis": "visual",
  "method": "siglip2_text_cosine_v1",
  "pageSize": 10,
  "maxCandidates": 2000,
  "includeUnscored": true
}

Do not carry a lexical cursor into that visual request or silently fall back to lexical ranking on encoder failure.

The ordinary schema, tool index, and checked-in examples give machine-readable shapes and nested examples. Deployments that advertise paired text ranking use the optional schema and contract delta. Current runtime output also includes the rightsEvidence and imageEvidence object fields described above. The response schemas/examples are intentionally sparse, not exhaustive nested runtime specifications; consult models, tool signatures, and response construction for the remaining details. In particular, the ordinary schema omits the optional text-method enum, accepts neither top-level extras nor policies: null, and does not specify the current object evidence fields. Its minimal capabilities example is not a current availability report; discover capabilities from the server you are using.

Minimal Python SDK call

This uses the official MCP Python SDK against a generic loopback endpoint and reads the bearer credential from a protected file:

import asyncio
import os
from pathlib import Path

import httpx2
from mcp.client.session import ClientSession
from mcp.client.streamable_http import streamable_http_client


async def main():
    url = os.getenv("MET_MCP_URL", "http://127.0.0.1:8000/mcp")
    token = Path(os.environ["MET_MCP_BEARER_TOKEN_FILE"]).read_text().strip()
    headers = {"Authorization": f"Bearer {token}"}
    async with httpx2.AsyncClient(headers=headers, timeout=20) as client:
        async with streamable_http_client(url, http_client=client) as streams:
            async with ClientSession(*streams) as session:
                await session.initialize()
                result = await session.call_tool("collection_capabilities", {})
                if result.is_error:
                    raise RuntimeError(result.content)
                print(result.structured_content)


asyncio.run(main())

Related MCP server: governed-rag-mcp

Install

Python 3.11 is required. requirements.lock contains the exact dependency versions used for the ordinary server, imports, and tests; it is a version pin set, not a hash-locked supply-chain attestation.

python3.11 -m venv .venv
.venv/bin/python -m pip install -r requirements.lock

For a network-isolated install, first obtain all wheels through your own reviewed process, then use the same pins from a local wheelhouse:

.venv/bin/python -m pip install --no-index --find-links /path/to/wheelhouse \
  -r requirements.lock

The source tree runs directly with PYTHONPATH=src; installing this project as a package is optional.

Supply and import data

No data/ directory is included. A serving directory must contain catalogue.sqlite, manifest.json, vector_ids.npy, and vectors.npy as produced by the importer. Metadata records may greatly outnumber available vectors; the generated manifest and collection_capabilities expose the actual denominators.

The primary bounded fixture import uses two publicly obtainable, checksum-verified inputs: the pinned Met Open Access CSV and the pinned first published SigLIP2 shard.

export MET_CSV=/path/to/MetObjects.csv
export MET_VECTOR_SHARD=/path/to/siglip2-00000-of-00052.parquet
export MET_OUTPUT=/path/to/generated/met-mcp-dataset

sha256sum "$MET_CSV" "$MET_VECTOR_SHARD"
PYTHONPATH=src .venv/bin/python -m met_mcp.build_fixture \
  --csv "$MET_CSV" \
  --vectors "$MET_VECTOR_SHARD" \
  --output "$MET_OUTPUT"

With the pinned inputs, this default fixture deterministically selects 1,000 CSV records and stores 200 vectors from the supplied first shard that intersect that selection. It does not provide all 484,956 CSV records or full-vector coverage. The expected hashes, revisions, upstream paths, licenses, and limits are recorded in DATA-SOURCES.md. The importer performs no fetches and rejects the wrong CSV or vector shard.

The public-input test builds that no-seed fixture twice in temporary directories, compares its snapshot, selection, and every derived receipt, then serves it over loopback and exercises all six tools through the official MCP Python SDK:

PYTHONPATH=src:. .venv/bin/pytest -q tests/test_public_import.py \
  --external-csv "$MET_CSV" \
  --external-vectors "$MET_VECTOR_SHARD"

--api-seeds is optional, separately dated enrichment from captures supplied by the user. The importer verifies every declared response hash and keeps API evidence separate from CSV evidence. The historical nine-object capture used by regression tests is not distributed publicly and cannot be reproduced by GETting today's API: current responses are new observations, not the original dated bytes or state. See DATA-SOURCES.md for the manifest format and limitation.

For full pinned-CSV metadata with only the 4,996 compatible vectors joined from the first shard, use the existing larger mode (not exercised by the bounded quickstart or public-input test):

PYTHONPATH=src .venv/bin/python -m met_mcp.build_fixture \
  --dataset-mode csv_baseline \
  --csv "$MET_CSV" \
  --vectors "$MET_VECTOR_SHARD" \
  --output /path/to/generated/met-mcp-csv-baseline

The ancillary met_mcp.full_vectors path is not a portable three-input complete release rebuild. Its import requires all 52 pinned shards plus project-specific census/ID-column receipts, a completed download checkpoint, and the specifically accepted csv-baseline-efac7fc7083c9ee44eb6 base dataset. It is retained for the historical release workflow, not presented as part of this public quickstart.

Run locally

The bearer secret is required. Prefer a protected token file, and keep the default loopback bind unless you have separately designed the network boundary.

export MET_MCP_TOKEN_FILE=/path/to/private/met-mcp-token
install -m 600 /dev/null "$MET_MCP_TOKEN_FILE"
.venv/bin/python -c 'import secrets; print(secrets.token_urlsafe(32))' > "$MET_MCP_TOKEN_FILE"
export MET_MCP_DATA_DIR=/path/to/generated/met-mcp-dataset
export MET_MCP_BEARER_TOKEN_FILE="$MET_MCP_TOKEN_FILE"
export MET_MCP_ALLOWED_HOSTS='127.0.0.1:8000,localhost:8000'
export MET_MCP_ALLOWED_ORIGINS='http://127.0.0.1:3000,http://localhost:3000'
PYTHONPATH=src .venv/bin/python -m met_mcp.cli serve

/healthz and /readyz disclose only status. /mcp requires the exact bearer token, an allowed Host, and (when present) an allowed Origin. The server defaults to one expensive query at a time, a 10-second query budget, bounded bodies, concurrency, and backlog. Generic local-only container configuration is in deploy/README.md.

Optional paired SigLIP2 text encoder

This path requires separately obtained files for google/siglip2-so400m-patch14-384 at revision e8e487298228002f3d8a82e0cd5c8ea9c567f57f. Verify every source file listed in DATA-SOURCES.md, then derive the text-only artifact without network access or modification of the source directory:

PYTHONPATH=src .venv/bin/python scripts/build_siglip2_text_assets.py \
  --source-dir /path/to/verified/full-checkpoint \
  --derived-dir /path/to/derived/text-tower

Install requirements-siglip2-text.lock in a separate Python 3.11 environment from the MCP server and start the sidecar from the derived directory. Plan a separate 4 GiB memory envelope for this optional process and validate capacity on your own host; this is a planning limit, not a host acceptance result. The sidecar forces offline library modes, validates all serving-file hashes against encoder-identity.json, serves on loopback by default, and admits one inference at a time.

SIGLIP2_TEXT_MODEL_DIR=/path/to/derived/text-tower \
  /path/to/text-venv/bin/python scripts/siglip2_text_encoder.py

export MET_MCP_SIGLIP2_ENCODER_URL=http://127.0.0.1:8080
export MET_MCP_SIGLIP2_ENCODER_IDENTITY_FILE=/path/to/derived/text-tower/encoder-identity.json
export MET_MCP_SIGLIP2_ENCODER_TIMEOUT_SECONDS=6

Set those encoder variables in the environment used to launch the MCP server, then start or restart the MCP process. Exporting them in another shell does not modify an already running process. The sidecar and MCP server remain separate environments.

The exact recipe—including lowercasing, 64-token padded input, no attention mask, and L2-normalized SiglipTextModel.pooler_output—is part of the identity. Cosine scores are discovery signals, not calibrated relevance or factual evidence. The published image vectors do not identify their original image-generation checkpoint revision or preprocessing, so cross-modal compatibility remains a documented limit.

Test

Asset-free coverage uses synthetic records only for mocked API capture state:

PYTHONPATH=src:. .venv/bin/pytest -q -m 'not external_data'

That command currently reports 57 passed and 23 deselected. An unqualified run with no external inputs reports 57 passed and 23 skipped. External modes are explicit:

  • --external-data-dir (or MET_MCP_TEST_DATA_DIR) supplies the historical seeded fixture expected by existing service, ranking, boundary, and SDK regression assertions. Those assertions refer to specific historical objects and are not a generic validator for an arbitrary generated dataset.

  • --external-api-seeds (or MET_MCP_TEST_API_SEEDS) supplies the exact dated set for historical manifest/hash regression validation.

  • --external-csv and --external-vectors (or MET_MCP_TEST_CSV and MET_MCP_TEST_VECTORS) enable the public no-seed two-build/SDK proof. Together with --external-api-seeds, they also enable the historical seeded importer regression. pyarrow from requirements.lock is required.

For example, the first two modes run with:

PYTHONPATH=src:. .venv/bin/pytest -q tests \
  --external-data-dir /path/to/historical-seeded-fixture \
  --external-api-seeds /path/to/api-seeds

Tests skip—not silently substitute—any external mode whose inputs are absent. The public-input command above is the portable generated-dataset check; the historical suite requires the undistributed dated fixture and seeds explicitly.

The six-tool schemas and examples are under contract/. Examples that show Met records retain their source labels and are illustrative rather than bundled data. Review THIRD-PARTY-NOTICES.md before redistributing external inputs or generated datasets.

Related MCP Connectors

Related MCP Servers

  • F
    license
    Not graded
    quality
    F
    maintenance
    Enables searching and researching document collections through hybrid semantic search and agentic research queries with grounded, cited answers. It allows users to list collections, scan document sections, and retrieve full Markdown content via MCP-compatible agents.
    34 npm
    -
  • A
    license
    A
    quality
    B
    maintenance
    Provides governed retrieval over MCP with hybrid search, strict confidence gating, and access control, exposing three read-only tools.
    3
    Apache 2.0
  • A
    license
    Not graded
    quality
    C
    maintenance
    Provides read-only, provenance-first repository navigation for agents and humans, with ranked lexical retrieval, exact query, document handles, symbol context, and change impact analysis.
    75 npm
    Apache 2.0
  • F
    license
    Not graded
    quality
    B
    maintenance
    Provides read-only MCP tools to list archived snapshots, retrieve methodology and proof bundles, and verify supplied evidence bundles.
    -