Skip to main content
Glama
lokah1945

AI Model Catalog MCP Server

by lokah1945
README.md
# AI Model Catalog MCP Server — NVIDIA NIM + Multi-Provider

> **Status: production-ready v2.3.0** — audited end to end on 2026-08-01
> ([`AUDIT_2026-08-01.md`](AUDIT_2026-08-01.md)): 170 unit tests passing,
> ruff clean, `pip-audit` clean, CI green, live smoke-tested on both
> transports. Enterprise hardening: MCP management-tool isolation on HTTP,
> Prometheus label sanitization, constant-time auth, request correlation,
> rate limiting + TRUST_PROXY, body/thread caps, security headers everywhere.

An MCP (Model Context Protocol) server that gives any MCP client (Claude
Desktop / Claude Code, Cursor, IDEs) a live, audited catalog of AI models:

- **NVIDIA NIM** — canonical source `https://build.nvidia.com/models`; fetched
  without an API key, deprecated models excluded, OpenAPI/templates resolved,
  cURL + Node.js + Python examples generated, tiered verification evidence
  stored. The hosted NIM catalog is **100% free to call** with an NVIDIA API key.
- **OpenRouter** — `https://openrouter.ai/api/v1/models` (official API)
- **Nous Portal** — `https://inference-api.nousresearch.com/v1/models` (official API)
- **opencode Zen** — `https://models.dev/api.json` (opencode's official catalog; see `https://opencode.ai/docs/zen`)
- **Blackbox AI** — official docs pricing table (`docs.blackbox.ai`); authenticated `https://api.blackbox.ai/v1/models` when `BLACKBOX_API_KEY` is set
- **Vercel AI Gateway** — `https://ai-gateway.vercel.sh/v1/models` (official API)

## FREE_ONLY mode

`.env` (tracked, no secrets) ships with:

```bash
FREE_ONLY=yes
```

With `FREE_ONLY=yes`, every provider listing — MCP tools, MCP resources, and
HTTP routes — only surfaces **free** models. Paid models cannot be listed *or*
looked up directly (server-side enforcement, not client-side filtering).
NVIDIA NIM is exempt because its hosted catalog is 100% free to call.
Free detection per provider uses each official source: `:free`/`-free` id
suffixes and zero prompt+completion pricing (OpenRouter/Nous/Vercel), zero/absent
cost (opencode Zen), and "Free" pricing rows (Blackbox docs).

API keys never go into `.env` — put them in `.env.local` (gitignored) or real
environment variables. See `.env.example`.

## Quick start (install.sh)

```bash
./install.sh --install    # venv + deps + tests + build catalog + start services
./install.sh --status     # service state, ports, catalog counts
./install.sh --restart    # restart both services
./install.sh --update     # refresh catalog data, then --restart to serve it
./install.sh --logs       # tail service logs
./install.sh --systemd    # systemd units (auto-start on boot, auto-restart)
./install.sh --stop / --uninstall
```

Ports live in `.env` (tracked): `MCP_PORT=9100` (MCP + dashboard),
`PORT=8787` (JSON API), `HOST=127.0.0.1`. Real env vars always win;
put secrets only in `.env.local`.

## MCP server (primary interface)

```bash
make install          # deps
make catalog          # build the NIM + provider catalog (key-less)
make mcp              # stdio (Claude Desktop / Claude Code)
make mcp-http         # Streamable HTTP + dashboard on 127.0.0.1:9100
#   http://127.0.0.1:9100/mcp        MCP protocol endpoint (Streamable HTTP)
#   http://127.0.0.1:9100/dashboard  live HTML dashboard (also served at /)
#   http://127.0.0.1:9100/health     liveness probe
#   http://127.0.0.1:9100/ready      readiness probe
```

Claude Desktop / Code config (`.mcp.json` is included in the repo):

```json
{
  "mcpServers": {
    "nim-catalog": {
      "command": "python",
      "args": ["mcp_server.py", "--database", "data/active_nvidia_nim.sqlite3"]
    }
  }
}
```

Tools: `search_models`, `get_model`, `get_usage_example`, `get_endpoints`,
`recommend_models`, `get_stats`, `list_tiers`, `list_publishers` (NVIDIA NIM) +
`list_providers`, `search_provider_models`, `get_provider_model` (multi-provider).

**Security note:** provider-key management tools (`openrouter_list_keys`,
`openrouter_create_key`, `openrouter_get_key`, `openrouter_delete_key`,
`openrouter_key_usage`, `openrouter_rotate_key`, `get_provider_keys_status`)
exist on the **stdio** transport only. With `--transport http` they are removed
from the registry, so the network-facing MCP endpoint never advertises or
honors key administration. The same routes on the JSON API additionally
require `CATALOG_API_TOKEN` to be set. `X-Forwarded-For` is ignored by default
— set `TRUST_PROXY=1` only when the service sits behind a proxy you control.

## Production deployment

Version **2.3.0** — enterprise-grade serving, stdlib-only hardening:

```bash
docker compose up -d                 # API :8787 + MCP HTTP :9100 (hardened containers)
docker compose run --rm refresher    # scheduled catalog rebuild
```

| Feature | How |
|---|---|
| Liveness / readiness | `GET /health` / `GET /ready` (DB present **and** populated) |
| Prometheus metrics | `GET /metrics` — requests, latency histogram, catalog gauges; label values sanitized against injection |
| Bearer auth | `CATALOG_API_TOKEN=…` (probes stay open; constant-time comparison; key-management routes require the token) |
| Request correlation | `X-Request-ID` (echoed/generated) + `X-Process-Time` headers on every response, mirrored into JSON logs |
| Rate limiting | `RATE_LIMIT_RPM=300` (429 + Retry-After; 0 disables); client identity from socket peer unless `TRUST_PROXY=1` |
| Structured logs | JSON access logs to stderr (`CATALOG_LOG=off` to silence) |
| DoS guards | 1 MiB request-body cap (413), bounded handler-thread pool (64) |
| Graceful shutdown | SIGTERM/SIGINT finishes in-flight requests |
| Security headers | nosniff, DENY, CSP, no-store on every response (JSON API **and** MCP HTTP) |

Operations guide: [`docs/OPERATIONS.md`](docs/OPERATIONS.md) · Security policy:
[`SECURITY.md`](SECURITY.md) · Release notes: [`CHANGELOG.md`](CHANGELOG.md)

## Dashboard

`make dashboard` generates `exports/dashboard.html` — a self-contained
(no-CDN) HTML dashboard with KPIs, NIM verification tiers/modalities/
publishers/transports, per-provider totals vs free with fetch health,
a searchable free-model explorer, a searchable NIM model explorer, and a
data-quality panel mirroring `validate_project.py`. The HTTP service serves
it live at `GET /dashboard` (also `/`):

```bash
make service   # then open http://127.0.0.1:8787/dashboard
```
Resources: `nim://stats`, `nim://models`, `nim://tiers`,
`nim://models/{publisher}/{slug}[/{language}]`, `providers://catalog`,
`providers://{provider}/models`.

## Current audited result

Counts drift as providers add/remove models; the numbers below are from the
last audited run (2026-07-28). Re-run `make catalog` for fresh numbers — the
`/stats` endpoint and `exports/tiered-audit.json` always reflect your local DB.

- NVIDIA NIM: Build index 139 → active non-deprecated **124**
  (57 free-hosted · 96 downloadable · 36 partner)
- OpenRouter: 341 models (18 free) · Nous Portal: 289 (5 free) ·
  opencode Zen: 85 (24 free) · Blackbox AI: 47 (2 free) ·
  Vercel AI Gateway: 307 (3 free)
- Live inference verification requires `NVIDIA_API_KEY` (`make verify`); gateway
  reconciliation requires the same key (`python ingest_gateway.py`).

`data/active_nvidia_nim.sqlite3` contains all active official models, the
multi-provider catalog (`provider_models`), and audit evidence.
`exports/verified-working.sqlite3` contains only latest live-verified hosted models.

## Gateway `/v1/models` reconciliation

```bash
export NVIDIA_API_KEY='nvapi-...'
python ingest_gateway.py --database data/active_nvidia_nim.sqlite3 --output exports
python export_unified.py --database data/active_nvidia_nim.sqlite3 --output unified-output
```

The gateway response currently exposes only four fields: `id`, `object`, `created`, and `owned_by`. The database preserves all of them in `gateway_models`, maps aliases to current Build entries, keeps gateway-only IDs, and enriches mapped records with Build OpenAPI/templates/docs. Query the `all_model_inventory` view for the full union.

## Build without API key

```bash
python -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt
python audit_catalog.py --database data/active_nvidia_nim.sqlite3 build
python providers.py --database data/active_nvidia_nim.sqlite3 --output exports
python audit_catalog.py --database data/active_nvidia_nim.sqlite3 export --output exports
python finalize_audit.py --database data/active_nvidia_nim.sqlite3 --output exports
```

## Optional live verification

```bash
export NVIDIA_API_KEY='nvapi-...'
python audit_catalog.py --database data/active_nvidia_nim.sqlite3 verify --rpm 20 --timeout 120
```

The live-test loop runs under the RPM cap, so a full pass can take a while. To
**resume an interrupted run** (skip models already tested, no re-spent credit):

```bash
python audit_catalog.py --database data/active_nvidia_nim.sqlite3 verify --resume --rpm 20 --timeout 120
```

One Build catalog ID:

```bash
python audit_catalog.py --database data/active_nvidia_nim.sqlite3 verify \
  --model z-ai/glm-5.2 --rpm 20 --timeout 120
```

## Automated update

```bash
chmod +x bin/update-catalog.sh
NIM_VERIFY=0 bin/update-catalog.sh
```

Set `NIM_VERIFY=1` only when live tests and credit use are intended.

The same pipeline is available as Make targets (`make help` for the full list):

```bash
make install      # pip install -r requirements.txt -r requirements-dev.txt -r requirements-mcp.txt
make test         # pytest unit suite (no network/API key needed)
make lint         # ruff lint gate (same as CI)
make catalog      # full key-less pipeline: build -> providers -> upstreams -> finalize -> export -> unified -> validate
make providers    # refresh only the multi-provider catalogs
make dashboard    # regenerate exports/dashboard.html
make service      # read-only catalog HTTP API on 127.0.0.1:8787
make mcp          # MCP server over stdio (Claude Desktop / Claude Code)
make mcp-http     # MCP server + dashboard over Streamable HTTP on 127.0.0.1:9100
make image        # build the catalog_service Docker image
make ci-local     # reproduce the CI key-less pipeline locally
```

## Testing

`tests/` is a fast, deterministic pytest suite (no network, no API key) that
pins the data-correctness invariants: model-ID validation/de-dup, page
classification, served-id selection, smoke-payload construction, upstream
joining, and deprecation date semantics — plus the security regression
surface (FREE_ONLY floor, management-tool isolation, metrics label
sanitization, X-Forwarded-For handling, LIKE escaping, request correlation,
auth/rate-limit paths). **170 tests**, run with `make test`.

## Continuous integration

`.github/workflows/ci.yml` runs the unit tests (Python 3.11/3.12/3.13), the
`pip-audit` dependency CVE gate, then the key-less pipeline (build →
upstreams → finalize → export → unified → validate) on every push/PR and
daily, so the build.nvidia.com contract and the pipeline runtime stay
healthy. The serving smoke test asserts FREE_ONLY over the wire.
No API key is required for CI; live verification stays opt-in via
`NIM_VERIFY=1` locally or `bin/build-strict.sh` with an `NVIDIA_API_KEY`.
Dependabot (`.github/dependabot.yml`) opens weekly PRs for pip and
GitHub Actions updates, each gated by the same CI.

Two additional CI smoke assertions are proposed (management gate → 403
without token; `X-Request-ID` header present) — see `AUDIT_2026-08-01.md` §7.
Apply them with a maintainer token that has the GitHub `workflows`
permission.

## Deployment

`catalog_service.py` ships as a read-only HTTP API. The included `Dockerfile`
serves a pre-built catalog over the API (non-root, healthchecked):

```bash
make catalog                                   # build data/active_nvidia_nim.sqlite3 first
docker build -t nvidia-nim-catalog .
docker run -p 8787:8787 -v "$PWD/data:/app/data:ro" nvidia-nim-catalog
```

## MCP server

`mcp_server.py` exposes the catalog as a **Model Context Protocol** server, so
MCP clients (Claude Desktop, Claude Code, Cursor, IDEs) can discover models,
read served IDs/endpoints, and get runnable cURL/Node.js examples on demand —
no copy-pasting from build.nvidia.com. It is a read-only adapter over the same
shared query layer (`catalog_queries.py`) as the HTTP API, so both transports
behave identically. Build the catalog first (`make catalog`), then run the
server.

Install the MCP dependency (extra):

```bash
pip install -r requirements-mcp.txt   # adds the `mcp` SDK
```

Run (the server reads `data/active_nvidia_nim.sqlite3` read-only):

```bash
make mcp           # stdio transport (default; for Claude Desktop / Claude Code)
make mcp-http      # Streamable HTTP transport on http://127.0.0.1:9100
# or explicitly:
python mcp_server.py --database data/active_nvidia_nim.sqlite3                      # stdio
python mcp_server.py --database data/active_nvidia_nim.sqlite3 --transport http --port 9100
```

### Connect from Claude Desktop (stdio)

Add an entry to `claude_desktop_config.json` (Settings → Developer → Edit Config).
Use the **absolute** path to the repo and (if needed) the absolute path to your
Python interpreter:

```jsonc
{
  "mcpServers": {
    "nim-catalog": {
      "command": "python",
      "args": ["/absolute/path/to/model_fetcher/mcp_server.py",
               "--database", "/absolute/path/to/model_fetcher/data/active_nvidia_nim.sqlite3"]
    }
  }
}
```

### Connect from Claude Code (`.mcp.json` in the project root)

```jsonc
{
  "mcpServers": {
    "nim-catalog": {
      "command": "python",
      "args": ["mcp_server.py", "--database", "data/active_nvidia_nim.sqlite3"]
    }
  }
}
```

### Inspect interactively

```bash
mcp dev mcp_server.py        # opens the MCP Inspector UI
```

### What the server exposes

**Tools** (the assistant calls these):

| Tool | Description |
|---|---|
| `search_models` | Search by text/modality/tier; filters for `working_only`/`free_only`. |
| `get_model` | Full detail: served_id, operations, endpoints, examples. |
| `get_usage_example` | Runnable cURL / Node.js code for a model. |
| `get_endpoints` | Every endpoint with host, transport, and auth requirement. |
| `recommend_models` | Steers to `live_inference_verified` models for production. |
| `get_stats` | Catalog counts, per-tier breakdown, last build time. |
| `list_tiers` | Verification tiers with counts and guarantees. |
| `list_providers` / `search_provider_models` / `get_provider_model` | Multi-provider catalog (OpenRouter, Nous, opencode, Blackbox, Vercel). |
| `openrouter_list_keys` / `create` / `get` / `delete` / `key_usage` / `rotate`, `get_provider_keys_status` | Provider-key administration — **stdio transport only**, never on `--transport http`. |

`serverInfo.version` in the MCP handshake reports the application version
(2.3.0), matching `GET /version` and the `catalog_info` Prometheus metric.

**Resources** (URI-addressable): `nim://stats`, `nim://models`, `nim://tiers`,
`nim://models/{publisher}/{slug}` (full model), and
`nim://models/{publisher}/{slug}/{language}` (a code example). The catalog id is
`publisher/slug`; because it contains a slash, the template splits it into two
segments — e.g. `nim://models/meta/llama-3.1-8b-instruct/curl`.

**Prompt**: `use_nim_model(catalog_id, task)` returns a ready-to-run integration
plan.

### Notes

- The server is **read-only**; it never writes and never runs the build pipeline.
  Refresh the catalog by rebuilding the DB on a schedule (`make catalog` /
  CI). `get_stats` reports the last build time so freshness is transparent.
- `recommend_models` defaults to `production_only=true`, returning only models
  that actually responded over HTTP — set it `false` to also see official
  downloadable/specialized models, clearly tagged.
- **Partial-build safety:** if the DB was produced by `audit_catalog build`
  alone (the tiered audit from `finalize_audit.py` not yet run),
  `get_stats`/`list_tiers` report `audit_populated: false` with a clear warning,
  and `recommend_models` falls back to free-endpoint models instead of silently
  returning an empty list. Run `make catalog` (which includes `finalize`) for
  full tiered data.
- A missing database returns a clear startup error rather than failing silently.

For a complete strict rebuild, including the gateway and gRPC stages:

```bash
chmod +x bin/build-strict.sh
NIM_RPM=20 bin/build-strict.sh
```

The strict pipeline now runs Build discovery, authenticated gateway ingestion and gateway-only tests, Build free-endpoint inference tests, gRPC/NVCF health audit, upstream audit, final audit, unified export, and strict export. Gateway-only records remain visible in the unified inventory but enter strict output only after a real inference pass.

## Verification policy

- `live_inference_verified`: received HTTP 2xx and eligible for working DB.
- `official_contract_downloadable`: active official NIM, but requires local NVIDIA GPU/container deployment.
- `official_contract_specialized_transport`: official Build/API contract verified; live testing needs gRPC, media, private access, or special infrastructure.
- `live_test_failed`: timeout/404/5xx; retained in audit but excluded from working.

Every model has cURL and Node.js records. For gRPC/private/custom transports, records explicitly state the official alternative instead of fabricating an HTTP request.

## Upstream endpoint audit

```bash
python audit_upstreams.py --database data/active_nvidia_nim.sqlite3 --output exports
```

The audit cross-checks stored endpoints against each model's Build OpenAPI servers, resolved code templates, Build guides, and official NVIDIA docs. Supported upstream families include:

- `integrate.api.nvidia.com` — hosted LLM/chat/embedding and selected APIs;
- `ai.api.nvidia.com` — legacy/specialized VLM, retrieval, Cosmos, and AV APIs;
- `health.api.nvidia.com` — BioNeMo/biology APIs;
- `optimize.api.nvidia.com` — optimization APIs such as cuOpt;
- `climate.api.nvidia.com` — climate/weather APIs;
- `grpc.nvcf.nvidia.com:443` — hosted NVCF gRPC requiring API key and function ID;
- `localhost` — self-hosted NIM management/inference endpoints.

`exports/upstream-audit.*` records the result. A Build URL with `documentation-only` means NVIDIA publishes no fixed public upstream (for example private/early-access or internally brokered playgrounds).