AI Model Catalog MCP Server
by lokah1945
README.md
# AI Model Catalog MCP Server — NVIDIA NIM + Multi-Provider
> **Status: production-ready v2.3.0** — audited end to end on 2026-08-01
> ([`AUDIT_2026-08-01.md`](AUDIT_2026-08-01.md)): 170 unit tests passing,
> ruff clean, `pip-audit` clean, CI green, live smoke-tested on both
> transports. Enterprise hardening: MCP management-tool isolation on HTTP,
> Prometheus label sanitization, constant-time auth, request correlation,
> rate limiting + TRUST_PROXY, body/thread caps, security headers everywhere.
An MCP (Model Context Protocol) server that gives any MCP client (Claude
Desktop / Claude Code, Cursor, IDEs) a live, audited catalog of AI models:
- **NVIDIA NIM** — canonical source `https://build.nvidia.com/models`; fetched
without an API key, deprecated models excluded, OpenAPI/templates resolved,
cURL + Node.js + Python examples generated, tiered verification evidence
stored. The hosted NIM catalog is **100% free to call** with an NVIDIA API key.
- **OpenRouter** — `https://openrouter.ai/api/v1/models` (official API)
- **Nous Portal** — `https://inference-api.nousresearch.com/v1/models` (official API)
- **opencode Zen** — `https://models.dev/api.json` (opencode's official catalog; see `https://opencode.ai/docs/zen`)
- **Blackbox AI** — official docs pricing table (`docs.blackbox.ai`); authenticated `https://api.blackbox.ai/v1/models` when `BLACKBOX_API_KEY` is set
- **Vercel AI Gateway** — `https://ai-gateway.vercel.sh/v1/models` (official API)
## FREE_ONLY mode
`.env` (tracked, no secrets) ships with:
```bash
FREE_ONLY=yes
```
With `FREE_ONLY=yes`, every provider listing — MCP tools, MCP resources, and
HTTP routes — only surfaces **free** models. Paid models cannot be listed *or*
looked up directly (server-side enforcement, not client-side filtering).
NVIDIA NIM is exempt because its hosted catalog is 100% free to call.
Free detection per provider uses each official source: `:free`/`-free` id
suffixes and zero prompt+completion pricing (OpenRouter/Nous/Vercel), zero/absent
cost (opencode Zen), and "Free" pricing rows (Blackbox docs).
API keys never go into `.env` — put them in `.env.local` (gitignored) or real
environment variables. See `.env.example`.
## Quick start (install.sh)
```bash
./install.sh --install # venv + deps + tests + build catalog + start services
./install.sh --status # service state, ports, catalog counts
./install.sh --restart # restart both services
./install.sh --update # refresh catalog data, then --restart to serve it
./install.sh --logs # tail service logs
./install.sh --systemd # systemd units (auto-start on boot, auto-restart)
./install.sh --stop / --uninstall
```
Ports live in `.env` (tracked): `MCP_PORT=9100` (MCP + dashboard),
`PORT=8787` (JSON API), `HOST=127.0.0.1`. Real env vars always win;
put secrets only in `.env.local`.
## MCP server (primary interface)
```bash
make install # deps
make catalog # build the NIM + provider catalog (key-less)
make mcp # stdio (Claude Desktop / Claude Code)
make mcp-http # Streamable HTTP + dashboard on 127.0.0.1:9100
# http://127.0.0.1:9100/mcp MCP protocol endpoint (Streamable HTTP)
# http://127.0.0.1:9100/dashboard live HTML dashboard (also served at /)
# http://127.0.0.1:9100/health liveness probe
# http://127.0.0.1:9100/ready readiness probe
```
Claude Desktop / Code config (`.mcp.json` is included in the repo):
```json
{
"mcpServers": {
"nim-catalog": {
"command": "python",
"args": ["mcp_server.py", "--database", "data/active_nvidia_nim.sqlite3"]
}
}
}
```
Tools: `search_models`, `get_model`, `get_usage_example`, `get_endpoints`,
`recommend_models`, `get_stats`, `list_tiers`, `list_publishers` (NVIDIA NIM) +
`list_providers`, `search_provider_models`, `get_provider_model` (multi-provider).
**Security note:** provider-key management tools (`openrouter_list_keys`,
`openrouter_create_key`, `openrouter_get_key`, `openrouter_delete_key`,
`openrouter_key_usage`, `openrouter_rotate_key`, `get_provider_keys_status`)
exist on the **stdio** transport only. With `--transport http` they are removed
from the registry, so the network-facing MCP endpoint never advertises or
honors key administration. The same routes on the JSON API additionally
require `CATALOG_API_TOKEN` to be set. `X-Forwarded-For` is ignored by default
— set `TRUST_PROXY=1` only when the service sits behind a proxy you control.
## Production deployment
Version **2.3.0** — enterprise-grade serving, stdlib-only hardening:
```bash
docker compose up -d # API :8787 + MCP HTTP :9100 (hardened containers)
docker compose run --rm refresher # scheduled catalog rebuild
```
| Feature | How |
|---|---|
| Liveness / readiness | `GET /health` / `GET /ready` (DB present **and** populated) |
| Prometheus metrics | `GET /metrics` — requests, latency histogram, catalog gauges; label values sanitized against injection |
| Bearer auth | `CATALOG_API_TOKEN=…` (probes stay open; constant-time comparison; key-management routes require the token) |
| Request correlation | `X-Request-ID` (echoed/generated) + `X-Process-Time` headers on every response, mirrored into JSON logs |
| Rate limiting | `RATE_LIMIT_RPM=300` (429 + Retry-After; 0 disables); client identity from socket peer unless `TRUST_PROXY=1` |
| Structured logs | JSON access logs to stderr (`CATALOG_LOG=off` to silence) |
| DoS guards | 1 MiB request-body cap (413), bounded handler-thread pool (64) |
| Graceful shutdown | SIGTERM/SIGINT finishes in-flight requests |
| Security headers | nosniff, DENY, CSP, no-store on every response (JSON API **and** MCP HTTP) |
Operations guide: [`docs/OPERATIONS.md`](docs/OPERATIONS.md) · Security policy:
[`SECURITY.md`](SECURITY.md) · Release notes: [`CHANGELOG.md`](CHANGELOG.md)
## Dashboard
`make dashboard` generates `exports/dashboard.html` — a self-contained
(no-CDN) HTML dashboard with KPIs, NIM verification tiers/modalities/
publishers/transports, per-provider totals vs free with fetch health,
a searchable free-model explorer, a searchable NIM model explorer, and a
data-quality panel mirroring `validate_project.py`. The HTTP service serves
it live at `GET /dashboard` (also `/`):
```bash
make service # then open http://127.0.0.1:8787/dashboard
```
Resources: `nim://stats`, `nim://models`, `nim://tiers`,
`nim://models/{publisher}/{slug}[/{language}]`, `providers://catalog`,
`providers://{provider}/models`.
## Current audited result
Counts drift as providers add/remove models; the numbers below are from the
last audited run (2026-07-28). Re-run `make catalog` for fresh numbers — the
`/stats` endpoint and `exports/tiered-audit.json` always reflect your local DB.
- NVIDIA NIM: Build index 139 → active non-deprecated **124**
(57 free-hosted · 96 downloadable · 36 partner)
- OpenRouter: 341 models (18 free) · Nous Portal: 289 (5 free) ·
opencode Zen: 85 (24 free) · Blackbox AI: 47 (2 free) ·
Vercel AI Gateway: 307 (3 free)
- Live inference verification requires `NVIDIA_API_KEY` (`make verify`); gateway
reconciliation requires the same key (`python ingest_gateway.py`).
`data/active_nvidia_nim.sqlite3` contains all active official models, the
multi-provider catalog (`provider_models`), and audit evidence.
`exports/verified-working.sqlite3` contains only latest live-verified hosted models.
## Gateway `/v1/models` reconciliation
```bash
export NVIDIA_API_KEY='nvapi-...'
python ingest_gateway.py --database data/active_nvidia_nim.sqlite3 --output exports
python export_unified.py --database data/active_nvidia_nim.sqlite3 --output unified-output
```
The gateway response currently exposes only four fields: `id`, `object`, `created`, and `owned_by`. The database preserves all of them in `gateway_models`, maps aliases to current Build entries, keeps gateway-only IDs, and enriches mapped records with Build OpenAPI/templates/docs. Query the `all_model_inventory` view for the full union.
## Build without API key
```bash
python -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt
python audit_catalog.py --database data/active_nvidia_nim.sqlite3 build
python providers.py --database data/active_nvidia_nim.sqlite3 --output exports
python audit_catalog.py --database data/active_nvidia_nim.sqlite3 export --output exports
python finalize_audit.py --database data/active_nvidia_nim.sqlite3 --output exports
```
## Optional live verification
```bash
export NVIDIA_API_KEY='nvapi-...'
python audit_catalog.py --database data/active_nvidia_nim.sqlite3 verify --rpm 20 --timeout 120
```
The live-test loop runs under the RPM cap, so a full pass can take a while. To
**resume an interrupted run** (skip models already tested, no re-spent credit):
```bash
python audit_catalog.py --database data/active_nvidia_nim.sqlite3 verify --resume --rpm 20 --timeout 120
```
One Build catalog ID:
```bash
python audit_catalog.py --database data/active_nvidia_nim.sqlite3 verify \
--model z-ai/glm-5.2 --rpm 20 --timeout 120
```
## Automated update
```bash
chmod +x bin/update-catalog.sh
NIM_VERIFY=0 bin/update-catalog.sh
```
Set `NIM_VERIFY=1` only when live tests and credit use are intended.
The same pipeline is available as Make targets (`make help` for the full list):
```bash
make install # pip install -r requirements.txt -r requirements-dev.txt -r requirements-mcp.txt
make test # pytest unit suite (no network/API key needed)
make lint # ruff lint gate (same as CI)
make catalog # full key-less pipeline: build -> providers -> upstreams -> finalize -> export -> unified -> validate
make providers # refresh only the multi-provider catalogs
make dashboard # regenerate exports/dashboard.html
make service # read-only catalog HTTP API on 127.0.0.1:8787
make mcp # MCP server over stdio (Claude Desktop / Claude Code)
make mcp-http # MCP server + dashboard over Streamable HTTP on 127.0.0.1:9100
make image # build the catalog_service Docker image
make ci-local # reproduce the CI key-less pipeline locally
```
## Testing
`tests/` is a fast, deterministic pytest suite (no network, no API key) that
pins the data-correctness invariants: model-ID validation/de-dup, page
classification, served-id selection, smoke-payload construction, upstream
joining, and deprecation date semantics — plus the security regression
surface (FREE_ONLY floor, management-tool isolation, metrics label
sanitization, X-Forwarded-For handling, LIKE escaping, request correlation,
auth/rate-limit paths). **170 tests**, run with `make test`.
## Continuous integration
`.github/workflows/ci.yml` runs the unit tests (Python 3.11/3.12/3.13), the
`pip-audit` dependency CVE gate, then the key-less pipeline (build →
upstreams → finalize → export → unified → validate) on every push/PR and
daily, so the build.nvidia.com contract and the pipeline runtime stay
healthy. The serving smoke test asserts FREE_ONLY over the wire.
No API key is required for CI; live verification stays opt-in via
`NIM_VERIFY=1` locally or `bin/build-strict.sh` with an `NVIDIA_API_KEY`.
Dependabot (`.github/dependabot.yml`) opens weekly PRs for pip and
GitHub Actions updates, each gated by the same CI.
Two additional CI smoke assertions are proposed (management gate → 403
without token; `X-Request-ID` header present) — see `AUDIT_2026-08-01.md` §7.
Apply them with a maintainer token that has the GitHub `workflows`
permission.
## Deployment
`catalog_service.py` ships as a read-only HTTP API. The included `Dockerfile`
serves a pre-built catalog over the API (non-root, healthchecked):
```bash
make catalog # build data/active_nvidia_nim.sqlite3 first
docker build -t nvidia-nim-catalog .
docker run -p 8787:8787 -v "$PWD/data:/app/data:ro" nvidia-nim-catalog
```
## MCP server
`mcp_server.py` exposes the catalog as a **Model Context Protocol** server, so
MCP clients (Claude Desktop, Claude Code, Cursor, IDEs) can discover models,
read served IDs/endpoints, and get runnable cURL/Node.js examples on demand —
no copy-pasting from build.nvidia.com. It is a read-only adapter over the same
shared query layer (`catalog_queries.py`) as the HTTP API, so both transports
behave identically. Build the catalog first (`make catalog`), then run the
server.
Install the MCP dependency (extra):
```bash
pip install -r requirements-mcp.txt # adds the `mcp` SDK
```
Run (the server reads `data/active_nvidia_nim.sqlite3` read-only):
```bash
make mcp # stdio transport (default; for Claude Desktop / Claude Code)
make mcp-http # Streamable HTTP transport on http://127.0.0.1:9100
# or explicitly:
python mcp_server.py --database data/active_nvidia_nim.sqlite3 # stdio
python mcp_server.py --database data/active_nvidia_nim.sqlite3 --transport http --port 9100
```
### Connect from Claude Desktop (stdio)
Add an entry to `claude_desktop_config.json` (Settings → Developer → Edit Config).
Use the **absolute** path to the repo and (if needed) the absolute path to your
Python interpreter:
```jsonc
{
"mcpServers": {
"nim-catalog": {
"command": "python",
"args": ["/absolute/path/to/model_fetcher/mcp_server.py",
"--database", "/absolute/path/to/model_fetcher/data/active_nvidia_nim.sqlite3"]
}
}
}
```
### Connect from Claude Code (`.mcp.json` in the project root)
```jsonc
{
"mcpServers": {
"nim-catalog": {
"command": "python",
"args": ["mcp_server.py", "--database", "data/active_nvidia_nim.sqlite3"]
}
}
}
```
### Inspect interactively
```bash
mcp dev mcp_server.py # opens the MCP Inspector UI
```
### What the server exposes
**Tools** (the assistant calls these):
| Tool | Description |
|---|---|
| `search_models` | Search by text/modality/tier; filters for `working_only`/`free_only`. |
| `get_model` | Full detail: served_id, operations, endpoints, examples. |
| `get_usage_example` | Runnable cURL / Node.js code for a model. |
| `get_endpoints` | Every endpoint with host, transport, and auth requirement. |
| `recommend_models` | Steers to `live_inference_verified` models for production. |
| `get_stats` | Catalog counts, per-tier breakdown, last build time. |
| `list_tiers` | Verification tiers with counts and guarantees. |
| `list_providers` / `search_provider_models` / `get_provider_model` | Multi-provider catalog (OpenRouter, Nous, opencode, Blackbox, Vercel). |
| `openrouter_list_keys` / `create` / `get` / `delete` / `key_usage` / `rotate`, `get_provider_keys_status` | Provider-key administration — **stdio transport only**, never on `--transport http`. |
`serverInfo.version` in the MCP handshake reports the application version
(2.3.0), matching `GET /version` and the `catalog_info` Prometheus metric.
**Resources** (URI-addressable): `nim://stats`, `nim://models`, `nim://tiers`,
`nim://models/{publisher}/{slug}` (full model), and
`nim://models/{publisher}/{slug}/{language}` (a code example). The catalog id is
`publisher/slug`; because it contains a slash, the template splits it into two
segments — e.g. `nim://models/meta/llama-3.1-8b-instruct/curl`.
**Prompt**: `use_nim_model(catalog_id, task)` returns a ready-to-run integration
plan.
### Notes
- The server is **read-only**; it never writes and never runs the build pipeline.
Refresh the catalog by rebuilding the DB on a schedule (`make catalog` /
CI). `get_stats` reports the last build time so freshness is transparent.
- `recommend_models` defaults to `production_only=true`, returning only models
that actually responded over HTTP — set it `false` to also see official
downloadable/specialized models, clearly tagged.
- **Partial-build safety:** if the DB was produced by `audit_catalog build`
alone (the tiered audit from `finalize_audit.py` not yet run),
`get_stats`/`list_tiers` report `audit_populated: false` with a clear warning,
and `recommend_models` falls back to free-endpoint models instead of silently
returning an empty list. Run `make catalog` (which includes `finalize`) for
full tiered data.
- A missing database returns a clear startup error rather than failing silently.
For a complete strict rebuild, including the gateway and gRPC stages:
```bash
chmod +x bin/build-strict.sh
NIM_RPM=20 bin/build-strict.sh
```
The strict pipeline now runs Build discovery, authenticated gateway ingestion and gateway-only tests, Build free-endpoint inference tests, gRPC/NVCF health audit, upstream audit, final audit, unified export, and strict export. Gateway-only records remain visible in the unified inventory but enter strict output only after a real inference pass.
## Verification policy
- `live_inference_verified`: received HTTP 2xx and eligible for working DB.
- `official_contract_downloadable`: active official NIM, but requires local NVIDIA GPU/container deployment.
- `official_contract_specialized_transport`: official Build/API contract verified; live testing needs gRPC, media, private access, or special infrastructure.
- `live_test_failed`: timeout/404/5xx; retained in audit but excluded from working.
Every model has cURL and Node.js records. For gRPC/private/custom transports, records explicitly state the official alternative instead of fabricating an HTTP request.
## Upstream endpoint audit
```bash
python audit_upstreams.py --database data/active_nvidia_nim.sqlite3 --output exports
```
The audit cross-checks stored endpoints against each model's Build OpenAPI servers, resolved code templates, Build guides, and official NVIDIA docs. Supported upstream families include:
- `integrate.api.nvidia.com` — hosted LLM/chat/embedding and selected APIs;
- `ai.api.nvidia.com` — legacy/specialized VLM, retrieval, Cosmos, and AV APIs;
- `health.api.nvidia.com` — BioNeMo/biology APIs;
- `optimize.api.nvidia.com` — optimization APIs such as cuOpt;
- `climate.api.nvidia.com` — climate/weather APIs;
- `grpc.nvcf.nvidia.com:443` — hosted NVCF gRPC requiring API key and function ID;
- `localhost` — self-hosted NIM management/inference endpoints.
`exports/upstream-audit.*` records the result. A Build URL with `documentation-only` means NVIDIA publishes no fixed public upstream (for example private/early-access or internally brokered playgrounds).
This server cannot be deployed
Maintenance
ActivityStale
ResponsivenessNo issues