apinav
Searches Google's public API directory as a live plugin, merging API listings into the candidate pool for semantic search and reranking.
Provides live search over the RapidAPI marketplace, fetching API listings per query and merging them into the candidate pool before reranking without persisting results.
apinav
Agent-driven API discovery: a local SQLite catalog of public APIs with keyword (FTS5) and semantic (embeddings + LLM rerank) search, exposed to AI agents as an MCP (Model Context Protocol) server.
The catalog has no bulk ingestion step. Live directory plugins fetch results per search query; every result that passes the quality gates is upserted into the database, embedded, and immediately searchable — the index grows organically during usage.
Built and maintained by an autonomous AI agent as part of a home-lab agent platform (the repo itself is agent-authored, from schema to this README).
Architecture
Plugins query API sources. Every live source is a uniform
search(query, limit)plugin (apis.io, marketplace, Smithery, Apify, Google APIs directory, HF Spaces) — polite pacing,Retry-Afterhandling, one dead plugin never takes the others down. No batch jobs, no crawls: plugins only fetch what a search actually asks for.Results cached in SQLite. After live results are merged into the candidate pool for the current query, novel ids (never seen before) are persisted to the local database via
ingest.py— spam/junk-gated, deduplicated, upserted with per-source provenance in thesourcecolumn. Re-sights are cheap id skips; a locked or failed persist never blocks the query.Embeddings for semantic search. Persisted rows are embedded with NVIDIA
nemotron-3-embed-1bin the same call (batch of 32) and appended to a numpy vector cache, so they are semantically searchable immediately. FTS5 triggers keep the keyword index current.
Related MCP server: APIClaw
Search workflow
Semantic candidates — embed the query, rank stored vectors by cosine similarity (numpy matrix fast path), plus an FTS5 pre-pass.
Live merge — every plugin runs its per-query
search()in the background; results are merged into the candidate pool before reranking so they compete on relevance instead of being appended to the tail.Persist — novel live rows are written to the database (organic growth).
Rerank last — LLM cross-encoder over the merged pool; results are returned with per-source provenance, similarity scores, and doc links.
Repo layout
File | Purpose |
| MCP server: keyword search, semantic search (live merge + persist), live search, record fetch, stats |
| Organic-growth persistence: quality gates, dedup, upsert + NVIDIA embed + vector-cache append for novel live results |
| SQLite schema (catalog + FTS5 + embeddings + source registry) |
| Central configuration: paths, endpoints, models, tuning parameters |
| Config loader used by all modules |
| Plugin registry config: base URLs, limits, cadence, drift notes |
| Source registry CLI: init, report, set |
| Live-search plugins (one module per source, uniform |
| Shared plugin plumbing: polite pacing, |
| Shared spam/junk detection battery |
| NVIDIA embedding backfill + numpy vector cache for fast ranking |
| CLI and gateway registration helper |
Plugin contract
Every live plugin has the same shape:
def search(query: str, limit: int = 5) -> list[dict]:
# plain HTTP GET (or optional local client), polite pacing via plugins/base
# returns rows normalized to the apinav node shape with a `source` tag
def build_links(row: dict) -> dict:
# optional: return {"url_docs": ..., "url_spec_json": ..., "url_spec_yaml": ...}
# for the row's source. The MCP server dispatches through plugins.build_links().Polite pacing (default 0.5 s between calls per source)
RateLimited(seconds)on HTTP 429 withRetry-Afterparsed and honoredOne dead plugin never takes the others down
Results are overlays: fetched per query, merged before rerank — and novel ids among them are persisted by
ingest.py(idempotent, lock-safe, best-effort: a locked or failed persist never blocks the query)Link resolution is owned by the plugin that owns the
sourcetag
A plugin that needs a non-public client (e.g. a session-based marketplace
client) loads it as an optional local module named marketplace_client.py
(with open_tab/close_tab/check_rate_limit/fetch_all); if it is missing the
plugin degrades gracefully to an empty result.
Requirements
Python 3.11+
openai-compatible client for embeddings/rerank — API keys are sourced viaconfig.secret()from the per-host.envfile (NVIDIA_API_KEY,OPENROUTER_API_KEY); nothing is hardcoded in code orconfig.yamlSQLite with FTS5
Coding standards
This repo is maintained by an autonomous agent, so the bar for readability is explicit. Follow these when editing:
Comments say why, not what. The what belongs in the code (names, control flow). A comment earns its place when it captures tribal knowledge a reader would otherwise have to dig out of git history or reverse-engineer:
Non-obvious data invariants (e.g.
endpoint_count-1= "unknown class", vs0= junk).Gotchas that bite (
INSERT OR REPLACEwipes a row, so preserve missingendpoint_count/sourcefrom the existing row).Rationale for a surprising choice (why FTS pre-pass exists, why live results merge before rerank, why the audio-direction regex judges the name only).
Generic restatements like # batch by 32 are noise — delete them.
Names over comments. Prefer a descriptive identifier over a comment that explains a vague one.
Docstrings are contracts, not history. State what a function/module does, its inputs/outputs, and its invariants. Do not put dates or attributions there — that is what git is for.
Section banners sparingly. Use # --- Section --- only for big logical
blocks in a long module (the server). Skip them for 3-line helpers.
Secrets never in config. Operational settings live in config.yaml;
credentials stay in .env and are reached only through config.secret().
Status
Personal home-lab project, published as-is. The catalog database and embedding matrix are multi-GB build artifacts and are not included — they accumulate from plugin results as you search.
This server cannot be deployed
Maintenance
Related MCP Connectors
Turn any task into the right API calls: discover, evaluate, and integrate public APIs.
Discover, compare, route, and execute machine-accessible capabilities for AI agents.
31Multi-engine search for AI agents. Trust scoring, local corpus, MCP-native. Self-hostable, BYOK.
AI service marketplace — agents discover, call, and pay for API services automatically.
Related MCP Servers
- AlicenseAqualityCmaintenanceEnables semantic search and discovery of free public APIs from an extensive catalog. Provides embedding-based search over API names and descriptions, plus detailed API information retrieval.210MIT
- AlicenseAqualityBmaintenanceThe API layer for AI agents. World's biggest API index with 22,000+ APIs and growing. Agents discover and call APIs at runtime with semantic search, structured metadata, and 18 Direct Call APIs including AI providers.14349 npm8MIT
- FlicenseNot gradedqualityDmaintenanceEnables AI agents to index and search across SQLite databases and CSV files to discover table schemas and column metadata. It provides a unified MCP API for data source management and structural exploration through natural language.-
- AlicenseNot gradedqualityDmaintenanceEnables AI agents to search and discover data across SQLite and CSV sources through an MCP interface, with metadata indexing and fuzzy search capabilities.MIT