Similarity Search MCP Server
Click on "Install Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Similarity Search MCP ServerRank these vectors against the query and show top 5 results"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
# Similarity Search API
Stateless similarity search over pre-computed vectors — NMI (normalized mutual information) + cosine fusion, with an entropy-calibrated blending weight computed per request. No vector database, no index to maintain, no infrastructure to run.
Available both as a plain HTTP API and as an MCP server (5 tools) for AI agents.
Important: this operates on vectors, not raw text
This API does not embed text for you. query and corpus entries are pre-computed numeric vectors (e.g. from your own embedding model). If you need text-to-vector embedding first, run that upstream and pass the resulting vectors here.
Related MCP server: smriti
Base URL
https://similarity-search-api-production.up.railway.app
Authentication
All business endpoints require an X-API-Key header:
X-API-Key:
/health requires no authentication.
Pricing
Two ways to pay, same endpoints:
x402 (pay-per-call, USDC on Base) — currently on Base Sepolia testnet, $0.01/call, no account or API key required beyond the x402 payment flow itself. A request without payment gets
402 Payment Requiredwith the payment details in thepayment-requiredresponse header.Stripe (metered billing) — for callers provisioned with an API key and a Stripe customer on the account.
Endpoints
POST /similarity/search
Rank a corpus against a query vector using the composite score.
{
"query": { "id": "q1", "vector": [0.12, -0.4, 0.91, "..."] },
"corpus": [
{ "id": "doc1", "vector": [0.10, -0.35, 0.88, "..."] },
{ "id": "doc2", "vector": [0.55, 0.02, -0.14, "..."] }
],
"top_k": 10,
"nmi_bins": 10,
"alpha_override": null
}All vectors in query and corpus must share the same dimensionality (2-4096 dims). top_k up to 1000. alpha_override (optional) pins the cosine/NMI blend weight instead of calibrating it from corpus entropy.
Response:
{
"results": [
{ "id": "doc1", "composite_score": 0.91, "cosine_similarity": 0.89, "nmi_score": 0.94, "rank": 1 }
],
"calibrated_alpha": 0.73,
"corpus_entropy": 3.85,
"query_id": "q1",
"corpus_size": 2,
"latency_ms": 43,
"request_fingerprint": "..."
}POST /similarity/calibrate-alpha/v1
Compute the entropy-calibrated alpha for a corpus without running a full search - useful for inspecting/debugging calibration behavior before committing to a search call.
POST /similarity/batch-score
Score up to 10,000 (vector_a, vector_b) pairs with a fixed alpha - no corpus/entropy overhead.
{
"pairs": [[[0.1, 0.2], [0.15, 0.19]]],
"alpha": 0.5,
"nmi_bins": 10
}GET /health
Liveness probe. No auth required. Not billed (excluded from both Stripe and x402).
Note:
POST /similarity/calibrate-alpha(without/v1) is a deprecated alias kept for backward compatibility - use/similarity/calibrate-alpha/v1.
MCP tools
Connect an MCP-compatible client (Claude, Cursor, etc.) to the streamable HTTP endpoint at: https://similarity-search-api-production.up.railway.app/mcp
Exposes 5 tools: rank_items_by_nmi_cosine_fusion, estimate_corpus_entropy_profile, score_pair_nmi_cosine, find_outlier_vectors_by_nmi_deficit, calibrate_alpha_from_query_entropy.
The scoring method
composite_score = alpha * cosine(query, doc) + (1 - alpha) * NMI_normalized(query, doc)
alpha is calibrated per-request from the Shannon entropy of the submitted corpus (unless you pass alpha_override) - high-entropy (dispersed) corpora lean toward cosine; low-entropy (dense/narrow) corpora lean toward NMI, which captures statistical dependence that cosine's geometric angle misses.
Limits
Corpus size: up to 500,000 items per request
Vector dimensionality: 2-4,096
batch-scorepairs: up to 10,000 per request
This server cannot be installed
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Servers
- FlicenseAqualityDmaintenanceEnables AI agents to query a local knowledge graph built from document collections using hybrid search (BM25 + vector fusion) and entity-relationship extraction. Supports privacy-first, offline operation with tools for semantic search, entity graph exploration, and corpus statistics.Last updated3
- AlicenseBqualityAmaintenanceEnables building and querying vector-based knowledge graphs with node and edge management and semantic search.Last updated515MIT
- Alicense-qualityAmaintenanceEnables local semantic search over documents (PDFs, code, etc.) using local embedding models, allowing AI agents and users to find information by meaning without API keys or cloud services.Last updated3MIT
Related MCP Connectors
Multi-engine search for AI agents. Trust scoring, local corpus, MCP-native. Self-hostable, BYOK.
Persistent memory for AI agents. Search, store, and recall across sessions.
Universal memory for AI agents and tools. Save, organize and search context anywhere.
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/nexus-mcp-infra/similarity-search-api-sdk'
If you have feedback or need assistance with the MCP directory API, please join our Discord server