mcp-kb-server
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@mcp-kb-serversearch our docs and support tickets for how to reset a user's MFA"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
mcp-kb-server
A remote MCP server that puts a company's scattered product knowledge behind one search, so an AI assistant in claude.ai can answer platform questions and always say where the answer came from.
I built it for a SaaS team whose knowledge lived in four places nobody had open at the same time. This repo is a cleaned-up version of that internal tool: the company, customers and content are replaced with a fictional company ("Northwind Cloud") and a tiny made-up corpus, but the architecture and the lessons are the real ones.
Stack: Node.js · Vercel serverless functions · MCP SDK (Streamable HTTP) · BM25 · OAuth 2.0 + PKCE · Freshdesk REST API · Slack · Vercel Blob
The problem
The answer to "how does X work?" was always somewhere, just never in the same place:
the Help Center says how a feature works,
recorded training sessions explain why and show examples,
resolved support tickets hold how it was solved in practice,
and the best advice, what to actually do, was in nobody's documentation at all, only in people's heads and Slack threads.
Related MCP server: Rememberizer MCP Server
What I built
One search over all four sources, exposed as an MCP server, plus a loop that fills the fourth source by itself.
claude.ai (Project)
│ OAuth 2.0 + PKCE
▼
Vercel
├─ /api/mcp MCP server, 3 tools: search · view · report_gap
├─ /api/oauth/* authorize, token, discovery metadata
├─ /api/cron/harvest daily harvest of approved Slack answers
└─ /api/status diagnostics (deliberately NOT an MCP tool)
│
├─ data/index.json help center + training + tickets snapshot (deployed file)
├─ Vercel Blob team-curated answers (private, updated without redeploy)
├─ Freshdesk API fresh tickets, live, READ-ONLY
└─ Slack gap reports in, validated answers out
question not answered ──► report_gap ──► Slack ──► a person replies in the thread
▲ │ reacts ✅
└──── searchable on the next query ◄── daily cron ◄─┘The improvement loop is the part I like most: when the assistant can't answer, it logs the question to Slack. A teammate answers in the thread and adds a ✅. A daily job turns the thread into a curated answer with provenance (slack-curated, ai-verified or mixed-curated) and the next person who asks gets it. The ✅ is the human validation.
Key decisions
Decision | Why |
BM25 instead of embeddings | The corpus is small and full of exact product terms ("playbook", "workflow", "template"), where lexical search is accurate. No extra API key, no vector store, no added latency. The index is built in memory at cold start in milliseconds. |
Three tools, not eight | Tool definitions travel to the model on every turn. I measured the cost: the fixed part dominated a typical conversation, so every extra tool is paid again each turn. Per-source search tools folded into |
Diagnostics outside MCP | A status tool cost tokens on every turn and its output (index dates, snapshots) only confused end users. It became a plain HTTP endpoint behind a secret. |
A reserved slot per source | By raw score the help center took most results and drowned the other sources. Each source with a decent result (at least a third of the best score) gets a slot; the rest is filled by score. |
Progressive disclosure |
|
"Similar cases" runs on the server |
|
Say when a result is only a lexical coincidence | BM25 always returns something. If the best result covers under a third of the significant query terms, the response says it's probably not documented instead of presenting it with false confidence. |
Read-only Freshdesk, enforced in code | One |
Mask PII before indexing | The index is a deployed file loaded whole into memory. Emails and phone numbers become |
Stateless OAuth with signed tokens | Serverless has nowhere to store codes. Authorization codes and tokens carry their own payload, signed with HMAC-SHA256, compared in constant time. Rotating one signing key revokes every token. |
Curated answers in a private Blob | A deployed file would need a redeploy for every answered question. A Blob store is read hot with a 5-minute in-memory TTL. |
Ticket labels as indexed text, not as a filter | Classification only covers recent tickets. Filtering by category would hide most of the history, where the old solutions live; as text, labels help where they exist and cost nothing where they don't. |
What went wrong, and what I learned
The fixed token cost is the real cost. Instructions plus tool definitions are paid every turn, whether or not a tool is used. Auditing a conversation (
npm run audit:tokens) showed it dominated the total. The cheapest optimization wasn't smarter retrieval, it was fewer and shorter tools.A "not documented" answer needs a next step. An empty result used to just say "no matches". Now it tells the assistant to rephrase with product terms and then call
report_gap, so the loop closes on every path without relying on project instructions being up to date.CDN caching hid new answers. Reading the Blob went through a CDN by default and served the old version after it was updated; a documented answer could take hours to appear. Reading with
useCache: falsefixed it, since there's already an in-memory TTL.Speech-to-text mangles the product name. Most mentions of the product in the transcripts were transcribed wrong, so people searching for it missed most of the content. Names are normalized at index time, and a speaker's own name embedded in their text is stripped. Long monologues are split, with timestamps interpolated so "min 12:30" stays correct.
Email footers define the search if you let them. tf-idf rewards rare words, and legal disclaimers are rare. Without trimming each message's signature and footer (per message, never on the whole ticket, or the solution in later replies is cut off), the "key terms" of a sync problem came from the confidentiality notice.
BM25 score doesn't tell you if a ticket is relevant. A bad result can outscore a good one depending on length. Relevance is judged with a vocabulary overlap coefficient instead (not Jaccard, which punishes texts of different length).
The claude.ai connector caches the tool schema. After changing tools or instructions you have to reload the connector, or it keeps using the old definitions.
A shared team password is a known limitation. There is no per-person revocation. The fix would be signing in through the company's identity provider.
Known limitations of this version
The phone-number mask is greedy on purpose (it errs toward masking): any run of 9+ digits with separators becomes
[tel], dates like2026-03-14included.The stemmer only handles English plurals and is intentionally conservative.
Help-center content comes from local Markdown files; in a real deployment you'd plug in a crawler or export, and nothing else in the pipeline changes.
The ticket snapshot is built from fixtures by default; pass
--live-ticketsto pull from a real Freshdesk account.
Run it yourself
npm install
npm test # 40+ unit and integration tests
npm run build:index # rebuild data/index.json from corpus/
npm run dev # http://localhost:3000/api/mcp (open, no auth: dev only)Smoke test with curl:
curl -s -X POST localhost:3000/api/mcp \
-H 'Content-Type: application/json' -H 'Accept: application/json, text/event-stream' \
-d '{"jsonrpc":"2.0","id":1,"method":"tools/call","params":{"name":"search","arguments":{"query":"rep on vacation still receiving leads"}}}'Try asking about assignment, WhatsApp templates, call quality or CRM sync: the fictional corpus covers all four sources for those topics.
Deploy:
Copy
.env.exampleand set the variables in Vercel (OAuth client, signing key, access password, Freshdesk, Slack, Blob token).vercel --prod.In claude.ai: Settings → Connectors → Add custom connector →
https://<your-deployment>/api/mcp, with the OAuth client id and secret.Check the whole flow without a browser:
BASE=https://<your-deployment> npm run e2e:oauth.
Refresh or extend the corpus by dropping files into corpus/ and running the matching npm run index:* script.
Project layout
api/mcp.js MCP server: tools, instructions, auth, CORS
api/oauth/ authorize (login + PKCE), token, discovery metadata
api/cron/harvest.js Slack ✅ threads -> private Blob
api/status.js HTTP diagnostics
lib/bm25.js BM25 + one-chunk-per-parent + per-source slots
lib/text.js tokenizer, stopwords, plural stemmer
lib/similar.js "similar cases": key terms, relevance, coverage
lib/freshdesk.js read-only client, delta cache, PII masking
lib/oauth.js signed tokens, PKCE, redirect_uri allow-list
lib/curated-store.js private Blob reader/writer for team answers
lib/slack.js gap reports (webhook)
lib/slack-harvest.js thread -> knowledge unit with provenance
scripts/build-index.mjs corpus/ -> data/index.json
scripts/sources/ help center, transcripts, curated builders
scripts/dev-server.mjs local Vercel-like server (+ OAuth routes)
scripts/e2e-oauth.mjs full OAuth flow against a running server
scripts/audit-tokens.mjs token-cost audit of a typical conversation
corpus/ fictional help articles, transcript, tickets, curated answers
test/ node:test suiteLicense
MIT
This server cannot be deployed
Maintenance
Related MCP Connectors
Shared company knowledge, workflows, and connected apps for the AIs your team already uses.
Your company's brain for AI agents. Cited, permission-aware knowledge across every system.
- KumbukaOAuthai.kumbuka
Governed, auditable knowledge your team curates for its AI assistants, self-hostable
Connect your team's living knowledge base — docs, data, issues, CRM — to Claude and ChatGPT.
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceEnables AI assistants to search through structured databases and unstructured content (documents, videos, files) using natural language queries with semantic understanding.MIT
- AlicenseNot gradedqualityDmaintenanceEnables semantic search and retrieval of personal and team knowledge from connected sources like Slack, Gmail, Google Drive, and Dropbox, with the ability to save new information for future recall.Apache 2.0
- FlicenseAqualityNot gradedmaintenanceExposes an internal engineering knowledge base to AI assistants, allowing users to search and retrieve standards, runbooks, and architecture decisions. It supports RAG-enhanced search, document scraping, and specialized prompts for incident investigation and code reviews.5-
- AlicenseNot gradedqualityCmaintenanceEnables AI assistants to crawl, index, and retrieve information from technical documentation using semantic search, with optional knowledge graph validation for code hallucination detection.MIT