Skip to main content
Glama

mcp-kb-server

A remote MCP server that puts a company's scattered product knowledge behind one search, so an AI assistant in claude.ai can answer platform questions and always say where the answer came from.

I built it for a SaaS team whose knowledge lived in four places nobody had open at the same time. This repo is a cleaned-up version of that internal tool: the company, customers and content are replaced with a fictional company ("Northwind Cloud") and a tiny made-up corpus, but the architecture and the lessons are the real ones.

Stack: Node.js · Vercel serverless functions · MCP SDK (Streamable HTTP) · BM25 · OAuth 2.0 + PKCE · Freshdesk REST API · Slack · Vercel Blob


The problem

The answer to "how does X work?" was always somewhere, just never in the same place:

  • the Help Center says how a feature works,

  • recorded training sessions explain why and show examples,

  • resolved support tickets hold how it was solved in practice,

  • and the best advice, what to actually do, was in nobody's documentation at all, only in people's heads and Slack threads.

Related MCP server: Rememberizer MCP Server

What I built

One search over all four sources, exposed as an MCP server, plus a loop that fills the fourth source by itself.

claude.ai (Project)
   │  OAuth 2.0 + PKCE
   ▼
Vercel
 ├─ /api/mcp             MCP server, 3 tools: search · view · report_gap
 ├─ /api/oauth/*         authorize, token, discovery metadata
 ├─ /api/cron/harvest    daily harvest of approved Slack answers
 └─ /api/status          diagnostics (deliberately NOT an MCP tool)
      │
      ├─ data/index.json   help center + training + tickets snapshot (deployed file)
      ├─ Vercel Blob       team-curated answers (private, updated without redeploy)
      ├─ Freshdesk API     fresh tickets, live, READ-ONLY
      └─ Slack             gap reports in, validated answers out

   question not answered ──► report_gap ──► Slack ──► a person replies in the thread
         ▲                                                   │ reacts ✅
         └──── searchable on the next query ◄── daily cron ◄─┘

The improvement loop is the part I like most: when the assistant can't answer, it logs the question to Slack. A teammate answers in the thread and adds a ✅. A daily job turns the thread into a curated answer with provenance (slack-curated, ai-verified or mixed-curated) and the next person who asks gets it. The ✅ is the human validation.

Key decisions

Decision

Why

BM25 instead of embeddings

The corpus is small and full of exact product terms ("playbook", "workflow", "template"), where lexical search is accurate. No extra API key, no vector store, no added latency. The index is built in memory at cold start in milliseconds.

Three tools, not eight

Tool definitions travel to the model on every turn. I measured the cost: the fixed part dominated a typical conversation, so every extra tool is paid again each turn. Per-source search tools folded into search + a sources filter, and the three view_* tools folded into one, since the reference prefix tells them apart.

Diagnostics outside MCP

A status tool cost tokens on every turn and its output (index dates, snapshots) only confused end users. It became a plain HTTP endpoint behind a secret.

A reserved slot per source

By raw score the help center took most results and drowned the other sources. Each source with a decent result (at least a third of the best score) gets a slot; the rest is filled by score.

Progressive disclosure

search returns ~55-word snippets with a link; the full text is requested separately with view. A typical search stays cheap.

"Similar cases" runs on the server

search with similar_to: <ticket> turns a ticket into a query using its most informative terms (tf-idf against the index, with a boost for product vocabulary, mail footers and courtesy phrases removed). The ticket body never enters the model's context.

Say when a result is only a lexical coincidence

BM25 always returns something. If the best result covers under a third of the significant query terms, the response says it's probably not documented instead of presenting it with false confidence.

Read-only Freshdesk, enforced in code

One fetch in the whole codebase, method pinned to GET, any other verb throws before reaching the network. A unit test proves no request leaves the process.

Mask PII before indexing

The index is a deployed file loaded whole into memory. Emails and phone numbers become [email] / [tel] at build time; full data is only available through view, live and authenticated.

Stateless OAuth with signed tokens

Serverless has nowhere to store codes. Authorization codes and tokens carry their own payload, signed with HMAC-SHA256, compared in constant time. Rotating one signing key revokes every token. redirect_uri is validated against claude.ai and subdomains, otherwise the login screen is an open redirect usable for phishing.

Curated answers in a private Blob

A deployed file would need a redeploy for every answered question. A Blob store is read hot with a 5-minute in-memory TTL.

Ticket labels as indexed text, not as a filter

Classification only covers recent tickets. Filtering by category would hide most of the history, where the old solutions live; as text, labels help where they exist and cost nothing where they don't.

What went wrong, and what I learned

  • The fixed token cost is the real cost. Instructions plus tool definitions are paid every turn, whether or not a tool is used. Auditing a conversation (npm run audit:tokens) showed it dominated the total. The cheapest optimization wasn't smarter retrieval, it was fewer and shorter tools.

  • A "not documented" answer needs a next step. An empty result used to just say "no matches". Now it tells the assistant to rephrase with product terms and then call report_gap, so the loop closes on every path without relying on project instructions being up to date.

  • CDN caching hid new answers. Reading the Blob went through a CDN by default and served the old version after it was updated; a documented answer could take hours to appear. Reading with useCache: false fixed it, since there's already an in-memory TTL.

  • Speech-to-text mangles the product name. Most mentions of the product in the transcripts were transcribed wrong, so people searching for it missed most of the content. Names are normalized at index time, and a speaker's own name embedded in their text is stripped. Long monologues are split, with timestamps interpolated so "min 12:30" stays correct.

  • Email footers define the search if you let them. tf-idf rewards rare words, and legal disclaimers are rare. Without trimming each message's signature and footer (per message, never on the whole ticket, or the solution in later replies is cut off), the "key terms" of a sync problem came from the confidentiality notice.

  • BM25 score doesn't tell you if a ticket is relevant. A bad result can outscore a good one depending on length. Relevance is judged with a vocabulary overlap coefficient instead (not Jaccard, which punishes texts of different length).

  • The claude.ai connector caches the tool schema. After changing tools or instructions you have to reload the connector, or it keeps using the old definitions.

  • A shared team password is a known limitation. There is no per-person revocation. The fix would be signing in through the company's identity provider.

Known limitations of this version

  • The phone-number mask is greedy on purpose (it errs toward masking): any run of 9+ digits with separators becomes [tel], dates like 2026-03-14 included.

  • The stemmer only handles English plurals and is intentionally conservative.

  • Help-center content comes from local Markdown files; in a real deployment you'd plug in a crawler or export, and nothing else in the pipeline changes.

  • The ticket snapshot is built from fixtures by default; pass --live-tickets to pull from a real Freshdesk account.

Run it yourself

npm install
npm test                  # 40+ unit and integration tests
npm run build:index       # rebuild data/index.json from corpus/
npm run dev               # http://localhost:3000/api/mcp (open, no auth: dev only)

Smoke test with curl:

curl -s -X POST localhost:3000/api/mcp \
  -H 'Content-Type: application/json' -H 'Accept: application/json, text/event-stream' \
  -d '{"jsonrpc":"2.0","id":1,"method":"tools/call","params":{"name":"search","arguments":{"query":"rep on vacation still receiving leads"}}}'

Try asking about assignment, WhatsApp templates, call quality or CRM sync: the fictional corpus covers all four sources for those topics.

Deploy:

  1. Copy .env.example and set the variables in Vercel (OAuth client, signing key, access password, Freshdesk, Slack, Blob token).

  2. vercel --prod.

  3. In claude.ai: Settings → Connectors → Add custom connector → https://<your-deployment>/api/mcp, with the OAuth client id and secret.

  4. Check the whole flow without a browser: BASE=https://<your-deployment> npm run e2e:oauth.

Refresh or extend the corpus by dropping files into corpus/ and running the matching npm run index:* script.

Project layout

api/mcp.js                 MCP server: tools, instructions, auth, CORS
api/oauth/                 authorize (login + PKCE), token, discovery metadata
api/cron/harvest.js        Slack ✅ threads -> private Blob
api/status.js              HTTP diagnostics
lib/bm25.js                BM25 + one-chunk-per-parent + per-source slots
lib/text.js                tokenizer, stopwords, plural stemmer
lib/similar.js             "similar cases": key terms, relevance, coverage
lib/freshdesk.js           read-only client, delta cache, PII masking
lib/oauth.js               signed tokens, PKCE, redirect_uri allow-list
lib/curated-store.js       private Blob reader/writer for team answers
lib/slack.js               gap reports (webhook)
lib/slack-harvest.js       thread -> knowledge unit with provenance
scripts/build-index.mjs    corpus/ -> data/index.json
scripts/sources/           help center, transcripts, curated builders
scripts/dev-server.mjs     local Vercel-like server (+ OAuth routes)
scripts/e2e-oauth.mjs      full OAuth flow against a running server
scripts/audit-tokens.mjs   token-cost audit of a typical conversation
corpus/                    fictional help articles, transcript, tickets, curated answers
test/                      node:test suite

License

MIT

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    D
    maintenance
    Enables semantic search and retrieval of personal and team knowledge from connected sources like Slack, Gmail, Google Drive, and Dropbox, with the ability to save new information for future recall.
    Apache 2.0
  • F
    license
    A
    quality
    Not graded
    maintenance
    Exposes an internal engineering knowledge base to AI assistants, allowing users to search and retrieve standards, runbooks, and architecture decisions. It supports RAG-enhanced search, document scraping, and specialized prompts for incident investigation and code reviews.
    5
    -
  • A
    license
    Not graded
    quality
    C
    maintenance
    Enables AI assistants to crawl, index, and retrieve information from technical documentation using semantic search, with optional knowledge graph validation for code hallucination detection.
    MIT