Skip to main content
Glama
README.md
# mcp-kb-server

A remote [MCP](https://modelcontextprotocol.io) server that puts a company's scattered product knowledge behind one search, so an AI assistant in claude.ai can answer platform questions and always say where the answer came from.

I built it for a SaaS team whose knowledge lived in four places nobody had open at the same time. This repo is a cleaned-up version of that internal tool: the company, customers and content are replaced with a fictional company ("Northwind Cloud") and a tiny made-up corpus, but the architecture and the lessons are the real ones.

**Stack:** Node.js · Vercel serverless functions · MCP SDK (Streamable HTTP) · BM25 · OAuth 2.0 + PKCE · Freshdesk REST API · Slack · Vercel Blob

---

## The problem

The answer to "how does X work?" was always somewhere, just never in the same place:

- the **Help Center** says how a feature works,
- recorded **training sessions** explain why and show examples,
- resolved **support tickets** hold how it was solved in practice,
- and the best advice, **what to actually do**, was in nobody's documentation at all, only in people's heads and Slack threads.

## What I built

One search over all four sources, exposed as an MCP server, plus a loop that fills the fourth source by itself.

```
claude.ai (Project)
   │  OAuth 2.0 + PKCE
   ▼
Vercel
 ├─ /api/mcp             MCP server, 3 tools: search · view · report_gap
 ├─ /api/oauth/*         authorize, token, discovery metadata
 ├─ /api/cron/harvest    daily harvest of approved Slack answers
 └─ /api/status          diagnostics (deliberately NOT an MCP tool)
      │
      ├─ data/index.json   help center + training + tickets snapshot (deployed file)
      ├─ Vercel Blob       team-curated answers (private, updated without redeploy)
      ├─ Freshdesk API     fresh tickets, live, READ-ONLY
      └─ Slack             gap reports in, validated answers out

   question not answered ──► report_gap ──► Slack ──► a person replies in the thread
         ▲                                                   │ reacts ✅
         └──── searchable on the next query ◄── daily cron ◄─┘
```

The improvement loop is the part I like most: when the assistant can't answer, it logs the question to Slack. A teammate answers in the thread and adds a ✅. A daily job turns the thread into a curated answer with provenance (`slack-curated`, `ai-verified` or `mixed-curated`) and the next person who asks gets it. The ✅ *is* the human validation.

## Key decisions

| Decision | Why |
|---|---|
| **BM25 instead of embeddings** | The corpus is small and full of exact product terms ("playbook", "workflow", "template"), where lexical search is accurate. No extra API key, no vector store, no added latency. The index is built in memory at cold start in milliseconds. |
| **Three tools, not eight** | Tool definitions travel to the model on *every* turn. I measured the cost: the fixed part dominated a typical conversation, so every extra tool is paid again each turn. Per-source search tools folded into `search` + a `sources` filter, and the three `view_*` tools folded into one, since the reference prefix tells them apart. |
| **Diagnostics outside MCP** | A status tool cost tokens on every turn and its output (index dates, snapshots) only confused end users. It became a plain HTTP endpoint behind a secret. |
| **A reserved slot per source** | By raw score the help center took most results and drowned the other sources. Each source with a decent result (at least a third of the best score) gets a slot; the rest is filled by score. |
| **Progressive disclosure** | `search` returns ~55-word snippets with a link; the full text is requested separately with `view`. A typical search stays cheap. |
| **"Similar cases" runs on the server** | `search` with `similar_to: <ticket>` turns a ticket into a query using its most informative terms (tf-idf against the index, with a boost for product vocabulary, mail footers and courtesy phrases removed). The ticket body never enters the model's context. |
| **Say when a result is only a lexical coincidence** | BM25 always returns *something*. If the best result covers under a third of the significant query terms, the response says it's probably not documented instead of presenting it with false confidence. |
| **Read-only Freshdesk, enforced in code** | One `fetch` in the whole codebase, method pinned to `GET`, any other verb throws before reaching the network. A unit test proves no request leaves the process. |
| **Mask PII before indexing** | The index is a deployed file loaded whole into memory. Emails and phone numbers become `[email]` / `[tel]` at build time; full data is only available through `view`, live and authenticated. |
| **Stateless OAuth with signed tokens** | Serverless has nowhere to store codes. Authorization codes and tokens carry their own payload, signed with HMAC-SHA256, compared in constant time. Rotating one signing key revokes every token. `redirect_uri` is validated against `claude.ai` and subdomains, otherwise the login screen is an open redirect usable for phishing. |
| **Curated answers in a private Blob** | A deployed file would need a redeploy for every answered question. A Blob store is read hot with a 5-minute in-memory TTL. |
| **Ticket labels as indexed text, not as a filter** | Classification only covers recent tickets. Filtering by category would hide most of the history, where the old solutions live; as text, labels help where they exist and cost nothing where they don't. |

## What went wrong, and what I learned

- **The fixed token cost is the real cost.** Instructions plus tool definitions are paid every turn, whether or not a tool is used. Auditing a conversation (`npm run audit:tokens`) showed it dominated the total. The cheapest optimization wasn't smarter retrieval, it was fewer and shorter tools.
- **A "not documented" answer needs a next step.** An empty result used to just say "no matches". Now it tells the assistant to rephrase with product terms and then call `report_gap`, so the loop closes on every path without relying on project instructions being up to date.
- **CDN caching hid new answers.** Reading the Blob went through a CDN by default and served the old version after it was updated; a documented answer could take hours to appear. Reading with `useCache: false` fixed it, since there's already an in-memory TTL.
- **Speech-to-text mangles the product name.** Most mentions of the product in the transcripts were transcribed wrong, so people searching for it missed most of the content. Names are normalized at index time, and a speaker's own name embedded in their text is stripped. Long monologues are split, with timestamps interpolated so "min 12:30" stays correct.
- **Email footers define the search if you let them.** tf-idf rewards rare words, and legal disclaimers are rare. Without trimming each message's signature and footer (per message, never on the whole ticket, or the solution in later replies is cut off), the "key terms" of a sync problem came from the confidentiality notice.
- **BM25 score doesn't tell you if a ticket is relevant.** A bad result can outscore a good one depending on length. Relevance is judged with a vocabulary overlap coefficient instead (not Jaccard, which punishes texts of different length).
- **The claude.ai connector caches the tool schema.** After changing tools or instructions you have to reload the connector, or it keeps using the old definitions.
- **A shared team password is a known limitation.** There is no per-person revocation. The fix would be signing in through the company's identity provider.

### Known limitations of this version

- The phone-number mask is greedy on purpose (it errs toward masking): any run of 9+ digits with separators becomes `[tel]`, dates like `2026-03-14` included.
- The stemmer only handles English plurals and is intentionally conservative.
- Help-center content comes from local Markdown files; in a real deployment you'd plug in a crawler or export, and nothing else in the pipeline changes.
- The ticket snapshot is built from fixtures by default; pass `--live-tickets` to pull from a real Freshdesk account.

## Run it yourself

```bash
npm install
npm test                  # 40+ unit and integration tests
npm run build:index       # rebuild data/index.json from corpus/
npm run dev               # http://localhost:3000/api/mcp (open, no auth: dev only)
```

Smoke test with curl:

```bash
curl -s -X POST localhost:3000/api/mcp \
  -H 'Content-Type: application/json' -H 'Accept: application/json, text/event-stream' \
  -d '{"jsonrpc":"2.0","id":1,"method":"tools/call","params":{"name":"search","arguments":{"query":"rep on vacation still receiving leads"}}}'
```

Try asking about assignment, WhatsApp templates, call quality or CRM sync: the fictional corpus covers all four sources for those topics.

**Deploy:**

1. Copy `.env.example` and set the variables in Vercel (OAuth client, signing key, access password, Freshdesk, Slack, Blob token).
2. `vercel --prod`.
3. In claude.ai: Settings → Connectors → Add custom connector → `https://<your-deployment>/api/mcp`, with the OAuth client id and secret.
4. Check the whole flow without a browser: `BASE=https://<your-deployment> npm run e2e:oauth`.

Refresh or extend the corpus by dropping files into `corpus/` and running the matching `npm run index:*` script.

## Project layout

```
api/mcp.js                 MCP server: tools, instructions, auth, CORS
api/oauth/                 authorize (login + PKCE), token, discovery metadata
api/cron/harvest.js        Slack ✅ threads -> private Blob
api/status.js              HTTP diagnostics
lib/bm25.js                BM25 + one-chunk-per-parent + per-source slots
lib/text.js                tokenizer, stopwords, plural stemmer
lib/similar.js             "similar cases": key terms, relevance, coverage
lib/freshdesk.js           read-only client, delta cache, PII masking
lib/oauth.js               signed tokens, PKCE, redirect_uri allow-list
lib/curated-store.js       private Blob reader/writer for team answers
lib/slack.js               gap reports (webhook)
lib/slack-harvest.js       thread -> knowledge unit with provenance
scripts/build-index.mjs    corpus/ -> data/index.json
scripts/sources/           help center, transcripts, curated builders
scripts/dev-server.mjs     local Vercel-like server (+ OAuth routes)
scripts/e2e-oauth.mjs      full OAuth flow against a running server
scripts/audit-tokens.mjs   token-cost audit of a typical conversation
corpus/                    fictional help articles, transcript, tickets, curated answers
test/                      node:test suite
```

## License

MIT