Skip to main content
Glama
FeRepos

Employee Knowledge Assistant

by FeRepos

Employee Knowledge Assistant

RAG + MCP proof-of-concept: an internal assistant that answers employee policy questions using retrieval-augmented generation exposed through an MCP server, optionally summarized by Claude.

What this is

This repo is a TypeScript proof of concept for an Employee Knowledge Assistant at a fictional company, Nimbus Retail Inc. It retrieves passages from markdown policy docs, exposes that retrieval as an MCP search_knowledge tool (plus company://policies/* resources), and can have Claude answer only from those passages. A small Next.js App Router chat UI in apps/web is the human front end; the same MCP server can also be run standalone. The project is a growth exercise in wiring RAG, MCP, and Claude together—not a production HR system.

Claude is optional. Search, indexing, and MCP all run locally with no API key. Without ANTHROPIC_API_KEY, the chat UI still works in retrieval-only mode: it shows the matched policy snippets instead of a Claude summary.

Related MCP server: mcp-business-bot

Prerequisites

  • Node.js 20+ and npm (the MCP SDK requires Node 18+; the Next.js 16 app in apps/web is happiest on 20+)

  • Commands must be run from the employee-knowledge-assistant/ directory (the folder that contains this package.json), not from Documents/MCP or your home directory

  • Anthropic API key — optional. Retrieval and npm run mcp:dev never call Anthropic. Add a key only if you want Claude to write a short answer. Create one at console.anthropic.com → API keys (paid / trial credits; there is no lasting free Claude API)

Setup

  1. Clone this repository and cd into employee-knowledge-assistant.

  2. Install root dependencies: npm install.

  3. If you will run the chat UI, also install the Next.js app: npm --prefix apps/web install.

  4. Copy environment defaults:

    cp .env.example .env

    Leave ANTHROPIC_API_KEY empty for free retrieval-only chat, or paste a key for Claude summaries. Optional: CLAUDE_MODEL (defaults to claude-sonnet-4-6) and VECTOR_STORE_PATH (defaults to ./data/vector-store.json).

  5. Build the local TF-IDF index from knowledge/:

    npm run build:index

    You should see a chunk count and Index built successfully at ….

Then run it in one of two ways:

(a) MCP server only — useful with an MCP inspector or any MCP client:

npm run mcp:dev

This starts tsx mcp/server.ts on stdio. Success looks like:

employee-knowledge-assistant MCP server listening on stdio (37 chunks, 7 resources)

The process then waits silently for a client. That is expected — it is not a website and will not open a browser. Stop it with Ctrl+C. It loads the vector store from VECTOR_STORE_PATH and markdown from knowledge/. It does not call Claude.

(b) Full chat UI — Next.js spawns the MCP server itself:

npm run dev:web

(equivalent: npm --prefix apps/web run dev)

Open http://localhost:3000. Do not start mcp:dev in parallel for this path. The first POST /api/ask calls getMcpClient(), which runs npx tsx mcp/server.ts once per Next.js process and reuses that subprocess.

Keep the root .env at the repository root. The API route also loads ../../.env when the Next.js cwd is apps/web. There is a copy of the template at apps/web/.env.example. After changing .env, restart the Next.js process.

Project structure

employee-knowledge-assistant/
├── apps/web/          # Next.js App Router chat UI and POST /api/ask
├── server/
│   ├── rag/           # Load, chunk, TF-IDF embed, vector store, retrieve
│   ├── claude/        # Anthropic client + grounded system/user prompts
│   └── orchestrator/  # MCP client + answerQuestion() pipeline
├── mcp/
│   ├── server.ts      # Stdio MCP server (tool + resources)
│   ├── tools/         # search_knowledge Zod schemas and handler
│   └── resources/     # company://policies/* URI map and readers
├── knowledge/         # Nimbus Retail policy markdown (the corpus)
├── scripts/           # build-index.ts — write the JSON vector store
├── tests/             # Vitest unit tests + gated e2e
├── docs/              # Architecture, data flow, interview notes
├── .env.example
├── package.json
└── tsconfig.json

Example questions

These match facts in knowledge/. With a Claude key, wording varies but the numbers should not. Without a key, you get the raw retrieved passages (same facts).

Question

Expected

How many WFH days are allowed?

Answered: 2 WFH days per week, requested in WorkSync at least 24 hours ahead. Sources typically include knowledge/wfh-policy.md.

How many annual leave days do employees get?

Answered: 18 days per year for full-time staff.

What happens if I submit an expense claim late?

Answered: submit in ExpenseFlow within 30 days; late reports are rejected unless Finance grants an exception.

What is the domestic hotel allowance for business travel?

Answered: $180 per night domestic, booked through Nimbus Travel Desk.

How long does a standard insurance claim take to process?

Answered: 15 business days after a complete HealthPortal submission.

What is the company dress code policy?

Not in the knowledge base. Retrieval should fail the 0.05 score floor (found: false / sufficientContext: false). The assistant should say it does not have that information rather than invent a dress code.

What is the cafeteria menu?

Same as above — out of scope for these policies.

Running tests

npm test          # vitest run — fast, mocked; e2e file self-skips
npm run test:watch
npm run test:e2e  # RUN_E2E_TESTS=true; real Claude + real MCP (needs a key; costs a little)

npm test covers loaders, chunking, TF-IDF, vector-store round-trip, retriever ranking, MCP tool/resource helpers, orchestrator assembly, retrieval-only mode, and error-handling fallbacks. tests/e2e.test.ts is skipped unless RUN_E2E_TESTS=true. It also skips (does not fail) if ANTHROPIC_API_KEY is unset. The e2e suite builds a temporary index so it does not overwrite ./data/vector-store.json.

Architecture

A question hits the Next.js UI, then POST /api/ask, which reuses one MCP stdio client. The orchestrator calls the MCP search_knowledge tool, which runs TF-IDF retrieval against the JSON store. If ANTHROPIC_API_KEY is set, Claude sees only those chunks (or an explicit “no context” user message). If the key is missing, answerQuestion() returns the retrieved passages directly. Either path yields an AnswerResult with answer, sources, and sufficientContext. Tools vs full-document resources, layering, and why TF-IDF is used here are in docs/architecture.md. A numbered request trace is in docs/data-flow.md.

Troubleshooting

Symptom

What the code is doing

Fix

npm error enoent … /Users/…/package.json

npm run mcp:dev (or npm test) was run from the wrong folder

cd employee-knowledge-assistant first — the directory that contains this README and package.json.

MCP prints listening on stdio then “does nothing”

Stdio server is idle waiting for an MCP client

Expected. Use an inspector, or use the chat UI (npm run dev:web) instead of treating this as a website.

MCP process exits immediately: Failed to load vector store at … and “Run npm run build:index first” on stderr

mcp/server.ts calls loadRetriever(storePath) and process.exit(1) if the JSON file cannot be loaded

From the repo root, run npm run build:index. Confirm VECTOR_STORE_PATH (default ./data/vector-store.json).

Chat answer: “I'm having trouble accessing the knowledge base right now…”

answerQuestion() caught a failed callSearchKnowledge() (MCP tool error, spawn/stdio failure after connect, invalid tool JSON, etc.), logged it with console.error, and returned a fallback AnswerResult

Ensure build:index succeeded. Restart the Next.js process so getMcpClient() can spawn a fresh npx tsx mcp/server.ts from the repo root. Check stderr for MCP tool call failed / Failed to connect to the MCP knowledge server.

Chat starts with “No Anthropic API key is set, so this is the retrieved policy text…”

Retrieval-only mode (hasClaudeApiKey() is false)

Optional. Add a key to .env and restart dev:web if you want a Claude summary. Leave it blank to stay free.

Chat answer: “I couldn't generate a response right now. Please try again.”

Claude messages.create threw or returned no text (invalid key, billing, or model name). Empty key no longer takes this path

Check the key at console.anthropic.com, billing/credits, and CLAUDE_MODEL. Or clear the key to use retrieval-only mode.

HTTP 400 Malformed JSON request body. or Question must be a non-empty string.

/api/ask rejected the POST body

Send JSON { "question": "…" } with a non-empty string.

HTTP 504 The request took too long. Please try again.

Promise.race in /api/ask fired after 30 seconds. The underlying MCP/Claude work is not cancelled

Retry. For a hung MCP child, restart next dev. This timeout is client-facing only.

HTTP 500 Something went wrong while answering. Please try again.

Uncaught error in the route (often MCP connect failing before answerQuestion), logged server-side without a stack in the JSON body

Same as spawn issues: index present, tsx available, cwd such that mcp/server.ts exists (the client walks up to the repo root).

Next.js Module not found: Can't resolve './mcpClient.js'

Turbopack does not map NodeNext .js specifiers to .ts files

apps/web is configured to use next dev --webpack with a .js → .ts extension alias. Use npm run dev:web (not a plain next dev without --webpack).

Answer says the knowledge base has no information; UI shows “No sources”

Retriever minScore default 0.05 filtered everything; found: false

Rephrase toward words that appear in knowledge/ (policy names, WorkSync, ExpenseFlow). Rebuild the index after editing markdown. TF-IDF is lexical—synonyms may miss.

Low-quality or off-topic snippets

Cosine similarity over TF-IDF bags of words, topK 3 from the orchestrator

Rebuild index after corpus changes. This POC does not use neural embeddings.

What I'd change for production

  • Replace local TF-IDF with a real embeddings model (API or local) so retrieval is semantic, not token-overlap.

  • Replace data/vector-store.json with a real vector database (indexes, filters, concurrent writers).

  • Add authentication and authorization on POST /api/ask (there is none today).

  • Wire real end-to-end cancellation (the 30s Promise.race does not abort Claude or the MCP child).

  • Structured logging and monitoring (tool latency, retrieval scores, Claude errors) instead of console.error only.

  • Run the MCP server as a long-lived process (or pool), not one stdio subprocess per Next.js server instance, so it can scale independently of the web tier.

Related MCP Connectors

Related MCP Servers

  • F
    license
    Not graded
    quality
    B
    maintenance
    Enables grounded question-answering over internal documents via a single MCP tool that retrieves relevant passages and generates answers with citations, returning sources and diagnostics.
    -
  • F
    license
    Not graded
    quality
    B
    maintenance
    Enables MCP clients to ask plain-language questions and receive answers grounded only in documents the configured role is cleared to read, with the same access-controlled tools available across any client.
    -