nexla-mcp-doc-qa
by ArundhatiCat
README.md
# nexla-mcp-doc-qa
An MCP server that exposes grounded, source-attributed question-answering over
a collection of PDF documents. Built for the Nexla Software Engineer take-home
assignment.
> **A note on the data:** the assignment references "4–5 PDF files" that were
> supposed to arrive by email. I didn't have those in hand while building this,
> so `data/pdfs/` ships with 4 synthetic sample PDFs (`scripts/generate_sample_pdfs.py`)
> covering a fictional company, "Acme Corp," across a whitepaper, a security
> doc, a customer case study, and a product roadmap — deliberately written so
> some questions require pulling facts from more than one document. **To use
> the real assignment PDFs, just drop them into `data/pdfs/` (removing or
> keeping the samples, your call) and run `python scripts/build_index.py`.**
> Nothing else needs to change — ingestion is format/content-agnostic.
---
## 1. Quick Start
```bash
# 1. Clone and enter the repo
git clone <this-repo-url>
cd nexla-mcp-doc-qa
# 2. Create a virtualenv (recommended) and install dependencies
python3 -m venv .venv
source .venv/bin/activate # Windows: .venv\Scripts\activate
pip install -r requirements.txt
# 3. (Optional) put your own PDFs in data/pdfs/, replacing the samples.
# If you skip this, the 4 sample PDFs generated for this repo are used.
# 4. Build the search index (parses PDFs, chunks, vectorizes, persists to data/index.pkl)
python scripts/build_index.py
# 5. Run the MCP server (stdio transport, for local/dev clients like Claude Desktop)
python src/server.py
```
The server builds its index lazily on first tool call if `data/index.pkl`
doesn't exist yet, so step 4 is technically optional — but running it
explicitly is faster for iteration and lets you inspect the chunk count
before wiring up a client.
### Connecting a real MCP client (e.g. Claude Desktop)
Add this to your client's MCP config (e.g. `claude_desktop_config.json`):
```json
{
"mcpServers": {
"nexla-doc-qa": {
"command": "/absolute/path/to/.venv/bin/python",
"args": ["/absolute/path/to/nexla-mcp-doc-qa/src/server.py"]
}
}
}
```
### Optional: LLM-synthesized answers instead of extractive excerpts
By default the server works fully offline with zero external API calls
(see [Architecture](#3-architecture) for why). If you export an Anthropic
API key, it will automatically switch to using Claude to synthesize a
direct, cited answer from the retrieved passages instead of returning raw
excerpts:
```bash
export ANTHROPIC_API_KEY=sk-ant-...
python src/server.py
```
### Running the smoke test
`tests/test_mcp_protocol.py` launches the server as a real subprocess over
stdio (the actual MCP transport, not a direct Python function call) and
exercises `list_tools` and `call_tool`:
```bash
python tests/test_mcp_protocol.py
```
---
## 2. Tool Documentation
### `query_documents(question: str, top_k: int = 5) -> dict`
The core Q&A tool. Retrieves the `top_k` most relevant chunks across **all**
indexed documents (not scoped to one file), then produces a grounded answer.
**Input:**
| field | type | required | description |
|---|---|---|---|
| `question` | string | yes | Natural-language question |
| `top_k` | integer | no (default 5) | Number of source passages to retrieve/consider |
**Output:**
```json
{
"answer": "Here is what the documents say, most relevant first: ...",
"sources": [
{"document": "acme_security_and_compliance.pdf", "page": 1, "section": "1. Certifications", "relevance_score": 0.2248}
],
"mode": "extractive"
}
```
`mode` is one of `"llm"`, `"extractive"`, `"extractive_fallback"` (LLM call failed,
fell back gracefully), or `"none"` (nothing relevant found).
**Example query:** `"What SLA does Acme guarantee for platform uptime?"`
→ see `examples/interaction_log.md` for full real outputs.
### `list_documents() -> list[dict]`
Lists every PDF currently indexed, with how many distinct pages were
extracted from each. Lets an agent scope its questions before querying.
```json
[{"document": "acme_data_platform_whitepaper.pdf", "pages_indexed": 4}]
```
### `rebuild_index() -> dict`
Re-parses everything in `data/pdfs/` and rebuilds `data/index.pkl` from
scratch. Call this after adding/removing/replacing PDFs without restarting
the server.
```json
{"status": "ok", "documents_indexed": 4, "chunks_indexed": 9}
```
---
## 3. Architecture
```
data/pdfs/*.pdf
│
▼
┌─────────────────┐ pdfplumber: per-page text + best-effort heading
│ src/ingest.py │ extraction, then overlapping character-window
│ │ chunking (900 chars, 150 overlap, sentence-aware)
└────────┬────────┘
│ List[Chunk] (doc_name, page_number, section, text)
▼
┌───────────────────┐ TF-IDF vectorization (scikit-learn) over all
│ src/vector_store.py│ chunks; cosine similarity for retrieval.
│ │ Pluggable: swap backend="sentence-transformers"
└────────┬───────────┘ for dense embeddings if you have model-download access.
│ persisted → data/index.pkl
▼
┌────────────────┐ Takes retrieved chunks + question. If
│ src/qa.py │ ANTHROPIC_API_KEY is set, asks Claude to
│ │ synthesize a cited answer strictly from context.
└────────┬───────┘ Otherwise, returns attributed excerpts directly.
│
▼
┌─────────────────┐ FastMCP server exposing query_documents,
│ src/server.py │ list_documents, rebuild_index as MCP tools
└─────────────────┘ over stdio transport.
```
**Design decisions worth explaining:**
- **TF-IDF over dense embeddings, by default.** This assignment's corpus is
4–5 documents, not 4–5 million. At this scale, TF-IDF + cosine similarity
retrieves relevant passages just as reliably as a sentence-transformer
model, with three practical wins: zero model download (works in
network-restricted environments), zero inference cost, and zero
non-determinism to debug. The `VectorStore` class is written with a
`backend` parameter specifically so this is a one-line swap
(`backend="sentence-transformers"`) if the corpus grows or semantic
matching (synonyms, paraphrase) becomes the bottleneck — it isn't at this
scale, since the sample queries in `examples/interaction_log.md` all
retrieve the correct passage.
- **Extractive-by-default, LLM-optional answer synthesis.** The assignment
requires the server to be runnable locally without assuming any particular
paid API is available. Rather than hard-requiring an LLM call, `qa.py`
returns clearly attributed excerpts by default, and *upgrades* to
Claude-synthesized prose only if `ANTHROPIC_API_KEY` is present — with a
try/except that falls back to extractive mode if the API call fails for
any reason. This also makes `mode` in every response an honest signal of
how the answer was produced, which matters for a "grounded answer" tool.
- **Section/page-level attribution at chunk creation time**, not
reconstructed after retrieval. Each `Chunk` carries `doc_name`,
`page_number`, and a best-effort `section` heading guessed from the first
heading-like line on that PDF page. This means attribution is never
approximate or inferred after the fact — it's a property of the chunk
itself.
- **Persisted index, not rebuilt per-query.** `build_index.py` is a
separate step from `server.py` so the (cheap, but non-zero) parsing +
vectorization cost is paid once, not on every tool call. `rebuild_index`
is still exposed as a tool for when documents change without a server
restart.
---
## 4. Vibe Coding Setup
**Tools used:** Claude (via an agentic coding environment with file/bash
access), for essentially the entire build — scaffolding, the ingestion/chunking
logic, the TF-IDF vector store, the MCP server, the stdio protocol test, and
this README.
**How I directed it:** I didn't prompt this as one big "build the assignment"
request. I worked through it in the order a reviewer would actually care
about, checking real output at each step rather than trusting a large diff:
1. Set up the repo skeleton and got a *minimal* PDF → text → chunk pipeline
running and printing real output before writing anything downstream of it.
2. Only after chunking looked right did I add the vector store — and I
specifically asked for a pluggable backend rather than hard-wiring one
embedding provider, because I didn't want the whole submission to be
unrunnable if a reviewer's machine can't reach a model host.
3. Built `qa.py` with the extractive path first and the LLM path as a
strict opt-in on top of it, then had it run through 4 real questions
against the sample PDFs so I could read the actual answers, not just see
"tests passed."
4. For the MCP server itself, I hit a real, non-obvious bug: `from __future__
import annotations` in `server.py` broke FastMCP's runtime type
inspection (`issubclass(param.annotation, Context)` blew up because the
annotation was still a string). The AI's first instinct was to add type:
ignore comments around it; I overrode that and had it actually explain
*why* FastMCP needs real runtime types, then removed the future-import
instead of papering over the symptom. I wanted to understand the fix,
not just apply it.
5. I insisted on an actual stdio subprocess test (`tests/test_mcp_protocol.py`)
rather than accepting "the functions work when I call them directly in
Python" as evidence of MCP compliance — those are not the same claim,
and the assignment explicitly asks for a demonstrably working MCP
interaction, not a working set of Python functions.
**Where I leaned on the AI vs. overrode it:**
- Leaned on it heavily for: boilerplate (dataclasses, argparse-free CLI
scripts, the sample-PDF generator), the sentence-boundary-aware chunking
logic, and drafting this README's architecture section from the actual
code rather than from scratch.
- Overrode/corrected it on: the FastMCP annotation bug above; an early draft
that used FAISS for the vector store, which I cut in favor of TF-IDF
because it added a dependency and an index-format decision for no
retrieval-quality benefit at this document count — YAGNI; and an early
version of `qa.py` that silently returned an empty answer on retrieval
miss instead of an explicit "not found" message, which I flagged because a
document Q&A tool should never look confidently wrong.
**Overall view on AI tooling in a software engineering workflow:** it's
fastest and most reliable when treated as a very quick first-draft generator
for pieces I already have a clear spec for in my head — chunking logic,
protocol plumbing, boilerplate — and weakest when I let it make silent
architectural calls (embedding provider, vector store, error-handling
behavior) without pushing back and asking it to justify them against the
actual constraints of the task. The failure mode to watch for isn't broken
code, it's *plausible* code — the FAISS vector store worked fine, it just
wasn't the right call for a 4-document corpus. The discipline that mattered
most was running real inputs through every layer before moving to the next
one, so that a wrong architectural default surfaced as "why did I just add a
dependency I don't need" rather than as a bug three layers downstream.
---
## 5. Repository Layout
```
nexla-mcp-doc-qa/
├── README.md
├── requirements.txt
├── data/
│ ├── pdfs/ # source PDFs (sample docs included, see note above)
│ └── index.pkl # generated by scripts/build_index.py (gitignored)
├── src/
│ ├── ingest.py # PDF parsing + chunking
│ ├── vector_store.py # embedding + retrieval (pluggable backend)
│ ├── qa.py # grounded answer synthesis (LLM or extractive)
│ └── server.py # MCP server (FastMCP, stdio transport)
├── scripts/
│ ├── build_index.py # one-shot: parse all PDFs → data/index.pkl
│ └── generate_sample_pdfs.py # generates the sample PDF corpus
├── tests/
│ └── test_mcp_protocol.py # real stdio MCP client smoke test
└── examples/
└── interaction_log.md # sample Q&A transcript with source attribution
```
This server cannot be deployed
Maintenance
ActivityStale
ResponsivenessNo issues