Skip to main content
Glama
mecoskun

Sourcebook

by mecoskun

Sourcebook — Document Q&A

Ask questions about uploaded documents and inspect the passages behind each answer. Sourcebook combines MCP retrieval, CPU embeddings, and a self-hosted open-source language model. No paid LLM API is used.

Status: local application implemented. CPU inference and the MCP workflow have been exercised locally; evaluation results are in docs/evaluation-results.json. Public backend hosting is not provisioned. Docker configuration is supplied; see validation notes below.

What you can try

  • Upload a text-based PDF, DOCX, or UTF-8 TXT, or load the fictional employee handbook.

  • Ask an independent question and get an answer with source citations.

  • Expand each source to read its exact extracted passage.

  • Remove documents or clear your temporary workspace.

  • If generation is unavailable or fails validation, inspect clearly labeled retrieved excerpts instead.

Limits: 10 MB/file, five documents/workspace, 100 PDF pages, 150,000 extracted characters and 250 chunks/document. Scanned PDFs/OCR, password-protected PDFs, and image understanding are excluded. DOCX/TXT cite paragraph/section locations because they have no stable page layout. Extraction of complex tables, headers, and unusual PDF layouts is imperfect.

Related MCP server: ragi

Run locally

Install Python 3.12 and uv, then run from this repository:

uv sync --frozen --python 3.12
uv run python scripts/prepare_models.py

Preparation downloads about 2.7 GB of model assets into ignored local folders. It verifies the pinned model and embedding checksums in models.lock.json. On Windows it also downloads and checks the official llama.cpp CPU runtime. On Linux/macOS, install a compatible llama-server or use Docker.

Start the shared CPU model in one terminal:

uv run python scripts/start_model.py

Start the application in another terminal:

uv run uvicorn app.main:app --host 127.0.0.1 --port 8000

Open http://127.0.0.1:8000. Load the sample and ask “How many vacation days do employees get?” The application also works in explicitly labeled excerpt mode if the model service is stopped.

Copy .env.example to .env only if you need to change settings. The app loads it relative to the project directory. Never commit .env, tokens, uploaded documents, indexes, or model binaries.

Model choices

  • Generation: Qwen3 4B GGUF distributed by Ollama, pinned by SHA-256. Qwen3 is Apache-2.0 licensed. The compact model is a CPU baseline, not a claim of enterprise-grade answer quality. The initial 1.7B comparison is retained in docs/evaluation-1.7b.json.

  • Embeddings: BAAI/bge-small-en, ONNX via FastEmbed, MIT licensed. English-first retrieval, loaded once by the MCP server. Assets come from Qdrant's published download.

  • Serving: llama.cpp, CPU only, four threads, 8,192-token context, one generation slot. A small Qwen3 ChatML template explicitly closes the thinking prefix so all output tokens are spent on the answer. Native Windows runtime version and Docker image digest are pinned separately.

Model provenance: Qwen model card, Ollama Qwen3, FastEmbed, llama.cpp.

Architecture

flowchart LR
    UI[Static web interface] --> API[FastAPI backend]
    API --> Parser[Bounded parser subprocess]
    API --> Bridge[MCP client / stdio]
    Bridge --> Tools[MCP document tools]
    Tools --> Index[Per-session in-memory vector index]
    API --> Model[Shared llama.cpp CPU service]

Uploads are parsed into text sections and indexed by the MCP index_document tool. Questions always trigger the MCP search_documents tool; the language model receives only retrieved passages. This is a deterministic retrieval workflow, not an autonomous tool-selecting agent. MCP stays internal, and the model cannot choose another user's session or invoke upload/deletion tools.

The model returns structured statements with supporting source IDs. The backend rejects unknown/missing IDs and attaches citation markers itself. Valid IDs do not prove that a claim is supported: model faithfulness remains an evaluation concern. Similarity is a ranking signal, not a calibrated confidence percentage.

Retrieval combines semantic similarity with rare query-word overlap and lightweight English word-ending normalization. The semantic threshold is 0.80; candidates down to 0.75 also need two matching content terms. A second model call checks whether drafted statements follow from cited evidence, address the whole question, and avoid ignoring contradictions in other retrieved passages. This is an additional guard, not an independent factual guarantee: the checker uses the same small model. These ranking constants were exercised on the handbook and the public NIST fixture; they are not calibrated probabilities and need broader validation for other domains.

app/model.py is the reusable inference adapter for the future transcript-only YouTube project. Its endpoint can point to one shared private model service. No OpenAI account/key is needed despite the compatible HTTP request format.

Isolation, retention, and resource limits

  • The backend creates unguessable bearer tokens. Only token hashes are held in backend memory; the browser keeps its token in session storage.

  • The backend supplies the MCP workspace ID, never a user-selected ID. Documents cannot be listed, searched, or deleted across workspaces.

  • Extracted text and vectors stay in memory. Original uploads may briefly use the framework's temporary upload spool and are closed after parsing. No user documents are committed or deliberately retained on disk.

  • Workspaces expire after one hour; a cleanup loop removes their server-side indexes within approximately 30 seconds after expiry, once any active operation finishes. Explicit deletion is also available. Browser-displayed answers remain until cleared or the page is reloaded.

  • One expensive operation runs at a time; excess requests receive a retry message. Defaults: 20 questions/workspace, 200 questions/day, 100 uploads/day, 100 new sessions/day, 20 active workspaces.

  • Counters/indexes are process-local and reset on restart. Run one API worker. Multi-instance deployment needs shared quotas and session storage.

  • Request bodies are byte-capped even without Content-Length. Document parsing has a 20-second timeout in a separate process and a 768 MB address-space limit on Linux. The Docker API container adds a memory/process limit. Windows local parsing does not have the Linux memory cap.

  • No chat history is sent to the model. Each question stands alone. The UI renders document/model text without interpreting HTML.

This is a bounded portfolio demo, not an audited confidential-document service. Use non-sensitive demo files. Public hosting needs HTTPS, explicit allowed origins, and a reverse proxy with upload/concurrency limits. Prompt injection and incorrect model answers remain possible; see the evaluation record.

Tests and evaluation

uv run pytest -q
uv run ruff check --config pyproject.toml app tests scripts
node --check frontend/app.js
uv run python scripts/evaluate.py

Unit/integration tests cover malformed inputs, PDF/DOCX/TXT extraction, chunk provenance, request caps, authentication/expiry, cross-workspace retrieval/deletion, model output validation, and actual MCP subprocess initialization/cleanup. These tests need no model download.

The live evaluation needs the app and model running. It uses a small fictional handbook and saves outputs/timings in docs/evaluation-results.json. Qwen3 4B passed 19 of 20 fixture checks, with a median response time of 6.8 seconds. The missed case was an unnecessary abstention on remote-work days. Its keyword checks are smoke tests, not a general answer-accuracy or citation-faithfulness score. Do not extrapolate local CPU speed to a shared hosting plan.

Docker and hosting

Prepare the model files first, then:

docker compose up --build

The API is bound to localhost:8000; the model has no exposed host port. Both containers use CPU. The compose file budgets 6 GB for inference and 2 GB for the API. The native model used approximately 5.34 GB locally; allow additional memory for the operating system (12 GB total is a safer starting point). Benchmark the chosen host before purchase. The Docker daemon was unavailable during initial Windows development, so container execution must be verified before deployment.

For deployment, place an HTTPS reverse proxy in front of the API, configure QA_ALLOWED_ORIGINS to your exact GitHub Pages origin, and keep model access private. The shared USD 20–30/month target is a planning estimate, not a purchased hosting plan or guarantee.

GitHub Pages publishes only frontend/. The manual Publish interface workflow accepts a public HTTPS backend origin and writes it to config.js safely. GitHub Actions secrets cannot make a secret embedded in static JavaScript private. Use backend environment configuration for credentials.

Without a backend origin, the public Pages interface enters a clearly labeled prepared sample demo. It serves two recorded model answers from frontend/demo.json; live uploads are disabled and other questions are not answered. Export these examples with uv run python scripts/build_demo.py after a successful evaluation. This preview is not the live hosted RAG backend.

Repository map

Location

Responsibility

frontend/

Responsive, dependency-free interface; GitHub Pages compatible

app/main.py

HTTP API, session lifecycle, quotas, upload orchestration

app/extraction.py, app/parse_worker.py

Bounded extraction and provenance

app/mcp_client.py, app/mcp_server.py

MCP lifecycle and document tools

app/retrieval.py

CPU embeddings and isolated vector search

app/model.py

Shared local-model adapter and citation validation

tests/

Offline regression tests and MCP integration checks

scripts/

Model preparation, serving, and live evaluation

docs/PROJECT_BRIEF.md

Agreed scope and handoff context for both projects

Credits

Inspired by the MCP teaching examples in Dave Ebbelaar's AI Cookbook. This is a separate implementation with an upload pipeline, isolated sessions, a web interface, local generation, and tests. Third-party libraries and model assets retain their respective licenses; model binaries are downloaded separately, not redistributed in this repository.

Informal question wording regression

A follow-up fix restricts verification to cited evidence and distinguishes informal eligibility questions from explicit accrual questions. The exact question "when does employee earn vacation" now returns the documented six-month availability rule. All eight focused live checks passed; see docs/wording-regression.json. These supplement the earlier 20-case baseline; they do not replace it or establish general accuracy.

Broader format reliability pass

Run uv run python scripts/evaluate_reliability.py with the backend/model running. It creates synthetic DOCX, two-page PDF, and TXT uploads and checks answers and citation locations through the real API/MCP path. The 12 recorded cases cover tables, informal wording, missing facts, and a document instruction attempting to alter a fact. All 12 are accepted after review of one exact cited missing-policy answer; the original 11/12 automated count and unchanged outputs are retained in docs/reliability-results.json. Offline suite: 25 tests.

DOCX extraction preserves paragraph/table order and repeats each table's first row as context for subsequent rows. Citations identify table and row. This assumes the first row is useful header context. Merged and nested DOCX tables are rejected with an explanation because the current parser cannot interpret them reliably. Unusually large tables and complex PDF layouts remain limitations. These synthetic fixtures do not substitute for representative real documents. Future phases and resume notes are in docs/ROADMAP.md.

Pre-hosting acceptance checks

Run uv run python scripts/evaluate_prehosting.py. The runner downloads one checksum-pinned public NIST SP 1300 PDF into ignored .runtime/, then exercises it alongside a long synthetic DOCX and conflicting policy fixtures. Reports are retained under docs/prehosting-*.json. Source: https://nvlpubs.nist.gov/nistpubs/SpecialPublications/NIST.SP.1300.pdf . The nine-page guide is a real public document; the long DOCX and conflict cases are controlled synthetic fixtures.

Short context-dependent questions such as “When does it open?” now ask the user to name the subject. This is a deliberately narrow grammar guard, not a general ambiguity detector. Questions remain independent; chat history is not supplied to the model. Conflicting policies may produce a cautious refusal rather than a complete comparison. Verification can detect only conflicts in retrieved passages, not every passage in every uploaded document.

Pre-hosting review completed

See docs/PREHOSTING_REVIEW.md for results and known limits. Offline tests: 41 passed; pre-hosting acceptance: 13/13; format cases: 12/12. A later handbook rerun exposed a founding-date error, now covered by a deterministic missing-origin-evidence guard and four passing targeted API cases. Failed runs remain in the repository; these separate suites are not a single general accuracy score. Next phase: CPU hosting benchmarking.

Available Tools

4 tools
delete_documentA

Delete one document, or all documents in this workspace when no ID is provided.

ParametersJSON Schema
NameRequiredDescriptionDefault
session_idYes
document_idNo

TDQS

A3.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the behavioral disclosure burden. It usefully reveals the dangerous default that omitting document_id deletes all documents in the workspace. However, it does not mention irreversibility, permission requirements, or what happens after deletion, which are important for a destructive operation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence with no filler. It efficiently conveys the core action and the critical all-documents branch without redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given this is a destructive tool with no annotations and no output schema, the description is not complete enough for safe invocation. It warns about the all-documents case but omits session_id semantics, postconditions, and any safety or error context.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It explains document_id semantics well: present means one document, absent means all documents. But it does not explain the required session_id parameter at all, leaving half the parameters unexplained despite the low schema coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly names the action ('Delete') and the resource ('document'), and distinguishes two modes: deleting one document or all documents in the workspace when no ID is provided. This makes its purpose distinct from siblings like index_document, list_documents, and search_documents.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit usage guidance for both parameter cases: provide a document_id to delete a single document, omit it to delete all documents in the workspace. It does not name alternatives or exclusion conditions, but no sibling is a deletion tool, so the context is reasonably clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

index_documentB

Index extracted document passages for one private workspace.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYes
sectionsYes
session_idYes

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description must carry the burden of disclosing side effects and constraints. It only mentions the private-workspace scope but does not explain whether re-indexing overwrites existing data, whether duplicates are rejected, what permissions are needed, or what response to expect.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is short and front-loaded with the core verb, which is good for scanning. However, the brevity reflects missing substance rather than efficient coverage of all relevant information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with three required parameters, no annotations, and no output schema, this description is incomplete. It provides useful context about private-workspace scoping but omits parameter semantics, operation behavior, and return conventions.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema has zero description coverage for all three required parameters, so the description must compensate. It weakly hints that 'sections' correspond to extracted passages, but it never explains session_id, name, or the expected structure of sections, leaving the agent without enough information.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Index'), a resource type ('extracted document passages'), and a clear scope ('one private workspace'). It is immediately distinguishable from sibling tools like list_documents and delete_document because the operation type is explicit and different.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit statement about when to choose this tool over alternatives. However, the intended usage is reasonably implied by the verb 'Index' and the sibling tool names, which makes the guidance adequate but not explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_documentsB

List source documents belonging to one workspace.

ParametersJSON Schema
NameRequiredDescriptionDefault
session_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. It only names the action and scope; it does not mention read-only guarantees, pagination, ordering, result limits, or any other behavior beyond the basic listing action.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is one short, front-loaded sentence with no filler. Every word contributes meaning, which is ideal for an agent scanning tool definitions.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

While the tool is simple and has an output schema, the description does not sufficiently connect the required session_id parameter to the workspace context, and it omits behavioral details that would be needed without annotations. The purpose is clear, but the invocation context is incomplete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, and the description never names or explains session_id. The phrase 'one workspace' gives an indirect hint, but the agent still has to guess how session_id maps to a workspace and what format it should take.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('List'), a specific resource ('source documents'), and a clear scope ('belonging to one workspace'). This distinguishes it from the sibling tools index_document, delete_document, and search_documents.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives a clear context for when to use the tool – enumerating source documents in a workspace – but it does not explicitly mention alternatives or state when to prefer search_documents over this tool. The usage boundary is mostly implied rather than stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_documentsA

Find relevant passages with source locations from this workspace only.

ParametersJSON Schema
NameRequiredDescriptionDefault
questionYes
session_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the behavioral disclosure burden. It does reveal that results include source locations and that search is limited to the current workspace, which is useful behavioral context. Still, it does not explicitly state that the operation is read-only or describe any side effects, permissions, or limits, though 'find' implies a non-mutating search.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single front-loaded sentence with no filler. It states the action, the output characteristic (source locations), and the scope ('this workspace only') efficiently, making it easy to scan and parse.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given that an output schema exists, the return shape is covered elsewhere. The description provides a minimal but adequate purpose and scope for a simple search tool, but it omits parameter semantics and explicit guidance on when to use it versus siblings, leaving noticeable gaps for an agent to resolve.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate for undocumented parameters, but it does not. It only indirectly hints that session_id scopes the search to a workspace; it gives no guidance on what constitutes a well-formed 'question' or how to obtain/supply the session_id.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Find') and resource ('relevant passages with source locations'), and scopes the operation to the current workspace. This clearly distinguishes it from sibling tools like index_document, list_documents, and delete_document, which perform different operations.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The tool's purpose implies when to use it: when the agent needs to find relevant passages rather than list, index, or delete documents. However, it does not explicitly name alternatives or state constraints like 'use instead of list_documents when content-level search is needed,' leaving some inference to the agent.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 4 tool updatesv0.1.0
    • First observeddelete_document
    • First observedindex_document
    • First observedlist_documents
    • First observedsearch_documents

TDQS

A3.7/5.0

Scored across 4 tools

Disambiguation5/5

Each tool targets a distinct operation: indexing, listing, deleting, and searching documents. There is no meaningful overlap in purpose, so an agent should have no trouble selecting the right tool.

Naming Consistency4/5

All names follow a clear verb_document(s) snake_case pattern, with verbs positioned first. The only minor inconsistency is the singular/plural noun form depending on whether the operation targets one document or a collection.

Tool Count5/5

Four tools is a well-scoped surface for a source-document workspace: add, list, search, and delete. Each tool serves a distinct core need without unnecessary bloat.

Completeness4/5

Core document lifecycle coverage exists: index, list, search, and delete. The main gaps are the lack of an explicit update/refresh operation and a full-document retrieval tool, though these can be worked around by deleting and re-indexing or by using search results.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    D
    maintenance
    An MCP server implementation that provides tools for retrieving and processing documentation through vector search, enabling AI assistants to augment their responses with relevant documentation context
    20 npm
    265
    MIT
  • A
    license
    A
    quality
    D
    maintenance
    Local-first RAG indexing and semantic search MCP server. Enables document retrieval and context-aware queries using local embedding models.
    3
    5 npm
    MIT
  • A
    license
    Not graded
    quality
    D
    maintenance
    An MCP server that indexes documents and serves relevant context to LLMs via Retrieval Augmented Generation (RAG).
    22 npm
    36
    MIT
  • F
    license
    Not graded
    quality
    D
    maintenance
    An MCP server that exposes document retrieval as tools (semantic search and source listing) for any LLM, using a vector index built from DocPilot's ingestion pipeline.
    -