case-chat
Provides integration with OpenAI-compatible endpoints (vLLM's /v1) for chat completion and tool use, enabling the server to use DiffusionGemma for conversational responses.
Stores a structured fake-case dataset in SQLite, allowing exact queries against ground-truth data such as timelines, entities, facts, flags, and observations.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@case-chatWhat was the guardianship petition filing date?"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
case-chat
A proof-of-concept conversational chat interface backed by DiffusionGemma served
via vLLM (OpenAI-compatible /v1) on a cloud-rented RTX 5090. It answers
questions using RAG over two read-only knowledge sources:
the synthetic test corpus of raw source documents (fictional Holcomb family Arkansas guardianship case), and
the existing domain-knowledge corpus (legal / behavioral / scripture), re-embedded into Qdrant and reachable via MCP.
It also exposes a structured fake-case dataset built from the synthetic corpus ground-truth (timeline / entities / facts / flags / observations) so exact questions like "when was the guardianship petition filed?" resolve against structured data — a stand-in for what case-project's extraction pipeline will eventually provide.
Hard boundaries
No Athena. LLM + embeddings go through vLLM / a Qwen3-Embedding-4B
/embeddingsendpoint, never the Athena daemon.No extracted data. RAG reads only raw source documents + the domain-knowledge corpus. case-project's
casedb(timeline events, evidence, observations, resolved participants, …) is off-limits. The fake-case dataset here is synthesized from the synthetic corpus's ground-truth, which is fictional — not the realcasedb.Data sovereignty. Only the fictional synthetic corpus and the non-sensitive domain-knowledge reference text leave local hardware. Real
case-data/never does.
Related MCP server: MCP-RAG-Assistant
Architecture
[Web app + chat orchestrator] ──MCP stdio──▶ [MCP retrieval server] ──▶ Qdrant
│ │
│ OpenAI /v1 (tools) ├─ /embeddings ─▶ Qwen3-Embedding-4B
▼ │ (TEI on box / local fallback)
DiffusionGemma (vLLM, 4-bit) └─ SQLite (fake-case dataset)Embedding contract (load-bearing)
Qwen/Qwen3-Embedding-4B · 2560-dim · cosine · L2-normalized · asymmetric
(queries wrapped Instruct: …\nQuery: …, documents bare). Kept identical to
domain-knowledge's build side so vectors converge. See
case_chat/embeddings/client.py.
Decisions
Architecture decisions are recorded under docs/decisions/.
Status
POC under construction. See the implementation plan / todo list.
Related MCP Connectors
Cloud or self-hosted knowledge for AI agents: hybrid search, reranking, GraphRAG, scoped MCP tools.
The CustomGPT.ai MCP server is a fully managed, RAG-powered endpoint that connects large language models with private knowledge bases and external data sources. It provides tools for retrieval-augmented generation queries (send_message), data ingestion (upload_file), and source listing, enabling AI agents to query private documents like PDFs with high accuracy and real-time citations.
Medical RAG: semantic search for clinical guidelines, drug interactions, diagnoses & EHR data.
Medical RAG: semantic search for clinical guidelines, drug interactions, diagnoses & EHR data.
Related MCP Servers
- FlicenseNot gradedqualityBmaintenanceMCP server for a modular RAG system that enables natural language question answering over enterprise documents with intent-aware routing, adaptive retrieval, and citation-backed responses.-
- AlicenseNot gradedqualityCmaintenanceHybrid RAG pipeline that indexes documents and exposes them via an MCP server, enabling natural language queries to retrieve relevant context chunks for LLMs.MIT
- FlicenseNot gradedqualityBmaintenanceEnables enterprise knowledge search through a chat interface, comparing traditional RAG with MCP-driven retrieval using hybrid BM25 and dense vector search, Ollama-powered answer generation, and question routing to domain-specific retrieval tools.-
- AlicenseNot gradedqualityCmaintenanceA modular RAG framework exposing knowledge retrieval tools via MCP, enabling AI assistants to perform hybrid search, reranking, and multimodal document queries with full observability and evaluation.MIT