ResearchMind MCP
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@ResearchMind MCPWhat do my uploaded papers say about retrieval-augmented generation?"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
ResearchMind MCP
An AI research assistant for working with a personal corpus of academic papers, exposed both as a REST API and as a Model Context Protocol (MCP) server.
Development status: pre-alpha. The system does not yet run. This repository is an architectural foundation with an approved delivery plan. Most capabilities described below are not implemented. The Development Status section is exact about what exists. Nothing here is production-ready or safe to expose to a network.
What It Does
The goal is a research assistant for an individual academic or a small research group. A researcher uploads papers, and the system makes their own corpus answerable:
Semantic search across everything they have uploaded.
Grounded answers to research questions, drawn only from their documents and returned with citations that resolve to a real page and section.
Document-scoped analysis — summarise a paper, compare several.
MCP is the integration seam. The same capabilities are reachable from the project's own web client and from an external MCP host such as Claude Desktop, because both are thin adapters over one service core.
What it is not. Not a literature search engine — it works on documents you supply. Not multi-tenant SaaS. Not a chatbot with general knowledge: if the answer is not in your corpus, the correct response is "no relevant sources", and the system is built to say so rather than to improvise.
Related MCP server: Athena
Architecture
A modular monolith. REST and MCP are sibling adapters over a shared Service Core; neither calls the other.
┌──────────────────┐ ┌────────────────────────┐
│ Next.js client │ │ MCP host │
│ (browser) │ │ (e.g. Claude Desktop) │
└────────┬─────────┘ └───────────┬────────────┘
│ HTTPS + Bearer JWT │ stdio subprocess
┌────────▼─────────────────┐ ┌────────────▼─────────────┐
│ REST ADAPTER │ │ MCP ADAPTER │
│ backend/api/ │ │ mcp_server/ │
│ routers · schemas │ │ tool registry+dispatch │
│ └────────────────┼───┼──► one identity resolver │
└────────┬─────────────────┘ └────────────┬─────────────┘
└───────────────┬──────────────────┘
┌────────────────────────▼───────────────────────────────┐
│ SERVICE CORE — backend/services/ │
│ Auth · Document · Ingestion · Search · Research │
│ every method takes an authenticated Principal │
└──┬──────────┬───────────┬───────────┬──────────────┬───┘
▼ ▼ ▼ ▼ ▼
┌──────┐ ┌─────────┐ ┌────────┐ ┌──────────┐ ┌────────────┐
│Repos │ │ Object │ │Embed │ │ Vector │ │ LLM │
│ │ │ storage │ │Provider│ │ Index │ │ Provider │
└──┬───┘ └────┬────┘ └───┬────┘ └────┬─────┘ └─────┬──────┘
▼ ▼ ▼ ▼ ▼
PostgreSQL volume FastEmbed Qdrant Anthropic
(SOURCE OF (content- (local, (INDEX (Claude)
TRUTH) addressed) 384-d) ONLY)
▲
┌────┴─────┐
│ Redis │ ARQ job queue + rate limits
└──────────┘Storage responsibilities
Store | Owns | Never |
PostgreSQL | Sole authority for what exists and who owns it | Vectors |
Qdrant | An index: vectors + | A source of truth |
Redis | ARQ job queue, job status, rate-limit counters | Anything whose loss is unrecoverable |
Object storage | Original uploaded bytes, content-addressed | Anything derivable |
Isolation invariant. Retrieval filters on
user_idin Qdrant (fast path) and re-validates every chunk against PostgreSQL ownership before any content reaches a prompt (correct path). If the two ever disagree, the system returns fewer results — never another user's document.
Rationale for every structural choice is in docs/adr/.
Technology Stack
Layer | Technology |
Language / runtime | Python 3.12 · Poetry |
API | FastAPI · Uvicorn · Pydantic v2 · pydantic-settings |
System of record | PostgreSQL 16 · SQLAlchemy 2 (async) · asyncpg · Alembic |
Vector index | Qdrant (cosine) |
Jobs & cache | Redis 7 · ARQ |
LLM | Anthropic Claude ( |
Embeddings | FastEmbed · |
Document parsing | PyMuPDF (block mode, thread-offloaded) |
MCP |
|
Auth | JWT ( |
Frontend | Next.js 14 (App Router) · React 18 · TypeScript · Tailwind · TanStack Query · Zustand · axios |
Testing | pytest · pytest-asyncio · testcontainers · httpx · gitleaks |
Quality | ruff · black (both enforced in CI) · mypy strict (enforced in CI over the M2 security surface; not yet repository-wide) |
Infrastructure | Docker · Docker Compose · GitHub Actions · Dependabot |
Embeddings run locally, so a full stack needs exactly one secret:
ANTHROPIC_API_KEY. See ADR-0004.
Repository Structure
Path | Contents |
| REST adapter ( |
| MCP adapter — server, tools, resources, prompts. Importable; handlers stubbed until M6 |
| Agent layer (collapsed to a single |
| Domain models, interfaces (ABCs), utilities — the layer everything depends on |
| RAG ingestion: parse → chunk → embed |
| Qdrant adapter |
| Redis adapter |
| Next.js web client |
|
|
| Dockerfiles and infrastructure configuration |
| Developer and CI helper scripts |
|
Development Status
Verified by execution, not by assumption. The full evidence base is in docs/architecture/AUDIT-2026-09.md.
What genuinely works
Component | State |
Domain models ( | ✅ Complete — 12 Pydantic v2 models |
Interfaces ( | ✅ Complete — 4 ABCs |
API schemas ( | ✅ Complete — 9 DTOs |
Settings ( | ✅ Complete |
FastAPI app construction | ✅ Builds and mounts routers |
Repository foundation | ✅ Docs, ADRs, CI hygiene, workflow |
What does not work
Area | State |
REST endpoints | ⚠️ Auth, |
Authentication & authorisation | ✅ M2 complete. Register and log in over HTTP; a forged, expired or unresolvable token is 401 on every protected route; a role without the permission is 403; another user's document is 404 whatever your role |
MCP layer | ⚠️ Imports correctly and lists its 7 tools; no tool handler is implemented yet (M6) |
RAG pipeline | ⚠️ PDF text extraction (M3/S3.3), section-aware chunking (M3/S3.4), embedding into Qdrant (M3/S3.5) and owner-validated retrieval reachable at |
Persistence | ⚠️ Schema (migrations |
Agents | ❌ Return |
Container builds | ⚠️ Two images, not three — the MCP container was removed in M0/S0.2 (ADR-0002 settled on stdio). Both were made to build in M0/S0.4–S0.5; not re-verified since |
Test suite | ✅ 941 passed, 2 xfailed, against real PostgreSQL and Redis in CI |
Roughly 10% complete by the pre-M0 audit's count. That figure has not been re-measured since, and is left as the audit stated it rather than revised by guess — M0 through M2/S2.4 have since replaced a good deal of declaration with behaviour, but no one has counted again.
Prerequisites
Python 3.12 · Poetry 1.8+
Node.js 20+ · npm
Docker + Docker Compose v2
An Anthropic API key
Git
Local Development
The full stack is not runnable yet. The frontend half now is: its image
builds and serves every route (Sprint M0/S0.4), and poetry check passes
(M0/S0.1). What remains unproven is the backend half of docker compose up
— finishing it is the rest of Milestone M0.
What works today
git clone <repository-url>
cd researchmind-mcp
cp .env.example .env # then set ANTHROPIC_API_KEY
cd frontend && cp .env.example .env.local && cd ..
./scripts/check-hygiene.sh # repository hygiene checksThe frontend runs on its own (verified in Sprint M0/S0.4):
cd frontend
npm ci # reproducible: installs from package-lock.json
npm run dev # http://localhost:3000
npm run build # production build, emits .next/standalone
npm run lint # ESLint via next/core-web-vitals
npm run type-check # tsc --noEmitNEXT_PUBLIC_* variables are inlined into the browser bundle at build time,
so they must be set before npm run build, not on the running container. See
docs/development/environment.md.
After Milestone M0 (not yet available)
docker compose up --build # full stack — backend half still unproven
poetry install && poetry run python main.py # backend onlyThis section is updated as each milestone makes a workflow genuinely functional. Commands are not documented here before they work.
Testing
poetry run pytest # ✅ 941 passed, 2 xfailed (needs postgres + redis; see testing.md)
poetry run ruff check . # ✅ enforced in CI
poetry run black --check . # ✅ enforced in CI
poetry run mypy . # ⚠️ 113 errors repo-wide; strict-clean and CI-enforced over the M2 and M3 surfaces
./scripts/check-hygiene.sh # ✅ enforced in CIThe two xfails are deliberate and strict: they pin behaviour that does not
exist yet (chunking, M3; readiness probing, M9) and will fail the build the day
it starts working, rather than passing silently.
Strategy, test levels and the blocking release-gate suites are in docs/development/testing.md.
Git Workflow
main (protected, validated states only) ← milestone/* ← sprint/*.
Conventional Commits. Tags mark validated milestones, never aspirational ones.
Full detail: docs/development/workflow.md.
Roadmap
Milestone | Outcome |
M0 | Build integrity — images build, imports resolve, CI green |
M1 | System of record — Postgres, repositories, migrations |
M2 | Fail-closed authentication and authorisation (complete — S2.1–S2.5) |
M3 | Document ingestion — PDF to owned, searchable vectors (S3.1 upload & storage, S3.2 ingestion worker, S3.3 PDF text extraction, S3.4 section-aware chunking, S3.5 embeddings & vector store done) |
M4 | Vector index & isolated retrieval — done (S4.1 retrieval foundation, S4.2 search API & composition root) |
M4 | Tenant-isolated retrieval |
M5 | Grounded answering — first working end-to-end flow |
M6 | MCP adapter |
M7 | RAG evaluation harness |
M8 | Web client |
M9 | Hardening and observability |
M10 | Release validation |
docs/roadmap/MILESTONES.md · docs/roadmap/COMPLETION_PLAN.md
Security
Authentication and authorisation are fail-closed as of M2: accounts are created and signed into over HTTP, every protected route refuses a token it cannot resolve to an active user (401), refuses a role that lacks the permission (403), and scopes every document to its owner in the SQL itself (404). There is still no rate limiting (M9), and the system is pre-alpha, so it must not be exposed to an untrusted network. Principles and invariants: docs/security/principles.md. To report a vulnerability, see SECURITY.md — please report privately.
Contributing
See CONTRIBUTING.md.
License
This server cannot be deployed
Maintenance
Related MCP Connectors
- docs2mcpOAuthcom.docs2mcp
Query your own PDFs and documents from any MCP client. Every answer cites the page it came from.
Personal knowledge base MCP server with semantic search, auto-categorization, metadata extraction
Academic literature search, retrieval, and private library management on top of OpenAlex.
Academic research MCP server for paper search, citation checks, graphs, and deep research.
Related MCP Servers
- FlicenseNot gradedqualityDmaintenanceEnables semantic search and conversational querying across a personal research library of PDFs, DOCX, and other documents using a vector database. It provides tools for document summarization, finding related papers, and high-accuracy retrieval for AI clients like Claude Desktop.-
- FlicenseNot gradedqualityDmaintenanceA local academic research assistant that indexes PDFs into a searchable vector library and exposes MCP tools for semantic search, claim extraction, contradiction detection, and multi-step research synthesis.-
- AlicenseAqualityFmaintenanceEnables searching arXiv and top AI conferences, finding related papers, generating research insights, and managing a personal library via MCP tools.541 npmMIT
- FlicenseAqualityBmaintenanceEnables semantic search across personal PDF paper collections with page-level citations, allowing users to query their library from any MCP-capable client.9-