Personal Knowledge-Base MCP Server
Provides semantic search over a personal knowledge base by generating embeddings for documents and queries using Google Gemini.
Click on "Install Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Personal Knowledge-Base MCP Serversearch my notes for Cauchy-Riemann equations"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Personal Knowledge-Base MCP Server
A Personal Knowledge-Base MCP Server that provides semantic search over a student-owned document collection using the Model Context Protocol (MCP), Gemini embeddings, and Qdrant.
Project Overview
This project exposes a personal knowledge base as callable MCP tools.
Instead of relying on keyword matching, the system converts user queries into vector embeddings and retrieves semantically relevant document chunks from Qdrant.
Related MCP server: Solarium
Architecture
User / MCP Client
|
v
MCP Server (FastMCP)
|
+----------------------+
| |
v v
search_notes() get_document()
|
v
Gemini Embedding API
|
v
Qdrant Vector Database
|
v
Ranked Chunks
|
v
Source + Page + Score + Text
## Features
* PDF document ingestion
* Page-by-page text extraction
* Recursive text chunking
* Gemini `gemini-embedding-001` embeddings
* Qdrant vector storage
* Semantic similarity search
* Source and page citations
* Confidence threshold for low-relevance queries
* Full-document retrieval
* Indexed-source listing
* MCP Inspector support
## MCP Tools
### `search_notes`
Searches the knowledge base using semantic similarity.
Arguments:
* `query`: search question or topic
* `top_k`: maximum number of results
Returns:
* similarity score
* source filename
* page number
* relevant text chunk
### `get_document`
Returns the complete text of an indexed PDF document.
Argument:
* `doc_id`: document filename
Example:
```text
Complex_Variables_Project_Report.pdflist_source_documents
Lists all indexed source documents.
Example output:
1. Complex_Variables_Project_Report.pdfProject Structure
Personal-Knowledge-MCP/
├── documents/
│ └── Complex_Variables_Project_Report.pdf
├── services/
│ ├── chunking.py
│ ├── embedding.py
│ ├── pdf_reader.py
│ └── qdrant_service.py
├── .env
├── .gitignore
├── evaluation.py
├── ingest.py
├── requirements.txt
└── server.pySetup
1. Create and activate virtual environment
python -m venv .venv
.venv\Scripts\Activate.ps12. Install dependencies
pip install -r requirements.txt3. Configure Gemini API key
Create a .env file in the project root:
GEMINI_API_KEY=your_api_key_hereNever commit .env to Git.
4. Start Qdrant
The project uses local Qdrant at:
http://localhost:6333Example Docker command:
docker run -d --name qdrant -p 6333:6333 -p 6334:6334 qdrant/qdrantDocument Ingestion
Place the PDF inside:
documents/Run:
python ingest.pyThe ingestion pipeline performs:
PDF
↓
Page extraction
↓
Chunking
↓
Gemini embeddings
↓
Qdrant storageEach stored chunk contains:
text
page
sourceRunning the MCP Server
Start the MCP Inspector:
mcp dev server.pyThe MCP server uses STDIO transport.
Available tools:
search_notes
get_document
list_source_documentsRetrieval Evaluation
A small evaluation set of five queries was used to check whether at least one expected relevant page appeared within the top three retrieved results.
Evaluation result:
Tests: 5
Successful hits: 5
Hit@3: 100%Example evaluation queries included:
What is a complex variable?
What are the Cauchy-Riemann equations?
How does the Laplace transform help engineering systems?
What is the difference between Laplace and Fourier transforms?
How is FFT used for audio noise reduction?
Confidence Filtering
The search tool uses an initial similarity threshold of:
0.60For example, relevant queries produced scores around:
0.79
0.76
0.75while an unrelated query produced scores around:
0.52Therefore low-scoring results are filtered and the tool returns:
No confident match found.Technologies
Python
FastMCP
Model Context Protocol (MCP)
Google Gemini Embeddings
Qdrant
PyMuPDF
LangChain Text Splitters
Docker
MCP Inspector
Current Knowledge Source
The current demonstration corpus is:
Complex_Variables_Project_Report.pdfThe document contains 7 pages and was split into 30 chunks for the indexed personal_knowledge collection.
Security
API keys are stored in
.env.envis excluded through.gitignoreSecrets should never be committed to source control
Future Improvements
Support Markdown and TXT documents
Add document-level persistent IDs
Improve duplicate-chunk handling
Expand the evaluation dataset
Add more retrieval metrics
Support multiple document collections
Add optional Qdrant Cloud deployment
This server cannot be installed
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Servers
- Flicense-qualityCmaintenanceMCP server that indexes a knowledge base into Chroma and provides search tools for retrieving document fragments via vector embeddings.
- Alicense-qualityDmaintenanceA knowledge base MCP server backed by Qdrant vector database with local embeddings for semantic search and document management.51ISC
- Flicense-qualityCmaintenanceMCP server providing RAG tools (search_notes, answer_from_notes) and resources for grounded answers over a local knowledge base.
- Flicense-qualityAmaintenanceA local knowledge base MCP server that enables retrieval and evidence-based Q&A over Obsidian Markdown notes, with high-recall embedding search, chunked indexing, hybrid retrieval, and three STDIO MCP tools for agent-driven recollection and quality-gated recall.
Related MCP Connectors
Agent-native MCP server over the public saagarpatel.dev corpus. Read-only, stateless.
Agentic search over your Dewey document collections from any MCP-compatible client.
Serve a folder of Markdown notes as an MCP server: hybrid search, reading, and sourced answers.
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/AmnaSarwar522/Personal-Knowledge-MCP'
If you have feedback or need assistance with the MCP directory API, please join our Discord server