docs_server
Enables ingesting and reviewing GitHub repositories, including indexing source code and identifying bugs, security vulnerabilities, and code-quality issues in the repository.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@docs_serverReview the repository for security vulnerabilities in the login flow."
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Bug Buster AI
Bug Buster AI is an MCP-powered retrieval-augmented generation assistant with two capabilities: it answers questions over ingested documents and reviews GitHub repositories for bugs, security issues, and code quality problems. It uses Hugging Face embeddings, ChromaDB, LangChain and LangGraph orchestration, and a Groq-hosted LLM to retrieve relevant evidence before answering.
Architecture
Hugging Face embeddings: Sentence Transformers embed documents and source-code chunks for semantic retrieval.
ChromaDB vector store: Local persistent storage keeps document chunks in
docmind_documentsand repository code incode_chunks.MCP server:
mcp_server/docs_server.pyexposessearch_docsandsearch_codeover stdio.LangGraph ReAct agent: Chooses the appropriate MCP search tool and grounds responses in retrieved content.
Groq LLM: Provides the default tool-calling model, configured through
GROQ_API_KEYandGROQ_MODEL.Streamlit UI: Provides chat, provider selection, repository ingestion, and tool-use visibility.
Related MCP server: Semantic Code Search MCP Server
Setup
Python 3.11 or newer is required.
python -m venv .venv
.\.venv\Scripts\Activate.ps1
pip install -r requirements.txt
Copy-Item .env.example .envOpen .env and fill in your Groq credentials:
DOCMIND_LLM_PROVIDER=groq
GROQ_API_KEY=your-groq-api-key
GROQ_MODEL=openai/gpt-oss-120bThe default embedding model is BAAI/bge-small-en-v1.5. Set DOCMIND_EMBEDDING_MODEL=BAAI/bge-large-en-v1.5 when higher retrieval quality is worth the additional memory use.
Usage
Ingest documents
Place PDF, Markdown, or text files in data/, then run:
python -m ingestion.embed_and_storeThe document chunks are stored in the document collection used by search_docs.
Ingest a GitHub repository
Clone and index a repository into the separate code_chunks collection:
python -m ingestion.embed_and_store_repo --repo-url <url> --resetThe Streamlit sidebar also provides a GitHub repo URL field and Ingest Repository button that calls the same ingestion function directly. The ingestion process skips generated/vendor directories, unsupported extensions, and source files larger than 500 KB.
Launch the app
streamlit run ui/app.pyAsk questions about indexed documents or request a code review, for example:
What does the documentation say about authentication?Review the repository for security vulnerabilities in the login flow.Find error-handling gaps in the API client.
Example
Question:
Review the repository for bugs in the authentication code.A grounded response may look like:
Issue: User input is interpolated directly into the SQL query, allowing SQL injection.
File: src/auth.py
Chunk: 2
Recommendation: Use a parameterized query and validate the input before execution.Bug Buster AI reports only issues visible in retrieved code and cites the relevant file_path and chunk_index.
Tech Stack
Hugging Face Transformers and Sentence Transformers
LangChain
LangGraph
Model Context Protocol (MCP)
ChromaDB
Groq
Streamlit
This server cannot be deployed
Maintenance
Related MCP Connectors
Search GitHub, npm, PyPI, StackOverflow, ArXiv from one MCP — built for coding agents.
Token-efficient search for coding agents over public and private documentation.
Securely search and manage workspace context files for AI agents and teams.
Code intelligence for coding agents: semantic, AST, graph, and full-text search. 279+ languages.
Related MCP Servers
- AlicenseNot gradedqualityAmaintenanceEnables AI agents to search and retrieve context from GitHub issues, pull requests, releases, and documentation using hybrid semantic search and time-ordered activity scans.24 npmApache 2.0

Semantic Code Search MCPofficial
FlicenseNot gradedqualityCmaintenanceProvides AI coding agents with structured access to indexed codebases via semantic search, symbol analysis, and file reading tools.12-- AlicenseNot gradedqualityDmaintenanceEnables AI agents to index and search local files, websites, GitHub repos, and packages using hybrid retrieval with reranking, all through IDE chat.4Apache 2.0
- AlicenseAqualityAmaintenanceEnables coding agents to query local notes, decisions, docs, and code with hybrid retrieval (BM25 + embeddings + reranking) and get path:line citations. It provides tools like rag_query for full-corpus search and search_knowledge for project-scoped knowledge recall.21MIT