SourceTap
Allows AI agents to learn and search any library directly from its GitHub repository by downloading and indexing documentation files (Markdown/MDX) and performing keyword search.
Click on "Install Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@SourceTapsearch the docs of https://github.com/tiangolo/fastapi for 'dependency injection'"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
SourceTap
SourceTap is an MVP of a Model Context Protocol (MCP server that lets your AI assistant learn and search any library directly from its GitHub repository or documentation URL.
Features
This project provides two tools:
query_docs(url, query): A RAG (Retrieval-Augmented Generation) tool.Input: Takes a URL to a ZIP archive (e.g., a GitHub repo archive) and a search query.
Process:
Downloads the ZIP file (cached via SQLite to prevent redundant downloads).
Extracts
.mdand.mdxcontent.Indexes the content in-memory using
minsearch(TF-IDF/Keyword search).
Output: Returns the full content of the top 5 most relevant documentation files.
Use Case: Helps AI agents understand libraries that are too new, private, or obscure for their base references.
fetch_web_content(url): A reader tool.Input: Any webpage URL.
Process: Proxies the request through
r.jina.aito convert HTML to clean, LLM-friendly Markdown.Output: The text content of the page.
Use Case: Inspecting specific documentation pages, blog posts, or issue threads.
Related MCP server: Search Docs MCP
Installation
To use this tool with your AI assistant (e.g., Claude Desktop, Cline), add the following configuration to your MCP Settings file:
{
"mcpServers": {
"sourcetap": {
"command": "uv",
"args": [
"--directory",
"/absolute/path/to/sourcetap",
"run",
"python",
"main.py"
]
}
}
}Note: Replace
/absolute/path/to/sourcetapwith the actual path to this directory on your machine. Theuvcommand will automatically handle dependency installation and environment setup when the server starts.
Project Architecture
The Tech Stack
MCP Framework: FastMCP (Python)
Web Scraping: Jina Reader API (via httpx)
Search Engine: minsearch (TF-IDF/Keyword search)
Caching: SQLite with WAL mode
Used MCP: Context7
AI Assistant: Google Gemini 3 Flash + Antigravity IDE
Caching Strategy
The project uses SQLite for persistent caching of downloaded ZIP files.
WAL Mode: Write-Ahead Logging enabled for better concurrent read/write performance.
Search Implementation
Uses minsearch for in-memory document search.
Text Fields: Indexes both
contentandfilenamefor comprehensive search.TF-IDF Scoring: Ranks documents by term frequency-inverse document frequency.
Top-K Retrieval: Returns the 5 most relevant documents per query.
Memory Efficient: Index is rebuilt per query (no persistent index storage).
Limitations & Possible Improvements
Keyword-Only Search: Currently uses TF-IDF. Semantic search with embeddings (e.g.,
all-MiniLM-L6-v2) would enable conceptual matching.Full-File Retrieval: Returns entire files. Smart chunking by headers would improve precision.
Markdown-Only: Only indexes
.mdand.mdxfiles. Code parsing (.py,.ts) would enable technical implementation queries.ZIP Archives: Downloads full repositories. GitHub Tree API would enable sparse downloading of only needed files.
No Persistent Index: Index is rebuilt per query. Persistent indexing would improve performance for repeated queries.
Single-Threaded Cache: SQLite cache is synchronous. Async cache operations would improve throughput.
This server cannot be installed
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Servers
- FlicenseNot gradedqualityDmaintenanceScrapes and indexes documentation websites to provide AI assistants with searchable access to documentation content, API references, and code examples through configurable URL crawling.
- FlicenseAqualityDmaintenanceEnables semantic search across multiple AI library documentations to keep coding assistants up-to-date.11
- FlicenseAqualityBmaintenanceEnables LLMs to dynamically search, scrape, and query official documentation of libraries like uv, OpenAI, LangChain, and LlamaIndex via Google Serper and Groq.1
- FlicenseNot gradedqualityCmaintenanceEnables retrieval and cleaning of official documentation for AI/Python libraries using search and LLM-based HTML cleaning.
Related MCP Connectors
@latest documentation and code examples to 9000+ libraries for LLMs and AI code editors in a singl…
Search your knowledge bases from any AI assistant using hybrid RAG.
Provide your AI coding tools with token-efficient access to up-to-date technical documentation for…
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/maxvoltage/sourcetap'
If you have feedback or need assistance with the MCP directory API, please join our Discord server