reposniffer-mcp
RepoSniffer
There's a repo for that. Let RepoSniffer find it.

AI-first GitHub repo discovery. Describe a feature in plain language — "markdown editor with live preview" — and RepoSniffer returns a ranked, adoption-grade list of open-source projects that implement it, with evidence and quality signals (stars, activity, license, archived status).
Built to be the grounding layer for coding agents: stop hallucinating repos, get a verifiable "best of kind" answer with reasons.
Quickstart
# CLI
uvx reposniffer "markdown editor live preview" --language python --top-k 5
# MCP server (stdio) — wire into opencode, Claude Code, Codex, Cursor, ...
uvx --from reposniffer reposniffer-mcpSet GITHUB_TOKEN to raise search rate limits (authenticated = 30 req/min vs ~10).
Wire into your agent
opencode — add to opencode.json:
{
"mcp": {
"reposniffer": {
"type": "local",
"command": ["uvx", "--from", "reposniffer", "reposniffer-mcp"],
"environment": { "GITHUB_TOKEN": "ghp_..." }
}
}
}Claude Code: claude mcp add reposniffer -- uvx --from reposniffer reposniffer-mcp
Codex/Cursor: add an MCP server pointing at uvx --from reposniffer reposniffer-mcp (stdio).
MCP tools
Tool | Purpose |
| Feature query → ranked candidates with score breakdown, snippet evidence, license verdict, recommendation |
| Verify an existing repo (alive? licensed? best-of-kind?) + 2 alternatives |
| Embedding backend, model, auth status |
Every result carries as_of (a freshness timestamp agents can cite), a flags list
(archived, no-license, strong-copyleft, stale, ...), a license_category
(permissive / weak-copyleft / strong-copyleft / unknown), and a targeted
snippet showing why the repo matched.
Architecture
Coarse candidate fetch — GitHub Search API (
in:readme, language/license/stars filters).Hybrid rerank — embed each candidate's description + README front matter (not the whole README, to avoid dilution), cosine vs embedded query, plus a lexical-overlap boost for literal matches.
Quality scoring — popularity (log stars), activity (pushed_at half-life), license category, archived penalty; weights differ by
intent(adoptvsstudy).Adoption safety — permissive/weak/strong-copyleft classification flags GPL/AGPL repos before you depend on them.
Local SQLite cache — repos, READMEs, embeddings, query results → fast repeat queries, index grows over time.
Embeddings are pluggable: default is a zero-config local fastembed ONNX model
(no torch, no API key); set REPOSNIFFER_EMBED_BACKEND=api plus an OpenAI-compatible
endpoint for stronger quality.
Eval
Ground-truth queries live in eval/queries.py (feature → known-good repos). Run with a
token (each query fetches ~50 READMEs):
GITHUB_TOKEN=ghp_... uv run python -m eval.runReports hit@1 / hit@3 / hit@5. Current live result: hit@1 0.50, hit@3 0.83, hit@5 0.83.
Known limitation: the candidate stage depends on GitHub Search API relevance, which can
fail to recall canonical repos with weak descriptions/READMEs (e.g. Kozea/WeasyPrint
— description is just "The awesome document factory"). Semantic rerank can only rank
what the candidate fetch surfaces.
Project layout
src/reposniffer/
config.py # env-driven settings
cache.py # sqlite store (repos, readmes, embeddings, query cache)
engine/
github.py # GitHub REST client + search query builder + text/snippet utils
embed.py # Embedder protocol: local fastembed + OpenAI-compatible API
score.py # quality + license-category + lexical scoring
search.py # orchestration (Engine)
mcp/server.py # MCPServer (mcp 2.x)
cli.py # Typer CLI
eval/ # golden query → repo eval harness
tests/ # offline (fake transport + fake embedder)Development
uv sync
uv run ruff check .
uv run ruff format --check .
uv run pyright
uv run pytestCI (lint/format/type/tests) runs on every push and PR. Publishing to PyPI happens on v*
tags via trusted publishing — enable it once on the PyPI project settings, then:
git tag v0.1.0 && git push --tags.
Note: mcp 2.x is used — MCPServer (FastMCP was renamed in mcp 2.0). Pin mcp<2 if you need the v1 API.
Contributing
We welcome contributions — bugs, features, docs, and eval cases. Please read CONTRIBUTING.md first; it covers the dev setup, coding standards, tests, the eval harness, and the PR workflow.
Found a bug? Open an issue with the exact query and output.
Have a feature idea? Discuss it in an issue before writing code.
Reporting a vulnerability? See SECURITY.md — don't post it publicly.
This project follows a Code of Conduct.
License
MIT
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/nathan-hoche/RepoSniffer'
If you have feedback or need assistance with the MCP directory API, please join our Discord server