GitHub Code MCP
This server gives an AI agent read-only access to GitHub source code, letting it search, fetch, browse, and analyze Java code fragments without writing to GitHub.
Search code:
search_codefinds matching files/fragments with highlighted match context (requiresGITHUB_TOKEN).Read files:
get_file_contentsfetches a full file or a specific 1-indexed line range.Browse repository structure:
list_directorywalks the repo tree, andget_repository_info/get_readmeprovide project metadata and overview.List branches and commits:
list_branchesshows branch names with latest SHAs;list_commitsshows recent history, optionally scoped to one file.Analyze Java symbols:
list_symbolsenumerates classes/methods/constructors with line ranges.Extract Java fragments:
get_code_fragmentextracts a named class/method/constructor body, with AST-accurate parsing via tree-sitter or a regex fallback.Check pull request status:
get_pr_statusreturns whether a PR is open/merged/closed, including merge time and URL.Semantic code retrieval:
find_relevant_codefinds the most relevant code chunks for descriptive criteria using in-memory RAG (embeddings + BM25 fusion).
Provides read access to GitHub repositories, enabling code search, file and line-range retrieval, Java symbol and fragment extraction, directory browsing, and repository metadata, readme, branch, and commit listing.
Click on "Install Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@GitHub Code MCPshow me the getReport method in ReportService.java"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
GitHub Code MCP
An MCP server that gives the Rapid7 SI Triage agent read access to source code
on GitHub — by default, anands-bounteous/nexpose.
It fills the gap the other two servers don't cover: jira-confluence-mcp owns
tickets/KB, log-intelligence-mcp owns log retrieval, but a Phase 2
investigation also needs to pull the actual Java source that produced a stack
trace or defect. This server hits the real GitHub REST + Search APIs directly
(no mocking, no local cloning) and is scoped specifically to reading code
and returning fragments of it, plus one narrow read-only exception —
get_pr_status — so the orchestrator can tell a support engineer whether a
past fix's pull request has actually merged. Broader issue/PR management
(creating/commenting/merging) stays out of scope.
owner/repo are optional per-call overrides on every tool, on top of the
GITHUB_OWNER/GITHUB_REPO defaults, so the server can be pointed at another
repo without a restart.
Tools
Tool | Purpose |
| The core "fetch fragments for input text" tool — GitHub code search with highlighted match fragments showing exactly where the query hit. Requires |
| Fetch a file, or just a 1-indexed inclusive line range of it. |
| Enumerate the classes/interfaces/enums/records/methods/constructors declared in a Java file, with line ranges — use this to find a symbol name for |
| Java-aware extraction of one method/class/constructor body by name. Falls back to a plain text context window if the symbol isn't a recognisable declaration. |
| Browse the repo tree. |
| Description, default branch, language, topics, stars. |
| Project overview as plain text. |
| Branch names + latest commit sha. |
| Recent history, optionally scoped to one file. |
| Live merge status of a pull request — by full PR URL (e.g. a historical-kb-mcp |
| Semantic ("RAG") retrieval: given descriptive criteria, return the specific chunk(s) of code most relevant to it — see "Semantic code retrieval (RAG)" below. |
Related MCP server: mcp-server-github
Java-aware fragment extraction
get_code_fragment/list_symbols are backed by a pluggable FRAGMENT_BACKEND:
tree-sitter(AST-accurate) — parses withtree-sitter+tree-sitter-java. Correctly handles generics (Map<String, List<Foo>>), annotations, records, nested/anonymous classes, and text blocks — anything a hand-rolled brace-counter gets wrong on real Java. Optional dependency:pip install -e ".[java]".regex(dependency-free fallback) — matches Java declaration syntax to find a symbol's header line, then a string/char/comment-aware balanced-brace scanner (aware of",',//,/* */, and Java 15+"""text blocks, so a{/}inside a literal or comment can't throw off the count) finds the matching close.auto(default) — triestree-sitterfirst; if the optional dependency isn't installed, logs a warning and degrades toregex. Explicit choices (tree-sitter/regex) raise instead of silently degrading.
If a requested symbol isn't a recognisable Java declaration (e.g. it's a field
name, or doesn't exist), both backends fall back to a plain
±FRAGMENT_CONTEXT_LINES text window around its first literal occurrence, with
match_type="context_window" in the response so the caller can tell it's a
lower-confidence result rather than an exact definition.
Semantic code retrieval (RAG)
find_relevant_code answers a question neither search_code (literal/keyword
match) nor get_code_fragment (exact lookup by symbol name, which you must
already know) can: "given descriptive criteria, show me the specific code
that's relevant." It runs the same chunk → embed → hybrid-retrieve pattern
log-intelligence-mcp uses for logs, adapted to source code, entirely
in-memory per call — there's no local vector store or SI_DATA_DIR concept
here, since this is an ad-hoc per-query tool, not a persistent index.
Pipeline:
Candidate files — recursively walks the repo tree (
list_directory) and ranks files by path-token overlap withcriteria, refined using Java symbol names (list_symbols) for the top slice — deliberately not GitHub's code-search index, which can lag indefinitely on new/small/low-star repos.Chunking (
code_chunking.py) — each candidate file's full text is fetched and chunked. Wherelist_symbolsrecognises Java structure, leaf symbols (methods/constructors, or a class/interface/enum/record with no symbol nested inside its own range) become atomic chunking units, and the lines between them (imports, package decl, class signature) become "structural" gap units — so the file is tiled exactly once with nothing duplicated or dropped. Non-Java files (or files where symbol extraction fails) fall back to fixed-size line-window units. Units are greedily packed into token-budgeted chunks (CODE_CHUNK_TARGET_TOKENS/_MAX_TOKENS), with trailing-unit overlap between consecutive chunks (CODE_CHUNK_OVERLAP_TOKENS) and oversized single units emitted whole and flagged rather than split.Hybrid retrieval (
code_retrieval.py) — dense embedding similarity (EMBED_BACKEND:sentence-transformers, or the dependency-freehashingTF-IDF fallback;autoprefers the former, degrading with a logged warning if it isn't installed) fused with BM25 keyword matching via Reciprocal Rank Fusion (RRF_K/DENSE_WEIGHT/SPARSE_WEIGHT), returning the toptop_kchunks withmatch_type="hybrid_retrieval".
Install the optional sentence-transformers backend with
pip install -e ".[rag]"; without it (or with EMBED_BACKEND=hashing
explicit), retrieval still works fully offline via the hashing fallback.
Auth & configuration
Copy .env.example to .env:
GITHUB_TOKEN=<create at github.com/settings/tokens>
GITHUB_API_BASE_URL=https://api.github.com
GITHUB_OWNER=anands-bounteous
GITHUB_REPO=nexpose
GITHUB_DEFAULT_REF=
MAX_FILE_KB=500
FRAGMENT_CONTEXT_LINES=20
FRAGMENT_BACKEND=auto
HTTP_TIMEOUT=30
HTTP_MAX_RETRIES=4
MCP_HTTP_HOST=127.0.0.1
MCP_HTTP_PORT=8082
# find_relevant_code (RAG) — chunking
CODE_CHUNK_TARGET_TOKENS=400
CODE_CHUNK_MAX_TOKENS=800
CODE_CHUNK_OVERLAP_TOKENS=80
# find_relevant_code (RAG) — embeddings: auto | sentence-transformers | hashing
EMBED_BACKEND=auto
EMBED_MODEL=all-mpnet-base-v2
EMBED_DIM_FALLBACK=512
# find_relevant_code (RAG) — hybrid retrieval (RRF)
RRF_K=60
DENSE_WEIGHT=1.0
SPARSE_WEIGHT=1.0
CANDIDATE_POOL=30GITHUB_TOKEN is a GitHub personal access token
(github.com/settings/tokens). It's optional for reading public repos — but
required for search_code (GitHub's code search API rejects unauthenticated
requests outright) and strongly recommended for everything else (5,000
requests/hour authenticated vs. 60/hour anonymous). A fine-grained PAT with
read-only "Contents" access is enough; no repo write scope is needed since
this server never writes to GitHub.
GITHUB_API_BASE_URL is overridable for GitHub Enterprise Server. It's
normalised to just the scheme+host, same as any pasted API URL.
The HTTP client retries 429/5xx with exponential backoff (honouring
Retry-After), and additionally watches GitHub's primary rate-limit signal
(X-RateLimit-Remaining: 0 + X-RateLimit-Reset) to sleep until the limit
resets rather than blindly backing off — controlled by HTTP_MAX_RETRIES and
HTTP_TIMEOUT.
Install & run
cd github-mcp
python -m venv .venv && source .venv/bin/activate # .venv\Scripts\Activate.ps1 on Windows
pip install -e . # base install: mcp, httpx, uvicorn, numpy
pip install -e ".[java]" # + tree-sitter/tree-sitter-java for AST-accurate fragments
pip install -e ".[rag]" # + sentence-transformers for the find_relevant_code embedding backend
cp .env.example .env # fill in GITHUB_TOKEN
# stdio:
python -m github_mcp --transport stdio
# HTTP (streamable-http at http://127.0.0.1:8082/mcp):
python -m github_mcp --transport httpRegister with an MCP client (stdio example)
{
"mcpServers": {
"github": {
"command": "python",
"args": ["-m", "github_mcp", "--transport", "stdio"],
"env": {
"GITHUB_TOKEN": "…",
"GITHUB_OWNER": "anands-bounteous",
"GITHUB_REPO": "nexpose"
}
}
}
}Ports
This server's HTTP transport defaults to 8082 — jira-confluence-mcp uses
8080 and log-intelligence-mcp uses 8081, so all three can run simultaneously.
Tests
pytest # in an environment with pytest installed
python tests/_runner.py # offline harness when pytest isn't installedCovers base-URL normalisation, content decoding, line-slicing, directory/repo/
branch/commit normalisation, search-query building and result normalisation,
and — via a stub HTTP transport (FakeClient, no network needed) — file
fetch with the MAX_FILE_KB size guard, directory listing, repo info/readme/
branches/commits, the search_code no-token guard, and the fragment-backend
factory. The Java regex backend is exercised directly against a realistic
fixture source file (tests/fixtures/Sample.java) covering nested classes, an
interface, a generic method, an annotated method, and a string literal
containing {/} to prove brace-in-string masking works.
The find_relevant_code RAG pipeline has its own coverage: symbol-aware
chunking correctness (test_code_chunking.py — no duplication/gaps, oversized-
unit flagging, line-window fallback), the hashing embedding backend + BM25
tokenisation + RRF fusion against a hand-computed score
(test_embeddings_bm25.py), and end-to-end tool behaviour via the FakeClient
pattern (test_find_relevant_code.py — candidate-file discovery via the
list_directory/list_symbols tree walk, and hybrid-retrieval results). These
force EMBED_BACKEND=hashing explicitly, so they never need
sentence-transformers installed.
42 tests total, all offline — none need GITHUB_TOKEN, network access, or the
optional sentence-transformers/tree-sitter-java dependencies.
If
tree-sitter-javaisn't installed, tests targetRegexJavaBackendexplicitly rather than relying onFRAGMENT_BACKEND=autoresolution, so the suite stays runnable regardless of what's pip-installed.
Live GitHub API calls (a real
search_code/get_file_contentsagainstanands-bounteous/nexpose) need a realGITHUB_TOKENand network access, which the automated test suite doesn't exercise — see "Install & run" above to try them manually.
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Connectors
Access the GitHub API, enabling file operations, repository management, search functionality, and…
Code intelligence for LLMs. Analyze, search, and retrieve code from any public git repository.
Screens public GitHub repos and PRs to generate risk maps, findings, and merge-readiness signals.
Ask a codebase what calls what: search, blast radius, paths between symbols, and diffs.
Related MCP Servers
- FlicenseNot gradedqualityCmaintenanceEnables access to GitHub repositories and data through the GitHub API. Supports retrieving repositories, issues, pull requests, and searching code across GitHub with authentication via personal access tokens.
- FlicenseAqualityCmaintenanceEnables interaction with GitHub repositories, issues, pull requests, code search, branches, and GitHub Actions workflows.848
- FlicenseNot gradedqualityDmaintenanceProvides read-only access to company GitHub repositories, enabling code search, file retrieval, documentation search, and repo browsing via natural language.
- AlicenseNot gradedqualityBmaintenanceEnables searching GitHub code with filters and reading arbitrary file contents from any repository via MCP.GPL 3.0
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/hareeshbounteous/github-mcp'
If you have feedback or need assistance with the MCP directory API, please join our Discord server