Skip to main content
Glama
theunreal

codebase-mcp

by theunreal

codebase-mcp

Remote MCP server that indexes your repos and lets AI agents search them semantically.

Developers connect by adding a URL — no local install needed.

Quick Start

1. Configure repos

Edit config.yaml:

repos:
  - name: my-backend
    path: /path/to/local/repo
    extensions: [".py", ".ts", ".java"]

  - name: my-frontend
    url: https://github.com/your-org/frontend.git
    branch: main
    extensions: [".ts", ".tsx"]

2. Run

With Docker:

docker compose up --build

Without Docker:

python -m venv .venv && source .venv/bin/activate
pip install mcp chromadb sentence-transformers tree-sitter tree-sitter-python tree-sitter-javascript tree-sitter-typescript tree-sitter-java pyyaml gitpython pydantic pydantic-settings structlog uvicorn
python -m src.server

First run downloads the embedding model (~80MB) and indexes all repos. Subsequent runs are incremental (only changed files).

3. Connect your IDE

Claude Desktop — edit ~/Library/Application Support/Claude/claude_desktop_config.json:

{
  "mcpServers": {
    "codebase": {
      "url": "http://localhost:8080/mcp"
    }
  }
}

Cursor — edit .cursor/mcp.json:

{
  "mcpServers": {
    "codebase": {
      "url": "http://localhost:8080/mcp"
    }
  }
}

Claude Code — edit .mcp.json:

{
  "mcpServers": {
    "codebase": {
      "url": "http://localhost:8080/mcp"
    }
  }
}

Replace localhost:8080 with your server's actual address when hosting remotely.

Related MCP server: PAMPA

Tools

Tool

What it does

semantic_search(query)

Natural language search: "how does auth work"

get_entity(name)

Exact lookup: "show me UserService"

get_file_skeleton(file_path)

File structure without bodies (saves tokens)

list_repos()

What's indexed

reindex()

Trigger re-indexing

All tools accept an optional repo_name parameter to filter by repo.

Architecture

IDE (Claude/Cursor) → Streamable HTTP → MCP Server → ChromaDB
                                            ↑
                                    tree-sitter AST chunking
                                    + sentence-transformers embeddings
                                    + git clone/pull on schedule
  • Chunking: tree-sitter parses code into functions, classes, methods (not character splits)

  • Embeddings: sentence-transformers (default all-MiniLM-L6-v2, or jina-embeddings-v2-base-code for better code search)

  • Storage: ChromaDB with cosine similarity

  • Incremental: file MD5 hashes tracked, only changed files re-embedded

  • Scheduled: auto git pull + re-index every N minutes (configurable)

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    D
    maintenance
    Enables semantic code search across multiple repositories using natural language queries. Provides intelligent code discovery, symbol lookups, and cross-repo dependency analysis for AI coding agents.
    MIT
  • A
    license
    Not graded
    quality
    C
    maintenance
    Provides semantic code search and retrieval capabilities for AI agents, enabling them to query codebases using natural language with automatic learning, hybrid search, and intelligent chunking of functions and classes.
    13 npm
    30
    ISC
  • A
    license
    Not graded
    quality
    D
    maintenance
    Enables AI agents to intelligently navigate and understand codebases by providing instant file descriptions, semantic search, and context-aware recommendations, eliminating the need to repeatedly scan files.
    20
    MIT
  • A
    license
    Not graded
    quality
    B
    maintenance
    Adds semantic code search to AI coding agents, enabling natural language queries across entire codebases to retrieve relevant code chunks, saving tokens and providing deep context.
    16 npm
    1
    MIT