Skip to main content
Glama
lex-tools

@lex-tools/codebase-context-dumper

Official
by lex-tools

codebase-context-dumper MCP Server

npm version License: Apache-2.0

A Model Context Protocol (MCP) server designed to easily dump your codebase context into Large Language Models (LLMs).

Why Use This?

Large context windows in LLMs are powerful, but manually selecting and formatting files from a large codebase is tedious. This tool automates the process by:

  • Recursively scanning your project directory.

  • Including text files from the specified directory tree that are not excluded by .gitignore rules.

  • Automatically skipping binary files.

  • Concatenating the content with clear file path markers.

  • Supporting chunking to handle codebases larger than the LLM's context window.

  • Integrating seamlessly with MCP-compatible clients.

Related MCP server: code-index-mcp

The easiest way to use this tool is via npx, which runs the latest version without needing a local installation.

Configure your MCP client (e.g., Claude Desktop, VS Code extensions) to use the following command:

{
  "mcpServers": {
    "codebase-context-dumper": {
      "command": "npx",
      "args": [
        "-y",
        "@lex-tools/codebase-context-dumper"
      ]
    }
  }
}

The MCP client will then be able to invoke the dump_codebase_context tool provided by this server.

Features & Tool Details

Tool: dump_codebase_context

Recursively reads text files from a specified directory, respecting .gitignore rules and skipping binary files. Concatenates content with file path headers/footers. Supports chunking the output for large codebases.

Functionality:

  • Scans the directory provided in base_path.

  • Respects .gitignore files at all levels (including nested ones and .git by default).

  • Detects and skips binary files.

  • Reads the content of each valid text file.

  • Prepends a header (--- START: relative/path/to/file ---) and appends a footer (--- END: relative/path/to/file ---) to each file's content.

  • Concatenates all processed file contents into a single string.

Input Parameters:

  • base_path (string, required): The absolute path to the project directory to scan.

  • num_chunks (integer, optional, default: 1): The total number of chunks to divide the output into. Must be >= 1.

  • chunk_index (integer, optional, default: 1): The 1-based index of the chunk to return. Requires num_chunks > 1 and chunk_index <= num_chunks.

Output: Returns the concatenated (and potentially chunked) text content.

Local Installation & Usage (Advanced)

If you prefer to run a local version (e.g., for development):

  1. Clone the repository:

    git clone git@github.com:lex-tools/codebase-context-dumper.git
    cd codebase-context-dumper
  2. Install dependencies:

    npm install
  3. Build the server:

    npm run build
  4. Configure your MCP client to point to the local build output:

    {
      "mcpServers": {
        "codebase-context-dumper": {
          "command": "/path/to/your/local/codebase-context-dumper/build/index.js" // Adjust path
        }
      }
    }

Contributing

Contributions are welcome! Please see CONTRIBUTING.md for details on development, debugging, and releasing new versions.

License

This project is licensed under the Apache License 2.0. See the LICENSE file for details.

Available Tools

1 tool
dump_codebase_contextA

Recursively reads text files from a specified directory, respecting .gitignore rules and skipping binary files. Concatenates content with file path headers/footers. Supports chunking the output for large codebases.

ParametersJSON Schema
NameRequiredDescriptionDefault
base_pathYesThe absolute path to the project directory to scan.
num_chunksNoOptional total number of chunks to divide the output into (default: 1).
chunk_indexNoOptional 1-based index of the chunk to return (default: 1). Requires num_chunks > 1.

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It effectively describes key operational traits: recursive file reading, .gitignore respect, binary file skipping, output formatting with headers/footers, and chunking for large outputs. However, it doesn't mention potential limitations like file size constraints, permission requirements, or error handling, leaving some gaps.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is highly concise and well-structured in two sentences: the first covers core functionality and constraints, the second addresses scalability. Every phrase adds value (e.g., 'respecting .gitignore rules', 'skipping binary files', 'chunking the output'), with no wasted words or redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's moderate complexity (recursive file operations, chunking) and lack of annotations/output schema, the description does a good job covering core behavior and constraints. It explains what the tool does, key features, and output handling, but omits details like return format, error scenarios, or performance implications, which would enhance completeness for an agent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the input schema already fully documents all three parameters (base_path, num_chunks, chunk_index). The description adds no additional parameter-specific information beyond what's in the schema, such as examples or edge cases. The baseline score of 3 reflects adequate but minimal value addition from the description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the specific action ('recursively reads text files', 'concatenates content with file path headers/footers') and resource ('from a specified directory'), including key behavioral details like respecting .gitignore rules and skipping binary files. With no sibling tools, it fully defines the tool's unique purpose without redundancy.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage for scanning codebases ('large codebases') and mentions chunking for scalability, but provides no explicit guidance on when to use this tool versus alternatives or any prerequisites. Since there are no sibling tools, the lack of comparative guidance is less critical, but still leaves usage context somewhat open-ended.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

TDQS

A3.9/5.0
Disambiguation5/5

With only one tool, there is no possibility of ambiguity or overlap between tools. The tool has a single, clearly defined purpose.

Naming Consistency5/5

A single tool inherently has perfect naming consistency, as there are no other tools to compare it against for patterns or conventions.

Tool Count2/5

One tool is too few for most practical server purposes, as it severely limits functionality and interaction. While the tool is well-described, a single tool feels thin and incomplete for a codebase context server.

Completeness2/5

The server's purpose appears to be codebase context management, but with only a dump tool, there are significant gaps. Missing operations like search, filter, update, or delete context make the surface incomplete and likely insufficient for agent workflows.

Maintenance

ActivityInactive
ResponsivenessSyncing

Resources

Unclaimed servers have limited discoverability.

Looking for Admin?

If you are the server author, to access and configure the admin panel.

Related MCP Connectors

Related MCP Servers

Latest Blog Posts

MCP directory API

We provide all the information about MCP servers via our MCP API.

curl -X GET 'https://glama.ai/api/mcp/v1/servers/lex-tools/codebase-context-dumper'

If you have feedback or need assistance with the MCP directory API, please join our Discord server