Skip to main content
Glama
README.md
<div align="center">
  <img src="./logo/logo.svg" alt="ConTXT BOX" width="720">

  <p><strong>A local-first external context box for coding agents.</strong></p>

  <p>
    <a href="https://github.com/Oshadha345/contxt-box/actions"><img alt="CI" src="https://img.shields.io/github/actions/workflow/status/Oshadha345/contxt-box/ci.yml?style=flat-square"></a>
    <a href="https://pypi.org/project/contxt-box/"><img alt="Python" src="https://img.shields.io/badge/python-3.12%2B-3776AB?style=flat-square&logo=python&logoColor=white"></a>
    <a href="https://github.com/modelcontextprotocol/python-sdk"><img alt="MCP" src="https://img.shields.io/badge/MCP-ready-111827?style=flat-square"></a>
    <a href="https://github.com/microsoft/markitdown"><img alt="MarkItDown" src="https://img.shields.io/badge/MarkItDown-primary-2563EB?style=flat-square"></a>
    <a href="https://github.com/docling-project/docling"><img alt="Docling" src="https://img.shields.io/badge/Docling-primary-059669?style=flat-square"></a>
    <a href="./LICENSE"><img alt="License" src="https://img.shields.io/badge/license-MIT-black?style=flat-square"></a>
  </p>
</div>

---

## What Is It?

ConTXT BOX is a strict, local-first knowledge layer that sits beside any project or document folder. It gives coding agents such as Claude Code, Codex, Cursor, and other MCP clients a fast external memory: indexed filenames, folders, neighbors, summaries, cached document/image context, and durable chat preservation.

The design is intentionally narrow. Documents and images are the core path because they cover most real user context. Heavy extraction uses exactly one configured engine: [MarkItDown](https://github.com/microsoft/markitdown) or [Docling](https://github.com/docling-project/docling). No multi-tool fallback chain is used in core extraction.

## Features

- Lazy indexing with `rel_path`, filename, folder, mtime, size, type, neighbors, folder summaries, and cheap file summaries.
- On-demand extraction only through MarkItDown or Docling.
- Permanent Markdown sidecars under `.contextbox/history/media/`.
- MCP tools for coding agents.
- Watchdog-based `watch` command for continuous index updates.
- Preview-only smart reorganization.
- Auto preservation into `.contextbox/CONTEXT.md` plus JSONL history.

## Quick Start

```bash
uv sync
uv run contxtbox --help
uv run contxtbox init --root "S:\Papers"
uv run contxtbox config-show --root "S:\Papers"
uv run contxtbox index --root "S:\Papers"
uv run contxtbox health --root "S:\Papers"
uv run contxtbox search "computer vision" --root "S:\Papers"
```

When commands are run from inside the target workspace, `--root` can be omitted.

Install the document/image engines:

```bash
uv sync --extra media
```

Extract one file with the strict default engine:

```bash
uv run contxtbox extract-media "Computer Vision\paper.pdf" --root "S:\Papers"
```

Use Docling explicitly:

```bash
uv run contxtbox extract-media "Computer Vision\paper.pdf" --root "S:\Papers" --engine docling
```

Watch a folder:

```bash
uv run contxtbox watch --root "S:\Papers"
```

Run production readiness checks:

```bash
uv run contxtbox health --root "S:\Papers" --fail-on-error
```

Show the effective workspace config:

```bash
uv run contxtbox config-show --root "S:\Papers"
```

Production and MCP setup guides:

- [Installation](docs/INSTALL.md)
- [Production readiness](docs/PRODUCTION.md)
- [MCP client setup](docs/MCP_CLIENTS.md)
- [Client verification](docs/CLIENT_VERIFICATION.md)

## How It Works

```text
workspace/
`-- .contextbox/
    |-- index.json
    |-- config.toml
    |-- CONTEXT.md
    |-- preservation.jsonl
    `-- history/
        `-- media/
            `-- sanitized__file__path.context.md
```

### Indexing Rules

`index`, `update_index`, and `watch` always record:

- `rel_path`
- `filename`
- `folder_path`
- `mtime`
- `size`
- `file_type`
- `neighbors`
- `parent_folder_summary`
- `last_indexed`
- `context_summary`

The default summary is cheap and deterministic. It uses filename, folder name, and 5-7 nearby files. It does not open PDFs or images during indexing.

### Configuration

`init` creates `.contextbox/config.toml`:

```toml
extraction_engine = "markitdown"
max_inline_bytes = 512000
large_file_bytes = 50000000
max_neighbors = 10
debounce_seconds = 2.0
auto_watch = true

ignored_dirs = [
  ".git",
  ".venv",
  "node_modules",
]

priority_folders = [
  "codebases/",
  "research/",
  "specs/",
  "decisions/",
  "assets/images/",
]
```

Use `"docling"` when you want Docling as the strict extraction engine.

### Extraction Rules

Heavy extraction only happens when:

- `extract-media path` is called,
- or an MCP client calls `get_file(path, depth="full")`.

The result is cached as Markdown in `.contextbox/history/media/`, and `index.json` receives:

- `extracted_at`
- `context_ref`
- `extraction_method`
- `extraction_status`
- `extraction_warnings`
- `extraction_duration_seconds`

Sidecars include the same audit header before extracted content. Status values are conservative:
`success`, `partial`, `metadata-only`, or `cached`.

### MCP Tools

- `update_index()`
- `server_info()`
- `set_root(root, index=true)`
- `health()`
- `search(query, limit=10)`
- `get_file(path, depth="metadata" | "full")`
- `pull_context(task, limit=5)`
- `extract_media(path, force=false)`
- `reorganize(instruction)`
- `auto_preserve_context(summary, metadata=null)`

Start the MCP server:

```bash
uv run contxtbox mcp --root "S:\Papers"
```

## Attribution

- [Model Context Protocol Python SDK](https://github.com/modelcontextprotocol/python-sdk), MIT.
- [MarkItDown](https://github.com/microsoft/markitdown), MIT.
- [Docling](https://github.com/docling-project/docling), MIT.
- [watchdog](https://github.com/gorakhargosh/watchdog), Apache-2.0.
- [sentence-transformers](https://github.com/huggingface/sentence-transformers), Apache-2.0 library with model-specific licenses.
- [ChromaDB](https://github.com/chroma-core/chroma), Apache-2.0.
- [gstack](https://github.com/garrytan/gstack), MIT, as workflow inspiration.
- [Ponytail](https://github.com/DietrichGebert/ponytail), MIT, as minimal-agent behavior inspiration.

## Roadmap

- Stronger semantic search over sidecars.
- Reorganization scoring based on folder summaries and neighbor cues.
- MCP client recipes for Claude Code, Codex, Cursor, and others.
- Safe apply/undo flow for reorganization.
- Configurable ignore rules and extraction engine policy.

## Contributing

New ideas, bug fixes, documentation improvements, integration recipes, and production hardening work are welcome. Open an issue for discussion, or submit a focused pull request with a clear description, tests where relevant, and the verification commands you ran.

Useful contribution areas:

- MCP client setup recipes for more coding tools.
- Better document/image extraction quality checks.
- Faster indexing and retrieval for large workspaces.
- Safer reorganization previews and apply/undo flows.
- Clearer docs, examples, and real-world testing notes.

See [CONTRIBUTING.md](CONTRIBUTING.md) for the development checks.

## Connect

- Email: [samarakoonf@gmail.com](mailto:samarakoonf@gmail.com)
- LinkedIn: [Oshadha Samarakoon](https://www.linkedin.com/in/oshadha-samarakoon-488638341/)

## License

MIT. See [LICENSE](./LICENSE).

## Release

PyPI publishing is configured for Trusted Publishing through GitHub Actions. See
[Production readiness](docs/PRODUCTION.md#public-release).

TDQS

B3/5.0

Scored across 10 tools

Disambiguation4/5

Most tools have distinct purposes (e.g., auto_preserve_context vs. extract_media, pull_context), but some overlap exists between context preservation tools, potentially causing confusion.

Naming Consistency3/5

Naming is mixed: some tools follow verb_noun (get_file, set_root, update_index), while others are just verbs (health, reorganize) or use different styles (server_info). This inconsistency reduces predictability.

Tool Count5/5

With 10 tools, the count is well within the ideal 3-15 range and appropriate for a context management server, covering essential operations without bloat.

Completeness4/5

The tool set covers core context operations like preservation, extraction, search, and reorganization. Missing explicit update/delete for context items, but overall it provides a coherent workflow for the domain.

Maintenance

ActivityStale
ResponsivenessNo issues