context-usage-mcp
# context-usage-mcp
A tiny [Model Context Protocol](https://modelcontextprotocol.io) server (TypeScript/Node, stdio)
that exposes a single tool, **`get_context_usage`**, reporting the current session's **raw token
usage** — so an agent can check after each task whether context has grown large and branch its
behavior on an absolute threshold. Runs under both **Claude Code** and **OpenAI Codex CLI**, reading
each host's own session transcript/rollout behind an auto-detected host-adapter layer.
## What it does
Neither host hands an stdio MCP server the live token counts, so this server reads usage
**indirectly** from the host's own on-disk session file, behind an auto-detected host-adapter layer.
The output shape is identical across hosts.
**Claude Code** passes only `CLAUDE_PROJECT_DIR` (no session id, transcript path, or counts):
1. Derive the project folder from `CLAUDE_PROJECT_DIR` (non-alphanumerics → `-`).
2. Pick the **most-recently-modified `*.jsonl`** in `~/.claude/projects/<folder>/` as the current
session (its transcript is being written while the tool is called).
3. Scan **backward** for the last assistant message carrying a `usage` block.
4. Sum `input_tokens + cache_creation_input_tokens + cache_read_input_tokens` → `context_tokens`.
The `~/.claude` root can be relocated with `CLAUDE_CONFIG_DIR`.
**OpenAI Codex CLI** passes *no* project/session env var at all, but launches the server with its
working directory set to the session's project dir. So the server uses `process.cwd()`:
1. Resolve the Codex root (`CODEX_HOME`, else `~/.codex`); rollouts live under
`<root>/sessions/YYYY/MM/DD/rollout-*.jsonl`.
2. Scan recent day-folders (today + yesterday, widening to ~7 days), reading only each candidate's
first `session_meta` line to match its `cwd` against `process.cwd()`; pick the **freshest** match.
3. Scan **backward** for the last `token_count` event and read its per-turn `last_token_usage`.
4. Map to the shared, disjoint breakdown. Codex's `cached_input_tokens` is a *subset* of
`input_tokens` (verified in source), so `input` ← `input − cached`; then
`context_tokens = input + cache_read + cache_creation` — no double-count.
Output is `output_tokens + reasoning_output_tokens`. The Codex `model_context_window` is ignored
(raw counts only, strict parity with the Claude host).
## Install & run
Run the published package directly with `npx` — no clone or build required:
```bash
npx -y @beeltec/context-usage-mcp
```
This is the recommended way to wire it into a host (see the registration sections below). The `-y`
flag skips the install prompt on first run.
> **Run from source (local dev)** — for hacking on the server itself:
> ```bash
> npm install
> npm run build # tsc → dist/, sets +x on dist/index.js
> npm test # parser unit tests
> ```
> Then register the absolute path to `dist/index.js` (see the "from source" note in each
> registration section).
## Host detection
The server auto-detects its host: `CLAUDE_PROJECT_DIR` set → Claude Code; else `CODEX_HOME` set or
`~/.codex/sessions` present → Codex; else it falls back to Claude. Set the environment variable
`CONTEXT_USAGE_HOST=claude|codex` to force a host (the override wins over auto-detection).
## Register with Claude Code (user scope)
```bash
claude mcp add --scope user context-usage -- npx -y @beeltec/context-usage-mcp
```
Verify with `claude mcp list`. Because MCP servers are loaded at session start, start a **new**
Claude Code session before calling the tool.
> **From source:** register the absolute path to this repo's built binary instead:
> ```bash
> claude mcp add --scope user context-usage -- node /absolute/path/to/dist/index.js
> ```
## Register with Codex CLI
Add an entry to `~/.codex/config.toml` (root movable with `CODEX_HOME`):
```toml
[mcp_servers.context-usage]
command = "npx"
args = ["-y", "@beeltec/context-usage-mcp"]
```
Or use the CLI equivalent:
```bash
codex mcp add context-usage -- npx -y @beeltec/context-usage-mcp
```
**Do not set a custom `cwd` for the server**: Codex passes no project/session identifier, so
discovery relies on the server inheriting the session's working directory via `process.cwd()`.
Setting `cwd` breaks the cwd-match. MCP servers load at session start, so start a **new** Codex
session before calling the tool.
> **From source:** use `command = "node"` with `args = ["/absolute/path/to/dist/index.js"]` (or
> `codex mcp add context-usage -- node /absolute/path/to/dist/index.js`).
## Distribution
Published to npm as [`@beeltec/context-usage-mcp`](https://www.npmjs.com/package/@beeltec/context-usage-mcp)
and run via `npx` (see [Install & run](#install--run)). Releases are published from CI with
provenance on each GitHub Release; running from a local build is also supported for development.
## The tool: `get_context_usage`
No arguments. Returns raw counts only — no percentage, no context-window detection. It may be
**unavailable** early in a session before the first model response. Both a JSON `text` block and
typed `structuredContent` (declared `outputSchema`) are returned with the same object. The shape is
identical under Claude Code and Codex.
**Available:**
```json
{
"available": true,
"context_tokens": 86026,
"breakdown": {
"input_tokens": 2,
"cache_creation_input_tokens": 1001,
"cache_read_input_tokens": 85023,
"output_tokens": 225
},
"session_id": "…",
"model": "claude-opus-4-8",
"timestamp": "2026-07-20T14:17:50.751Z",
"reason": null
}
```
**Unavailable** (never throws): numeric fields are `null` and `reason` explains why (no
transcript/rollout found, empty file, or no usage-bearing message/`token_count` event yet).
| Field | Meaning |
|-------|---------|
| `available` | Whether a reading was obtained |
| `context_tokens` | `input + cache_creation + cache_read` (the context fed back to the model) |
| `breakdown` | Raw per-category counts, including `output_tokens` |
| `session_id` / `model` / `timestamp` | Metadata from the source message/event |
| `reason` | Why the reading is unavailable (`null` when available) |
## Intended usage
The server returns **raw numbers only**; the agent decides what "too long" means. The pattern:
call `get_context_usage` after each task and branch on an absolute token threshold defined in the
agent's prompt/CLAUDE.md (e.g. wrap up or hand off when `context_tokens` exceeds a limit).
## Known risks
- **Internal file format.** Both the Claude transcript JSONL and the Codex rollout JSONL are
undocumented internal formats that may change between host versions; parsing them directly is an
accepted fragility.
- **Concurrent-session race.** The freshest-file heuristic (freshest `*.jsonl` for Claude; freshest
cwd-matching `rollout-*.jsonl` for Codex) is racy if multiple sessions run in the same project at
once — it may pick another session's file.
TDQS
Scored across 1 tool
With only one tool, there is no possibility of ambiguity or misselection. The tool's purpose is clearly defined and distinct by default.
The single tool name 'get_context_usage' follows a clear verb_noun pattern. Consistency is trivially maintained as there are no other tools to contrast with.
The server has only one tool, which feels thin for a typical server but is arguably sufficient for its narrow purpose. It is not trivial, but the count is on the lower end of the acceptable range.
The tool covers the core raw token usage reporting, but it explicitly lacks percentage or context-window detection, leaving a notable gap for users who need derived insights. The domain is narrowly focused, but the surface is somewhat incomplete.