Skip to main content
Glama
beeltec

context-usage-mcp

by beeltec
README.md
# context-usage-mcp

A tiny [Model Context Protocol](https://modelcontextprotocol.io) server (TypeScript/Node, stdio)
that exposes a single tool, **`get_context_usage`**, reporting the current session's **raw token
usage** — so an agent can check after each task whether context has grown large and branch its
behavior on an absolute threshold. Runs under both **Claude Code** and **OpenAI Codex CLI**, reading
each host's own session transcript/rollout behind an auto-detected host-adapter layer.

## What it does

Neither host hands an stdio MCP server the live token counts, so this server reads usage
**indirectly** from the host's own on-disk session file, behind an auto-detected host-adapter layer.
The output shape is identical across hosts.

**Claude Code** passes only `CLAUDE_PROJECT_DIR` (no session id, transcript path, or counts):

1. Derive the project folder from `CLAUDE_PROJECT_DIR` (non-alphanumerics → `-`).
2. Pick the **most-recently-modified `*.jsonl`** in `~/.claude/projects/<folder>/` as the current
   session (its transcript is being written while the tool is called).
3. Scan **backward** for the last assistant message carrying a `usage` block.
4. Sum `input_tokens + cache_creation_input_tokens + cache_read_input_tokens` → `context_tokens`.

The `~/.claude` root can be relocated with `CLAUDE_CONFIG_DIR`.

**OpenAI Codex CLI** passes *no* project/session env var at all, but launches the server with its
working directory set to the session's project dir. So the server uses `process.cwd()`:

1. Resolve the Codex root (`CODEX_HOME`, else `~/.codex`); rollouts live under
   `<root>/sessions/YYYY/MM/DD/rollout-*.jsonl`.
2. Scan recent day-folders (today + yesterday, widening to ~7 days), reading only each candidate's
   first `session_meta` line to match its `cwd` against `process.cwd()`; pick the **freshest** match.
3. Scan **backward** for the last `token_count` event and read its per-turn `last_token_usage`.
4. Map to the shared, disjoint breakdown. Codex's `cached_input_tokens` is a *subset* of
   `input_tokens` (verified in source), so `input` ← `input − cached`; then
   `context_tokens = input + cache_read + cache_creation` — no double-count.

Output is `output_tokens + reasoning_output_tokens`. The Codex `model_context_window` is ignored
(raw counts only, strict parity with the Claude host).

## Install & run

Run the published package directly with `npx` — no clone or build required:

```bash
npx -y @beeltec/context-usage-mcp
```

This is the recommended way to wire it into a host (see the registration sections below). The `-y`
flag skips the install prompt on first run.

> **Run from source (local dev)** — for hacking on the server itself:
> ```bash
> npm install
> npm run build          # tsc → dist/, sets +x on dist/index.js
> npm test               # parser unit tests
> ```
> Then register the absolute path to `dist/index.js` (see the "from source" note in each
> registration section).

## Host detection

The server auto-detects its host: `CLAUDE_PROJECT_DIR` set → Claude Code; else `CODEX_HOME` set or
`~/.codex/sessions` present → Codex; else it falls back to Claude. Set the environment variable
`CONTEXT_USAGE_HOST=claude|codex` to force a host (the override wins over auto-detection).

## Register with Claude Code (user scope)

```bash
claude mcp add --scope user context-usage -- npx -y @beeltec/context-usage-mcp
```

Verify with `claude mcp list`. Because MCP servers are loaded at session start, start a **new**
Claude Code session before calling the tool.

> **From source:** register the absolute path to this repo's built binary instead:
> ```bash
> claude mcp add --scope user context-usage -- node /absolute/path/to/dist/index.js
> ```

## Register with Codex CLI

Add an entry to `~/.codex/config.toml` (root movable with `CODEX_HOME`):

```toml
[mcp_servers.context-usage]
command = "npx"
args = ["-y", "@beeltec/context-usage-mcp"]
```

Or use the CLI equivalent:

```bash
codex mcp add context-usage -- npx -y @beeltec/context-usage-mcp
```

**Do not set a custom `cwd` for the server**: Codex passes no project/session identifier, so
discovery relies on the server inheriting the session's working directory via `process.cwd()`.
Setting `cwd` breaks the cwd-match. MCP servers load at session start, so start a **new** Codex
session before calling the tool.

> **From source:** use `command = "node"` with `args = ["/absolute/path/to/dist/index.js"]` (or
> `codex mcp add context-usage -- node /absolute/path/to/dist/index.js`).

## Distribution

Published to npm as [`@beeltec/context-usage-mcp`](https://www.npmjs.com/package/@beeltec/context-usage-mcp)
and run via `npx` (see [Install & run](#install--run)). Releases are published from CI with
provenance on each GitHub Release; running from a local build is also supported for development.

## The tool: `get_context_usage`

No arguments. Returns raw counts only — no percentage, no context-window detection. It may be
**unavailable** early in a session before the first model response. Both a JSON `text` block and
typed `structuredContent` (declared `outputSchema`) are returned with the same object. The shape is
identical under Claude Code and Codex.

**Available:**

```json
{
  "available": true,
  "context_tokens": 86026,
  "breakdown": {
    "input_tokens": 2,
    "cache_creation_input_tokens": 1001,
    "cache_read_input_tokens": 85023,
    "output_tokens": 225
  },
  "session_id": "…",
  "model": "claude-opus-4-8",
  "timestamp": "2026-07-20T14:17:50.751Z",
  "reason": null
}
```

**Unavailable** (never throws): numeric fields are `null` and `reason` explains why (no
transcript/rollout found, empty file, or no usage-bearing message/`token_count` event yet).

| Field | Meaning |
|-------|---------|
| `available` | Whether a reading was obtained |
| `context_tokens` | `input + cache_creation + cache_read` (the context fed back to the model) |
| `breakdown` | Raw per-category counts, including `output_tokens` |
| `session_id` / `model` / `timestamp` | Metadata from the source message/event |
| `reason` | Why the reading is unavailable (`null` when available) |

## Intended usage

The server returns **raw numbers only**; the agent decides what "too long" means. The pattern:
call `get_context_usage` after each task and branch on an absolute token threshold defined in the
agent's prompt/CLAUDE.md (e.g. wrap up or hand off when `context_tokens` exceeds a limit).

## Known risks

- **Internal file format.** Both the Claude transcript JSONL and the Codex rollout JSONL are
  undocumented internal formats that may change between host versions; parsing them directly is an
  accepted fragility.
- **Concurrent-session race.** The freshest-file heuristic (freshest `*.jsonl` for Claude; freshest
  cwd-matching `rollout-*.jsonl` for Codex) is racy if multiple sessions run in the same project at
  once — it may pick another session's file.

TDQS

A4.5/5.0

Scored across 1 tool

Disambiguation5/5

With only one tool, there is no possibility of ambiguity or misselection. The tool's purpose is clearly defined and distinct by default.

Naming Consistency5/5

The single tool name 'get_context_usage' follows a clear verb_noun pattern. Consistency is trivially maintained as there are no other tools to contrast with.

Tool Count3/5

The server has only one tool, which feels thin for a typical server but is arguably sufficient for its narrow purpose. It is not trivial, but the count is on the lower end of the acceptable range.

Completeness3/5

The tool covers the core raw token usage reporting, but it explicitly lacks percentage or context-window detection, leaving a notable gap for users who need derived insights. The domain is narrowly focused, but the surface is somewhat incomplete.

Maintenance

ActivitySlowing
ResponsivenessNo issues