Skip to main content
Glama
PavolKum

claude-quota-mcp

by PavolKum

claude-quota-mcp

Let your agent check its own Claude usage limits, so it can plan long-running work around resets instead of dying mid-task.

Every Claude usage monitor renders for human eyeballs: statuslines, tray apps, terminal dashboards. This is the other direction: a single MCP tool the agent calls to answer questions like:

  • "The session window resets in 40 minutes and I'm only at 30%. Front-load the heavy work now."

  • "The weekly window is at 85%. Downshift to a cheaper model before the wall, not after."

  • "Quota resets before this migration would finish. Schedule the second half for the next window."

Where the numbers come from

The tool tries three sources, best first, and always says which one answered.

Source

What it is

Credential

Network

Windows

Statusline snapshot (default)

The documented rate_limits fields Claude Code pipes to its statusline command (five_hour and seven_day: used_percentage, resets_at), persisted locally by claude-quota-statusline

none

none

5-hour, weekly (account-wide)

OAuth usage endpoint (opt-in only)

The undocumented endpoint behind Claude Code's /usage, reached with your Claude Code OAuth token

your Claude Code token, read-only

api.anthropic.com, cached 180 s

5-hour, weekly, per-model weekly

Local estimate (fallback)

Reconstruction from ~/.claude/projects/**/*.jsonl using the ccusage "blocks" method

none

none

5-hour, times exact, percent estimated against this machine's own history

Why the endpoint is off by default. Anthropic's Claude Code legal page says the consumer OAuth token is intended exclusively to support ordinary use of Claude Code and other native Anthropic applications, and the endpoint is not documented. Anthropic publishes no usage or quota API for individual Pro or Max subscriptions. The statusline fields, by contrast, are documented at https://code.claude.com/docs/en/statusline, carry exact subscriber numbers, and need no credential at all. So that is the default. If you understand the terms question and still want the per-model windows, set CLAUDE_QUOTA_ALLOW_OAUTH_ENDPOINT=1 for the MCP server or hook, or pass --allow-oauth to the CLI. The request then identifies itself honestly as claude-quota-mcp; nothing here impersonates Claude Code.

Limits of the statusline source, stated plainly: account-wide windows only, no per-model weekly window; data exists only after Claude Code has made its first API response and rendered a statusline; freshness equals the last render, and the tool reports the snapshot age and labels it stale after fifteen minutes.

Related MCP server: mcp-token-saver

What it returns

| Window | Used | Resets in | Severity |
|---|---|---|---|
| 5-hour session | 17% | 4h 36m | normal |
| Weekly (all models) | 5% | 2d 3h | normal |

Why exact percentages matter: local-transcript estimators can only guess percent-used against your own historical maximum. On the machine this was built on, that estimate said 0.7% when the real figure was 13%. Reset times reconstruct fine locally; percentages don't.

Install

Source: https://github.com/PavolKum/claude-quota-mcp. PyPI release: pending (the uvx claude-quota-mcp one-liner will appear here once the first release is published). Until then install straight from GitHub, or from a local clone:

python -m pip install git+https://github.com/PavolKum/claude-quota-mcp
# or, inside a clone:
python -m pip install .

Then register the MCP server with your host. Claude Code:

claude mcp add quota -- claude-quota-mcp

Any MCP host's config:

{
  "mcpServers": {
    "quota": { "command": "claude-quota-mcp" }
  }
}

Non-MCP agents and scripts can shell out instead:

python -m claude_quota                # best available source, JSON
python -m claude_quota --local        # force the transcript estimate
python -m claude_quota --allow-oauth  # permit the undocumented endpoint for this call
python -m claude_quota --codex        # Codex snapshot only

Feed the snapshot: the statusline command

The default source needs Claude Code to hand its statusline payload to this package once per render. Two ways:

Standalone. Use the bundled command as your statusline. It writes the snapshot and prints a compact (5h) 53% (7d) 6%:

"statusLine": { "type": "command", "command": "claude-quota-statusline" }

Alongside your own statusline. Run it quietly first, then your script, by teeing stdin to both (any shell that has tee and process substitution):

tee >(claude-quota-statusline --quiet) | python my_statusline.py

Or call it from Python inside your statusline:

from claude_quota.statusline import write_snapshot
write_snapshot(stdin_data)  # the parsed JSON Claude Code piped in

Only the documented rate_limits fields are persisted. Session ids, working directories, model names and context-window data in the payload are not written anywhere. Snapshot location: %LOCALAPPDATA%\claude-quota-mcp\statusline-snapshot.json (override with CLAUDE_QUOTA_SNAPSHOT).

The hook: quota awareness the agent can't forget

A tool the agent must remember to call gets forgotten exactly when it matters, deep in a long task. The included quota-guard hook inverts that: quota state is pushed into context instead of pulled.

  • SessionStart: always injects a one-line baseline (5h session 23% (resets 4h 11m); weekly 5% ...).

  • UserPromptSubmit: silent in the common case; speaks only when a window crosses a threshold (default 70%), severity escalates, or there's an opportunity: the session window resets soon with usage to spare. Unused quota expires at reset, so it nudges the agent to front-load heavy work. Repeats are throttled per session.

Claude Code settings.json:

"hooks": {
  "SessionStart":     [{ "hooks": [{ "type": "command", "command": "claude-quota-hook" }] }],
  "UserPromptSubmit": [{ "hooks": [{ "type": "command", "command": "claude-quota-hook" }] }]
}

Tune with env vars: QUOTA_HOOK_WARN_PERCENT (70), QUOTA_HOOK_OPPORTUNITY_MINUTES (60), QUOTA_HOOK_OPPORTUNITY_MAX_PERCENT (50), QUOTA_HOOK_REMIND_MINUTES (30). The hook reads the statusline snapshot, so per-prompt frequency costs nothing.

Every emitted line carries the current local time (Now Fri 2026-07-18 15:04); agents have no clock, and "resets in 40m" is meaningless without one. Set QUOTA_HOOK_TIME_ALWAYS=1 to also emit a minimal <now> tag on every prompt even when quota has nothing to say (about 10 tokens of temporal grounding; default off).

Codex too

The same tool reports OpenAI Codex limits (provider: "codex"; the default "all" covers both). Codex CLI echoes the server's own rate-limit state into its local rollout logs (~/.codex/sessions/**/rollout-*.jsonl), so the exact numbers are read straight from disk: no token, no endpoint, no rate-limit risk. Freshness equals Codex's last activity and is always reported (snapshot_age_seconds). This is the same local-only pattern the Claude statusline source follows.

To let Codex check its own quota, register the server in ~/.codex/config.toml:

[mcp_servers.quota]
command = "claude-quota-mcp"

Design notes

  • One tool, tiny schema. MCP servers tax every message with their tool schemas; this one adds a single tool with two optional parameters.

  • Documented sources first. Both default sources (Claude statusline snapshot, Codex rollout logs) are local files written by the vendors' own tools. No credential leaves the machine unless you opt in.

  • 429-safe when opted in. The OAuth endpoint rate-limits aggressively per token. Responses are disk-cached (180 s TTL, shared across processes), and any network failure serves stale cache before falling back.

  • Graceful degradation. No snapshot, no opt-in, offline: falls back to the local transcript reconstruction (times exact, percent estimated) and says so in the output. It never fails silently into wrong numbers.

  • Stdlib core. quota.py and statusline.py are plain stdlib; mcp and pydantic are needed only by the MCP server itself.

Tests

python test_hook_bridge.py
python test_statusline.py

Why this exists

Built as a side quest of a multi-agent orchestration experiment: agents coordinating over a shared bus needed to decide when to push hard and when to wind down a lane, which turns out to require knowing your own quota.

MIT licensed.

Related MCP Connectors

Related MCP Servers

  • A
    license
    A
    quality
    Not graded
    maintenance
    Provides real-time visibility into Claude Pro and Max subscription usage limits directly within Claude Code by utilizing local OAuth tokens. It enables users to monitor session and weekly usage across different models and receive alerts regarding rate-limiting status.
    4
    -
  • A
    license
    A
    quality
    C
    maintenance
    Real-time Claude.ai subscription awareness for AI coding assistants. Surfaces live utilization, forecasts limits, gates expensive operations, and measures real per-task cost.
    5
    17 npm
    6
    MIT
  • A
    license
    Not graded
    quality
    C
    maintenance
    Surface Claude Code token usage, estimated cost, and plan-limit status in any MCP client. Enables agents to query usage data from local logs and Anthropic API.
    MIT
  • A
    license
    Not graded
    quality
    C
    maintenance
    Provides coding agents with pre-flight cost estimation by analyzing local logs to forecast token usage and quota impact before expensive work, exposing tools for remaining quota, per-task cost estimates, and attempt affordability.
    162 npm
    MIT