claude-quota-mcp
Reports OpenAI Codex usage limits by reading the exact rate-limit state echoed into local Codex CLI rollout logs (~/.codex/sessions/**/rollout-*.jsonl), with no token or network access required.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@claude-quota-mcpcheck my Claude usage and when the session window resets"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
claude-quota-mcp
Let your agent check its own Claude usage limits, so it can plan long-running work around resets instead of dying mid-task.
Every Claude usage monitor renders for human eyeballs: statuslines, tray apps, terminal dashboards. This is the other direction: a single MCP tool the agent calls to answer questions like:
"The session window resets in 40 minutes and I'm only at 30%. Front-load the heavy work now."
"The weekly window is at 85%. Downshift to a cheaper model before the wall, not after."
"Quota resets before this migration would finish. Schedule the second half for the next window."
Where the numbers come from
The tool tries three sources, best first, and always says which one answered.
Source | What it is | Credential | Network | Windows |
Statusline snapshot (default) | The documented | none | none | 5-hour, weekly (account-wide) |
OAuth usage endpoint (opt-in only) | The undocumented endpoint behind Claude Code's | your Claude Code token, read-only |
| 5-hour, weekly, per-model weekly |
Local estimate (fallback) | Reconstruction from | none | none | 5-hour, times exact, percent estimated against this machine's own history |
Why the endpoint is off by default. Anthropic's Claude Code legal page says
the consumer OAuth token is intended exclusively to support ordinary use of
Claude Code and other native Anthropic applications, and the endpoint is not
documented. Anthropic publishes no usage or quota API for individual Pro or Max
subscriptions. The statusline fields, by contrast, are documented at
https://code.claude.com/docs/en/statusline, carry exact subscriber numbers,
and need no credential at all. So that is the default. If you understand the
terms question and still want the per-model windows, set
CLAUDE_QUOTA_ALLOW_OAUTH_ENDPOINT=1 for the MCP server or hook, or pass
--allow-oauth to the CLI. The request then identifies itself honestly as
claude-quota-mcp; nothing here impersonates Claude Code.
Limits of the statusline source, stated plainly: account-wide windows only, no per-model weekly window; data exists only after Claude Code has made its first API response and rendered a statusline; freshness equals the last render, and the tool reports the snapshot age and labels it stale after fifteen minutes.
Related MCP server: mcp-token-saver
What it returns
| Window | Used | Resets in | Severity |
|---|---|---|---|
| 5-hour session | 17% | 4h 36m | normal |
| Weekly (all models) | 5% | 2d 3h | normal |Why exact percentages matter: local-transcript estimators can only guess percent-used against your own historical maximum. On the machine this was built on, that estimate said 0.7% when the real figure was 13%. Reset times reconstruct fine locally; percentages don't.
Install
Source: https://github.com/PavolKum/claude-quota-mcp. PyPI release: pending
(the uvx claude-quota-mcp one-liner will appear here once the first release
is published). Until then install straight from GitHub, or from a local clone:
python -m pip install git+https://github.com/PavolKum/claude-quota-mcp
# or, inside a clone:
python -m pip install .Then register the MCP server with your host. Claude Code:
claude mcp add quota -- claude-quota-mcpAny MCP host's config:
{
"mcpServers": {
"quota": { "command": "claude-quota-mcp" }
}
}Non-MCP agents and scripts can shell out instead:
python -m claude_quota # best available source, JSON
python -m claude_quota --local # force the transcript estimate
python -m claude_quota --allow-oauth # permit the undocumented endpoint for this call
python -m claude_quota --codex # Codex snapshot onlyFeed the snapshot: the statusline command
The default source needs Claude Code to hand its statusline payload to this package once per render. Two ways:
Standalone. Use the bundled command as your statusline. It writes the
snapshot and prints a compact (5h) 53% (7d) 6%:
"statusLine": { "type": "command", "command": "claude-quota-statusline" }Alongside your own statusline. Run it quietly first, then your script, by
teeing stdin to both (any shell that has tee and process substitution):
tee >(claude-quota-statusline --quiet) | python my_statusline.pyOr call it from Python inside your statusline:
from claude_quota.statusline import write_snapshot
write_snapshot(stdin_data) # the parsed JSON Claude Code piped inOnly the documented rate_limits fields are persisted. Session ids, working
directories, model names and context-window data in the payload are not
written anywhere. Snapshot location: %LOCALAPPDATA%\claude-quota-mcp\statusline-snapshot.json
(override with CLAUDE_QUOTA_SNAPSHOT).
The hook: quota awareness the agent can't forget
A tool the agent must remember to call gets forgotten exactly when it matters, deep in a long task. The included quota-guard hook inverts that: quota state is pushed into context instead of pulled.
SessionStart: always injects a one-line baseline (
5h session 23% (resets 4h 11m); weekly 5% ...).UserPromptSubmit: silent in the common case; speaks only when a window crosses a threshold (default 70%), severity escalates, or there's an opportunity: the session window resets soon with usage to spare. Unused quota expires at reset, so it nudges the agent to front-load heavy work. Repeats are throttled per session.
Claude Code settings.json:
"hooks": {
"SessionStart": [{ "hooks": [{ "type": "command", "command": "claude-quota-hook" }] }],
"UserPromptSubmit": [{ "hooks": [{ "type": "command", "command": "claude-quota-hook" }] }]
}Tune with env vars: QUOTA_HOOK_WARN_PERCENT (70), QUOTA_HOOK_OPPORTUNITY_MINUTES
(60), QUOTA_HOOK_OPPORTUNITY_MAX_PERCENT (50), QUOTA_HOOK_REMIND_MINUTES (30).
The hook reads the statusline snapshot, so per-prompt frequency costs nothing.
Every emitted line carries the current local time (Now Fri 2026-07-18 15:04);
agents have no clock, and "resets in 40m" is meaningless without one. Set
QUOTA_HOOK_TIME_ALWAYS=1 to also emit a minimal <now> tag on every prompt
even when quota has nothing to say (about 10 tokens of temporal grounding;
default off).
Codex too
The same tool reports OpenAI Codex limits (provider: "codex"; the default
"all" covers both). Codex CLI echoes the server's own rate-limit state into
its local rollout logs (~/.codex/sessions/**/rollout-*.jsonl), so the exact
numbers are read straight from disk: no token, no endpoint, no rate-limit risk.
Freshness equals Codex's last activity and is always reported
(snapshot_age_seconds). This is the same local-only pattern the Claude
statusline source follows.
To let Codex check its own quota, register the server in ~/.codex/config.toml:
[mcp_servers.quota]
command = "claude-quota-mcp"Design notes
One tool, tiny schema. MCP servers tax every message with their tool schemas; this one adds a single tool with two optional parameters.
Documented sources first. Both default sources (Claude statusline snapshot, Codex rollout logs) are local files written by the vendors' own tools. No credential leaves the machine unless you opt in.
429-safe when opted in. The OAuth endpoint rate-limits aggressively per token. Responses are disk-cached (180 s TTL, shared across processes), and any network failure serves stale cache before falling back.
Graceful degradation. No snapshot, no opt-in, offline: falls back to the local transcript reconstruction (times exact, percent estimated) and says so in the output. It never fails silently into wrong numbers.
Stdlib core.
quota.pyandstatusline.pyare plain stdlib;mcpandpydanticare needed only by the MCP server itself.
Tests
python test_hook_bridge.py
python test_statusline.pyWhy this exists
Built as a side quest of a multi-agent orchestration experiment: agents coordinating over a shared bus needed to decide when to push hard and when to wind down a lane, which turns out to require knowing your own quota.
MIT licensed.
This server cannot be deployed
Maintenance
Related MCP Connectors
Read-only Codex usage-limit reset data: 24/48h reset forecast, dated reset record, service status.
Token guard and rate limiter preventing runaway API cost spikes for OpenAI and Anthropic.
Live status, API pricing and rate limits for ChatGPT, Claude, Gemini, Cursor and 42+ AI tools.
Live status and health checks for AI coding providers: Claude, Cursor, Copilot, Codex and more.
Related MCP Servers
- AlicenseAqualityNot gradedmaintenanceProvides real-time visibility into Claude Pro and Max subscription usage limits directly within Claude Code by utilizing local OAuth tokens. It enables users to monitor session and weekly usage across different models and receive alerts regarding rate-limiting status.4-
- AlicenseAqualityCmaintenanceReal-time Claude.ai subscription awareness for AI coding assistants. Surfaces live utilization, forecasts limits, gates expensive operations, and measures real per-task cost.517 npm6MIT
- AlicenseNot gradedqualityCmaintenanceSurface Claude Code token usage, estimated cost, and plan-limit status in any MCP client. Enables agents to query usage data from local logs and Anthropic API.MIT
- AlicenseNot gradedqualityCmaintenanceProvides coding agents with pre-flight cost estimation by analyzing local logs to forecast token usage and quota impact before expensive work, exposing tools for remaining quota, per-task cost estimates, and attempt affordability.162 npmMIT