claudeusage
# claudeusage
Measure whether your Claude subscription is actually paying off — from local logs, with no upload and no Git integration.
> **Status: finished (October 2026).** This was a personal research project. It works and stays public, but it is no longer developed and issues may go unanswered. The measurements below are from one account over about five weeks (2026-09-01 → 10-08).
*[한국어 문서: README.ko.md](README.ko.md) — the Korean version is the working notebook, with the full measurements and the mistakes we made getting there.*
## Why this exists
On a subscription you don't pay per token, so nothing tells you whether you are using it well. Most personal tools count **how much** you used. That is the wrong number: burning tokens in circles scores high. Meta ran an internal leaderboard like that and shut it down — people started idling AI to climb it. A few free profilers do look for waste, but they price it in API dollars — and on a subscription the thing that runs out is the limit gauge, which weighs tokens very differently (see below).
Tools that measure quality do exist, but they are all built for teams: they need your Git repo, your PRs, your issue tracker. That leaves a gap.
| | Team tools | claudeusage |
| --- | --- | --- |
| Who is measured | A lead measures the team | You measure yourself |
| Needs | Git repo, PRs, tracker | One local log directory |
| Measures | Return on team investment | Whether your subscription pays off |
| Price | SaaS or self-hosted | Free, runs locally |
Git-based tools track code with `git blame`, so anything you never committed is invisible to them. This reconstructs the edit history itself, so experiments and code you rewrote before committing still count.
## What it measures
**Lines that survived.** Every edit is folded in time order, so only what was still there at the end of the session counts. Code you rewrote three turns later drops out by itself. You cannot game the score by idling.
**Waste, in money.** Rework, failures and rejections add up to 4.1% here. Money spent because conversations got long: **12.3%**. Both are API-dollar figures. Whether starting fresh conversations actually saves *limit* was never proven — a new conversation pays to re-read its context — so treat the 12.3% as a pointer, not a promise.
**What the rate limit actually eats.** Your real constraint is the limit gauge, not dollars. Cache reads are most of the cost but weigh about 1/70 of fresh input against the limit. **Saving money and saving limit are different skills.**
**What chat ate.** Limits apply to the whole account, so claude.ai chat and the mobile app drain the same gauge — while leaving no local trace. Any log-only tool undercounts. This one subtracts what the logs explain from what the gauge actually did, and checks that estimate against the per-product breakdown Anthropic returns.
## What one account's data showed
Measured from about 52,000 status-line samples and 33,000 logged requests. Every multiplier names what it is divided by.
- **Cache reads weigh about 1/70 of fresh input against the limit** (per token, fitted on 107 five-hour windows). They are 97% of all tokens and most of the dollar cost, yet the limit ends up split into similar shares between output, fresh input and cache reads. Dollar-based waste reports overstate them.
- **A weekly-limit change shows up in the gauges alone.** Weekly rise ÷ five-hour rise moved from 0.099 (40 windows, 09-01→13) to 0.125 (41 windows, 09-15→10-02): 1.26×. Anthropic's announced change (bonus +50% → +25%) predicts 1.20×. This only confirms a published change; one account gives about ±3%.
- **Model weight per output token, Opus 5 = 1:** Fable 5.1 ≈ 3.8×, Opus 4.8 ≈ 1.7× (few clean windows; direction solid, exact value not).
- **Chat outside Claude Code was about 10% of the limit**, matching the 8–9% per-product breakdown Anthropic returns.
## Install
Requires Python 3.9+ (standard library only, no dependencies) and macOS for the live-limit tool.
```bash
git clone https://github.com/syk8015/claudeusage && cd claudeusage
./install.sh # package + skill + MCP server (user scope)
claude mcp list # claudeusage … ✔ Connected
```
`install.sh` installs the package (trying `uv`, `pipx`, `pip`, then a dedicated venv — modern Pythons block system installs), drops the Claude Code skill in `~/.claude/skills/`, and registers the MCP server. `./install.sh --uninstall` reverses it.
That gives you the `claudeusage` command and the `claudeusage-mcp` server. There is no PyPI release; install from the clone.
### One more step for the limit tools
Claude Code shows your limit gauge and throws it away — nothing stores it. `claudeusage limit` and `claudeusage chat` need that history, so a collector has to sit on your status line. Add this to `~/.claude/settings.json`:
```json
"statusLine": {
"type": "command",
"command": "claudeusage statusline --print"
}
```
Already have a status line you like? Drop `--print` and pipe into it — the collector passes the JSON straight through:
```json
"command": "claudeusage statusline | bash ~/.claude/my-statusline.sh"
```
Nothing edits your settings for you — your status line is yours.
**`claudeusage value` and `claudeusage usage` work right away**, with no collector. The limit analysis gets useful after a few days of samples.
Samples live in `~/.claudeusage/` (override with `CLAUDEUSAGE_DATA`). Running from a clone keeps using the repo's own `data/` directory, so an existing history is never orphaned.
## Use it from Claude
Ask in plain language — "am I getting my money's worth?", "why is my limit draining so fast?", "how much of my limit did chat eat?" The skill picks the right tool.
| MCP tool | What it answers |
| --- | --- |
| `subscription_value` | Lines that survived, waste, cost by activity |
| `limit_breakdown` | What drives the 5-hour and weekly limits |
| `chat_share` | How much of the limit chat ate |
| `current_limits` | Limits right now, plus the per-product breakdown |
## Use it from the shell
```bash
claudeusage value --all # all projects
claudeusage value --project PATH --waste
claudeusage limit # limit burn per 5-hour window
claudeusage limit --fit # fit: what the gauge weighs
claudeusage chat # chat's share, weekly
claudeusage usage # live limits + product breakdown
claudeusage check # both consistency checkers
```
Output is English by default. `CLAUDEUSAGE_LANG=ko` switches it to Korean.
## How the chat estimate works
Three sources, each blind in a different way, so they are merged:
| Source | Gives | Blind when |
| --- | --- | --- |
| Status line samples (`claudeusage statusline`) | Gauge readings with reset times | Claude Code isn't running |
| Claude desktop app history | Gauge readings every 15 min | The app isn't running |
| OAuth usage endpoint | **Per-product truth** (Claude Code / chat / Cowork) | Only the current weekly window |
The estimate is a subtraction: `chat = gauge rise − what local logs explain`. The "explained" part uses per-model weights fitted by `claudeusage limit`, fitted only on windows with no suspected off-log usage — otherwise chat usage inflates the weights and the estimator goes blind to itself.
Two things had to be corrected, both found by running it on real data:
1. **Half-observed windows tilt the subtraction.** If the gauge is only watched for part of a window, the rise reads low while the attributed cost reads high. Only windows with a known reset time, watched from zero, are counted (55 of 107 here).
2. **Log-derived cost runs 6–9% low** — background calls never reach the logs. Left alone, that gap becomes fake chat usage. Each window is scaled against the status line's own total.
`claudeusage check` verifies 15 invariants, including an anchor: a real window where chat was confirmed by hand.
## Accuracy, honestly
- The per-product breakdown is ground truth, but only two weekly windows of it were collected. Calling the estimator "validated" needs more.
- The per-token fit is unstable for Opus 5.5: with 40 windows it drives the output-token weight to zero, which cannot be right. Inputs are correlated; per-dollar figures are steadier than per-token ones.
- Claude Code deletes conversation logs after 30 days by default (`cleanupPeriodDays`). Gauge samples older than that lose the request data they need, so the limit fit only ever sees the last month.
- Model weights come from 5–7 clean windows for the less-used models. Direction is solid; the exact multiplier still moves.
- Limit gauges are integers. One change point is coarse; answers are averages over ~1,200 of them.
- `claudeusage usage` reads Claude Code's credentials from the macOS keychain, sends them only to `api.anthropic.com`, and never prints or stores them. What it records is an allowlist: timestamp, limit kind, model name, percentage, reset time, product shares.
## Scope
Claude only. Merging ChatGPT and Gemini exports was the original plan for "combine coding and chat"; it was dropped on purpose. Everything here is built around one Claude account and its rate limits.
## Registry
Listed on the official MCP registry as **`io.github.syk8015/claudeusage`**.
<!-- mcp-name: io.github.syk8015/claudeusage -->
Also listed on [Glama](https://glama.ai/mcp/servers/syk8015/claudeusage).
## License
MIT
TDQS
Scored across 4 tools
subscription_value and current_limits are clearly distinct, and the detailed descriptions make limit_breakdown and chat_share distinguishable as mechanism analysis vs. cross-product attribution. An agent skimming names alone could conflate the two rate-limit tools, but the descriptions resolve the ambiguity.
All four names are lowercase snake_case compound nouns, which gives the set a consistent shape. The convention is not verb_noun like typical action tools, and chat_share is a bit ambiguous semantically, but there is no jarring style mixing.
Four tools is compact but each one earns its place: local usage value, limit drivers, chat attribution, and live limit status. The server scope is narrow enough that this feels well-scoped rather than thin.
The set covers the core workflows: measuring local value, explaining limit consumption, attributing non-local usage, and querying current limits. The main gap is that current_limits can record samples to a log, but no tool reads that history back for time-series analysis.