Skip to main content
Glama
README.md
# Agentrava

Strava, for coding agents. An MCP server that turns a finished session into a
bragging card — route map, climb profile, headline stats, badges, personal records.

<p align="center">
  <img src="docs/example.png" width="46%" alt="A single session card">
  <img src="docs/recap.png" width="46%" alt="A season recap card">
</p>
<p align="center">
  <img src="docs/weekly.png" width="46%" alt="A Weekly Snap card">
  <img src="docs/monthly.png" width="46%" alt="A Monthly Snap card">
</p>
<p align="center"><em>A session · a season · a week · a month. All four are generated from a synthetic season — no real usage data ships in this repo.</em></p>

**Nothing on a card is self-reported.** A hook parses the session transcript for
tool calls, tokens, diff hunks, recovered errors and moving time. The agent never
gets to describe its own workout. Works with Claude Code, Codex CLI and Cursor.

## Install

```bash
git clone https://github.com/lukisimi/agentrava ~/agentrava && cd ~/agentrava
npm run setup
```

Installs dependencies, registers the MCP server at user scope, and adds the Stop
hook. Idempotent, backs up every file it edits, and reversible:

```bash
npm run setup -- --manual      # keep the tools, stop logging every turn
npm run setup -- --auto        # put automatic logging back
npm run setup -- --cursor      # also install the Cursor probe
npm run setup -- --codex       # also register the MCP server with Codex CLI
npm run setup -- --uninstall   # remove everything (your data is left alone)
```

Restart Claude Code, then `node scripts/backfill.mjs` to log your history.

Any MCP client works — it speaks stdio:

```json
{ "mcpServers": { "agentrava": { "command": "node", "args": ["/path/to/agentrava/src/index.js"] } } }
```

## Tools

- **`agentrava`** — start here: every tool with an example, plus your totals and streak.
- **`snapshot`** — card for the session in progress, measured from the live transcript.
  Takes `photo` (`"chat"` uses an image you just pasted) and `title` to rename it inline.
- **`log_activity`** — log a session by hand; unreported fields count as zero.
- **`recap`** — one card for a whole period: totals, activity heatmap, hour-of-day
  histogram, trophy case, longest streak, biggest session. Optional `from` / `to`.
- **`weekly_snap`** · **`monthly_snap`** — a week or a month on one card, same size as a session card.
  Take `period` (`this`, `last`, `last7` / `last30`, or a date), `title`, `hide_projects`, and `pick`.
- **`get_profile`** — career totals, streak, personal records, trophy case.
- **`rename_session`** · **`rename_project`** · **`list_projects`** — your own names, kept across re-logs.
- **`list_activities`** · **`leaderboard`** · **`set_athlete`**

## The metaphor

| Strava | Agentrava | Formula |
|---|---|---|
| Distance | ground covered | `churn / 100 + tool_calls / 25` km |
| Elevation | the parts that hurt | `files × 37 + errors × 120 + tests_failed × 45` m |
| Moving time | session time, idle excluded | gaps over 5 min don't count |
| Pace | minutes per km | `time / distance` |
| Suffer score | Effort, 0–100 | cadence, elevation, tokens, retries |
| Calories | tokens burned | input + cache writes + output |
| Economy | tokens per km | lower is leaner |
| Gear | the model | measured, never assumed |
| — | API cost | priced per message at list rates |

Every weight is **fitted to real sessions**, not guessed. Churn alone left the
median session at 0.00 km — most sessions read and search far more than they
write — which is why tool calls carry distance too.

## What the numbers actually mean

The inputs are measured. The **scales are invented** — 100 lines = 1 km, an error
= 120 m — chosen so a median session lands near a plausible 4.4 km. That makes
cards comparable **between your own sessions**, which is what records and the
leaderboard rest on, and meaningless outside Agentrava.

Measured across 134 sessions:

| | correlates most with | r |
|---|---|---|
| Distance | tool calls | **0.94** |
| Distance | churn | 0.86 |
| Elevation | errors recovered | **0.91** |
| Elevation | files changed | 0.88 |

So distance is *volume of activity* — 69% of it from the tool-call term — and
elevation is *friction*, 58% of it from errors. They correlate 0.79 with each
other: overlapping, but about a third of elevation is information distance
doesn't carry, which is what separates a long easy session from a short brutal one.

### Not every failed tool call is friction

An error is a tool result the client flagged as failed. Sampled across 166 real
Claude Code errors: 45% a shell command exiting non-zero, 11% bad arguments or a
missing file, 7% a browser step, 19% app-level errors from the user's own tools —
and **12% nothing to do with the agent at all**: the human declining a tool call,
the permission layer blocking one, a model or MCP server briefly unavailable.

Those last ones are counted separately (`errors_environmental`) and **left out of
the climb**, because 120 m of elevation and a loop in the route should mean work
that had to be redone, not a moment where someone hit Escape. Across this store
that removed 89 errors and 8 km of climb, and demoted 4 cards from Debug.

The test, in [`src/errors.js`](src/errors.js), is a phrase list run against the
**first 300 characters** of the failure text — these failures announce themselves
in their first line, while a long command output can mention "rate limit"
anywhere. 609 *successful* Cursor tool results contain that phrase and not one of
them is a rate limit. Deliberately excluded from the list: `permission denied`
(a real filesystem failure the agent must work around), a bare `timed out`
(usually its own command hanging) and `connection refused` (usually a dev server
it forgot to start).

Each client hides the text somewhere different — Claude Code in the
`tool_result`, Cursor in `toolFormerData.result`, Codex in a failed item's
`result` for an MCP call or `stderr` for a shell one. Serialising a whole Codex
item instead tested the command line and the id, which counted a session that
merely *printed* the word "rejected" as a refusal.

**Errors are not a measure of inefficiency.** Their raw count is mostly a size
measurement — 0.87 against tool calls. As a rate (errors per 100 tool calls) they
correlate 0.08 with tokens per line changed and −0.01 with seconds per line
changed. Across 78 sessions, a 5× difference in error rate bought 20% more tokens
per line and 18% *less* time per line: no signal. Elevation honestly means
eventful, not wasteful.

**Raw tokens cannot rank efficiency.** They correlate 0.72 with distance, so the
number mostly says how big a session was. Economy (tokens per km) correlates 0.08
with distance — size-independent, and therefore actually comparable. It measures
token cost per unit of *volume*, not of *value*: a session that finds the right
answer in five calls scores badly on it.

**None of this measures whether the work was any good.** A session that flails for
800 tool calls outscores one that fixes the bug in five. Nothing in a transcript
reliably encodes outcome — that's a ceiling, not a tuning problem.

## The route and the climb profile

**The route map** is a random walk seeded by the activity id, so a card always
redraws identically. Two things in it are real: its length and density come from
tool calls, and **every error you recovered from draws as a loop** — the trace
shows where you went in circles.

**The climb profile** under it is cumulative elevation: flat where the session ran
smoothly, stepping up wherever a file was written or an error recovered, bucketed
by moving time so an idle gap doesn't collapse it. The area under the curve is the
elevation figure on the card.

That strip was decoration until recently — a seeded random walk reading no session
data at all, the same label over pure noise. Sessions with fewer than three climb
events now get **no strip at all** rather than an invented one. 81% have a profile.

## Weekly and Monthly Snap

```bash
node scripts/snap.mjs week                 # this calendar week so far
node scripts/snap.mjs week last            # the previous calendar week
node scripts/snap.mjs week last7           # rolling seven days, ending today
node scripts/snap.mjs month                # this calendar month so far
node scripts/snap.mjs month last30         # rolling thirty days (14d, 90d… also work)
node scripts/snap.mjs month 2026-08 --pick 086f7bf6   # feature a session you chose
node scripts/snap.mjs week --title "Shipped the new onboarding" --hide-projects
```

**Calendar or rolling.** A calendar week shared on a Wednesday is a stub — it
says "this week so far" and two days are empty. `last7` and `last30` are always
whole windows ending today, which is usually what you want to post. Calendar
periods stay the default because streaks, months and heatmaps are calendar
things. Rolling cards label themselves by range and by weekday, since a rolling
week does not start on a Monday.

Both are **1080×1350, the same frame as a session card**, so a snap and a card sit
side by side in a feed. A six-row month is the tightest case and still clears the
footer.

**Weekly** — agent time per day, sessions / active days / projects, tool calls,
estimated cost, and the longest session. **Monthly** — a Monday-first heatmap,
the same headline counts, the top three projects by time, and a featured session.

What makes the numbers hold together:

- **Daily buckets are recorded, not inferred.** Both parsers credit each stretch
  of moving time to the local day it happened, splitting at midnight, so a session
  from 23:50 to 00:05 counts 10 minutes on one day and 5 on the next. Every bar
  and heatmap cell comes from those buckets, and so does Active Days — seven bars
  can't sit next to a five-day headline.
- **A session crossing a period boundary appears in both**, with its time
  apportioned. Tool calls and cost have no per-day record, so they are apportioned
  by the same time share.
- **"Agent time" is summed across sessions.** Nine agents running in parallel for
  five hours is 45 hours of agent time on one day. The card says so under the
  chart; it is not your working hours.
- **Cost says when it is partial** — `partial · 38 of 40 priced` — and reads
  "not recorded" rather than $0 when nothing could be priced.
- **"Month's pick · Selected by you" appears only when you picked it.** Otherwise
  the feature is labelled "Longest session", which is what it is.
- **An unfinished period says so** — "This week so far".
- **Nothing is inferred about outcome.** A `--title` is yours to write; the card
  never claims anything shipped.

## Badges

Earnable, not participation trophies:

`Negative Splits` deleted more than you wrote · `Flawless` no errors, no failed tests ·
`Hill Repeats` climbed out of it 3+ times · `Marathon` 1h+ · `Ultra` 3h+ ·
`Sprint` under 3 minutes with a diff · `Yak Shave` 30+ tool calls, barely a diff ·
`All Green` full suite, zero red · `Furnace` 5M+ tokens · `Nocturnal` 11pm–5am ·
`Everest` 3000m+ · `10K Club` 10 km · `Gran Fondo` 40 km · `Polyglot` 3+ languages ·
`Red Zone` effort 90+ · `Sightseeing` all reading, no writing ·
`Signed Off` 10+ edits accepted, none sent back (Cursor only)

Measured frequency: `Hill Repeats` 47%, `Marathon` 44%, `Yak Shave` 36%,
`Polyglot` 31%, `10K Club` 25%, `Ultra` 22%, `Flawless` 22%, `Nocturnal` 19%,
`Red Zone` 14%, `Furnace` 8%, `Everest` 8%. Average 3.1 badges per card.

Personal records only fire once there is something to beat, so the first activity
never claims one.

## Logging

Three modes, in descending cost:

| | per-turn cost | logs sessions | keeps streaks honest |
|---|---|---|---|
| **auto** (Stop hook) | ~300 ms | automatically | yes |
| **manual + day stamp** (default of `--manual`) | ~10 ms | when you ask | yes |
| **manual only** (`--manual --no-stamp`) | none | when you ask | no |

The full hook re-parses the transcript and redraws the card every turn. The day
stamp is a two-line shell script that appends today's date and nothing else — it
never starts a Node process, and writes one line per *day*. Streaks count the
union of logged-activity days and stamped days, so both modes mean the same thing.

By hand, any time:

```bash
npm run log                       # the session you are in
node scripts/log-now.mjs --cursor # the Cursor conversation you are in
node scripts/log-now.mjs --codex  # the Codex CLI session you are in
node scripts/log-now.mjs --codex --list
node scripts/log-now.mjs --list   # 15 most recent, newest first
node scripts/log-now.mjs 9e22ccfa # one session by id prefix
```

Logging **upserts** — running it repeatedly on one session updates that activity
instead of stacking duplicates. Sessions under 8 tool calls or 2 minutes are
ignored.

### What the Claude Code hook measures

| Field | Source |
|---|---|
| Tool calls | `tool_use` blocks in assistant messages |
| Tokens | `usage` input / output / cache write / cache read, per model |
| Lines ± | `structuredPatch` hunks from Edit/Write results |
| Files | Edit/Write paths, **plus** shell redirect / `tee` / `sed -i` targets |
| Errors recovered | `tool_result.is_error` |
| Moving time | consecutive timestamp gaps, each capped at 5 min |
| Model | `message.model`, most frequent in the session |

Known limits: **cache reads are excluded from the token total** (replayed context
is not work done); **shell writes are detected by regex**, deliberately
conservative — it misses writes rather than inventing them, and line counts for
those files are not recovered, so churn under-reports on shell-heavy sessions;
**session type is a guess** from files, churn and error count.

## Backfill

```bash
node scripts/backfill.mjs --dry-run    # report only, writes nothing
node scripts/backfill.mjs              # log everything not yet logged
node scripts/backfill.mjs --force      # rebuild from empty (backs up first)
node scripts/cursor-backfill.mjs       # same, for Cursor
node scripts/codex-backfill.mjs        # same, for Codex CLI
```

Sorted by session time, because personal records are judged against prior history
— replaying out of order would award them to whichever session happened to be
processed first. Walks `~/.claude/projects/` recursively (git-worktree sessions
live several levels deep). Roughly 900 MB of transcripts takes ~25 s including
card rendering.

## Cursor

Cursor stores chat in SQLite at
`~/Library/Application Support/Cursor/User/globalStorage/state.vscdb`, one row per
message, keyed `bubbleId:<conversationId>:<bubbleId>`.

The database is **WAL-mode**, and while Cursor runs it usually has megabytes of
uncommitted log. Opening it with `immutable=1` makes SQLite ignore the WAL — which
hides the newest conversations entirely and throws "malformed" when a checkpoint
lands mid-read. Reads use `mode=ro`, falling back to a snapshot of the db plus its
`-wal` and `-shm`. The whole database is read in **one grouped pass** (~60 s for
289 conversations); per-conversation `LIKE` queries each scan a multi-GB table.

### What Cursor actually records

Measured across 222 logged conversations:

| signal | coverage | |
|---|---|---|
| tool calls | 100% | ✅ |
| moving time | 100% | ✅ |
| climb profile | 89% | ✅ |
| errors | 57% | ✅ a real `status` field, cleaner than Claude Code's boolean |
| `userDecision` | 40% | ✅ **accepted / rejected per edit** — no equivalent in Claude Code |
| tokens | 4% | ❌ `tokenCount` unpopulated since Jan 2026 |
| **files changed** | **6%** | ❌ see below |
| **line churn** | **0%** | ❌ see below |

Cursor stores arguments for only **477 of 15,142** `edit_file_v2` calls; the rest
have empty `rawArgs`, and the result holds content hashes rather than paths. So
**which file an edit touched is usually unrecoverable**, and `files_changed`
counts only the subset that is.

This was worse before: the parser took a path from *any* tool carrying one,
including `read_file_v2`, so files merely opened counted as changed and inflated
elevation (median Cursor elevation 555 m → 120 m once restricted to real edits).
Under-reporting something unmeasurable beats inflating it, so Cursor elevation
rests mainly on errors — which it does record reliably.

Cursor has a `stop` hook with the same stdio-JSON contract as Claude Code
(`conversation_id`, `transcript_path`, `workspace_roots`, `status`), and
`hooks/cursor-probe.mjs` captures one real payload. Live auto-logging is **not**
wired up: a full scan takes ~60 s, too slow for every turn.

## Codex CLI

Codex writes one JSONL rollout per session under
`~/.codex/sessions/YYYY/MM/DD/rollout-<timestamp>-<uuid>.jsonl`, with a timestamp
on every line, so moving time, daily buckets and the climb profile come out the
same way they do for Claude Code. 130 rollouts totalling 330 MB parse in about
two seconds; the largest single file, 119 MB, takes 566 ms.

**It records file changes better than the other two clients.** An `item_completed`
event of type `FileChange` carries a `changes` map: the full body for an added
file, a unified diff for an edited one. Churn is counted from those diffs rather
than inferred from shell commands (Claude Code) or given up on (Cursor).

Measured across the 24 sessions above the logging floor:

| signal | coverage | |
|---|---|---|
| tool calls | 100% | ✅ `custom_tool_call`, `function_call`, `tool_search_call` |
| moving time | 100% | ✅ |
| tokens in/out/cached | 100% | ✅ `token_count.info.total_token_usage` |
| files and line churn | 100% of sessions that changed a file | ✅ exact, from `FileChange` diffs |
| errors | 54% | ✅ `item_completed` with `status: "failed"` |
| model | 100% | ✅ from `turn_context` |
| est. API cost | 100% | ✅ OpenAI list prices, per model and tier |
| accepted / rejected edits | 0% | ❌ not recorded |

Token totals are cumulative in the log, so the last `token_count` wins rather
than being summed, and `input_tokens` includes the cached portion — the cached
tokens are subtracted back out so a Codex card's token count means what a Claude
Code card's does.

**Pricing a Codex session takes two passes.** The cumulative totals say how many
tokens there were; the per-request `last_token_usage` figures say which model and
which price tier they ran under. Summing the per-request figures instead would
overstate a session that replays part of its own history — one rollout in four
came out 5% high — so the totals stay authoritative and the per-request numbers
decide only how the tokens divide.

OpenAI charges a whole request at 2× input and 1.5× output once its prompt passes
272K input tokens. Across 3,329 requests here the largest prompt was 245K, and
Codex ran a 258K context window, so the higher tier never applied — but it is
implemented rather than assumed away.

Codex's own `codex-auto-review` model has no published rate. Its turns price at
zero rather than being guessed at, exactly as an unrecognised Claude model does.

Two things in the log are not what they look like:

- **The first user message is not the prompt.** Codex opens a session by feeding
  itself the environment block, `AGENTS.md`, the plugin list and a replayed
  approval history as user messages. Taken verbatim, sessions were titled
  `<recommended_plugins>`. Those are skipped.
- **Each Codex conversation gets its own directory** when Codex runs from the
  ChatGPT app: `<workspace>/Codex/<date>/<conversation-slug>`. One project per
  session, each named after the conversation — which would put chat titles on a
  card meant to be shared. They group under the Codex workspace instead. A Codex
  session started inside a real repository still reports that repository.

Of 130 rollouts, 105 fall below the 8-tool-call floor: most are chat-only threads
with no tools at all, which the same floor would reject in any client.

There is **no auto-logging hook**. Codex's `notify` takes a single program and is
usually already claimed by something else, so taking it over would break whatever
was there. Log from Codex with the `snapshot` MCP tool (`npm run setup -- --codex`
registers the server in `~/.codex/config.toml`) or with
`node scripts/log-now.mjs --codex`.

## Cards

### Athlete and gear

The athlete is **you**, not the model — Strava does not file your rides under the
bike. The model is gear, shown under the title with the client:
`Claude Opus 5 · Cursor`.

```bash
node scripts/whoami.mjs "Your Name"        # the name on every card, past and future
node scripts/whoami.mjs --avatar ~/me.jpg  # a picture for the circle
node scripts/whoami.mjs --avatar chat      # use an image you just pasted into the chat
node scripts/whoami.mjs --no-avatar        # back to the initial
node scripts/whoami.mjs                    # show both
```

The avatar replaces the initial in the circle on every card and snap, clipped to
the circle. It is copied into `~/.agentrava/avatars/` on set, so moving or
deleting the original does not blank your cards, and it is swappable any time —
setting a new one redraws everything. `set_athlete` takes `avatar` and
`avatar_reset` for the same thing from chat.

`set_athlete` does the same from chat. Unset, cards read "Athlete" — the safe
default for sharing.

### Renaming sessions and projects

```bash
node scripts/rename.mjs session 4e374e0a "The photo that kept vanishing"
node scripts/rename.mjs session 4e374e0a --reset          # back to the generated title
node scripts/rename.mjs projects                          # list, with paths
node scripts/rename.mjs project "acme-api-service" "API"
node scripts/rename.mjs project API --reset
node scripts/rename.mjs project API --hide                # keep it off every card
```

Or from chat: `rename_session`, `rename_project`, `list_projects`. Every card
result ends with the options that apply to it, phrased for the assistant to put
to you —
adding a photo, renaming, or hiding project names before you share. Affected cards
are redrawn immediately — changing a name in a table does nothing to a PNG
already on disk.

Names are presentation, stored beside the measured data rather than in it —
session titles in `overrides.json`, project names in `projects.json` — so a
re-log or a forced backfill keeps them, and a reset always restores the
generated name.

**Which project a session belongs to comes from the transcript**, not from where
you happened to run the logger: the most frequent working directory recorded in
the session that resolves to a repository. Logging the same session from `$HOME`
used to drop its project entirely, so a project could vanish from a snap with no
work having happened. Agent scaffolding — `~/.claude`, `~/.cursor`, Claude's
scratch workspaces — is never a project.

**A project is its repository path, not its name.** Two repositories both called
`web` stay separate, and giving two projects the same display name does not merge
them. A name that matches more than one project is refused with the paths listed,
rather than guessed; so is a session id prefix that matches more than one session.
Anything under a repository's `.claude/` or `.cursor/` — worktrees, skills —
counts as that repository.

Hiding a project removes its name from session cards and shows it as
"Project A" on snaps. The route is seeded by activity id only, so renaming a
session never redraws its route.

### Client names and logos

Cards name the client in the header. Known ids: `claude-code`, `claude`, `cursor`,
`openai`, `codex`, `grok`, `copilot`, `windsurf`, `zed`.

**With a logo installed, the mark sits large in the top-right corner** — 58px,
balancing the athlete's avatar on the left, with the activity type and streak
shifted in beside it. Without one there is nothing to put there, so the client
stays a **tinted chip** under the title instead — Claude Code terracotta, Codex
green, Cursor white — rather than trailing the gear line as grey text.

**No logo artwork ships with this repo.** Those marks are trademarks, and bundling
them into an MIT repo means redistributing brand assets that most brand guidelines
restrict. Naming a product is ordinary nominative use; shipping its logo is not.
Until you install one the chip shows a coloured dot.

```bash
node scripts/logo.mjs                      # what is installed, and for how many sessions
node scripts/logo.mjs codex ~/openai.svg   # install one
node scripts/logo.mjs codex chat           # use an image you just pasted into the chat
node scripts/logo.mjs codex --remove
```

The file lands in `~/.agentrava/logos/<client>.svg` (or `.png`, under 512 KB) and
every card for that client is redrawn. Whether you may put a given mark on a card
you post is between you and that owner's brand guidelines — this repo takes no
position beyond not shipping the artwork itself.

### Photos

Strava lets you put your ride photo behind the route. So does this.

```bash
node scripts/card.mjs <session> --photo ~/me-in-a-hammock.jpg
node scripts/card.mjs <session> --photo chat   # the image you just pasted
node scripts/card.mjs <session> --no-photo
```

`--photo chat` needs no file: an image pasted into Claude Code never becomes a
file on disk — it lives as base64 in the transcript — so this recovers the most
recent one. jpg/png/gif/webp under 8 MB, embedded so the card stays one
self-contained file.

A centred full-size route sits right on whoever is in the photo, so photo cards
default to a **smaller route on the left**. Move it yourself if the subject is
elsewhere — there is no face detection, on purpose:

```bash
node scripts/card.mjs <session> --route right --route-scale 0.6
node scripts/card.mjs <session> --route auto      # back to the default
```

Photo and route placement live in `~/.agentrava/overrides.json`, apart from the
measured data, so re-logging a session or a forced backfill keeps them. (Before
this, every re-log redrew the card from scratch and silently dropped the photo.)

### Before you share one

The subtitle is your first prompt, and prompts name customers, vendors and
internal projects.

```bash
node scripts/privacy.mjs              # list subtitles that look sensitive
node scripts/privacy.mjs --strip      # blank just those
node scripts/privacy.mjs --strip-all  # blank all, and stop recording them
node scripts/card.mjs <session> --no-summary
```

The detector flags company suffixes and capitalised proper names; **it will not
catch everything** — "fix the checkout bug for acme" reads as clean. Read the
subtitle before you post one, or turn summaries off entirely with
`{"summaries": "off"}` in `~/.agentrava/config.json`.

## Cost

Claude Code records four token classes per message, so a session is priced per
message at whatever model produced it. Codex is priced per request — see above.
Rates are Anthropic and OpenAI list prices (`src/pricing.js`, read from their
docs on 18 September 2026); both vendors bill cache writes at 1.25× input and
cache reads at 0.1×, so one table shape covers both.

A model with no published rate is left unpriced rather than guessed at, and a
card with no price shows `—` rather than `$0.00`. Snaps say how many of their
sessions were priced, so a gap is visible instead of silent.

**This is not a bill.** A subscription does not charge per token. The figure is
what the session *would* have cost on the API — useful for comparing sessions,
useless as an invoice.

The split is the interesting part. Across 136 priced sessions: 0.7M input, 48.5M
output, 373M cache writes and 17.2 **billion** cache reads. Cache reads are 98% of
all tokens, which is why they dominate cost even at a tenth of the input rate.

## Data

Activities live in `~/.agentrava/activities.json`, cards in `~/.agentrava/cards/`
as PNG and SVG. Override with `AGENTRAVA_HOME`. **Nothing leaves the machine** —
there is no network call anywhere in this server.

Writes are serialised with a `mkdir`-based cross-process lock: every session's hook
writes the same file, and without it a 20-way concurrent test lost 19 writes.

Activity ids are derived from the client and session id, so a forced rebuild
produces the same ids, the same routes and the same card files — it used to mint
new ones and orphan hundreds of cards each time.

Projects are resolved through agent worktrees: a session in
`<repo>/.claude/worktrees/<name>` belongs to `<repo>`, read straight off the path
so it still works after the worktree is deleted. Without this, one week counted
18 projects that were really 2.

## Development

```bash
npm run demo                      # render sample cards from a synthetic season
node scripts/e2e.js               # drive the server over real MCP stdio
node scripts/rerender.js --prune  # redraw stored cards, delete orphans
node scripts/recap.js 2026-08-01 2026-08-31
```

## License

MIT — see [LICENSE](LICENSE).

TDQS

B3.4/5.0

Scored across 7 tools

Disambiguation4/5

Most tools have clearly distinct roles: profile, feed, recap, leaderboard, and athlete naming are all unambiguous. The only real boundary issue is log_activity versus snapshot, which both log sessions and return cards, though snapshot's explicit preference for Claude Code work helps clarify the split.

Naming Consistency3/5

The set mixes verb_noun names like log_activity, set_athlete, get_profile, and list_activities with bare nouns like snapshot, recap, and leaderboard, so there is no consistent pattern. All names are readable and lowercase, but the convention is not uniform enough to be considered mostly consistent.

Tool Count5/5

Seven tools is well-scoped for a niche gamified activity tracker, and each tool covers a distinct part of logging, viewing, summarizing, and comparing activities. No tool feels redundant, and the count is appropriate for the server's purpose.

Completeness4/5

The core lifecycle is covered: log_activity and snapshot create activities, list_activities reads the feed, and profile, recap, and leaderboard provide aggregation and comparison. Minor gaps like no delete or update activity endpoint and no single-activity detail view are workable but not severe.

Maintenance

ActivityMaintained
ResponsivenessNo issues