agentrava
An MCP server that turns coding sessions into Strava-style achievement cards with stats, badges, and personal records.
Log activity: Manually report a session's stats (lines, errors, tokens, etc.) and get a card back.
Snapshot current session: Automatically measure the live transcript (tool calls, diffs, errors, moving time) and log/update the running session.
Set athlete: Change the display name shown on every card.
Get profile: View career totals, current streak, personal records, and trophy case.
List activities: Show recent logged sessions with headline stats.
Season recap: Generate a summary card for a period (heatmap, hour-of-day histogram, longest streak, biggest session).
Leaderboard: Rank sessions by distance, elevation, duration, effort, tokens, or tool calls.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@agentravalog my latest coding session and show me my achievement card"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Agentrava
Strava, for coding agents. An MCP server that turns a finished session into a bragging card — route map, climb profile, headline stats, badges, personal records.
Nothing on a card is self-reported. A hook parses the session transcript for tool calls, tokens, diff hunks, recovered errors and moving time. The agent never gets to describe its own workout. Works with Claude Code, Codex CLI and Cursor.
Install
git clone https://github.com/lukisimi/agentrava ~/agentrava && cd ~/agentrava
npm run setupInstalls dependencies, registers the MCP server at user scope, and adds the Stop hook. Idempotent, backs up every file it edits, and reversible:
npm run setup -- --manual # keep the tools, stop logging every turn
npm run setup -- --auto # put automatic logging back
npm run setup -- --cursor # also install the Cursor probe
npm run setup -- --codex # also register the MCP server with Codex CLI
npm run setup -- --uninstall # remove everything (your data is left alone)Restart Claude Code, then node scripts/backfill.mjs to log your history.
Any MCP client works — it speaks stdio:
{ "mcpServers": { "agentrava": { "command": "node", "args": ["/path/to/agentrava/src/index.js"] } } }Related MCP server: Strava MCP Server
Tools
agentrava— start here: every tool with an example, plus your totals and streak.snapshot— card for the session in progress, measured from the live transcript. Takesphoto("chat"uses an image you just pasted) andtitleto rename it inline.log_activity— log a session by hand; unreported fields count as zero.recap— one card for a whole period: totals, activity heatmap, hour-of-day histogram, trophy case, longest streak, biggest session. Optionalfrom/to.weekly_snap·monthly_snap— a week or a month on one card, same size as a session card. Takeperiod(this,last,last7/last30, or a date),title,hide_projects, andpick.get_profile— career totals, streak, personal records, trophy case.rename_session·rename_project·list_projects— your own names, kept across re-logs.list_activities·leaderboard·set_athlete
The metaphor
Strava | Agentrava | Formula |
Distance | ground covered |
|
Elevation | the parts that hurt |
|
Moving time | session time, idle excluded | gaps over 5 min don't count |
Pace | minutes per km |
|
Suffer score | Effort, 0–100 | cadence, elevation, tokens, retries |
Calories | tokens burned | input + cache writes + output |
Economy | tokens per km | lower is leaner |
Gear | the model | measured, never assumed |
— | API cost | priced per message at list rates |
Every weight is fitted to real sessions, not guessed. Churn alone left the median session at 0.00 km — most sessions read and search far more than they write — which is why tool calls carry distance too.
What the numbers actually mean
The inputs are measured. The scales are invented — 100 lines = 1 km, an error = 120 m — chosen so a median session lands near a plausible 4.4 km. That makes cards comparable between your own sessions, which is what records and the leaderboard rest on, and meaningless outside Agentrava.
Measured across 134 sessions:
correlates most with | r | |
Distance | tool calls | 0.94 |
Distance | churn | 0.86 |
Elevation | errors recovered | 0.91 |
Elevation | files changed | 0.88 |
So distance is volume of activity — 69% of it from the tool-call term — and elevation is friction, 58% of it from errors. They correlate 0.79 with each other: overlapping, but about a third of elevation is information distance doesn't carry, which is what separates a long easy session from a short brutal one.
Not every failed tool call is friction
An error is a tool result the client flagged as failed. Sampled across 166 real Claude Code errors: 45% a shell command exiting non-zero, 11% bad arguments or a missing file, 7% a browser step, 19% app-level errors from the user's own tools — and 12% nothing to do with the agent at all: the human declining a tool call, the permission layer blocking one, a model or MCP server briefly unavailable.
Those last ones are counted separately (errors_environmental) and left out of
the climb, because 120 m of elevation and a loop in the route should mean work
that had to be redone, not a moment where someone hit Escape. Across this store
that removed 89 errors and 8 km of climb, and demoted 4 cards from Debug.
The test, in src/errors.js, is a phrase list run against the
first 300 characters of the failure text — these failures announce themselves
in their first line, while a long command output can mention "rate limit"
anywhere. 609 successful Cursor tool results contain that phrase and not one of
them is a rate limit. Deliberately excluded from the list: permission denied
(a real filesystem failure the agent must work around), a bare timed out
(usually its own command hanging) and connection refused (usually a dev server
it forgot to start).
Each client hides the text somewhere different — Claude Code in the
tool_result, Cursor in toolFormerData.result, Codex in a failed item's
result for an MCP call or stderr for a shell one. Serialising a whole Codex
item instead tested the command line and the id, which counted a session that
merely printed the word "rejected" as a refusal.
Errors are not a measure of inefficiency. Their raw count is mostly a size measurement — 0.87 against tool calls. As a rate (errors per 100 tool calls) they correlate 0.08 with tokens per line changed and −0.01 with seconds per line changed. Across 78 sessions, a 5× difference in error rate bought 20% more tokens per line and 18% less time per line: no signal. Elevation honestly means eventful, not wasteful.
Raw tokens cannot rank efficiency. They correlate 0.72 with distance, so the number mostly says how big a session was. Economy (tokens per km) correlates 0.08 with distance — size-independent, and therefore actually comparable. It measures token cost per unit of volume, not of value: a session that finds the right answer in five calls scores badly on it.
None of this measures whether the work was any good. A session that flails for 800 tool calls outscores one that fixes the bug in five. Nothing in a transcript reliably encodes outcome — that's a ceiling, not a tuning problem.
The route and the climb profile
The route map is a random walk seeded by the activity id, so a card always redraws identically. Two things in it are real: its length and density come from tool calls, and every error you recovered from draws as a loop — the trace shows where you went in circles.
The climb profile under it is cumulative elevation: flat where the session ran smoothly, stepping up wherever a file was written or an error recovered, bucketed by moving time so an idle gap doesn't collapse it. The area under the curve is the elevation figure on the card.
That strip was decoration until recently — a seeded random walk reading no session data at all, the same label over pure noise. Sessions with fewer than three climb events now get no strip at all rather than an invented one. 81% have a profile.
Weekly and Monthly Snap
node scripts/snap.mjs week # this calendar week so far
node scripts/snap.mjs week last # the previous calendar week
node scripts/snap.mjs week last7 # rolling seven days, ending today
node scripts/snap.mjs month # this calendar month so far
node scripts/snap.mjs month last30 # rolling thirty days (14d, 90d… also work)
node scripts/snap.mjs month 2026-08 --pick 086f7bf6 # feature a session you chose
node scripts/snap.mjs week --title "Shipped the new onboarding" --hide-projectsCalendar or rolling. A calendar week shared on a Wednesday is a stub — it
says "this week so far" and two days are empty. last7 and last30 are always
whole windows ending today, which is usually what you want to post. Calendar
periods stay the default because streaks, months and heatmaps are calendar
things. Rolling cards label themselves by range and by weekday, since a rolling
week does not start on a Monday.
Both are 1080×1350, the same frame as a session card, so a snap and a card sit side by side in a feed. A six-row month is the tightest case and still clears the footer.
Weekly — agent time per day, sessions / active days / projects, tool calls, estimated cost, and the longest session. Monthly — a Monday-first heatmap, the same headline counts, the top three projects by time, and a featured session.
What makes the numbers hold together:
Daily buckets are recorded, not inferred. Both parsers credit each stretch of moving time to the local day it happened, splitting at midnight, so a session from 23:50 to 00:05 counts 10 minutes on one day and 5 on the next. Every bar and heatmap cell comes from those buckets, and so does Active Days — seven bars can't sit next to a five-day headline.
A session crossing a period boundary appears in both, with its time apportioned. Tool calls and cost have no per-day record, so they are apportioned by the same time share.
"Agent time" is summed across sessions. Nine agents running in parallel for five hours is 45 hours of agent time on one day. The card says so under the chart; it is not your working hours.
Cost says when it is partial —
partial · 38 of 40 priced— and reads "not recorded" rather than $0 when nothing could be priced."Month's pick · Selected by you" appears only when you picked it. Otherwise the feature is labelled "Longest session", which is what it is.
An unfinished period says so — "This week so far".
Nothing is inferred about outcome. A
--titleis yours to write; the card never claims anything shipped.
Badges
Earnable, not participation trophies:
Negative Splits deleted more than you wrote · Flawless no errors, no failed tests ·
Hill Repeats climbed out of it 3+ times · Marathon 1h+ · Ultra 3h+ ·
Sprint under 3 minutes with a diff · Yak Shave 30+ tool calls, barely a diff ·
All Green full suite, zero red · Furnace 5M+ tokens · Nocturnal 11pm–5am ·
Everest 3000m+ · 10K Club 10 km · Gran Fondo 40 km · Polyglot 3+ languages ·
Red Zone effort 90+ · Sightseeing all reading, no writing ·
Signed Off 10+ edits accepted, none sent back (Cursor only)
Measured frequency: Hill Repeats 47%, Marathon 44%, Yak Shave 36%,
Polyglot 31%, 10K Club 25%, Ultra 22%, Flawless 22%, Nocturnal 19%,
Red Zone 14%, Furnace 8%, Everest 8%. Average 3.1 badges per card.
Personal records only fire once there is something to beat, so the first activity never claims one.
Logging
Three modes, in descending cost:
per-turn cost | logs sessions | keeps streaks honest | |
auto (Stop hook) | ~300 ms | automatically | yes |
manual + day stamp (default of | ~10 ms | when you ask | yes |
manual only ( | none | when you ask | no |
The full hook re-parses the transcript and redraws the card every turn. The day stamp is a two-line shell script that appends today's date and nothing else — it never starts a Node process, and writes one line per day. Streaks count the union of logged-activity days and stamped days, so both modes mean the same thing.
By hand, any time:
npm run log # the session you are in
node scripts/log-now.mjs --cursor # the Cursor conversation you are in
node scripts/log-now.mjs --codex # the Codex CLI session you are in
node scripts/log-now.mjs --codex --list
node scripts/log-now.mjs --list # 15 most recent, newest first
node scripts/log-now.mjs 9e22ccfa # one session by id prefixLogging upserts — running it repeatedly on one session updates that activity instead of stacking duplicates. Sessions under 8 tool calls or 2 minutes are ignored.
What the Claude Code hook measures
Field | Source |
Tool calls |
|
Tokens |
|
Lines ± |
|
Files | Edit/Write paths, plus shell redirect / |
Errors recovered |
|
Moving time | consecutive timestamp gaps, each capped at 5 min |
Model |
|
Known limits: cache reads are excluded from the token total (replayed context is not work done); shell writes are detected by regex, deliberately conservative — it misses writes rather than inventing them, and line counts for those files are not recovered, so churn under-reports on shell-heavy sessions; session type is a guess from files, churn and error count.
Backfill
node scripts/backfill.mjs --dry-run # report only, writes nothing
node scripts/backfill.mjs # log everything not yet logged
node scripts/backfill.mjs --force # rebuild from empty (backs up first)
node scripts/cursor-backfill.mjs # same, for Cursor
node scripts/codex-backfill.mjs # same, for Codex CLISorted by session time, because personal records are judged against prior history
— replaying out of order would award them to whichever session happened to be
processed first. Walks ~/.claude/projects/ recursively (git-worktree sessions
live several levels deep). Roughly 900 MB of transcripts takes ~25 s including
card rendering.
Cursor
Cursor stores chat in SQLite at
~/Library/Application Support/Cursor/User/globalStorage/state.vscdb, one row per
message, keyed bubbleId:<conversationId>:<bubbleId>.
The database is WAL-mode, and while Cursor runs it usually has megabytes of
uncommitted log. Opening it with immutable=1 makes SQLite ignore the WAL — which
hides the newest conversations entirely and throws "malformed" when a checkpoint
lands mid-read. Reads use mode=ro, falling back to a snapshot of the db plus its
-wal and -shm. The whole database is read in one grouped pass (~60 s for
289 conversations); per-conversation LIKE queries each scan a multi-GB table.
What Cursor actually records
Measured across 222 logged conversations:
signal | coverage | |
tool calls | 100% | ✅ |
moving time | 100% | ✅ |
climb profile | 89% | ✅ |
errors | 57% | ✅ a real |
| 40% | ✅ accepted / rejected per edit — no equivalent in Claude Code |
tokens | 4% | ❌ |
files changed | 6% | ❌ see below |
line churn | 0% | ❌ see below |
Cursor stores arguments for only 477 of 15,142 edit_file_v2 calls; the rest
have empty rawArgs, and the result holds content hashes rather than paths. So
which file an edit touched is usually unrecoverable, and files_changed
counts only the subset that is.
This was worse before: the parser took a path from any tool carrying one,
including read_file_v2, so files merely opened counted as changed and inflated
elevation (median Cursor elevation 555 m → 120 m once restricted to real edits).
Under-reporting something unmeasurable beats inflating it, so Cursor elevation
rests mainly on errors — which it does record reliably.
Cursor has a stop hook with the same stdio-JSON contract as Claude Code
(conversation_id, transcript_path, workspace_roots, status), and
hooks/cursor-probe.mjs captures one real payload. Live auto-logging is not
wired up: a full scan takes ~60 s, too slow for every turn.
Codex CLI
Codex writes one JSONL rollout per session under
~/.codex/sessions/YYYY/MM/DD/rollout-<timestamp>-<uuid>.jsonl, with a timestamp
on every line, so moving time, daily buckets and the climb profile come out the
same way they do for Claude Code. 130 rollouts totalling 330 MB parse in about
two seconds; the largest single file, 119 MB, takes 566 ms.
It records file changes better than the other two clients. An item_completed
event of type FileChange carries a changes map: the full body for an added
file, a unified diff for an edited one. Churn is counted from those diffs rather
than inferred from shell commands (Claude Code) or given up on (Cursor).
Measured across the 24 sessions above the logging floor:
signal | coverage | |
tool calls | 100% | ✅ |
moving time | 100% | ✅ |
tokens in/out/cached | 100% | ✅ |
files and line churn | 100% of sessions that changed a file | ✅ exact, from |
errors | 54% | ✅ |
model | 100% | ✅ from |
est. API cost | 100% | ✅ OpenAI list prices, per model and tier |
accepted / rejected edits | 0% | ❌ not recorded |
Token totals are cumulative in the log, so the last token_count wins rather
than being summed, and input_tokens includes the cached portion — the cached
tokens are subtracted back out so a Codex card's token count means what a Claude
Code card's does.
Pricing a Codex session takes two passes. The cumulative totals say how many
tokens there were; the per-request last_token_usage figures say which model and
which price tier they ran under. Summing the per-request figures instead would
overstate a session that replays part of its own history — one rollout in four
came out 5% high — so the totals stay authoritative and the per-request numbers
decide only how the tokens divide.
OpenAI charges a whole request at 2× input and 1.5× output once its prompt passes 272K input tokens. Across 3,329 requests here the largest prompt was 245K, and Codex ran a 258K context window, so the higher tier never applied — but it is implemented rather than assumed away.
Codex's own codex-auto-review model has no published rate. Its turns price at
zero rather than being guessed at, exactly as an unrecognised Claude model does.
Two things in the log are not what they look like:
The first user message is not the prompt. Codex opens a session by feeding itself the environment block,
AGENTS.md, the plugin list and a replayed approval history as user messages. Taken verbatim, sessions were titled<recommended_plugins>. Those are skipped.Each Codex conversation gets its own directory when Codex runs from the ChatGPT app:
<workspace>/Codex/<date>/<conversation-slug>. One project per session, each named after the conversation — which would put chat titles on a card meant to be shared. They group under the Codex workspace instead. A Codex session started inside a real repository still reports that repository.
Of 130 rollouts, 105 fall below the 8-tool-call floor: most are chat-only threads with no tools at all, which the same floor would reject in any client.
There is no auto-logging hook. Codex's notify takes a single program and is
usually already claimed by something else, so taking it over would break whatever
was there. Log from Codex with the snapshot MCP tool (npm run setup -- --codex
registers the server in ~/.codex/config.toml) or with
node scripts/log-now.mjs --codex.
Cards
Athlete and gear
The athlete is you, not the model — Strava does not file your rides under the
bike. The model is gear, shown under the title with the client:
Claude Opus 5 · Cursor.
node scripts/whoami.mjs "Your Name" # the name on every card, past and future
node scripts/whoami.mjs --avatar ~/me.jpg # a picture for the circle
node scripts/whoami.mjs --avatar chat # use an image you just pasted into the chat
node scripts/whoami.mjs --no-avatar # back to the initial
node scripts/whoami.mjs # show bothThe avatar replaces the initial in the circle on every card and snap, clipped to
the circle. It is copied into ~/.agentrava/avatars/ on set, so moving or
deleting the original does not blank your cards, and it is swappable any time —
setting a new one redraws everything. set_athlete takes avatar and
avatar_reset for the same thing from chat.
set_athlete does the same from chat. Unset, cards read "Athlete" — the safe
default for sharing.
Renaming sessions and projects
node scripts/rename.mjs session 4e374e0a "The photo that kept vanishing"
node scripts/rename.mjs session 4e374e0a --reset # back to the generated title
node scripts/rename.mjs projects # list, with paths
node scripts/rename.mjs project "acme-api-service" "API"
node scripts/rename.mjs project API --reset
node scripts/rename.mjs project API --hide # keep it off every cardOr from chat: rename_session, rename_project, list_projects. Every card
result ends with the options that apply to it, phrased for the assistant to put
to you —
adding a photo, renaming, or hiding project names before you share. Affected cards
are redrawn immediately — changing a name in a table does nothing to a PNG
already on disk.
Names are presentation, stored beside the measured data rather than in it —
session titles in overrides.json, project names in projects.json — so a
re-log or a forced backfill keeps them, and a reset always restores the
generated name.
Which project a session belongs to comes from the transcript, not from where
you happened to run the logger: the most frequent working directory recorded in
the session that resolves to a repository. Logging the same session from $HOME
used to drop its project entirely, so a project could vanish from a snap with no
work having happened. Agent scaffolding — ~/.claude, ~/.cursor, Claude's
scratch workspaces — is never a project.
A project is its repository path, not its name. Two repositories both called
web stay separate, and giving two projects the same display name does not merge
them. A name that matches more than one project is refused with the paths listed,
rather than guessed; so is a session id prefix that matches more than one session.
Anything under a repository's .claude/ or .cursor/ — worktrees, skills —
counts as that repository.
Hiding a project removes its name from session cards and shows it as "Project A" on snaps. The route is seeded by activity id only, so renaming a session never redraws its route.
Client names and logos
Cards name the client in the header. Known ids: claude-code, claude, cursor,
openai, codex, grok, copilot, windsurf, zed.
With a logo installed, the mark sits large in the top-right corner — 58px, balancing the athlete's avatar on the left, with the activity type and streak shifted in beside it. Without one there is nothing to put there, so the client stays a tinted chip under the title instead — Claude Code terracotta, Codex green, Cursor white — rather than trailing the gear line as grey text.
No logo artwork ships with this repo. Those marks are trademarks, and bundling them into an MIT repo means redistributing brand assets that most brand guidelines restrict. Naming a product is ordinary nominative use; shipping its logo is not. Until you install one the chip shows a coloured dot.
node scripts/logo.mjs # what is installed, and for how many sessions
node scripts/logo.mjs codex ~/openai.svg # install one
node scripts/logo.mjs codex chat # use an image you just pasted into the chat
node scripts/logo.mjs codex --removeThe file lands in ~/.agentrava/logos/<client>.svg (or .png, under 512 KB) and
every card for that client is redrawn. Whether you may put a given mark on a card
you post is between you and that owner's brand guidelines — this repo takes no
position beyond not shipping the artwork itself.
Photos
Strava lets you put your ride photo behind the route. So does this.
node scripts/card.mjs <session> --photo ~/me-in-a-hammock.jpg
node scripts/card.mjs <session> --photo chat # the image you just pasted
node scripts/card.mjs <session> --no-photo--photo chat needs no file: an image pasted into Claude Code never becomes a
file on disk — it lives as base64 in the transcript — so this recovers the most
recent one. jpg/png/gif/webp under 8 MB, embedded so the card stays one
self-contained file.
A centred full-size route sits right on whoever is in the photo, so photo cards default to a smaller route on the left. Move it yourself if the subject is elsewhere — there is no face detection, on purpose:
node scripts/card.mjs <session> --route right --route-scale 0.6
node scripts/card.mjs <session> --route auto # back to the defaultPhoto and route placement live in ~/.agentrava/overrides.json, apart from the
measured data, so re-logging a session or a forced backfill keeps them. (Before
this, every re-log redrew the card from scratch and silently dropped the photo.)
Before you share one
The subtitle is your first prompt, and prompts name customers, vendors and internal projects.
node scripts/privacy.mjs # list subtitles that look sensitive
node scripts/privacy.mjs --strip # blank just those
node scripts/privacy.mjs --strip-all # blank all, and stop recording them
node scripts/card.mjs <session> --no-summaryThe detector flags company suffixes and capitalised proper names; it will not
catch everything — "fix the checkout bug for acme" reads as clean. Read the
subtitle before you post one, or turn summaries off entirely with
{"summaries": "off"} in ~/.agentrava/config.json.
Cost
Claude Code records four token classes per message, so a session is priced per
message at whatever model produced it. Codex is priced per request — see above.
Rates are Anthropic and OpenAI list prices (src/pricing.js, read from their
docs on 18 September 2026); both vendors bill cache writes at 1.25× input and
cache reads at 0.1×, so one table shape covers both.
A model with no published rate is left unpriced rather than guessed at, and a
card with no price shows — rather than $0.00. Snaps say how many of their
sessions were priced, so a gap is visible instead of silent.
This is not a bill. A subscription does not charge per token. The figure is what the session would have cost on the API — useful for comparing sessions, useless as an invoice.
The split is the interesting part. Across 136 priced sessions: 0.7M input, 48.5M output, 373M cache writes and 17.2 billion cache reads. Cache reads are 98% of all tokens, which is why they dominate cost even at a tenth of the input rate.
Data
Activities live in ~/.agentrava/activities.json, cards in ~/.agentrava/cards/
as PNG and SVG. Override with AGENTRAVA_HOME. Nothing leaves the machine —
there is no network call anywhere in this server.
Writes are serialised with a mkdir-based cross-process lock: every session's hook
writes the same file, and without it a 20-way concurrent test lost 19 writes.
Activity ids are derived from the client and session id, so a forced rebuild produces the same ids, the same routes and the same card files — it used to mint new ones and orphan hundreds of cards each time.
Projects are resolved through agent worktrees: a session in
<repo>/.claude/worktrees/<name> belongs to <repo>, read straight off the path
so it still works after the worktree is deleted. Without this, one week counted
18 projects that were really 2.
Development
npm run demo # render sample cards from a synthetic season
node scripts/e2e.js # drive the server over real MCP stdio
node scripts/rerender.js --prune # redraw stored cards, delete orphans
node scripts/recap.js 2026-08-01 2026-08-31License
MIT — see LICENSE.
Available Tools
7 toolsget_profileAthlete profileB
Career totals, current streak, personal records and the trophy case across every logged activity.
| Name | Required | Description | Default |
|---|---|---|---|
| athlete | No | Filter to one athlete. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are present, so the description must carry behavioral weight. It explains the scope ('across every logged activity') and the output categories, which gives some transparency about aggregation behavior. However, it does not clarify default behavior when the optional athlete parameter is omitted, response shape, permissions, or whether anything beyond reading occurs.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, compact sentence with no filler. It front-loads the main output categories and closes with the important scope qualifier ('across every logged activity'). Every phrase contributes meaning.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple tool with one optional parameter and no output schema, the description names the main content areas but leaves some gaps. The term 'trophy case' is not explained, and the behavior when athlete is omitted is not stated. Overall it is adeuate but not fully self-sufficient.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, and the schema already explains that athlete filters to one athlete. The tool description adds no detail about how the athlete filter interacts with 'every logged activity,' but because the schema is complete, a baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly communicates that this tool returns an athlete profile built from career totals, current streak, personal records, and trophy case information. It does not use an explicit verb like 'retrieves' or 'returns,' but the tool name plus the resource content make the purpose clear. It is distinct from raw activity listing, though it doesn't explicitly contrast with sibling tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is given on when to call this tool versus list_activities, recap, snapshot, or leaderboard. The description implies an aggregate/profile use case, but it never states exclusions, alternatives, or prerequisites such as needing a selected athlete.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
leaderboardLeaderboardB
Rank logged sessions by a metric: distance, elevation, duration, effort, tokens or tool_calls.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| metric | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It states the ranking behavior but omits important details such as whether the call is read-only, the sorting direction, default limit behavior, or the shape of the returned leaderboard. This is insufficient for an unannotated tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence that front-loads the core action and resource and lists the metric options without any filler. It is efficient and easy to scan.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no annotations and no output schema, the description is not complete enough for confident invocation. An agent is left without information about the return value format, default parameters, or how the ranking is ordered. The description covers only the basic function.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description lists the metric enum values, but these are already fully captured in the schema. It adds no explanation of what 'effort' means, and it says nothing about the 'limit' parameter. With 0% schema description coverage, the description fails to compensate for the missing parameter semantics.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Rank'), names the resource ('logged sessions'), and enumerates the valid metrics, making it easy for an agent to understand what the tool does. It clearly distinguishes this from sibling tools like list_activities or recap, which have different purposes.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus its siblings. It does not mention alternatives or scenarios where ranking would be inappropriate; the use case is only implicit in the name and the verb 'Rank'.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_activitiesRecent activitiesC
The feed: recent logged sessions with their headline stats.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | How many, newest first. Default 10. | |
| athlete | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden of explaining behavior. It says the tool returns recent logged sessions and headline stats, but does not state whether this is read-only, how results are ordered, whether pagination exists, or how the athlete parameter affects results.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single concise sentence and is front-loaded with the main purpose. It avoids bloat, though it may be too terse to compensate for missing paramdetails.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, no annotations, and one parameter is undocumented. The description conveys only a high-level view of the result and does not mention filtering, ordering, or what 'headline stats' concretely includes. For a simple two-parameter read tool it is close but still has important gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema documents limit ('How many, newest first. Default 10.') but athlete is left undescribed in both schema and description. The description adds no parameter meaning, so the agent has insufficient info to correctly use the athlete parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states this tool returns 'recent logged sessions with their headline stats,' which identifies the resource and output focus. It is not a tautology and can be distinguished from siblings like log_activity, but it does not explicitly contrast with them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance about when to prefer this tool over siblings such as leaderboard or recap, and no mention of alternatives or exclusions. The intended use is only inferred from 'recent' and 'feed,' so the agent does not receive explicit selection criteria.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
log_activityLog an activityA
Finish a coding session and get a Strava-style achievement card back as an image. Report the session honestly — line churn becomes distance, files and recovered errors become elevation, and the card awards badges and personal records against your own history. Call this when the user asks you to brag, or at the end of a session worth remembering.
| Name | Required | Description | Default |
|---|---|---|---|
| date | No | ISO timestamp. Defaults to now. | |
| repo | No | Repo or project name. | |
| type | No | What kind of session this was. Defaults to "feature". | |
| photo | No | Card background, with the route drawn over it — Strava-style. A local image path (jpg/png/gif/webp, under 8 MB), or "chat" to use the image the user most recently pasted into this conversation. Ask the user for one; do not invent a path. | |
| title | No | Optional. Left blank, it is auto-named Strava-style from the clock and type — "Morning Refactor", "Late Night Debug". | |
| tokens | No | Tokens burned, if you know it. | |
| athlete | No | Who did the work. Defaults to "Claude". | |
| summary | No | One line on what you actually did. Shown under the title. | |
| languages | No | Languages touched. | |
| tool_calls | No | How many tool calls you made. | |
| lines_added | No | Lines added. | |
| tests_failed | No | Tests that failed. | |
| tests_passed | No | Tests that passed. | |
| files_changed | No | Distinct files created or edited. | |
| lines_removed | No | Lines removed. | |
| duration_seconds | No | Wall-clock length of the session. | |
| errors_recovered | No | Times you hit an error and worked past it. These draw as loops on the route map — be honest, they are the best part. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the transparency burden and does disclose key behavior: the output is an image card, metrics are mapped gamification-style, and honesty is required. It implies persistence through 'your own history' but does not explicitly state that an activity record is saved.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, front-loaded with the tool's purpose and output, followed by the trigger condition. The playful wording earns its place by reinforcing the Strava-style behavior without adding fluff.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The output is described as an image card, schema documentation covers parameters, and trigger conditions are clear. It could be more complete by explicitly stating that the activity is persisted and what happens on return, but none of this is critical for selecting or invoking the tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema description coverage is 100%, so the schema already documents all 17 parameters. The description adds a helpful metaphor—line churn becomes distance, errors become elevation—but does not explain individual parameter details, which is acceptable given full schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific action and result: finish a coding session and receive a Strava-style achievement card as an image. It distinguishes itself by promising the card, badges, and personal records, though it does not explicitly contrast with sibling tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear call triggers: when the user asks to brag or at the end of a session worth remembering. It does not list excluded scenarios or alternative sibling tools, but the guidance is sufficiently clear for an agent to know when to invoke it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
recapSeason recapB
One card summarising a whole period: totals, a day-by-day activity heatmap, an hour-of-day histogram of when the work actually happened, the trophy case, longest streak and biggest session. Defaults to everything logged.
| Name | Required | Description | Default |
|---|---|---|---|
| to | No | End date, YYYY-MM-DD. Omit for today. | |
| from | No | Start date, YYYY-MM-DD. Omit for the beginning. | |
| title | No | Headline. Defaults to "N Activities". | |
| athlete | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden of behavioral disclosure, and it does reveal the output components and the default date range. However, it does not state whether the tool is read-only, whether it respects a selected athlete, or what happens when no data exists.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence with no wasted words, front-loaded with the core purpose and followed by concrete output details. It is appropriately sized for the tool's complexity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description is reasonably complete for a non-mutating summary tool, listing the main content of the output card and the default scope. However, with no output schema and no annotations, it leaves the athlete parameter's role ambiguous and does not clarify whether 'everything logged' is global or scoped to the current athlete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 75%, and the schema already documents to, from, and title defaults. The description's 'Defaults to everything logged' reinforces the date-range semantics but adds no new meaning for the undocumented athlete parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a clear purpose: produce a single card summarizing a whole period with totals, a heatmap, an hour-of-day histogram, trophy case, longest streak, and biggest session. This distinguishes it from list_activities and log_activity, though it does not explicitly differentiate it from snapshot.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no explicit guidance on when to use recap versus sibling tools like snapshot, leaderboard, or list_activities. It mentions a default scope ('Defaults to everything logged'), but this is a parameter behavior rather than usage routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
set_athleteSet the athlete nameA
Set the name shown on every card. The athlete is the person whose account this is — the model that did the work is recorded separately as gear. Applies to past cards too. Ask the user what they want; do not guess a name from their email or filesystem.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | Display name, e.g. "Luka" or "Luka Pecavar". Max 40 characters. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral burden and does well: it discloses that the change applies to past cards too, and warns against guessing. It could add more about overwriting behavior or lack of reversibility, but the core behavioral implications are covered.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, each earning its place: the main effect, the athlete/gear distinction, and a necessary user-interaction safeguard. Information is front-loaded and there is no fluff.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple one-parameter setter, the description covers what is set, the scope of the change, and how to obtain the value. It does not describe the return value or error behavior, but that is a minor gap for this mutating tool with no output schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema already fully documents the single parameter including format and max length, so the baseline is 3. The description adds useful context about the name being shown on cards, but does not add new parameter-level meaning beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific action ('Set') on a specific resource (the athlete name shown on every card), and clarifies that the athlete is the account owner, not the model/gear. This distinguishes it clearly from siblings like get_profile and list_activities.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit usage guidance: always ask the user for the name and never guess from email or filesystem. It does not explicitly name alternatives, but the guidance is clear enough for an agent to know when and how to invoke this tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
snapshotSnapshot the current sessionA
Log the session that is running right now and return its card — mid-session, without waiting for it to end. Numbers are measured from the transcript (tool calls, tokens, diffs, recovered errors, moving time), not reported by you, so prefer this over log_activity whenever the work happened in Claude Code. Safe to call repeatedly: it updates the same activity instead of adding duplicates. With no argument it guesses the current session (matching working directory, else most recently written) and names which it chose — check that before repeating the numbers.
| Name | Required | Description | Default |
|---|---|---|---|
| photo | No | Card background: a local image path, or "chat" to use the image the user most recently pasted into this conversation. | |
| session | No | Session id prefix. Omit for the session in progress. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries full behavioral burden and does so thoroughly: numbers come from the transcript rather than the agent, repeated calls update rather than duplicate, and session guessing has a defined fallback with a named choice.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Every sentence earns its place: core action first, then sourcing, then alternative preference, then idempotency, then default behavior. It is dense but not padded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with optional parameters and no output schema, the description covers the essential call-time decisions: when to use it, what it measures, repeat-safety, and default session selection. Nothing critical is left for the agent to infer.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds meaning beyond the schema by explaining the no-argument session-guessing behavior and warning the agent to check which session was chosen before repeating numbers.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Opens with a specific verb and resource: 'Log the session that is running right now and return its card.' It also explicitly distances itself from log_activity, making the tool's scope and identity unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a concrete selection rule: prefer this over log_activity when the work happened in Claude Code. It also clarifies when it is safe to call repeatedly and how the no-argument default behaves.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
7 tool updates
v0.1.0- First observed
get_profile - First observed
leaderboard - First observed
list_activities - First observed
log_activity - First observed
recap - First observed
set_athlete - First observed
snapshot
TDQS
Scored across 7 tools
Most tools have clearly distinct roles: profile, feed, recap, leaderboard, and athlete naming are all unambiguous. The only real boundary issue is log_activity versus snapshot, which both log sessions and return cards, though snapshot's explicit preference for Claude Code work helps clarify the split.
The set mixes verb_noun names like log_activity, set_athlete, get_profile, and list_activities with bare nouns like snapshot, recap, and leaderboard, so there is no consistent pattern. All names are readable and lowercase, but the convention is not uniform enough to be considered mostly consistent.
Seven tools is well-scoped for a niche gamified activity tracker, and each tool covers a distinct part of logging, viewing, summarizing, and comparing activities. No tool feels redundant, and the count is appropriate for the server's purpose.
The core lifecycle is covered: log_activity and snapshot create activities, list_activities reads the feed, and profile, recap, and leaderboard provide aggregation and comparison. Minor gaps like no delete or update activity endpoint and no single-activity detail view are workable but not severe.
Maintenance
Related MCP Connectors
Ask your AI about your fitness: activities and data from Garmin, COROS, Strava, GPX and more
Turn Claude or ChatGPT into a cycling coach that plans your week, grades it, and adapts. Free beta.
- MyoAmigoOAuthcom.myoamigo
Agent-first strength-training platform across iOS, Web & MCP: read & write workouts, PRs & plans.
- Coach MCPOAuthai.iamcoach
Your endurance training data in your AI assistant: activities, recovery, plan, workout edits.
Related MCP Servers
- AlicenseBqualityDmaintenanceIntegrates with the Strava API to allow AI assistants to access fitness data including athlete profiles, activity history, and segment statistics. It enables users to query detailed performance metrics and explore geographic segment data through natural language commands.8182 npmMIT
- FlicenseAqualityDmaintenanceEnables AI agents to interact with the Strava API to retrieve athlete statistics and activity data. It provides tools for listing recent activities and fetching detailed information for specific workout sessions.7-
- AlicenseBqualityDmaintenanceEnables users to interact with their Strava data through natural language to analyze workouts, track fitness progress, and explore routes. It supports retrieving detailed activity stats, heart rate data, and segment insights directly within AI assistants.26312 npmMIT
- AlicenseAqualityDmaintenanceEnables AI assistants to directly access and analyze Strava activity data, including runs, rides, and swims, through natural language queries.410 npmMIT