Skip to main content
Glama

Agentrava

Strava, for coding agents. An MCP server that turns a finished session into a bragging card — route map, climb profile, headline stats, badges, personal records.

Nothing on a card is self-reported. A hook parses the session transcript for tool calls, tokens, diff hunks, recovered errors and moving time. The agent never gets to describe its own workout. Works with Claude Code, Codex CLI and Cursor.

Install

git clone https://github.com/lukisimi/agentrava ~/agentrava && cd ~/agentrava
npm run setup

Installs dependencies, registers the MCP server at user scope, and adds the Stop hook. Idempotent, backs up every file it edits, and reversible:

npm run setup -- --manual      # keep the tools, stop logging every turn
npm run setup -- --auto        # put automatic logging back
npm run setup -- --cursor      # also install the Cursor probe
npm run setup -- --codex       # also register the MCP server with Codex CLI
npm run setup -- --uninstall   # remove everything (your data is left alone)

Restart Claude Code, then node scripts/backfill.mjs to log your history.

Any MCP client works — it speaks stdio:

{ "mcpServers": { "agentrava": { "command": "node", "args": ["/path/to/agentrava/src/index.js"] } } }

Related MCP server: Strava MCP Server

Tools

  • agentrava — start here: every tool with an example, plus your totals and streak.

  • snapshot — card for the session in progress, measured from the live transcript. Takes photo ("chat" uses an image you just pasted) and title to rename it inline.

  • log_activity — log a session by hand; unreported fields count as zero.

  • recap — one card for a whole period: totals, activity heatmap, hour-of-day histogram, trophy case, longest streak, biggest session. Optional from / to.

  • weekly_snap · monthly_snap — a week or a month on one card, same size as a session card. Take period (this, last, last7 / last30, or a date), title, hide_projects, and pick.

  • get_profile — career totals, streak, personal records, trophy case.

  • rename_session · rename_project · list_projects — your own names, kept across re-logs.

  • list_activities · leaderboard · set_athlete

The metaphor

Strava

Agentrava

Formula

Distance

ground covered

churn / 100 + tool_calls / 25 km

Elevation

the parts that hurt

files × 37 + errors × 120 + tests_failed × 45 m

Moving time

session time, idle excluded

gaps over 5 min don't count

Pace

minutes per km

time / distance

Suffer score

Effort, 0–100

cadence, elevation, tokens, retries

Calories

tokens burned

input + cache writes + output

Economy

tokens per km

lower is leaner

Gear

the model

measured, never assumed

—

API cost

priced per message at list rates

Every weight is fitted to real sessions, not guessed. Churn alone left the median session at 0.00 km — most sessions read and search far more than they write — which is why tool calls carry distance too.

What the numbers actually mean

The inputs are measured. The scales are invented — 100 lines = 1 km, an error = 120 m — chosen so a median session lands near a plausible 4.4 km. That makes cards comparable between your own sessions, which is what records and the leaderboard rest on, and meaningless outside Agentrava.

Measured across 134 sessions:

correlates most with

r

Distance

tool calls

0.94

Distance

churn

0.86

Elevation

errors recovered

0.91

Elevation

files changed

0.88

So distance is volume of activity — 69% of it from the tool-call term — and elevation is friction, 58% of it from errors. They correlate 0.79 with each other: overlapping, but about a third of elevation is information distance doesn't carry, which is what separates a long easy session from a short brutal one.

Not every failed tool call is friction

An error is a tool result the client flagged as failed. Sampled across 166 real Claude Code errors: 45% a shell command exiting non-zero, 11% bad arguments or a missing file, 7% a browser step, 19% app-level errors from the user's own tools — and 12% nothing to do with the agent at all: the human declining a tool call, the permission layer blocking one, a model or MCP server briefly unavailable.

Those last ones are counted separately (errors_environmental) and left out of the climb, because 120 m of elevation and a loop in the route should mean work that had to be redone, not a moment where someone hit Escape. Across this store that removed 89 errors and 8 km of climb, and demoted 4 cards from Debug.

The test, in src/errors.js, is a phrase list run against the first 300 characters of the failure text — these failures announce themselves in their first line, while a long command output can mention "rate limit" anywhere. 609 successful Cursor tool results contain that phrase and not one of them is a rate limit. Deliberately excluded from the list: permission denied (a real filesystem failure the agent must work around), a bare timed out (usually its own command hanging) and connection refused (usually a dev server it forgot to start).

Each client hides the text somewhere different — Claude Code in the tool_result, Cursor in toolFormerData.result, Codex in a failed item's result for an MCP call or stderr for a shell one. Serialising a whole Codex item instead tested the command line and the id, which counted a session that merely printed the word "rejected" as a refusal.

Errors are not a measure of inefficiency. Their raw count is mostly a size measurement — 0.87 against tool calls. As a rate (errors per 100 tool calls) they correlate 0.08 with tokens per line changed and −0.01 with seconds per line changed. Across 78 sessions, a 5× difference in error rate bought 20% more tokens per line and 18% less time per line: no signal. Elevation honestly means eventful, not wasteful.

Raw tokens cannot rank efficiency. They correlate 0.72 with distance, so the number mostly says how big a session was. Economy (tokens per km) correlates 0.08 with distance — size-independent, and therefore actually comparable. It measures token cost per unit of volume, not of value: a session that finds the right answer in five calls scores badly on it.

None of this measures whether the work was any good. A session that flails for 800 tool calls outscores one that fixes the bug in five. Nothing in a transcript reliably encodes outcome — that's a ceiling, not a tuning problem.

The route and the climb profile

The route map is a random walk seeded by the activity id, so a card always redraws identically. Two things in it are real: its length and density come from tool calls, and every error you recovered from draws as a loop — the trace shows where you went in circles.

The climb profile under it is cumulative elevation: flat where the session ran smoothly, stepping up wherever a file was written or an error recovered, bucketed by moving time so an idle gap doesn't collapse it. The area under the curve is the elevation figure on the card.

That strip was decoration until recently — a seeded random walk reading no session data at all, the same label over pure noise. Sessions with fewer than three climb events now get no strip at all rather than an invented one. 81% have a profile.

Weekly and Monthly Snap

node scripts/snap.mjs week                 # this calendar week so far
node scripts/snap.mjs week last            # the previous calendar week
node scripts/snap.mjs week last7           # rolling seven days, ending today
node scripts/snap.mjs month                # this calendar month so far
node scripts/snap.mjs month last30         # rolling thirty days (14d, 90d… also work)
node scripts/snap.mjs month 2026-08 --pick 086f7bf6   # feature a session you chose
node scripts/snap.mjs week --title "Shipped the new onboarding" --hide-projects

Calendar or rolling. A calendar week shared on a Wednesday is a stub — it says "this week so far" and two days are empty. last7 and last30 are always whole windows ending today, which is usually what you want to post. Calendar periods stay the default because streaks, months and heatmaps are calendar things. Rolling cards label themselves by range and by weekday, since a rolling week does not start on a Monday.

Both are 1080×1350, the same frame as a session card, so a snap and a card sit side by side in a feed. A six-row month is the tightest case and still clears the footer.

Weekly — agent time per day, sessions / active days / projects, tool calls, estimated cost, and the longest session. Monthly — a Monday-first heatmap, the same headline counts, the top three projects by time, and a featured session.

What makes the numbers hold together:

  • Daily buckets are recorded, not inferred. Both parsers credit each stretch of moving time to the local day it happened, splitting at midnight, so a session from 23:50 to 00:05 counts 10 minutes on one day and 5 on the next. Every bar and heatmap cell comes from those buckets, and so does Active Days — seven bars can't sit next to a five-day headline.

  • A session crossing a period boundary appears in both, with its time apportioned. Tool calls and cost have no per-day record, so they are apportioned by the same time share.

  • "Agent time" is summed across sessions. Nine agents running in parallel for five hours is 45 hours of agent time on one day. The card says so under the chart; it is not your working hours.

  • Cost says when it is partial — partial · 38 of 40 priced — and reads "not recorded" rather than $0 when nothing could be priced.

  • "Month's pick · Selected by you" appears only when you picked it. Otherwise the feature is labelled "Longest session", which is what it is.

  • An unfinished period says so — "This week so far".

  • Nothing is inferred about outcome. A --title is yours to write; the card never claims anything shipped.

Badges

Earnable, not participation trophies:

Negative Splits deleted more than you wrote · Flawless no errors, no failed tests · Hill Repeats climbed out of it 3+ times · Marathon 1h+ · Ultra 3h+ · Sprint under 3 minutes with a diff · Yak Shave 30+ tool calls, barely a diff · All Green full suite, zero red · Furnace 5M+ tokens · Nocturnal 11pm–5am · Everest 3000m+ · 10K Club 10 km · Gran Fondo 40 km · Polyglot 3+ languages · Red Zone effort 90+ · Sightseeing all reading, no writing · Signed Off 10+ edits accepted, none sent back (Cursor only)

Measured frequency: Hill Repeats 47%, Marathon 44%, Yak Shave 36%, Polyglot 31%, 10K Club 25%, Ultra 22%, Flawless 22%, Nocturnal 19%, Red Zone 14%, Furnace 8%, Everest 8%. Average 3.1 badges per card.

Personal records only fire once there is something to beat, so the first activity never claims one.

Logging

Three modes, in descending cost:

per-turn cost

logs sessions

keeps streaks honest

auto (Stop hook)

~300 ms

automatically

yes

manual + day stamp (default of --manual)

~10 ms

when you ask

yes

manual only (--manual --no-stamp)

none

when you ask

no

The full hook re-parses the transcript and redraws the card every turn. The day stamp is a two-line shell script that appends today's date and nothing else — it never starts a Node process, and writes one line per day. Streaks count the union of logged-activity days and stamped days, so both modes mean the same thing.

By hand, any time:

npm run log                       # the session you are in
node scripts/log-now.mjs --cursor # the Cursor conversation you are in
node scripts/log-now.mjs --codex  # the Codex CLI session you are in
node scripts/log-now.mjs --codex --list
node scripts/log-now.mjs --list   # 15 most recent, newest first
node scripts/log-now.mjs 9e22ccfa # one session by id prefix

Logging upserts — running it repeatedly on one session updates that activity instead of stacking duplicates. Sessions under 8 tool calls or 2 minutes are ignored.

What the Claude Code hook measures

Field

Source

Tool calls

tool_use blocks in assistant messages

Tokens

usage input / output / cache write / cache read, per model

Lines ±

structuredPatch hunks from Edit/Write results

Files

Edit/Write paths, plus shell redirect / tee / sed -i targets

Errors recovered

tool_result.is_error

Moving time

consecutive timestamp gaps, each capped at 5 min

Model

message.model, most frequent in the session

Known limits: cache reads are excluded from the token total (replayed context is not work done); shell writes are detected by regex, deliberately conservative — it misses writes rather than inventing them, and line counts for those files are not recovered, so churn under-reports on shell-heavy sessions; session type is a guess from files, churn and error count.

Backfill

node scripts/backfill.mjs --dry-run    # report only, writes nothing
node scripts/backfill.mjs              # log everything not yet logged
node scripts/backfill.mjs --force      # rebuild from empty (backs up first)
node scripts/cursor-backfill.mjs       # same, for Cursor
node scripts/codex-backfill.mjs        # same, for Codex CLI

Sorted by session time, because personal records are judged against prior history — replaying out of order would award them to whichever session happened to be processed first. Walks ~/.claude/projects/ recursively (git-worktree sessions live several levels deep). Roughly 900 MB of transcripts takes ~25 s including card rendering.

Cursor

Cursor stores chat in SQLite at ~/Library/Application Support/Cursor/User/globalStorage/state.vscdb, one row per message, keyed bubbleId:<conversationId>:<bubbleId>.

The database is WAL-mode, and while Cursor runs it usually has megabytes of uncommitted log. Opening it with immutable=1 makes SQLite ignore the WAL — which hides the newest conversations entirely and throws "malformed" when a checkpoint lands mid-read. Reads use mode=ro, falling back to a snapshot of the db plus its -wal and -shm. The whole database is read in one grouped pass (~60 s for 289 conversations); per-conversation LIKE queries each scan a multi-GB table.

What Cursor actually records

Measured across 222 logged conversations:

signal

coverage

tool calls

100%

✅

moving time

100%

✅

climb profile

89%

✅

errors

57%

✅ a real status field, cleaner than Claude Code's boolean

userDecision

40%

✅ accepted / rejected per edit — no equivalent in Claude Code

tokens

4%

❌ tokenCount unpopulated since Jan 2026

files changed

6%

❌ see below

line churn

0%

❌ see below

Cursor stores arguments for only 477 of 15,142 edit_file_v2 calls; the rest have empty rawArgs, and the result holds content hashes rather than paths. So which file an edit touched is usually unrecoverable, and files_changed counts only the subset that is.

This was worse before: the parser took a path from any tool carrying one, including read_file_v2, so files merely opened counted as changed and inflated elevation (median Cursor elevation 555 m → 120 m once restricted to real edits). Under-reporting something unmeasurable beats inflating it, so Cursor elevation rests mainly on errors — which it does record reliably.

Cursor has a stop hook with the same stdio-JSON contract as Claude Code (conversation_id, transcript_path, workspace_roots, status), and hooks/cursor-probe.mjs captures one real payload. Live auto-logging is not wired up: a full scan takes ~60 s, too slow for every turn.

Codex CLI

Codex writes one JSONL rollout per session under ~/.codex/sessions/YYYY/MM/DD/rollout-<timestamp>-<uuid>.jsonl, with a timestamp on every line, so moving time, daily buckets and the climb profile come out the same way they do for Claude Code. 130 rollouts totalling 330 MB parse in about two seconds; the largest single file, 119 MB, takes 566 ms.

It records file changes better than the other two clients. An item_completed event of type FileChange carries a changes map: the full body for an added file, a unified diff for an edited one. Churn is counted from those diffs rather than inferred from shell commands (Claude Code) or given up on (Cursor).

Measured across the 24 sessions above the logging floor:

signal

coverage

tool calls

100%

✅ custom_tool_call, function_call, tool_search_call

moving time

100%

✅

tokens in/out/cached

100%

✅ token_count.info.total_token_usage

files and line churn

100% of sessions that changed a file

✅ exact, from FileChange diffs

errors

54%

✅ item_completed with status: "failed"

model

100%

✅ from turn_context

est. API cost

100%

✅ OpenAI list prices, per model and tier

accepted / rejected edits

0%

❌ not recorded

Token totals are cumulative in the log, so the last token_count wins rather than being summed, and input_tokens includes the cached portion — the cached tokens are subtracted back out so a Codex card's token count means what a Claude Code card's does.

Pricing a Codex session takes two passes. The cumulative totals say how many tokens there were; the per-request last_token_usage figures say which model and which price tier they ran under. Summing the per-request figures instead would overstate a session that replays part of its own history — one rollout in four came out 5% high — so the totals stay authoritative and the per-request numbers decide only how the tokens divide.

OpenAI charges a whole request at 2× input and 1.5× output once its prompt passes 272K input tokens. Across 3,329 requests here the largest prompt was 245K, and Codex ran a 258K context window, so the higher tier never applied — but it is implemented rather than assumed away.

Codex's own codex-auto-review model has no published rate. Its turns price at zero rather than being guessed at, exactly as an unrecognised Claude model does.

Two things in the log are not what they look like:

  • The first user message is not the prompt. Codex opens a session by feeding itself the environment block, AGENTS.md, the plugin list and a replayed approval history as user messages. Taken verbatim, sessions were titled <recommended_plugins>. Those are skipped.

  • Each Codex conversation gets its own directory when Codex runs from the ChatGPT app: <workspace>/Codex/<date>/<conversation-slug>. One project per session, each named after the conversation — which would put chat titles on a card meant to be shared. They group under the Codex workspace instead. A Codex session started inside a real repository still reports that repository.

Of 130 rollouts, 105 fall below the 8-tool-call floor: most are chat-only threads with no tools at all, which the same floor would reject in any client.

There is no auto-logging hook. Codex's notify takes a single program and is usually already claimed by something else, so taking it over would break whatever was there. Log from Codex with the snapshot MCP tool (npm run setup -- --codex registers the server in ~/.codex/config.toml) or with node scripts/log-now.mjs --codex.

Cards

Athlete and gear

The athlete is you, not the model — Strava does not file your rides under the bike. The model is gear, shown under the title with the client: Claude Opus 5 · Cursor.

node scripts/whoami.mjs "Your Name"        # the name on every card, past and future
node scripts/whoami.mjs --avatar ~/me.jpg  # a picture for the circle
node scripts/whoami.mjs --avatar chat      # use an image you just pasted into the chat
node scripts/whoami.mjs --no-avatar        # back to the initial
node scripts/whoami.mjs                    # show both

The avatar replaces the initial in the circle on every card and snap, clipped to the circle. It is copied into ~/.agentrava/avatars/ on set, so moving or deleting the original does not blank your cards, and it is swappable any time — setting a new one redraws everything. set_athlete takes avatar and avatar_reset for the same thing from chat.

set_athlete does the same from chat. Unset, cards read "Athlete" — the safe default for sharing.

Renaming sessions and projects

node scripts/rename.mjs session 4e374e0a "The photo that kept vanishing"
node scripts/rename.mjs session 4e374e0a --reset          # back to the generated title
node scripts/rename.mjs projects                          # list, with paths
node scripts/rename.mjs project "acme-api-service" "API"
node scripts/rename.mjs project API --reset
node scripts/rename.mjs project API --hide                # keep it off every card

Or from chat: rename_session, rename_project, list_projects. Every card result ends with the options that apply to it, phrased for the assistant to put to you — adding a photo, renaming, or hiding project names before you share. Affected cards are redrawn immediately — changing a name in a table does nothing to a PNG already on disk.

Names are presentation, stored beside the measured data rather than in it — session titles in overrides.json, project names in projects.json — so a re-log or a forced backfill keeps them, and a reset always restores the generated name.

Which project a session belongs to comes from the transcript, not from where you happened to run the logger: the most frequent working directory recorded in the session that resolves to a repository. Logging the same session from $HOME used to drop its project entirely, so a project could vanish from a snap with no work having happened. Agent scaffolding — ~/.claude, ~/.cursor, Claude's scratch workspaces — is never a project.

A project is its repository path, not its name. Two repositories both called web stay separate, and giving two projects the same display name does not merge them. A name that matches more than one project is refused with the paths listed, rather than guessed; so is a session id prefix that matches more than one session. Anything under a repository's .claude/ or .cursor/ — worktrees, skills — counts as that repository.

Hiding a project removes its name from session cards and shows it as "Project A" on snaps. The route is seeded by activity id only, so renaming a session never redraws its route.

Client names and logos

Cards name the client in the header. Known ids: claude-code, claude, cursor, openai, codex, grok, copilot, windsurf, zed.

With a logo installed, the mark sits large in the top-right corner — 58px, balancing the athlete's avatar on the left, with the activity type and streak shifted in beside it. Without one there is nothing to put there, so the client stays a tinted chip under the title instead — Claude Code terracotta, Codex green, Cursor white — rather than trailing the gear line as grey text.

No logo artwork ships with this repo. Those marks are trademarks, and bundling them into an MIT repo means redistributing brand assets that most brand guidelines restrict. Naming a product is ordinary nominative use; shipping its logo is not. Until you install one the chip shows a coloured dot.

node scripts/logo.mjs                      # what is installed, and for how many sessions
node scripts/logo.mjs codex ~/openai.svg   # install one
node scripts/logo.mjs codex chat           # use an image you just pasted into the chat
node scripts/logo.mjs codex --remove

The file lands in ~/.agentrava/logos/<client>.svg (or .png, under 512 KB) and every card for that client is redrawn. Whether you may put a given mark on a card you post is between you and that owner's brand guidelines — this repo takes no position beyond not shipping the artwork itself.

Photos

Strava lets you put your ride photo behind the route. So does this.

node scripts/card.mjs <session> --photo ~/me-in-a-hammock.jpg
node scripts/card.mjs <session> --photo chat   # the image you just pasted
node scripts/card.mjs <session> --no-photo

--photo chat needs no file: an image pasted into Claude Code never becomes a file on disk — it lives as base64 in the transcript — so this recovers the most recent one. jpg/png/gif/webp under 8 MB, embedded so the card stays one self-contained file.

A centred full-size route sits right on whoever is in the photo, so photo cards default to a smaller route on the left. Move it yourself if the subject is elsewhere — there is no face detection, on purpose:

node scripts/card.mjs <session> --route right --route-scale 0.6
node scripts/card.mjs <session> --route auto      # back to the default

Photo and route placement live in ~/.agentrava/overrides.json, apart from the measured data, so re-logging a session or a forced backfill keeps them. (Before this, every re-log redrew the card from scratch and silently dropped the photo.)

Before you share one

The subtitle is your first prompt, and prompts name customers, vendors and internal projects.

node scripts/privacy.mjs              # list subtitles that look sensitive
node scripts/privacy.mjs --strip      # blank just those
node scripts/privacy.mjs --strip-all  # blank all, and stop recording them
node scripts/card.mjs <session> --no-summary

The detector flags company suffixes and capitalised proper names; it will not catch everything — "fix the checkout bug for acme" reads as clean. Read the subtitle before you post one, or turn summaries off entirely with {"summaries": "off"} in ~/.agentrava/config.json.

Cost

Claude Code records four token classes per message, so a session is priced per message at whatever model produced it. Codex is priced per request — see above. Rates are Anthropic and OpenAI list prices (src/pricing.js, read from their docs on 18 September 2026); both vendors bill cache writes at 1.25× input and cache reads at 0.1×, so one table shape covers both.

A model with no published rate is left unpriced rather than guessed at, and a card with no price shows — rather than $0.00. Snaps say how many of their sessions were priced, so a gap is visible instead of silent.

This is not a bill. A subscription does not charge per token. The figure is what the session would have cost on the API — useful for comparing sessions, useless as an invoice.

The split is the interesting part. Across 136 priced sessions: 0.7M input, 48.5M output, 373M cache writes and 17.2 billion cache reads. Cache reads are 98% of all tokens, which is why they dominate cost even at a tenth of the input rate.

Data

Activities live in ~/.agentrava/activities.json, cards in ~/.agentrava/cards/ as PNG and SVG. Override with AGENTRAVA_HOME. Nothing leaves the machine — there is no network call anywhere in this server.

Writes are serialised with a mkdir-based cross-process lock: every session's hook writes the same file, and without it a 20-way concurrent test lost 19 writes.

Activity ids are derived from the client and session id, so a forced rebuild produces the same ids, the same routes and the same card files — it used to mint new ones and orphan hundreds of cards each time.

Projects are resolved through agent worktrees: a session in <repo>/.claude/worktrees/<name> belongs to <repo>, read straight off the path so it still works after the worktree is deleted. Without this, one week counted 18 projects that were really 2.

Development

npm run demo                      # render sample cards from a synthetic season
node scripts/e2e.js               # drive the server over real MCP stdio
node scripts/rerender.js --prune  # redraw stored cards, delete orphans
node scripts/recap.js 2026-08-01 2026-08-31

License

MIT — see LICENSE.

Available Tools

7 tools
get_profileAthlete profileB

Career totals, current streak, personal records and the trophy case across every logged activity.

ParametersJSON Schema
NameRequiredDescriptionDefault
athleteNoFilter to one athlete.

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are present, so the description must carry behavioral weight. It explains the scope ('across every logged activity') and the output categories, which gives some transparency about aggregation behavior. However, it does not clarify default behavior when the optional athlete parameter is omitted, response shape, permissions, or whether anything beyond reading occurs.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, compact sentence with no filler. It front-loads the main output categories and closes with the important scope qualifier ('across every logged activity'). Every phrase contributes meaning.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple tool with one optional parameter and no output schema, the description names the main content areas but leaves some gaps. The term 'trophy case' is not explained, and the behavior when athlete is omitted is not stated. Overall it is adeuate but not fully self-sufficient.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and the schema already explains that athlete filters to one athlete. The tool description adds no detail about how the athlete filter interacts with 'every logged activity,' but because the schema is complete, a baseline of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly communicates that this tool returns an athlete profile built from career totals, current streak, personal records, and trophy case information. It does not use an explicit verb like 'retrieves' or 'returns,' but the tool name plus the resource content make the purpose clear. It is distinct from raw activity listing, though it doesn't explicitly contrast with sibling tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is given on when to call this tool versus list_activities, recap, snapshot, or leaderboard. The description implies an aggregate/profile use case, but it never states exclusions, alternatives, or prerequisites such as needing a selected athlete.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

leaderboardLeaderboardB

Rank logged sessions by a metric: distance, elevation, duration, effort, tokens or tool_calls.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
metricNo

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. It states the ranking behavior but omits important details such as whether the call is read-only, the sorting direction, default limit behavior, or the shape of the returned leaderboard. This is insufficient for an unannotated tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence that front-loads the core action and resource and lists the metric options without any filler. It is efficient and easy to scan.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no annotations and no output schema, the description is not complete enough for confident invocation. An agent is left without information about the return value format, default parameters, or how the ranking is ordered. The description covers only the basic function.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The description lists the metric enum values, but these are already fully captured in the schema. It adds no explanation of what 'effort' means, and it says nothing about the 'limit' parameter. With 0% schema description coverage, the description fails to compensate for the missing parameter semantics.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Rank'), names the resource ('logged sessions'), and enumerates the valid metrics, making it easy for an agent to understand what the tool does. It clearly distinguishes this from sibling tools like list_activities or recap, which have different purposes.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus its siblings. It does not mention alternatives or scenarios where ranking would be inappropriate; the use case is only implicit in the name and the verb 'Rank'.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_activitiesRecent activitiesC

The feed: recent logged sessions with their headline stats.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoHow many, newest first. Default 10.
athleteNo

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden of explaining behavior. It says the tool returns recent logged sessions and headline stats, but does not state whether this is read-only, how results are ordered, whether pagination exists, or how the athlete parameter affects results.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single concise sentence and is front-loaded with the main purpose. It avoids bloat, though it may be too terse to compensate for missing paramdetails.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, no annotations, and one parameter is undocumented. The description conveys only a high-level view of the result and does not mention filtering, ordering, or what 'headline stats' concretely includes. For a simple two-parameter read tool it is close but still has important gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema documents limit ('How many, newest first. Default 10.') but athlete is left undescribed in both schema and description. The description adds no parameter meaning, so the agent has insufficient info to correctly use the athlete parameter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states this tool returns 'recent logged sessions with their headline stats,' which identifies the resource and output focus. It is not a tautology and can be distinguished from siblings like log_activity, but it does not explicitly contrast with them.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance about when to prefer this tool over siblings such as leaderboard or recap, and no mention of alternatives or exclusions. The intended use is only inferred from 'recent' and 'feed,' so the agent does not receive explicit selection criteria.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

log_activityLog an activityA

Finish a coding session and get a Strava-style achievement card back as an image. Report the session honestly — line churn becomes distance, files and recovered errors become elevation, and the card awards badges and personal records against your own history. Call this when the user asks you to brag, or at the end of a session worth remembering.

ParametersJSON Schema
NameRequiredDescriptionDefault
dateNoISO timestamp. Defaults to now.
repoNoRepo or project name.
typeNoWhat kind of session this was. Defaults to "feature".
photoNoCard background, with the route drawn over it — Strava-style. A local image path (jpg/png/gif/webp, under 8 MB), or "chat" to use the image the user most recently pasted into this conversation. Ask the user for one; do not invent a path.
titleNoOptional. Left blank, it is auto-named Strava-style from the clock and type — "Morning Refactor", "Late Night Debug".
tokensNoTokens burned, if you know it.
athleteNoWho did the work. Defaults to "Claude".
summaryNoOne line on what you actually did. Shown under the title.
languagesNoLanguages touched.
tool_callsNoHow many tool calls you made.
lines_addedNoLines added.
tests_failedNoTests that failed.
tests_passedNoTests that passed.
files_changedNoDistinct files created or edited.
lines_removedNoLines removed.
duration_secondsNoWall-clock length of the session.
errors_recoveredNoTimes you hit an error and worked past it. These draw as loops on the route map — be honest, they are the best part.

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the transparency burden and does disclose key behavior: the output is an image card, metrics are mapped gamification-style, and honesty is required. It implies persistence through 'your own history' but does not explicitly state that an activity record is saved.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, front-loaded with the tool's purpose and output, followed by the trigger condition. The playful wording earns its place by reinforcing the Strava-style behavior without adding fluff.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The output is described as an image card, schema documentation covers parameters, and trigger conditions are clear. It could be more complete by explicitly stating that the activity is persisted and what happens on return, but none of this is critical for selecting or invoking the tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema description coverage is 100%, so the schema already documents all 17 parameters. The description adds a helpful metaphor—line churn becomes distance, errors become elevation—but does not explain individual parameter details, which is acceptable given full schema coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific action and result: finish a coding session and receive a Strava-style achievement card as an image. It distinguishes itself by promising the card, badges, and personal records, though it does not explicitly contrast with sibling tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear call triggers: when the user asks to brag or at the end of a session worth remembering. It does not list excluded scenarios or alternative sibling tools, but the guidance is sufficiently clear for an agent to know when to invoke it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

recapSeason recapB

One card summarising a whole period: totals, a day-by-day activity heatmap, an hour-of-day histogram of when the work actually happened, the trophy case, longest streak and biggest session. Defaults to everything logged.

ParametersJSON Schema
NameRequiredDescriptionDefault
toNoEnd date, YYYY-MM-DD. Omit for today.
fromNoStart date, YYYY-MM-DD. Omit for the beginning.
titleNoHeadline. Defaults to "N Activities".
athleteNo

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the burden of behavioral disclosure, and it does reveal the output components and the default date range. However, it does not state whether the tool is read-only, whether it respects a selected athlete, or what happens when no data exists.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence with no wasted words, front-loaded with the core purpose and followed by concrete output details. It is appropriately sized for the tool's complexity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description is reasonably complete for a non-mutating summary tool, listing the main content of the output card and the default scope. However, with no output schema and no annotations, it leaves the athlete parameter's role ambiguous and does not clarify whether 'everything logged' is global or scoped to the current athlete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 75%, and the schema already documents to, from, and title defaults. The description's 'Defaults to everything logged' reinforces the date-range semantics but adds no new meaning for the undocumented athlete parameter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a clear purpose: produce a single card summarizing a whole period with totals, a heatmap, an hour-of-day histogram, trophy case, longest streak, and biggest session. This distinguishes it from list_activities and log_activity, though it does not explicitly differentiate it from snapshot.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no explicit guidance on when to use recap versus sibling tools like snapshot, leaderboard, or list_activities. It mentions a default scope ('Defaults to everything logged'), but this is a parameter behavior rather than usage routing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

set_athleteSet the athlete nameA

Set the name shown on every card. The athlete is the person whose account this is — the model that did the work is recorded separately as gear. Applies to past cards too. Ask the user what they want; do not guess a name from their email or filesystem.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYesDisplay name, e.g. "Luka" or "Luka Pecavar". Max 40 characters.

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the behavioral burden and does well: it discloses that the change applies to past cards too, and warns against guessing. It could add more about overwriting behavior or lack of reversibility, but the core behavioral implications are covered.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, each earning its place: the main effect, the athlete/gear distinction, and a necessary user-interaction safeguard. Information is front-loaded and there is no fluff.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple one-parameter setter, the description covers what is set, the scope of the change, and how to obtain the value. It does not describe the return value or error behavior, but that is a minor gap for this mutating tool with no output schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema already fully documents the single parameter including format and max length, so the baseline is 3. The description adds useful context about the name being shown on cards, but does not add new parameter-level meaning beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific action ('Set') on a specific resource (the athlete name shown on every card), and clarifies that the athlete is the account owner, not the model/gear. This distinguishes it clearly from siblings like get_profile and list_activities.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit usage guidance: always ask the user for the name and never guess from email or filesystem. It does not explicitly name alternatives, but the guidance is clear enough for an agent to know when and how to invoke this tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

snapshotSnapshot the current sessionA

Log the session that is running right now and return its card — mid-session, without waiting for it to end. Numbers are measured from the transcript (tool calls, tokens, diffs, recovered errors, moving time), not reported by you, so prefer this over log_activity whenever the work happened in Claude Code. Safe to call repeatedly: it updates the same activity instead of adding duplicates. With no argument it guesses the current session (matching working directory, else most recently written) and names which it chose — check that before repeating the numbers.

ParametersJSON Schema
NameRequiredDescriptionDefault
photoNoCard background: a local image path, or "chat" to use the image the user most recently pasted into this conversation.
sessionNoSession id prefix. Omit for the session in progress.

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries full behavioral burden and does so thoroughly: numbers come from the transcript rather than the agent, repeated calls update rather than duplicate, and session guessing has a defined fallback with a named choice.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Every sentence earns its place: core action first, then sourcing, then alternative preference, then idempotency, then default behavior. It is dense but not padded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with optional parameters and no output schema, the description covers the essential call-time decisions: when to use it, what it measures, repeat-safety, and default session selection. Nothing critical is left for the agent to infer.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds meaning beyond the schema by explaining the no-argument session-guessing behavior and warning the agent to check which session was chosen before repeating numbers.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Opens with a specific verb and resource: 'Log the session that is running right now and return its card.' It also explicitly distances itself from log_activity, making the tool's scope and identity unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives a concrete selection rule: prefer this over log_activity when the work happened in Claude Code. It also clarifies when it is safe to call repeatedly and how the no-argument default behaves.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 7 tool updatesv0.1.0
    • First observedget_profile
    • First observedleaderboard
    • First observedlist_activities
    • First observedlog_activity
    • First observedrecap
    • First observedset_athlete
    • First observedsnapshot

TDQS

B3.4/5.0

Scored across 7 tools

Disambiguation4/5

Most tools have clearly distinct roles: profile, feed, recap, leaderboard, and athlete naming are all unambiguous. The only real boundary issue is log_activity versus snapshot, which both log sessions and return cards, though snapshot's explicit preference for Claude Code work helps clarify the split.

Naming Consistency3/5

The set mixes verb_noun names like log_activity, set_athlete, get_profile, and list_activities with bare nouns like snapshot, recap, and leaderboard, so there is no consistent pattern. All names are readable and lowercase, but the convention is not uniform enough to be considered mostly consistent.

Tool Count5/5

Seven tools is well-scoped for a niche gamified activity tracker, and each tool covers a distinct part of logging, viewing, summarizing, and comparing activities. No tool feels redundant, and the count is appropriate for the server's purpose.

Completeness4/5

The core lifecycle is covered: log_activity and snapshot create activities, list_activities reads the feed, and profile, recap, and leaderboard provide aggregation and comparison. Minor gaps like no delete or update activity endpoint and no single-activity detail view are workable but not severe.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    B
    quality
    D
    maintenance
    Integrates with the Strava API to allow AI assistants to access fitness data including athlete profiles, activity history, and segment statistics. It enables users to query detailed performance metrics and explore geographic segment data through natural language commands.
    8
    182 npm
    MIT
  • A
    license
    B
    quality
    D
    maintenance
    Enables users to interact with their Strava data through natural language to analyze workouts, track fitness progress, and explore routes. It supports retrieving detailed activity stats, heart rate data, and segment insights directly within AI assistants.
    26
    312 npm
    MIT