Skip to main content
Glama

Zone Two

Your training log, in conversation. An MCP server that reads your own Garmin Connect data — sleep, HRV, training readiness, activities, body composition — so Claude can answer questions about it directly.

zonetwo.vercel.app · not affiliated with Garmin.

Authentication, and the caveat

Garmin's official Health API is only issued under a commercial partner agreement, so this uses the same private endpoints the Garmin Connect app does, via garminconnect. That works well for your own account and is what every Garmin integration of this kind does, but it is unsupported by Garmin: endpoints can change without notice, and aggressive polling can get an account rate-limited.

You log in once, interactively. Credentials are exchanged for OAuth tokens that are cached on disk and last about a year; the server itself never sees a password and never logs in. That split is deliberate — MFA prompts cannot be answered over an MCP stdio connection, so a server that tried to log in would just hang.

Related MCP server: GC-MCP

Install (Claude Desktop)

Download zonetwo.mcpb from zonetwo.vercel.app or the latest release, and open it. Claude Desktop shows an install dialog asking for your Garmin email and password. Nothing else is needed — no Python, no terminal, no config files.

If your account uses two-step verification, Garmin emails you a code the first time Claude reads your data. Paste it into the conversation and Claude will finish signing in. The code lasts 30 minutes, and the saved token then keeps you signed in for about a year.

Your password is used once, to obtain that token. It is stored by Claude Desktop's own configuration, not by this server, and is only re-used if the token is ever rejected.

Setup (from source)

Requires uv and Python 3.11+.

git clone https://github.com/alchemist-s/zonetwo.git
cd zonetwo
uv sync
uv run zonetwo login     # prompts for email, password, and MFA code if enabled
uv run zonetwo status    # confirms the tokens work

login must run in a real terminal — it prompts for a password and, if your account has MFA, a code. It cannot run through a non-interactive shell (such as Claude Code's ! prefix), and says so rather than failing with a traceback.

With MFA off, credentials can come from the environment instead:

GARMIN_EMAIL=you@example.com GARMIN_PASSWORD=... uv run zonetwo login

Garmin rate-limits login attempts by IP and can answer 429 on the first strategy the client tries. The library falls back across several; if the whole attempt fails that way, wait a few minutes rather than retrying immediately.

Tokens land in ~/.garminconnect (override with GARMINTOKENS), written 0600 in a 0700 directory.

Wire it into Claude Code

claude mcp add zonetwo --scope user -- uv --directory /absolute/path/to/zonetwo run zonetwo

Use an absolute path — the stored config does not expand ~. --scope user makes the server available in every project; without it the default local scope binds it to whichever directory you ran the command from.

Wire it into Claude Desktop

In ~/Library/Application Support/Claude/claude_desktop_config.json:

{
  "mcpServers": {
    "garmin": {
      "command": "uv",
      "args": ["--directory", "/absolute/path/to/zonetwo", "run", "zonetwo"]
    }
  }
}

If uv is not on the launcher's PATH, use its absolute path — which uv.

Tools

Dates accept YYYY-MM-DD, today, yesterday, or -7d. Omitting a date means today; omitting a range means the last seven days.

Tool

What it returns

garmin_training_history

Weekly volume, longest run and pace over N weeks — the input for a plan

garmin_create_workout

Build a structured session (intervals, pace targets) and optionally schedule it

garmin_schedule_workout

Put an existing workout on a date

garmin_scheduled_workouts

What's on the calendar this month

garmin_unschedule_workout

Take a session off the calendar, keeping the workout

garmin_delete_workout

Delete a saved workout (plans only — never recorded activities)

garmin_list_workouts

Workouts saved on the account

garmin_export_activities

Activities in a date range as CSV

garmin_briefing

Sleep, HRV, stress, readiness, Body Battery and daily totals in one call

garmin_whoami

Account name and unit preferences

garmin_devices

Registered devices and last sync

garmin_daily_summary

Steps, calories, floors, stress, body battery for a day

garmin_sleep

Stages, sleep score, overnight SpO2 and HRV

garmin_heart_rate

Resting/min/max plus daily average

garmin_hrv

Overnight HRV against your baseline

garmin_stress

Time spent in each stress band

garmin_body_battery

Charge and drain per day over a range

garmin_steps

Daily totals, or 15-minute buckets via intraday_date

garmin_spo2 / garmin_respiration

Pulse ox and breathing rate

garmin_intensity_minutes

Moderate and vigorous minutes

garmin_training_readiness

Readiness score and its contributing factors

garmin_training_status

Status, acute/chronic load balance, VO2 max

garmin_vo2max

VO2 max, fitness age, heat/altitude acclimation

garmin_race_predictions

Predicted 5K / 10K / half / marathon times

garmin_personal_records

PRs by activity type

garmin_activities

Recent activities, newest first

garmin_activities_by_date

Activities in a date range

garmin_last_activity

The most recent activity

garmin_activity

One activity in detail (raw=true for chart samples)

garmin_activity_splits

Per-lap splits

garmin_activity_weather

Weather during an activity

garmin_weight

Weigh-ins and body composition

garmin_api_get

Any Garmin API path, for what the above do not cover

Units travel in the field name

Garmin reports distances in metres, durations in seconds, speeds in metres per second and body mass in grams, all unlabelled. A bare "weight": 74500 reads as kilograms to anything summarising it, and "averageSpeed": 2.75 reads as km/h. So fields are renamed to carry their unit — distanceMeters, durationSeconds, averageSpeedMetersPerSecond, weightKg — and grams are converted, because that one is wrong rather than merely ambiguous. Foot-based activities also get a derived pace ("5:59 min/km"), which is the number a runner actually reads.

Rate limits

Garmin answers 429 rather than queuing, per IP. Calls are capped at three in flight and retried with exponential backoff and jitter; the semaphore is released before sleeping, so one throttled call does not stall unrelated ones. A 429 that survives the retry budget reaches the caller as a message saying to wait, not as a generic failure.

Why the responses are trimmed

Garmin's payloads are built for a dashboard: one night of sleep carries thousands of per-minute samples, and a 20-activity list runs past 100 kB of mostly-null fields. Every tool returns a compact projection and keeps the raw payload behind raw=true. Anything still over ~60 kB comes back as an explicit truncated marker rather than cut-off JSON — truncated JSON reads as complete data that happens to end early, and gets summarised as fact.

What you get depends on your watch

Nothing here is tied to a particular account — log in with yours and it works. But Garmin computes different metrics on different hardware, and the API returns an empty response rather than an error for the ones your device does not support. On a Venu 3, for instance, garmin_training_readiness returns a literal [] and garmin_training_status comes back with every field null: those are Forerunner/Fenix features. zonetwo check reports that as none rather than ok, so you can tell "my watch doesn't do this" from "this is broken".

What it may change

Two kinds of write, with different rules, because they carry different risk.

Plans — creating, scheduling and unscheduling workouts — are always available. They are additive, reversible, and the reason the server exists. Deleting a saved workout is allowed for the same reason: a plan is a draft.

Records — renaming an activity, logging weight or hydration — are opt-in via ZONETWO_ENABLE_WRITES=1, because they edit history rather than intent.

Recorded activities can never be deleted. No such tool exists at any setting. Plans are drafts; history is history.

Writing workouts

garmin_create_workout takes steps a model can plausibly write, and translates them into Garmin's nested format:

{
  "name": "6 x 800m",
  "steps": [
    {"kind": "warmup", "length": "10min"},
    {"repeat": 6, "steps": [
      {"kind": "interval", "length": "800m", "pace": "4:30-4:20"},
      {"kind": "recovery", "length": "90s"}
    ]},
    {"kind": "cooldown", "length": "10min"}
  ],
  "schedule_date": "2026-09-25"
}

Lengths are distances (800m, 5km, 3mi) or times (10min, 90s, 1h); min is parsed before m, so a ten minute warmup is never ten metres. Paces are minutes per kilometre and become the speed bands Garmin expects.

Privacy

Everything runs on your own computer. Requests go straight from your machine to Garmin: there is no server operated by the author, no telemetry, no analytics and no error reporting. (The project website counts page views with Vercel Web Analytics; the server you install reports nothing.)

Your password is used once, to obtain an access token. The token is saved at ~/.garminconnect (mode 0600 inside a 0700 directory) and the password is not used again unless that token is rejected — after about a year, or if you change your Garmin password. The password is never written to disk by this server, never logged, and goes nowhere except Garmin's own sign-in service.

Your health data is fetched on demand and never cached. Note that anything you ask about becomes part of your Claude conversation, which Anthropic processes under their privacy policy.

Scope. Read-only unless GARMIN_MCP_ENABLE_WRITES=1 (or the equivalent install option). Even then it can only rename an activity and record weight or hydration — no tool for deleting anything exists in this server.

Removing it. Uninstall the extension, delete ~/.garminconnect, or change your Garmin password to invalidate the token from Garmin's side.

Not intended for use by anyone under 16.

Troubleshooting

  • "No Garmin tokens…" — run zonetwo login.

  • Tokens rejected — they expire after about a year, and a password change invalidates them. zonetwo login --force.

  • Rate-limited — Garmin throttles per account. Wait a few minutes; avoid looping over long date ranges a day at a time.

  • Today's numbers look wrong — Garmin only has what the watch last synced to your phone.

  • China accounts — zonetwo login --china.

Testing

Four levels, cheapest first.

1. Offline suite — no account, no network.

uv run pytest

Runs against a stub shaped like real Garmin payloads. Covers date parsing, response shaping, every tool's projection logic, and the error paths.

2. Live smoke test — one command, hits your real account.

uv run zonetwo login     # once
uv run zonetwo check

check calls every read-only tool through the real MCP dispatch path and prints one line each:

  garmin_sleep                 ok    {"date": "2026-09-20", "sleepTimeSeconds": 27000, …}
  garmin_hrv                   none  (no data for this date)
  garmin_activity_weather      FAIL  Garmin returned 500

23/24 tools responded (1 with no data), 0 failed.

none is not a failure — it means your device does not record that metric, or has not synced it for that date. Use --date 2026-09-19 to test against a day that has definitely synced; today is often partial.

3. MCP Inspector — poke individual tools in a browser.

uv run mcp dev src/zonetwo/server.py:mcp --with-editable .

Opens a UI where you can list tools, read their schemas, and call them with your own arguments. Needs npx. The --with-editable . matters: the Inspector runs the server in its own environment, which otherwise lacks garminconnect.

4. End to end in Claude Code.

claude mcp add zonetwo --scope user -- uv --directory /absolute/path/to/zonetwo run zonetwo
claude mcp list          # should show garmin as connected

Then ask something that needs real data — "how did I sleep last night?", "what was my longest run this month?", "is my HRV trending down?" — and check the numbers against the Garmin Connect app.

To test the write tools, add -e GARMIN_MCP_ENABLE_WRITES=1 to the claude mcp add command.

Available Tools

35 tools
garmin_activitiesRecent activitiesB
Read-only

Most recent activities, newest first. activity_type filters e.g. running, cycling.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
offsetNo
activity_typeNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The readOnlyHint annotation already establishes that this is a safe read operation, so the description does not need to restate that. It does add useful ordering context ('newest first') and filter behavior, but it does not mention pagination handling or what happens when no activities are found, so the added transparency is moderate.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is very short, front-loads the main purpose, and contains no filler words. It could be slightly more structured by starting with an explicit verb, but every included phrase earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is simple, has an output schema, and all parameters are optional, so the description does not need to explain return values. However, it omits any differentiation from sibling activity tools and provides no indication of default pagination or how to combine filters, leaving noticeable gaps for an agent trying to choose the right call.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description carries the burden of explaining parameters. It only explains activity_type with examples ('running, cycling'); limit and offset are left entirely to their names and defaults. This is only partial compensation for the missing schema descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the resource ('activities') and the scope ('most recent, newest first'), so an agent can tell it is a listing operation. It lacks an explicit verb like 'list' or 'get', and it does not distinguish itself from sibling tools such as garmin_last_activity or garmin_activities_by_date, so it misses the top score.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance about when to use this tool versus closely related siblings like garmin_activity, garmin_last_activity, or garmin_activities_by_date. The description only tells what the tool does; it does not provide exclusions, prerequisites, or selection criteria.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

garmin_activities_by_dateActivities in a date rangeB
Read-only

Activities within a date range (default: last 7 days), oldest first.

ParametersJSON Schema
NameRequiredDescriptionDefault
endNo
startNo
activity_typeNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already mark this as read-only, so the description's main added value is the default window and ordering ('oldest first'). It does not address pagination, limits, or what happens when only one bound is provided, but given the readOnlyHint and simple nature, this is acceptable.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One sentence, front-loaded with the resource and scope, no filler. Every word earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

While an output schema exists and annotations cover safety, the description omits critical invocation details (date string format, meaning of null params, activity_type values). For a tool with 0% schema description coverage, this is too sparse to enable confident correct usage without external assumptions.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description carries the burden of explaining `start`, `end`, and `activity_type`. It only vaguely references a 'date range' and a default, leaving date formats, inclusivity, and activity-type values undocumented. The agent is not told how to format `start`/`end` strings or what `activity_type` accepts.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly identifies the resource ('Activities') and the filtering dimension ('date range'), and adds ordering information. However, it does not differentiate from siblings like garmin_activities or garmin_last_activity, so an agent might have to inspect those to choose correctly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage for retrieving activities within a specified date range and notes the default (last 7 days), giving context. But it does not explicitly state when to use this over sibling tools, nor any exclusions such as 'for a single activity use garmin_activity'.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

garmin_activityActivity detailA
Read-only

One activity in detail. raw=true returns the full payload including chart samples.

ParametersJSON Schema
NameRequiredDescriptionDefault
rawNo
activity_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description adds useful behavioral context beyond the readOnlyHint and openWorldHint annotations by explaining that raw=true returns the full payload including chart samples. This helps the agent understand the tool's optional behavior without contradicting the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence that immediately states the core purpose and then adds the key raw=true behavior. Every word earns its place; no filler or redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a two-parameter read-only tool with an output schema, the description covers the essential call details: what the tool returns and the optional raw behavior. The main gap is usage differentiation, which is already penalized under usage guidelines.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It explains the raw parameter's effect, but it does not clarify the activity_id parameter beyond its self-evident name and title. Partial compensation for one of the two parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool returns 'one activity in detail,' using a specific verb and resource. It distinguishes the singular focus from the sibling list tools, though it does not explicitly name an alternative.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus siblings like garmin_last_activity, garmin_activity_splits, or garmin_activities. The phrase 'one activity in detail' implies a use case, but no explicit context or exclusions are provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

garmin_activity_splitsActivity splitsC
Read-only

Per-lap splits for an activity.

ParametersJSON Schema
NameRequiredDescriptionDefault
activity_idYes

TDQS

C2.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=true, covering the safety profile. The description adds that the output is per-lap splits, which is useful context, but it doesn't describe the response format, potential size, or any special behavior like pagination. Given the annotations, this is acceptable but not rich.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, clear sentence with no redundancy. The core information is front-loaded, and there is no wasted text.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with one parameter and no output schema, the description gives a basic idea of the response (per-lap splits) but doesn't elaborate on what fields are included or any prerequisites. It is sufficient for a simple tool but could benefit from a hint about the data structure.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema has 0% description coverage, so the description must compensate for the activity_id parameter. It does mention that the splits are for an activity, implying activity_id is the identifier, but it doesn't provide format, source, or requirements beyond the name. This is minimal compensation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states that this tool provides per-lap splits for an activity, clearly distinguishing it from siblings like garmin_activity (full activity details) and garmin_activities (list). It specifies the resource and granularity, though it lacks an explicit verb like 'retrieve' or 'get'. The purpose is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives. It does not mention that this is for getting split data specifically, nor does it compare with similar tools like garmin_activity or garmin_last_activity. An agent would have to infer when splits are needed.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

garmin_activity_weatherActivity weatherC
Read-only

Weather recorded during an activity.

ParametersJSON Schema
NameRequiredDescriptionDefault
activity_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=true, covering the safety profile and potential for unexpected fields. The description adds minimal behavioral context ('weather recorded during an activity') but does not disclose edge cases such as missing weather data or how it handles activities without weather. It does not contradict annotations, so a 3 is appropriate.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence with no filler. It communicates the core purpose efficiently. While it could be slightly more informative, it is appropriately concise for a simple tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With only one parameter and an existing output schema, the tool is relatively simple. However, the description lacks essential context: it does not explain what activity_id refers to, how to obtain it, or when to use this tool. The openWorldHint annotation suggests unknown return fields, but the description offers no guidance on expected data shape or usage scenarios.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, meaning the schema provides no explanation for activity_id beyond its title 'Activity Id'. The description fails to compensate: it never mentions the parameter or clarifies how to obtain a valid activity_id. The agent is left with no added meaning beyond the bare schema field.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: it returns weather data recorded during an activity. This is specific and distinguishes it from siblings like garmin_activity (activity details) and garmin_activity_splits (split data). It is not a tautology, though it could be more explicit by adding a verb like 'retrieves'.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is given on when to use this tool versus alternatives. It does not mention prerequisites (e.g., needing an activity_id from a prior call), nor does it exclude cases where weather might be unavailable. The usage context is entirely implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

garmin_api_getDirect Garmin API requestA
Read-only

Call any Garmin Connect API path directly (read-only GET).

For endpoints the dedicated tools do not cover. Paths look like '/usersummary-service/usersummary/daily/{displayName}?calendarDate=2026-09-19'. Prefer a dedicated tool when one exists — this returns unshaped payloads.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathYes

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint and openWorldHint, and the description adds value by specifying the method is GET, that it returns unshaped payloads, and that it is a direct raw API call. This goes beyond the structured hints and sets accurate expectations about the output format.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact, front-loaded with the core behavior, and every sentence earns its place: scope, path format example, and the prefer-dedicated-tool guidance. There is no filler or repetition of schema details.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For an open-ended escape-hatch tool with a single parameter and no output schema, this is nearly complete: it explains when to use it, what it returns (unshaped payloads), and gives a realistic path example. The only minor gap is not detailing exactly whether the full base URL must be included, but the example implies the path starts at the service root.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With only one parameter (path) and 0% schema description coverage, the description compensates by showing a concrete path example with placeholder syntax and query string. It clearly conveys that the tool expects an arbitrary API path beginning with a service segment, even though it does not exhaustively document path construction rules.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'Call any Garmin Connect API path directly (read-only GET)'. It clearly differentiates itself from the many dedicated sibling tools by explicitly saying it covers endpoints those tools do not, and mentions the alternative behavior of preferring a dedicated tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives explicit when-to-use guidance: for endpoints the dedicated tools do not cover. It also tells the agent to prefer a dedicated tool when one exists, and restricts usage to read-only GET operations, which is a clear exclusion.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

garmin_body_batteryBody BatteryB
Read-only

Body Battery charge and drain per day over a date range (default: last 7 days).

ParametersJSON Schema
NameRequiredDescriptionDefault
endNo
startNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=true and openWorldHint=true, so the agent knows it's a safe read. The description adds that it returns charge and drain per day and the default date range, which is helpful but doesn't cover behaviors like pagination, rate limits, or error conditions. Given the annotations, a score of 3 is appropriate.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence with no filler. It front-loads the core purpose and appends the date-range default, which is efficient and easy to parse.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has only two parameters and an output schema, so the missing piece is parameter semantics. The description says 'over a date range' but doesn't explain how to specify start/end, what null means, or the date format. Since schema coverage is 0%, this is a significant gap that leaves the agent guessing on custom ranges.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must explain the parameters. It mentions a date range and a default of last 7 days, implying start and end control the range, but it doesn't specify date formats, what null means, or how the default is applied. This is insufficient compensation for the lack of schema descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly identifies the resource (Body Battery) and the type of data (charge and drain per day). It distinguishes from sibling tools by naming the specific metric, though it lacks an explicit verb like 'retrieve' or 'get'. The default date range adds clarity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus the many sibling health-data tools (e.g., garmin_heart_rate, garmin_stress). It only states what it does and the default range, without exclusions or alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

garmin_briefingMorning briefingA
Read-only

Morning snapshot in one call: sleep, HRV, Body Battery, readiness, stress and RHR.

Fetches each section concurrently so answering "how am I doing today?" costs
one round trip instead of six. A section that fails is named in
``sectionsUnavailable`` rather than failing the whole briefing, and a
section the device does not record simply reports its own note.

Returns the numbers only — no training advice. Read readiness, HRV and Body
Battery together rather than any one alone.
ParametersJSON Schema
NameRequiredDescriptionDefault
dateNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=true, so the description need not restate that. It adds value by disclosing concurrent fetching, per-section error handling (sectionsUnavailable), and behavior for unrecorded sections. It also clarifies that it returns numbers only, which is a behavioral trait. No contradiction with annotations; the description complements them well.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is three short paragraphs. The first sentence immediately states the purpose and scope. The second paragraph explains concurrency and error handling, the third clarifies return content and advice. It is front-loaded and each sentence earns its place, though the lack of parameter explanation is a structural omission that prevents a 5.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For an aggregate read tool with an output schema, the description covers the main behavior: what it returns (numbers only), concurrency, and error handling. However, it omits any explanation of the optional 'date' parameter, which is essential for correct invocation. Since the parameter is optional and has a default, an agent might infer it defaults to today, but this is not explicit. Overall, the description is adequate for the primary use but incomplete regarding parameters.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has a single optional 'date' parameter with 0% schema description coverage. The description does not mention the date parameter at all, leaving its meaning and format entirely to the schema (which only specifies type string/null and default null). An agent has no guidance on how to format the date or what it controls (presumably which day's briefing). This is a significant gap given low schema coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: 'Morning snapshot in one call' listing the specific metrics (sleep, HRV, Body Battery, readiness, stress, RHR). It distinguishes itself from siblings by aggregating six data types into a single call, and explicitly notes it returns numbers only, no training advice. This is a precise, non-tautological statement.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explains when to use it: for a quick 'how am I doing today?' question, and highlights the efficiency benefit of one round trip vs six. It also advises reading readiness, HRV, and Body Battery together, which is domain guidance. It does not explicitly name alternative tools for individual metrics, but the context implies that separate tools exist for each metric. This is clear usage context, though not exhaustive.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

garmin_create_workoutCreate a workoutA

Create a structured running workout on Garmin, optionally scheduling it.

Steps are plain objects, in order. Each needs a ``kind`` (warmup, run,
interval, recovery, rest, cooldown) and a ``length`` — a distance like
"800m", "5km" or a time like "10min", "90s". Optionally add ``pace``
("4:30" or a range "4:30-4:20", minutes per kilometre) or
``heart_rate_zone`` (1-5), and a ``note``.

Repeats nest: ``{"repeat": 6, "steps": [{"kind": "interval", "length":
"800m", "pace": "4:30-4:20"}, {"kind": "recovery", "length": "90s"}]}``.

Once scheduled it appears on the watch on that date.
ParametersJSON Schema
NameRequiredDescriptionDefault
nameYes
stepsYes
descriptionNo
schedule_dateNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description substantially exceeds the minimal annotations by disclosing the execution model: ordered step objects, accepted kind/length/pace formats, heart-rate zones, repeat nesting, and the fact that scheduling makes the workout appear on the watch. This is meaningful behavioral context and is consistent with readOnlyHint=false.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description opens with the core action and then uses compact, concrete examples to define the nested step structure. It is detailed but every sentence earns its place, and the examples are immediately interpretable.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need no explanation, and the description covers all non-obvious invocation details: step kinds, length units, pace formats, heart-rate zones, repeats, and scheduling side effects. An agent can construct a valid call without additional clarification.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description carries the burden for parameter meaning. It thoroughly explains steps (kind, length, pace, heart_rate_zone, note, repeat) and schedule_date via the watch behavior, but name and description are left to their schema titles rather than explicitly described.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb (Create), a resource (structured running workout on Garmin), and an optional behavior (scheduling). It clearly distinguishes this tool from siblings like list_workouts and delete_workout by describing what the workout contains.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The primary use case is clear: create a structured running workout and optionally schedule it. However, the description never names alternatives such as garmin_schedule_workout or states when not to use this tool, so the distinction from scheduling an existing workout is only implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

garmin_daily_summaryDaily summaryB
Read-only

Whole-day wellness rollup: steps, calories, distance, floors, stress, body battery.

ParametersJSON Schema
NameRequiredDescriptionDefault
dateNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=true, so the safety profile is covered. The description adds the list of wellness metrics included, which is useful, but it does not disclose behavioral details such as how the optional date is interpreted, what happens with no date, or timezone handling.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single concise sentence that front-loads the core concept ('whole-day wellness rollup') and then lists included metrics without wasted words. It is appropriately brief for a simple tool, though it sacrifices parameter clarity for brevity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is low-complexity with one optional parameter and an output schema, so the description need not explain return values. However, it leaves the date parameter completely unexplained, including accepted format and what a null default means, which is a notable gap for the only input an agent must provide.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate by explaining the date parameter, but it never mentions date, its format, or its default behavior. The parameter name and default null provide minimal meaning, but the description adds no semantic value for the only input.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description identifies a specific resource ('whole-day wellness rollup') and lists the included metrics (steps, calories, distance, floors, stress, body battery), making the tool's purpose reasonably clear. It does not explicitly contrast itself with sibling tools like garmin_steps or garmin_stress, but the 'rollup' framing distinguishes it as an aggregate view.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies this tool is for a daily aggregate wellness overview, which suggests when an agent might prefer it over metric-specific siblings. However, it does not explicitly state when not to use it or name alternatives for single-metric queries, leaving usage selection mostly to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

garmin_delete_workoutDelete a saved workoutA
Destructive

Delete a saved workout. Only affects plans — recorded activities cannot be deleted.

ParametersJSON Schema
NameRequiredDescriptionDefault
workout_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true and readOnlyHint=false, covering the destructive nature. The description adds the specific behavioral scoping to plans vs. activities, which is valuable context beyond the annotations. No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is only two sentences, front-loaded with the action and key constraint. Every sentence adds necessary information, with no redundant phrasing.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple delete tool with one parameter and an output schema, the description covers the main behavioral nuance (plans only). However, it does not clarify how to obtain the workout_id or differentiate plan IDs from activity IDs, which may require cross-referencing sibling tools like garmin_list_workouts. This is a minor gap, keeping it from a perfect score.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 0% schema description coverage, the description should compensate by explaining the parameter. It does not mention workout_id at all, leaving the agent to infer its meaning from the tool name and title. The parameter name is self-explanatory, but the description fails to explicitly state that workout_id refers to a plan ID, not an activity ID, which is a notable gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Delete') and resource ('saved workout'), and clarifies scope ('Only affects plans — recorded activities cannot be deleted'). This clearly differentiates from related siblings like garmin_unschedule_workout (unscheduling) and activity-related tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly states the tool only affects plans, not recorded activities, giving a clear boundary on when to use it. While it doesn't name alternatives or exclusion conditions, the distinction between plans and activities is the key usage constraint and is well-communicated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

garmin_devicesYour devicesA
Read-only

List the Garmin devices registered to the account, newest sync first.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The readOnlyHint and openWorldHint annotations already establish the safe read-only nature. The description adds useful behavioral context by specifying that results are account-scoped and ordered by most recent sync. With an output schema present, no additional return-format disclosure is needed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, front-loaded sentence conveys the action, resource, scope, and ordering with no filler. The title is slightly redundant but the description itself is optimally concise.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter read-only listing tool with an output schema, the description fully covers what an agent needs: what is listed, whose devices, and the order. No additional prerequisites, auth notes, or side-effect warnings are relevant.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has zero parameters, so there is nothing for the description to clarify about arguments. The baseline of 4 applies; the description appropriately focuses on the operation and result ordering instead.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('List'), a clear resource ('Garmin devices'), and the account scope, plus an ordering detail ('newest sync first'). This clearly distinguishes it from the many sibling tools focused on activities, metrics, and workouts.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description makes the use case plain: retrieve the account's registered Garmin devices. It does not explicitly name alternatives or exclusions, but the tool's scope is clear enough that an agent can infer when to select it over the metric-focused siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

garmin_export_activitiesExport activities as CSVA
Read-only

Export activities in a date range as CSV — the export Garmin makes awkward.

One header row then one row per activity, newest last. Narrow the range if the response comes back truncated.

ParametersJSON Schema
NameRequiredDescriptionDefault
endNo
startNo
activity_typeNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description adds behavioral detail beyond the read-only and open-world annotations: one header row, one row per activity, newest-last ordering, and possible truncation. It does not contradict the annotations and usefully warns the agent to narrow the range on truncation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two compact sentences front-load the core operation and then add format and truncation behavior without padding. Every clause carries useful information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present and read-only/open-world annotations, the description covers the CSV shape and truncation behavior. The main gap is parameter value semantics, especially activity_type and date format, which weakens end-to-end callability.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema descriptions are absent, so the description must compensate. It communicates that start/end define a date range, but it never explains the expected date format or what activity_type accepts, leaving the parameters underspecified.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names the exact operation ('Export activities in a date range as CSV'), making the resource and output format unambiguous. The CSV format distinguishes it from the activity tools, but it does not explicitly name alternatives like garmin_activities or garmin_activities_by_date.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No explicit when-to-use or alternative-routing guidance is given; the date-range and truncation tip imply the intended use but don't exclude sibling tools. A sentence pointing to raw-activity tools for non-CSV data would make it more actionable.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

garmin_heart_rateHeart rateA
Read-only

Heart rate for a day. Summary gives min/max/resting; raw adds the 2-minute series.

ParametersJSON Schema
NameRequiredDescriptionDefault
rawNo
dateNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Description adds useful behavior beyond annotations by explaining that summary returns min/max/resting and raw adds the 2-minute series. Annotations already cover readOnly and openWorld hints, so the safety profile is complete. No edge cases, availability caveats, or rate-limit issues are mentioned, but these are less critical given the read-only nature.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences with no filler or redundancy. The core information—day scope, summary mode, and raw mode—is front-loaded and immediately actionable.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only, single-day metric with an output schema, the description covers the essential facets: time scope, data granularity, and mode toggle. It omits date formatting and default behavior, but these are minor given the minimal parameter count. Overall it provides enough for a competent agent to invoke the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It clarifies the 'raw' boolean by defining what it adds, and it implies 'date' through 'for a day.' However, it does not specify the expected date format or the behavior of the null default, leaving some ambiguity.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States the resource clearly as 'heart rate for a day' and distinguishes between summary and raw output. It lacks an explicit verb like 'Get' or 'List,' which slightly weakens the action clarity. The topic is clearly distinct from sibling Garmin metric tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides context for a single day and explains the two modes (summary vs raw), which helps select between them. Does not explicitly state when to use this tool over sibling tools, though the metric's specificity makes the choice fairly unambiguous. No exclusionary guidance or alternative tool names are given.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

garmin_hrvHeart rate variabilityB
Read-only

Overnight heart rate variability: last-night average, baseline and status.

ParametersJSON Schema
NameRequiredDescriptionDefault
dateNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint and openWorldHint, so the safety profile is established. The description adds some context by scoping to overnight data and naming the three output dimensions, but it does not disclose behavioral traits such as data availability limitations, how 'status' is derived, or how missing nights are handled. This is acceptable given annotation coverage but not rich.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence with no filler. Every word earns its place by conveying the data scope and the specific metrics, making it easy for an agent to parse quickly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the low complexity (one optional parameter), the annotations, and the presence of an output schema, the description is reasonably positioned but not fully complete. The main gap is the undefined date parameter and the absence of any usage guidance, which are not compensated by the existing structured metadata.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the description never mentions the 'date' parameter. The phrase 'last-night' loosely implies the date refers to the night of interest, but there is no explicit explanation of what date means, the expected format, or what null/default behavior does. With a single optional parameter, the description should compensate by clarifying its semantics.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific resource and scope: overnight heart rate variability, with the relevant metrics (last-night average, baseline, status). It clearly identifies what the tool is about, though it uses a noun phrase rather than an explicit verb like 'get' or 'retrieve' and does not explicitly differentiate it from sibling tools such as garmin_heart_rate.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives, no exclusions, and no conditions. While the metric name implies a use case, the agent is given no explicit information about when to choose garmin_hrv over garmin_heart_rate, garmin_training_readiness, or garmin_stress.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

garmin_intensity_minutesIntensity minutesB
Read-only

Moderate and vigorous intensity minutes for a day, against the weekly goal.

ParametersJSON Schema
NameRequiredDescriptionDefault
dateNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already provide readOnlyHint=true and openWorldHint=true, so the safety profile is covered. The description adds useful behavioral context by clarifying that the data is measured against the weekly goal, which goes beyond the annotations without contradicting them.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single compact sentence that front-loads the core metric and includes the key goal context. Every word earns its place, and there is no filler or repetition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

This is a low-complexity read-only tool with an output schema and one optional parameter. The description covers what is returned—daily intensity minutes against the weekly goal—and the annotations cover safety. The main gap is lack of date parameter usage detail, but the overall context is otherwise sufficient.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate for the undocumented date parameter. It says 'for a day,' which hints that the date parameter selects the day, but it does not explain date format, optional/default behavior, or what null means, leaving significant ambiguity for a parameter with no schema description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description identifies a specific resource—moderate and vigorous intensity minutes—and scopes it to a day against the weekly goal, which distinguishes it from sibling tools like garmin_heart_rate or garmin_steps. It lacks an explicit verb but is not a tautology and tells an agent what data this tool provides.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is given about when to use this tool versus alternatives such as garmin_daily_summary or garmin_activities_by_date. The description implies it is for intensity minutes for a single day, but it never states exclusion criteria or points to a sibling for related use cases.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

garmin_last_activityMost recent activityD
Read-only

The most recently recorded activity.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

D1.7/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=true. The description adds no behavioral context beyond a restatement of the name—no mention of data format, limitations, or any quirk. Since it adds zero value over annotations, it scores at the floor.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely concise—a single short phrase—but it under-delivers on content. While there is no word waste, the phrase is so generic that it fails to structure any meaningful information. It is not a case of efficient brevity but rather of under-specification.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Although the tool has an output schema and annotations cover read-only/open-world, the description says almost nothing about the scope of the returned data. It does not clarify whether this returns a full activity object or a summary, nor does it hint at any filters. For a simple tool, this may be marginally acceptable, but it remains too vague to be considered complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so the baseline for this dimension is 4 per the rubric. There is nothing that the description could add to parameter semantics; the schema already covers everything (100% coverage). The absence of parameter info is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose1/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'The most recently recorded activity' is a near-tautology of the tool name 'garmin_last_activity'. It does not specify a distinct verb, resource scope, or differentiate among siblings like 'garmin_activity' or 'garmin_activities'. An agent cannot infer what makes this tool unique beyond the name.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives such as 'garmin_activities' (list all) or 'garmin_activity' (a specific activity). There is no mention of context, exclusions, or preferred scenarios.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

garmin_list_workoutsSaved workoutsA
Read-only

List the workouts saved on the Garmin account, newest first.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The annotations already communicate read-only safety and open-world behavior. The description adds one behavioral trait — newest-first ordering — but discloses nothing about pagination, rate limits, or how broad the 'saved workouts' scope is. This is a reasonable baseline with annotations present.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One short sentence, front-loaded with the action and resource, and it includes a useful ordering detail. There is no wasted text or repetition of schema fields.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple one-parameter read-only list tool, the description covers the core operation and ordering. The output schema handles return-value details, and annotations cover safety. The only minor gap is not explaining how the optional limit affects result count or pagination.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema exposes one optional 'limit' parameter with a default of 25, but the description never mentions it. With 0% schema description coverage, the description was expected to compensate, and it does not. The parameter name is fairly self-explanatory, which prevents a score of 1, but the gap remains.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb ('List'), a concrete resource ('workouts saved on the Garmin account'), and an ordering guarantee ('newest first'). The qualifying word 'saved' distinguishes this from sibling tools like garmin_scheduled_workouts or garmin_create_workout without requiring the agent to open schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The intended use is implied: list saved workouts rather than create, schedule, or delete them. However, the description does not explicitly name alternatives, exclusions, or conditions for when to prefer this tool over garmin_scheduled_workouts or garmin_activities. This is adequate but not explicitly guiding.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

garmin_personal_recordsPersonal recordsB
Read-only

Personal records across activity types.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=true. The description adds a scoping detail ('across activity types') but does not disclose any other behavioral aspects such as return format, rate limits, or data freshness. This adds a small amount of context beyond annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence with no wasted words. It is front-loaded with the core subject and is appropriately sized for a parameterless tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple tool with no parameters, the description is mostly complete in stating what it does. However, it lacks any description of the output structure or typical usage, and it does not differentiate from similar tools in the sibling set, leaving the agent to infer when to use it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, so the input schema fully covers them. Baseline of 4 applies because there is nothing to add. The description's mention of 'across activity types' hints at the scope of data but not parameter syntax, which is irrelevant here.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description identifies the resource (personal records) and scope (across activity types), which clearly distinguishes it from sibling tools like garmin_activities or garmin_activity. However, it lacks an explicit verb (e.g., 'retrieve' or 'list'), so it's clear but not maximally specific.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives. It does not mention alternatives, prerequisites, or scenarios where another tool would be more appropriate.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

garmin_race_predictionsRace time predictionsA
Read-only

Predicted race times for 5K, 10K, half and full marathon.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=true, covering the safety profile. The description adds the specific race distances returned, which is useful context, but it does not disclose behavioral traits such as whether predictions are model estimates from current fitness data or how they update over time. The description does not contradict the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single sentence that front-loads the resource and enumerates the full set of distances. Every word earns its place, with no filler or redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 0-parameter read-only tool with an output schema present, this description is nearly complete. It could clarify whether all four distances return in a single call, but since there are no parameters, that ambiguity is minor and the output schema carries the return-value detail.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters with 100% schema coverage, so the baseline is 4 and there are no inputs to document. The description appropriately focuses on what the call returns rather than inputs.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description identifies the resource (predicted race times) and enumerates the exact scope (5K, 10K, half, full marathon). The word 'predicted' implicitly distinguishes it from garmin_personal_records, though no explicit verb is used and sibling differentiation is not stated.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus closely related alternatives such as garmin_personal_records (actual achievements), garmin_vo2max, or garmin_training_status. With 35 siblings, the absence of any differentiation or context leaves the agent to infer usage from the name alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

garmin_respirationBreathing rateB
Read-only

Breathing rate for a day: waking, sleeping, highest and lowest breaths per minute.

ParametersJSON Schema
NameRequiredDescriptionDefault
dateNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=true, so the agent knows this is a safe read operation. The description adds that the tool returns waking, sleeping, highest, and lowest values, but it does not explain how an omitted or null date behaves or any data-availability caveats. No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single front-loaded sentence with no filler. It states the resource and scope first, then compactly lists the returned metrics, so every word earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple read-only tool with an output schema and one optional parameter, the description adequately explains that it returns daily respiration metrics. The main gap is the underspecified date parameter semantics, which matters because the schema itself provides no description for it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate for the undocumented date parameter. Saying 'for a day' implies date selects the day, but it does not specify accepted date formats, the meaning of null, or what happens when the parameter is omitted. This is insufficient for the tool's only parameter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly identifies the resource (breathing rate) and the temporal scope (a day), and enumerates the specific measures returned (waking, sleeping, highest, lowest breaths per minute). It lacks an explicit verb like 'retrieves' and does not differentiate from siblings, but the intent is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The daily scope and the enumerated metrics imply this tool is for per-day respiration summaries, and there is no dedicated respiration sibling among the listed tools. However, there is no explicit guidance on when to use it versus alternatives such as garmin_daily_summary or garmin_heart_rate, so usage is only implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

garmin_scheduled_workoutsScheduled workoutsB
Read-only

List workouts scheduled in a given month (defaults to the current one).

ParametersJSON Schema
NameRequiredDescriptionDefault
yearNo
monthNo

TDQS

B3.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The annotations already declare readOnlyHint=true and openWorldHint=true, so the safety profile is covered. The description adds the default-to-current-month behavior and monthly scoping, but it does not disclose output shape, pagination, timezone handling, or how scheduled workouts differ from ordinary workouts.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence that states the action, the scope, and the default. Every word earns its place and there is no redundant detail.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple read-only list with two optional parameters, the description is minimally viable. Still, it lacks any relationship guidance to garmin_list_workouts and does not mention what a scheduled workout entry looks like in the response.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description carries the burden of explaining the two parameters. It clarifies that the month defaults to the current one, but it never explains the year parameter, valid ranges, or how null year/month combinations behave.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a concrete operation ('List workouts') and a bounded resource ('scheduled in a given month'), and it adds the default-month behavior. However, it does not explicitly differentiate itself from garmin_list_workouts or garmin_schedule_workout, so it stops short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance about when to use this tool versus garmin_list_workouts, garmin_schedule_workout, garmin_unschedule_workout, or other siblings. The only usage hint is the month default, which is a parameter behavior rather than an explicit selection rule.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

garmin_schedule_workoutSchedule a workoutB

Put an existing workout on the calendar for a given date.

ParametersJSON Schema
NameRequiredDescriptionDefault
dateYes
workout_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate this is a write operation that is not destructive. The description adds that it places an existing workout on a calendar, which is a helpful side-effect clarification, but it does not disclose potential overwrite behavior, timezone handling, or other scheduling nuances.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence that is front-loaded with the core action and includes the key constraint ('existing'). No filler or redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple two-parameter tool with an output schema, the description covers the basic operation. However, it omits date format details and any guidance about conflicts with existing scheduled workouts, making it only minimally complete for reliable invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description carries the burden of explaining parameters. It mentions 'existing workout' and 'given date', but does not specify the date format, how to obtain the workout_id, or what values are valid. This is insufficient for a tool whose schema provides no descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action: put an existing workout on the calendar for a date. The word 'existing' helps distinguish this from garmin_create_workout, and the scheduling verb distinguishes it from garmin_unschedule_workout. It does not name a sibling explicitly, but the scope is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies the tool is for scheduling an already-created workout on a specific date. It gives no explicit guidance about when not to use it or when to prefer related tools like create_workout, list_workouts, or scheduled_workouts.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

garmin_sleepSleepA
Read-only

Sleep for the night ending on the given date: stages, score and overnight vitals.

ParametersJSON Schema
NameRequiredDescriptionDefault
rawNo
dateNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover the read-only and open-world behavior, lowering the burden on the description. The description adds useful context about date interpretation and return contents, but it does not disclose behavior when date is omitted or what the raw flag changes. There is no contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single focused sentence with no filler. The resource, date semantics, and return categories are front-loaded, making the description efficient and scannable.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers the core resource, date scope, and return categories, and the output schema plus annotations cover return values and safety. However, the raw parameter is undocumented everywhere, and the behavior when no date is supplied is unstated, leaving a notable gap for a tool that accepts zero required parameters.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, but it only clarifies the date parameter ('given date') and leaves raw completely unexplained. An agent cannot infer what raw=true does, and the default behavior for a null date is also not addressed.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly identifies the resource (sleep) and a precise scope ('night ending on the given date'), and names returned content (stages, score, overnight vitals). It is distinguishable from sibling metric tools like garmin_heart_rate or garmin_steps, though it does not explicitly contrast with garmin_daily_summary.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'for the night ending on the given date' gives clear context for when to call this tool. It does not state exclusions or alternatives, but the resource and date semantics make the intended use understandable among the sibling Garmin tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

garmin_spo2Blood oxygenA
Read-only

Pulse oximetry for a day: average and lowest overnight SpO2.

ParametersJSON Schema
NameRequiredDescriptionDefault
dateNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=true, so the agent knows this is a safe read operation. The description adds that it provides average and lowest overnight SpO2, which is useful context. However, it doesn't disclose details like whether the date is required, what happens if no data exists for the date, or the format of the response. With annotations covering the safety profile, a 3 is appropriate.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence with no waste. It front-loads the core purpose ('Pulse oximetry for a day') and then specifies the key outputs ('average and lowest overnight SpO2'). Every word earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has a single optional parameter, an output schema, and annotations covering safety. The description is adequate for a simple read-only tool, but it doesn't clarify the meaning of the null default or the date format. Given the simplicity, this is a minor gap, but the description could be more complete by stating what null means.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. The description mentions 'for a day' which implies the date parameter, but it doesn't explain the date format, what null means, or how the date affects the returned data. The parameter is optional with a default of null, but the description doesn't clarify whether null means today or all data. This is a gap, but the single parameter is relatively self-explanatory.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'Pulse oximetry for a day' with 'average and lowest overnight SpO2.' This clearly identifies what the tool does and distinguishes it from sibling tools like garmin_heart_rate or garmin_respiration, though it doesn't explicitly name a sibling.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage for a single day's SpO2 data, and the date parameter is optional with a default of null, suggesting it may default to today. However, it doesn't explicitly state when to use this tool versus alternatives like garmin_daily_summary or garmin_heart_rate, nor does it mention any exclusions or prerequisites.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

garmin_stepsStepsA
Read-only

Daily step totals over a range, or 15-minute buckets for one day via intraday_date.

ParametersJSON Schema
NameRequiredDescriptionDefault
endNo
startNo
intraday_dateNo

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=true, so the safety profile is known. The description adds behavioral context by explaining the intraday_date parameter and the two output modes (daily totals vs. 15-minute buckets), which is not obvious from the schema alone. No contradiction exists between description and annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, information-dense sentence. It front-loads the core purpose and clearly separates the two modes with 'or'. There is no filler or redundancy. Every word contributes to understanding the tool's function.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's simplicity (no required parameters, no output schema, read-only), the description is largely sufficient. It covers the main usage patterns and the role of intraday_date. However, it omits details like date format expectations or the response structure, which might be helpful but are not critical for a basic data retrieval tool. The annotations handle safety, so the description is complete enough for typical usage.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It explains that start/end define a range for daily totals, and intraday_date selects a single day for 15-minute buckets. This gives meaning to the three parameters beyond their names and types. It doesn't specify date formats or inclusivity, but for a simple read-only tool, it covers the essential semantics.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states exactly what the tool does: returns daily step totals over a date range or 15-minute buckets for a single day via intraday_date. It uses specific verbs and resources, clearly distinguishing between the two modes. It is not a tautology and stands apart from sibling tools like heart rate or stress.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explains the two usage scenarios (range vs. intraday) but does not explicitly say when to choose this tool over alternatives among the many Garmin siblings. It implies that you use this for step data, but no exclusions or comparisons are provided. It gives enough context for the primary use case but lacks cross-tool routing guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

garmin_stressStressB
Read-only

All-day stress: average, max and time spent in each stress band.

ParametersJSON Schema
NameRequiredDescriptionDefault
dateNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, so the safe-read behavior is covered. The description adds useful context that this is a daily aggregation (average, max, band durations), but it does not disclose date-default behavior, timezone assumptions, or how missing data is handled. That is acceptable given the read-only annotation, but not rich.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single concise phrase with the key output content front-loaded. There is no filler, repetition, or unnecessary detail.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Although an output schema exists and read-only annotations are present, the only input parameter (`date`) is completely unexplained. An agent does not know whether to pass a date, what format to use, or what the default behavior is. This is a meaningful gap for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the description never mentions the `date` parameter at all. It does not explain whether null means today, what date format is expected, or how date selection affects the returned summary. The description carries none of the parameter documentation burden.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly identifies the resource (Garmin stress) and the exact data provided: average, max, and time spent in each stress band for the day. This distinguishes it from sibling tools like garmin_heart_rate or garmin_hrv by naming a specific metric and its output format.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is given for when to choose this tool over siblings or what conditions make it appropriate. The word 'stress' implies the metric, but there is no mention of alternatives, exclusions, prerequisites, or typical call context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

garmin_submit_mfa_codeFinish signing in to GarminA

Finish signing in to Garmin with the one-time code it emailed.

Only needed when a Garmin tool has just reported that a code was sent. The code is valid for 30 minutes; afterwards the saved tokens last about a year and this is not asked for again.

ParametersJSON Schema
NameRequiredDescriptionDefault
codeYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only indicate readOnlyHint=false and destructiveHint=false. The description adds meaningful behavioral context: the code is emailed, valid for 30 minutes, and successful submission leads to saved tokens lasting about a year. It doesn't detail what happens on an invalid code, but the provided auth lifecycle context is valuable.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and front-loaded: the first sentence states the core purpose, and the second supplies the trigger condition and validity details. No filler or redundant restatement of the tool name or schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-parameter auth-completion tool, the description covers the essential context: when to call, what the code is, how long it is valid, and what to expect logistically. The presence of an output schema means the description doesn't need to explain return values.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must clarify the code parameter. It does so by identifying the code as the one-time emailed code and adding the 30-minute validity constraint, which goes beyond the schema's bare 'Code' label. This is adequate for a single, simple parameter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific action ('Finish signing in to Garmin') with a specific method ('the one-time code it emailed'), making the tool's purpose immediately clear. The trigger condition ('Only needed when a Garmin tool has just reported that a code was sent') distinguishes it from the many data-retrieval sibling tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives an explicit when-to-use condition: only after a Garmin tool reports that a code was sent. It also provides the code's 30-minute validity window and notes that this step won't be needed again for about a year, which helps the agent decide whether to call this tool now or avoid it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

garmin_training_historyTraining history by weekA
Read-only

Weekly training volume over recent weeks — the input for writing a plan.

Returns one row per week (distance, sessions, longest run, average pace) plus overall totals, in a single call. Use this before designing a plan rather than fetching activities one range at a time.

ParametersJSON Schema
NameRequiredDescriptionDefault
weeksNo
activity_typeNorunning

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=true, so the safety profile is covered. The description adds useful behavioral context: it returns aggregated weekly rows plus overall totals in a single call, and it is scoped to 'recent weeks' with a default of 12 weeks. It does not disclose pagination or exact date-range behavior, but the output schema likely covers the return shape. The description adds value beyond annotations without contradicting them.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences with no filler. The first sentence states the resource and purpose, the second details the output shape, and the third gives usage guidance. Every sentence earns its place, and the key scoping information ('recent weeks', 'single call') is front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is a read-only summary with an output schema, so the description does not need to explain return values in detail. It covers the main use case, the output granularity, and the alternative approach. The only minor gap is that it does not explicitly describe the activity_type parameter's effect or allowed values, but the default and the running-specific output fields make this inferable. For a read-only aggregation tool, this is nearly complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It does partially: it explains the 'weeks' concept by saying 'recent weeks' and 'one row per week', and it implies activity_type by saying 'running' is the default and the output includes distance/sessions/longest run/average pace. However, it does not explicitly explain the activity_type parameter or its allowed values, and it does not state the range of the weeks parameter. Still, the description gives enough context to infer both parameters' roles.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Returns'), a clear resource ('weekly training volume over recent weeks'), and the exact output shape (one row per week with distance, sessions, longest run, average pace, plus overall totals). It also distinguishes itself from sibling tools by framing the output as the input for writing a plan, which separates it from garmin_activities, garmin_activities_by_date, and garmin_last_activity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly says when to use it: 'Use this before designing a plan rather than fetching activities one range at a time.' This names the alternative pattern (fetching activities range by range) and gives a clear condition for choosing this tool. It also implies the tool is a high-level summary, not a detailed activity fetch.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

garmin_training_readinessTraining readinessC
Read-only

Training readiness score for a day, with the factors that drove it.

ParametersJSON Schema
NameRequiredDescriptionDefault
dateNo

TDQS

C2.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=true, which signal a safe read operation that may return dynamic data. The description adds the detail that it returns the score 'for a day' and the 'factors that drove it,' which gives some context about the response scope. However, it does not disclose defaults (e.g., what happens when date is null), timezone handling, or behavior when no data exists, leaving gaps beyond what annotations cover.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, concise sentence with no wasted words. It front-loads the core purpose and includes a meaningful detail (factors) without redundancy. It is appropriately minimal for a simple read tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has one optional parameter with zero schema coverage and no output schema, the description is insufficient. It omits any explanation of the date parameter, default behavior, or how to interpret the score. While the tool is simple, the lack of parameter semantics and usage guidance makes it incomplete for an agent to call correctly without external knowledge.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the description makes no mention of the 'date' parameter at all. The only parameter is entirely undocumented, and the description does not compensate by explaining how date affects results, its format, or default behavior. This leaves the agent guessing about a critical input.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('score') and resource ('training readiness for a day'), and mentions it includes the driving factors. It is clear about the core purpose but does not explicitly contrast with siblings like garmin_training_status or garmin_hrv, which could cause some ambiguity for an agent deciding among them.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives. It does not state when to use training readiness over training status or other metrics, nor does it mention any prerequisites or context like needing prior data. The agent must infer usage from the name alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

garmin_training_statusTraining statusB
Read-only

Training status, acute/chronic load balance and VO2 max as of a date.

ParametersJSON Schema
NameRequiredDescriptionDefault
dateNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The annotations already declare readOnlyHint=true, so no safety contradiction exists. The description adds a temporal behavior ('as of a date'), which clarifies that results are a snapshot for a particular date, but it does not disclose more about data availability, units, or how missing data is handled.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One short sentence, no filler, no redundancy. The core output is front-loaded and the date scoping is stated efficiently.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the low complexity (one optional parameter, read-only annotations, and an output schema present), the description is mostly complete for selecting and invoking the tool. The main gaps—date format and null behavior—are meaningful but minor for such a simple tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description carries the full burden for explaining the sole `date` parameter. It does clarify that the date is the as-of date, but it does not explain expected format, the meaning of null/default, or how the date affects the returned metrics. This is thin compensation for an undocumented parameter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly identifies the resource (training status) and the specific metrics returned (acute/chronic load balance and VO2 max), scoped by date. It lacks an explicit verb and does not directly differentiate from sibling tools like garmin_vo2max, but the combined metric list makes its purpose reasonably clear.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives such as garmin_vo2max, garmin_training_readiness, or garmin_daily_summary. The description implies a metrics-query use case but provides no explicit context, exclusion, or preferred-alternative guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

garmin_unschedule_workoutUnschedule a workoutA

Remove a scheduled workout from the calendar. The workout itself is kept.

ParametersJSON Schema
NameRequiredDescriptionDefault
scheduled_workout_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already mark this as mutating and non-destructive. The description adds value by clarifying that the workout itself is kept, which goes beyond the generic annotations and prevents the misconception that unscheduling deletes the workout. No further behavioral disclosure (permissions, reversibility) is given, but for this simple operation the added context is sufficient.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences with zero filler. The first sentence names the action and target, the second sentence preempts confusion with deletion. Every word earns its place, and the key nuance is front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-parameter tool with an output schema, the description covers the core behavior and the key non-destructive nuance. It doesn't mention where scheduled_workout_id comes from (e.g., garmin_scheduled_workouts), but that can be inferred from sibling tools. The main gap is the missing parameter guidance, which is a completeness issue but does not undermine the overall clarity for a simple tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has one required parameter with zero description coverage, and the description provides no information about scheduled_workout_id. It does not clarify whether this is the ID of the schedule entry or the workout ID, leaving a critical ambiguity. With 0% schema coverage, the description was expected to compensate but does not mention the parameter at all.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Remove') and resource ('scheduled workout from the calendar') and explicitly adds that the workout itself is kept, which distinguishes it from the sibling garmin_delete_workout. An agent can immediately tell this removes a calendar entry without deleting the underlying workout.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear context for when to use the tool (removing a schedule while preserving the workout) and implicitly differentiates it from deletion. However, it does not explicitly name alternative tools or state when not to use it, leaving the exclusion to inference rather than direct guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

garmin_vo2maxVO2 max and fitness ageA
Read-only

VO2 max, fitness age and heat/altitude acclimation as of a date.

ParametersJSON Schema
NameRequiredDescriptionDefault
dateNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already provide readOnlyHint=trueable, so the read-only nature is covered without contradiction. The description adds only the metric list and temporal scoping, but does not disclose behavioral details such as data freshness, null-date handling, or whether all metrics are always present.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single concise fragment with no filler words. It front-loads the main metric names and the date context, making it easy to scan.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With only one optional parameter, no required fields, annotations covering read-only behavior, and an output schema present, the description provides enough context for basic invocation. The main gap is missing date format/null semantics, but the default null in the schema and output schema mitigate this.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate for parameter meaning. 'As of a date' does clarify that the optional date parameter sets the temporal reference point, but it does not specify date format, what null means, or whether partial dates are accepted.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description identifies the specific resource (VO2 max, fitness age, heat/altitude acclimation) and adds a temporal scope ('as of a date'), which distinguishes it by topic from sibling tools. It lacks an explicit verb like 'retrieve' or 'get,' but the intent is still clear.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'as of a date' implies the tool is used to fetch these metrics for a particular date, providing some usage context. However, there is no explicit guidance on when to choose this over related siblings such as daily_summary or training_status, and no mention of what omitting the date means.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

garmin_weightWeight and body compositionA
Read-only

Weigh-ins over a date range (default: last 7 days), with body composition where recorded.

ParametersJSON Schema
NameRequiredDescriptionDefault
endNo
startNo

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations include readOnlyHint=true and openWorldHint=true, and the description does not contradict them. The description adds the default behavior of 'last 7 days' and clarifies that body composition is only included 'where recorded,' which is useful context beyond annotations. However, it doesn't disclose pagination, rate limits, or how date range is interpreted when not specified. With annotations covering safety, a 3 is appropriate.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence, front-loading the core resource ('Weigh-ins') and scope ('date range'), with the default unpacked immediately. It also adds the caveat about body composition. There is zero waste, and every word earns its place, making it highly concise and well-structured.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's simplicity (2 params, no required, no output schema), the description is largely sufficient, but it lacks details on parameter format, pagination, and error handling. The annotations cover safety (readOnly, openWorld), so the description's job is to clarify usage, which it partially does. It could mention that start and end are optional and how they interact, but the overall picture is adequate for a basic call.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate for the undocumented parameters. It mentions 'date range' and 'default: last 7 days,' which implies that start and end are date range boundaries, but it doesn't specify format (e.g., ISO 8601), null handling, or inclusive/exclusive bounds. The description adds some meaning but not comprehensive guidance, so baseline 3 is fitting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool retrieves weigh-ins over a date range with body composition where recorded, using a specific verb ('Weigh-ins') and resource ('date range'). It is distinct from siblings like garmin_heart_rate or garmin_activity, though it doesn't explicitly differentiate from other health metrics tools, so it's clear but not perfect.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies when to use it (to get weight data) but provides no explicit guidance on when not to use it or alternatives. It doesn't mention that this is the only weight-specific tool among siblings, but the purpose is clear enough that an agent can infer usage. Lacks explicit exclusions or alternative routing, so it's adequate but not strong.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

garmin_whoamiGarmin accountA
Read-only

Identify the signed-in Garmin account and its unit preferences.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already mark the tool as read-only and open-world. The description adds the meaningful detail that the tool returns both the signed-in account and unit preferences. It does not contradict the safe read-only behavior implied by annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, efficient sentence with no filler. It front-loads the primary purpose and adds the useful secondary output detail ('unit preferences') in a compact way.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a no-parameter identity lookup with read-only annotations and an output schema, the description is sufficient. It tells the agent exactly what information the tool provides, and nothing else is required to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There are zero parameters and the schema coverage is 100%, so there is no parameter documentation burden for the description. The description fully covers what the tool does without needing parameter elaboration.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb 'Identify' and the resource: the signed-in Garmin account plus its unit preferences. It is immediately distinguishable from all sibling tools, which focus on activities, metrics, workouts, and so on.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The intended usage is clear: use this tool when you need to know the currently signed-in Garmin account or its unit preferences. It does not explicitly name alternatives or exclusions, but among the sibling list this tool is uniquely self-describing as an identity/account tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 35 tool updatesv0.4.0
    • First observedgarmin_activities
    • First observedgarmin_activities_by_date
    • First observedgarmin_activity
    • First observedgarmin_activity_splits
    • First observedgarmin_activity_weather
    • First observedgarmin_api_get
    • First observedgarmin_body_battery
    • First observedgarmin_briefing
    • First observedgarmin_create_workout
    • First observedgarmin_daily_summary
    • First observedgarmin_delete_workout
    • First observedgarmin_devices
    • First observedgarmin_export_activities
    • First observedgarmin_heart_rate
    • First observedgarmin_hrv
    • First observedgarmin_intensity_minutes
    • First observedgarmin_last_activity
    • First observedgarmin_list_workouts
    • First observedgarmin_personal_records
    • First observedgarmin_race_predictions
    • First observedgarmin_respiration
    • First observedgarmin_schedule_workout
    • First observedgarmin_scheduled_workouts
    • First observedgarmin_sleep
    • First observedgarmin_spo2
    • First observedgarmin_steps
    • First observedgarmin_stress
    • First observedgarmin_submit_mfa_code
    • First observedgarmin_training_history
    • First observedgarmin_training_readiness
    • First observedgarmin_training_status
    • First observedgarmin_unschedule_workout
    • First observedgarmin_vo2max
    • First observedgarmin_weight
    • First observedgarmin_whoami

TDQS

B3/5.0

Scored across 35 tools

Disambiguation4/5

Each tool targets a distinct Garmin data source or action, and the metric tools are clearly separated by health domain. The main ambiguity risk is between rollups like garmin_briefing and garmin_daily_summary, and between activity list variants, but the descriptions are specific enough to resolve those.

Naming Consistency4/5

The garmin_ prefix and snake_case style are consistent throughout. Reader tools use noun-ish names like garmin_hrv and garmin_steps while action tools use verb_noun names like create_workout and delete_workout; a few outliers like garmin_api_get and garmin_whoami break the pattern but are still predictable.

Tool Count2/5

35 tools is well over the 25+ threshold and feels heavy for an agent to navigate, even though the domain is broad. Many health metrics and workout operations are legitimately distinct, but the surface would benefit from consolidation, such as grouping daily wellness metrics into fewer tools.

Completeness4/5

The set covers the core Garmin wellness and activity domains thoroughly: health metrics, activity history/detail, personal records, training status, and workout lifecycle management. Minor gaps exist, like no update_workout or activity deletion, but garmin_api_get provides a read-only escape hatch and the main workflows have no dead ends.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • F
    license
    Not graded
    quality
    B
    maintenance
    Enables AI assistants to access and query Garmin Connect health and fitness data, including sleep, HRV, training load, and activities, with an optional coaching plugin for personalized training plans.
    4 npm
    4
    -
  • A
    license
    Not graded
    quality
    D
    maintenance
    Enables AI assistants to query live Garmin Connect health and fitness data, including daily metrics, activities, sleep analysis, and trends via natural language.
    MIT