Skip to main content
Glama
pluton74mac

garmin-mcp-triathlon

by pluton74mac

Garmin_MCP_Triathlon

Installs two commands: garmin-mcp-triathlon (the server your MCP client launches) and garmin-mcp-triathlon-auth (one-time Garmin authentication).

Fork of Taxuspt/garmin_mcp (749★, MIT) — a Garmin Connect MCP server purpose-built for triathlon & endurance coaching.

171 tools in total. 140 come from upstream — Garmin health data, activities, workouts, devices, gear and the rest. 31 are new: 18 triathlon workout builders and 13 coaching tools across analytics, composite views, bulk data retrieval and plan automation.

The coaching tools return measurements only. Thresholds, verdicts and recommendations live in a separate coaching skill — see Design Principle.


What's Different from Upstream

31 New Tools

Module

Tools

What They Do

Workout Builders

18

Cycling, running, swimming, brick/multi-sport and strength — natural params → verified Garmin JSON → upload. (The module holds 23; five are upstream builders kept as-is, listed below.)

Bulk Data

3

get_health_series (7 metrics × N days, one call), get_activity_series, get_athlete_context (LTHR, FTP, HR zone floors)

Coaching Analytics

7

Readiness score + factors, load breakdown, zone distribution, scheduled-vs-completed pairing, performance trend, cardiac drift, weekly minutes

Composite Views

2

Morning brief (5 endpoints in one round trip), athlete status snapshot with baseline deviations

Plan Execution

1

Weekly plan creator — YAML → built, uploaded and scheduled on the Garmin calendar

Every one of them returns measurements. What the numbers mean is the coaching skill's job — see Design Principle.

3 Critical Mapping Bugs Fixed

The upstream workout_builders.py and the old json_encoder.py had silent bugs. Our builder tools fix them:

Target

Bug (Old)

Fix (Our Fork)

Device Display

Exact cycling watts

ID 6 power.between → Garmin stores it as pace.zone

ID 2 power.zone + targetValueOne/Two

"240-270W" not a pace target

Custom HR range (e.g. 130-145 bpm)

Silent drop (.heart.rate key crash)

ID 4 heart.rate.zone + targetValueOne/Two

"130-145 bpm" not blank

No validation pipeline

Raw JSON, silent upload failures

All builders go through upload_workout validation

Catches errors before upload

On cycling watt targets. Earlier versions of this table recommended target ID 6 with power.between, following upstream's docstrings. A live upload/read-back against a real account shows Garmin silently rewrites that to pace.zone on a cycling workout — the watt bounds survive but are reinterpreted as pace. Target ID 2 (power.zone) with the bounds in targetValueOne/targetValueTwo round trips intact and is what the builders now emit.

Every builder passes integration tests, and the coaching tools are verified against a live Garmin account — mocks alone hid several payload shape bugs (see normalize_sleep, normalize_readiness).

Verified on device

Uploaded from these builders, synced to an Instinct 2X Solar, and read off the watch. Anything not in this table is verified against the API only.

Target

Encoding

Watch shows

Cycling watts

ID 2 power.zone + targetValueOne/Two

240-270W

Cycling power zone

ID 2 power.zone + zoneNumber

Pwr. Zone 3

Custom HR range

ID 4 heart.rate.zone + targetValueOne/Two

130-145 bpm

Named HR zone

ID 4 heart.rate.zone + zoneNumber

HR zone 4

Repeat groups

RepeatGroupDTO + numberOfIterations

4×, recovery 1:30

Swim pace

ID 6 pace.zone, bounds in m/s

1:20/100m

Swim pace band

ID 6 pace.zone, two bounds

1:15-1:25/100m

The one that does not work: ID 6 power.between on a cycling workout. It uploads without error and reads back with the watt bounds intact, but Garmin stores it as pace.zone and the watch renders 240 m/s as 864.00km/h. That is what motivated the ID 2 correction above — see upstream issue #245.

power.between is now rejected on upload, whatever workoutTargetTypeId it arrives with, and the error names power.zone as the replacement. Garmin treats the id as authoritative and ignores the key string, so a wrong pairing cannot be caught by an id/key cross-check — the key has to be refused by name. Accepting it for backwards compatibility only preserved a silent wrong answer. Upstream PR #194 reached the same conclusion independently, by its own live round trip.


Related MCP server: Garmin Workouts MCP Server

Quick Start (Hermes Agent)

git clone https://github.com/pluton74mac/Garmin_MCP_Triathlon.git
cd Garmin_MCP_Triathlon

One-Command Setup

./scripts/hermes-setup.sh

This handles everything: installs uv standalone (required — pip-installed won't work), creates the wrapper script, and writes the MCP config to ~/.hermes/config.yaml. Then type /reload-mcp in your Hermes chat.

If you run garmin-mcp-triathlon under a dedicated Hermes profile rather than the default one, pass --profile <name> so the config is written to ~/.hermes/profiles/<name>/config.yaml instead of the global config:

./scripts/hermes-setup.sh --profile triathlon-coach

Manual Setup (if you prefer step-by-step)

1. Install uv standalone

# Required — pip-installed uv is NOT in Hermes' PATH
curl -LsSf https://astral.sh/uv/install.sh | sh

2. Authenticate

cd Garmin_MCP_Triathlon
uv run garmin-mcp-triathlon-auth

Enter your Garmin email, password, and MFA code. Tokens saved to ~/.garminconnect/.

3. Create wrapper script

REPO_DIR="$(cd Garmin_MCP_Triathlon && pwd)"  # absolute path to your clone
cat > ~/.local/bin/garmin-mcp-triathlon << EOF
#!/usr/bin/env bash
cd "$REPO_DIR"
exec "\$HOME/.local/bin/uv" run garmin-mcp-triathlon
EOF
chmod +x ~/.local/bin/garmin-mcp-triathlon

Use the actual absolute path to your clone here, not a placeholder — a wrapper that can't resolve REPO_DIR will silently no-op the cd and uv run will fail to find pyproject.toml.

Why a wrapper? Hermes Agent may not parse command + args arrays in config.yaml correctly — it can spawn the server name as the command instead of uv. The wrapper bundles the cd + exec uv run into one executable, bypassing this bug entirely.

4. Configure Hermes

Write the MCP config with Python YAML (Hermes guards config.yaml from file tools). Target the global config only if you're not using a dedicated Hermes profile — if you are, write to ~/.hermes/profiles/<name>/config.yaml instead, or the global default profile gets silently reconfigured:

import os
import yaml
p = '~/.hermes/config.yaml'  # or ~/.hermes/profiles/<name>/config.yaml for a dedicated profile
p = os.path.expanduser(p)
with open(p) as f:
    c = yaml.safe_load(f)
c['mcp_servers']['garmin-mcp-triathlon'] = {
    'command': os.path.expanduser('~/.local/bin/garmin-mcp-triathlon'),
    'timeout': 300  # cold start: Garmin auth takes 10-15s
}
with open(p, 'w') as f:
    yaml.safe_dump(c, f, default_flow_style=False, allow_unicode=True, sort_keys=False)

5. Load Tools

In a Hermes chat session: /reload-mcp

Verify with: get_user_profile — should return your Garmin profile.

Common Pitfalls

Symptom

Cause

Fix

Failed to spawn: garmin-mcp-triathlon

Hermes misparsing command+args

Use wrapper script (Step 3)

Connection closed in hermes mcp test

Cold-start timeout (Garmin auth takes 10-15s)

Bump timeout to 300, retry

uv: command not found

pip-installed uv, not standalone

Install standalone (Step 1)

Other MCP servers disappeared

Overwriting mcp_servers block

Use Python YAML to merge, not replace

Tools not visible after /reload-mcp

Config cached or parsing error

Verify with hermes mcp list and hermes mcp test

Nutrition tools return 403 Forbidden

Garmin nutrition/food-log API not enabled for the account

Account-level, not a bug here — reads and writes both fail before this code runs

Brick workout says not compatible on the watch

Device does not support multi-sport structured workouts

See the note under Brick / Multi-Sport builders

A workout step shows no target on the watch

Step list often omits it

Press into the step — the target is usually there


Workout Builder Catalog

Cycling (6 builders)

create_cycling_endurance_workout(name, duration_min, hr_zone="Z2", warmup_min=15, cooldown_min=15)
create_cycling_tempo_workout(name, duration_min, hr_zone="Z3", warmup_min=15, cooldown_min=15)
create_cycling_sweet_spot_workout(name, reps=3, work_min=20, rest_min=5, warmup_min=15, cooldown_min=10)
create_cycling_interval_workout(name, reps=5, work_sec=180, rest_sec=180, power_low=250, power_high=270, ...)
create_cycling_over_under_workout(name, reps=3, over_sec=60, under_sec=120, over_pct=105, under_pct=90, ...)
create_cycling_ftp_test_workout(name="FTP Test", warmup_min=20, test_min=20, cooldown_min=15)

Running (6 builders)

create_run_easy_workout(name, duration_min, hr_zone="Z2", warmup_min=10, cooldown_min=10)
create_run_tempo_workout(name, duration_min, hr_zone="Z4", warmup_min=10, cooldown_min=10)
create_run_long_workout(name, duration_min, hr_min=130, hr_max=145, ...)  # custom BPM range!
create_run_intervals_workout(name, reps=6, distance_m=400, rest_sec=120, hr_zone="Z5", ...)
create_run_hills_workout(name, reps=8, hill_sec=60, jog_down_sec=90, ...)
create_run_progression_workout(name, blocks=[...], warmup_min=15, cooldown_min=10)

Swimming (4 builders)

create_swim_endurance_workout(name, distance_m=1500, pace="1:45/100m", stroke="freestyle", pool_length=25)
create_swim_intervals_workout(name, reps=4, distance_m=200, rest_sec=30, pace="1:40/100m", ...)
create_swim_threshold_workout(name, distance_m=800, pace="1:42/100m", ...)
create_swim_drills_workout(name, drills=[{name, distance_m, equipment, stroke}, ...], ...)

Brick / Multi-Sport (2 builders)

create_brick_bike_run_workout(name, bike_duration_min=60, run_duration_min=20, bike_hr_zone="Z2", run_hr_zone="Z3")
create_brick_swim_bike_workout(name, swim_distance_m=1500, bike_duration_min=60, swim_pace="1:45/100m", ...)

Check your watch supports multi-sport workouts before relying on these. Both builders upload valid multi_sport workouts, but many Garmin watches cannot run a structured multi-sport workout and will report the workout as not compatible when you try to send it to the device. Confirmed on an Instinct 2X Solar and a Forerunner 245 Music — and a multi-sport workout created natively in the Garmin Connect app is rejected identically, so this is a device limitation rather than an encoding fault. Multi-sport structured workouts are generally a higher-tier feature (Forerunner 745/945/955/965, Fenix 6 and later, Enduro).

Upstream builders (preserved)

create_walk_run_workout, create_z2_walk_workout, create_strength_workout, create_run_workout, upload_workout, schedule_week


Coaching Analytics Catalog

All 13 coaching tools, and only the 13 that exist. Each returns measurements; none returns a verdict. Thresholds live in the coaching skill — see Design Principle.

Bulk Data (3 tools)

Tool

Returns

get_health_series(start, end, metrics=None)

Per-day body battery (4 values), HRV, resting HR, sleep, stress, training load, readiness — plus errors[], api_calls and a rate_limited flag

get_activity_series(start, end)

Per-activity date, sport, duration, distance, HR, power, training effect, optional HR-zone seconds

get_athlete_context()

LTHR, cycling/running FTP with as_of + is_stale, per-sport HR zone floors, VO2max, physical data, preferences, not_available

Individual Analytics (7 tools)

Tool

Returns

get_training_readiness_composite(date)

Garmin's readiness score, its level, six factor percentages

get_training_load_breakdown(start, end)

Minutes per sport plus Garmin's acute/chronic load, ACWR and TSB

get_zone_distribution(start, end)

Seconds per HR zone as percentages, by sport

get_workout_compliance(start, end)

Scheduled workouts paired with same-day activities

get_performance_trend(metric, sport, days)

Per-activity pace or power + avg HR, regression slope

get_cardiac_drift_analysis(activity_id)

hr_drift_pct — needs power and ≥60 min at 1 s sampling

get_weekly_load_progression(weeks=12)

Minutes per ISO week, week-over-week change

Composite Views (2 tools)

Tool

Impact

get_morning_brief(date)

5 calls → 1 — sleep, recovery, readiness, today's workout

get_athlete_status_snapshot(date)

Current values, Garmin baselines, deviations

Plan Automation (1 tool)

Tool

What It Does

create_weekly_plan(plan_yaml_path)

Reads YAML/JSON → creates all workouts → schedules each on its own date in the Garmin Calendar

Nine tools were removed, not relocated

run_safety_check, check_overtraining_risk, get_injury_risk_assessment, get_reds_risk_assessment, get_load_adjustment_recommendation, get_recovery_trend, get_weekly_health_summary, generate_taper_plan and validate_weekly_plan no longer exist, and neither does src/garmin_mcp/coaching_safety.py. Each of them encoded a coaching judgement — a threshold, a load curve, a gate — inside the data layer. Two of them returned opposite verdicts on identical data. That reasoning now lives in skills/triathlon-coaching/, where every threshold is one line of rules.yaml with its provenance recorded.

If you are looking for a safety gate, injury screen or taper, it is in the skill, not here. See docs/coaching-split-audit.md.


Architecture

Garmin_MCP_Triathlon/
├── src/garmin_mcp/                # Preserved upstream namespace
│   ├── *.py                       # Upstream modules (UNCHANGED)
│   ├── workout_builders.py        # EXTENDED: +18 triathlon builders
│   │
│   ├── coaching_data.py           # NEW: 3 bulk retrieval tools
│   ├── coaching_analytics.py      # NEW: 7 measurement tools
│   ├── coaching_composite.py      # NEW: 2 aggregated view tools
│   └── coaching_planning.py       # NEW: 1 plan execution tool
│
├── skills/triathlon-coaching/     # NEW: the judgement layer
│   ├── rules.yaml                 #   every threshold, one file
│   ├── scripts/evaluate.py        #   contains no numbers
│   └── references/                #   provenance for each threshold
│
├── tests/
│   ├── unit/                      # Unit tests for builders
│   ├── integration/               # Integration tests (mocked Garmin API)
│   │   ├── test_workout_builders_tools.py # EXTENDED
│   │   ├── test_coaching_data_tools.py    # NEW
│   │   ├── test_thinned_surface.py        # NEW: no verdicts leak
│   │   ├── test_fetch_failures_surface.py # NEW
│   │   └── test_*_reads.py                # NEW: payload-shape guards
│   └── e2e/                       # End-to-end (real Garmin creds)

Every new module follows the upstream pattern: configure(client) + register_tools(app).

Design Principle

┌────────────────────────────────────────┐
│  garmin-mcp-triathlon (DATA LAYER)      │
│  "What does the data say?"             │
│  Raw Garmin data → structured JSON     │
└────────────┬───────────────────────────┘
             │ MCP tool calls
             ▼
┌────────────────────────────────────────┐
│  Hermes Coaching Skills (INTELLIGENCE) │
│  "What should we do about it?"         │
│  Interpret, recommend, plan            │
└────────────────────────────────────────┘

The MCP returns data. The coach decides what to do.

This is enforced, not aspirational. No coaching tool returns a threshold, a severity, a gate or a sentence of advice; a test walks every tool's output looking for that vocabulary. Nine tools that did were removed and seven were thinned — docs/coaching-split-audit.md records what each one encoded and why.

The judgement lives in skills/triathlon-coaching/, where every threshold sits in one editable rules.yaml and the evaluator contains no numbers at all. references/rationale.md records where each number came from and what live data says about it.

Two rules the data layer keeps:

  • A failed fetch is never silence. Every tool that walks a date range returns an errors[] array. A day with no data is absent from the results; a day whose request raised is in errors. Collapsing those two is how an expired token used to produce a confident all-clear.

  • Nothing is substituted for a missing reading. No zeros, no plausible defaults. An absent measurement is absent.


The errors[] contract

Every tool that walks a date range or a list of activities returns an errors array. It is part of the tool's output contract, not a debugging aid, and the coaching skill depends on it.

Schema

{"date": "2026-08-02", "metric": "hrv", "error": "429 Too Many Requests"}

Field

Type

Meaning

date

YYYY-MM-DD

the day whose request failed

metric

string

the metric that was lost, never the endpoint

error

string

the exception text, unedited

get_activity_series adds activity_id for per-activity failures (HR-zone lookups) and omits date when the failure is not day-scoped. get_athlete_context uses {"source": ..., "error": ...} — its calls are not per-day.

The three states it exists to separate

State

results

errors

Everything worked

full

[]

Athlete has genuine gaps — watch not worn

short

[]

Fetch failed — expired token, 429, outage

short

populated

Rows two and three are byte-identical in the results. Without errors they are indistinguishable, and that is precisely how an expired token used to produce a confident all-clear from the safety gate.

Rules

  1. A day with no data is absent from the results. A day whose request raised is in errors. Never both, never neither.

  2. metric, not endpoint. body_battery and stress share get_stats; when that call fails, both metric names appear. A caller should not have to know Garmin's endpoint topology to understand what it just lost.

  3. Nothing is substituted. No zeros, no plausible defaults, no backfilling a missing value from a neighbouring field.

  4. A non-empty errors adds a warning string saying in prose that the gap is not a negative finding. The consumer is usually a language model, and a sentence is harder to skip than an integer.

  5. api_calls reports the real cost, so the price of a wide date range is visible rather than inferred.

  6. A rate limit aborts the walk. A 429 sets rate_limited: true and stops immediately rather than working through the rest of the range. See below.

Rate limiting

Garmin publishes no limits for this API. What is known from the community is that the aggressive limiting sits on the login/SSO endpoints and is keyed per account — not per IP or user agent — with reported blocks lasting from about an hour to 48+ hours. Token-based auth keeps this server off that path almost entirely: it resumes from ~/.garminconnect/ rather than signing in.

garminconnect 0.3.2 paces and retries login only — both of its anti-WAF sleeps live inside the SSO functions. Data calls have no backoff whatsoever; a 429 raises straight through. A 60-day, 6-source get_health_series walk is roughly 360 unpaced requests, so on hitting a limit the walk stops at the first refusal instead of firing hundreds more. Days already retrieved are returned and are complete; everything after the stop is unknown, and the warning says so.

If you see rate_limited: true, wait before retrying and ask for a shorter range or fewer metrics. Do not loop.

For consumers

Never draw a negative conclusion from a short result set while errors is non-empty. "No overtraining signals" and "we could not look" are different statements. The coaching skill turns its safety gate to unknown — never green — whenever errors is populated, and a real trigger still outranks it so a red gate is not downgraded by an unrelated 429.

Emitted by

get_health_series, get_activity_series, get_athlete_context, get_morning_brief and get_athlete_status_snapshot (the last two as fetch_errors, since they are single-date tools rather than range walks).


Garmin payload shapes worth knowing

These cost real debugging time. Each was found by reading a live payload, never by inferring from a plausible key name — and each one, before it was found, produced a confidently wrong number rather than an error. Most are handled by a named helper (in coaching_analytics.py unless noted); use the helper rather than reading the field directly.

Endpoint

Actual shape

Helper

get_sleep_data

summary nested under dailySleepDTO; sleepScores.overall is a dict with .value

normalize_sleep, sleep_score, sleep_hours

get_training_readiness

one-element list; no maxPossible; factors are <thing>FactorPercent

normalize_readiness, readiness_level

get_rhr_day

allMetrics.metricsMap.WELLNESS_RESTING_HEART_RATE[].value, not a flat restingHeartRate

extract_resting_hr

get_training_status

acuteTrainingLoadDTO under mostRecentTrainingStatus.latestTrainingStatusData.<deviceId>

_extract_acute_load_dto

get_stats

bodyBatteryMostRecentValue is the end-of-day drain, not the overnight charge

body_battery_at_wake

get_hrv_data

the seven-day figure is weeklyAvg; there is no lastSevenDaysAvg

read weeklyAvg

download_activity

takes an ActivityDownloadFormat enum with no FIT member — the FIT arrives inside the ORIGINAL zip

_extract_fit_bytes (activity_analysis.py)

get_max_metrics

returns [] on some accounts; VO2max is in get_user_profile().userData

—

HR zone floors

live at /biometric-service/heartRateZones, which garminconnect does not wrap

raw connectapi

get_activity

no hrInTimezones — HR zones come from get_activity_hr_in_timezones, one call per activity

—

get_body_battery_events

event series, not a daily summary; use get_stats

—

get_activities_by_date

list, newest first — sort before treating position as time

sport_family

workoutScheduleSummariesScalar

a JSON scalar taking Date args; sub-selecting fields is rejected

fetch_scheduled_workouts

Sport keys must be enumerated explicitly. activityType.parentTypeId is not a usable grouping key — 17 is shared by running, cycling, hiking and walking, while trail_running reports 1 and road_biking reports 2. The discipline lists live in SPORT_TYPE_KEYS; note that outdoor rides are road_biking, not cycling, and open water is open_water_swimming. Missing those two silently dropped activities from load, zone and injury analysis.

pace.zone bounds are metres per second. For swim paces use _pace_to_mps (100 / seconds_per_100m). Inverting this is easy to miss because the common default 1:40/100m is exactly 100 s — the one value where the correct and inverted expressions agree.

Multi-segment workouts need workout-unique stepOrder. Restarting at 1 per segment makes Garmin reject the upload outright; _renumber_steps_across_segments numbers continuously and descends into repeat groups.


Tool Filtering

171 tools is a lot of context. Filter per skill with GARMIN_ENABLED_TOOLS:

Skill

Enable these

Health Dashboard

get_stats, get_sleep_data, get_hrv_data, get_body_battery, get_stress_data, get_training_readiness_composite, get_morning_brief

Workout Review

get_activity_splits, get_activity, get_activity_details, get_training_effect, get_activity_fit_data, get_cardiac_drift_analysis

Workout Manager

All create_*_workout builders + schedule_week + create_weekly_plan

Weekly Insights

get_zone_distribution, get_workout_compliance, get_training_load_breakdown, get_performance_trend, get_weekly_load_progression

Coaching Skill

get_health_series, get_activity_series, get_athlete_context — the three bulk tools are all the skill's evaluator needs

Names are checked at startup: anything in GARMIN_ENABLED_TOOLS that matches no registered tool is reported on stderr rather than silently ignored.

Set via MCP server env:

"env": {
  "GARMIN_ENABLED_TOOLS": "get_morning_brief,get_sleep_data,get_training_readiness_composite,..."
}

Testing

# All tests (unit + integration) — 634 tests
uv run pytest tests/unit/ tests/integration/ -v

# Specific module
uv run pytest tests/integration/test_workout_builders_tools.py -v

# End-to-end (requires real Garmin credentials)
uv run pytest tests/e2e/ -m e2e -v

634 tests pass across tests/unit and tests/integration; 664 including the coaching skill's own suite (pytest -m "not e2e"). pytest -m e2e is 10 passed, 6 skipped — the skips are the nutrition tests, gated on a live probe because that API is 403 on accounts without the feature. Zero regressions on upstream tests.

The mock is specced against the real client

tests/conftest.py builds the Garmin client with create_autospec against a real Garmin instance, so a call with the wrong arity — or to a method that does not exist — fails immediately. This matters: an earlier revision used a bare Mock(), which accepts anything, and the suite was fully green while seven tools were calling the API incorrectly and failing on every invocation.

Two rules when extending the fixtures:

  • Spec against an instance, not the class. Garmin.__init__ assigns .client and the garmin_connect_* URLs, which a class-level autospec cannot see.

  • Set defaults with client.method.return_value = ..., never client.method = Mock(...) — the latter replaces the autospec'd child and silently discards signature checking.

Autospec is necessary but not sufficient

Autospec constrains call shapes; it says nothing about whether the payload you assert on matches what Garmin actually returns. Several bugs survived a green suite because the fixtures encoded shapes the API does not produce — sleep summaries nested under dailySleepDTO, training readiness returned as a one-element list, HR zones served from a separate endpoint. Fixtures in this repo are kept faithful to live payloads for that reason.

Manual Display Test (Required for Builders)

Uploading successfully is not the same as displaying correctly — Garmin silently rewrites some targets on save. To verify:

  1. Call a builder via MCP (e.g. create_cycling_interval_workout)

  2. Open Garmin Connect → Workouts → verify name, sport, steps, targets

  3. Sync to device → start workout → press into each step to see its target; the step list alone often does not show it

  4. Confirm the target reads in the units you asked for (watts, bpm, min/100m)


Critical Mapping Reference

Target Type

Correct ID

Correct Key

Extra Fields

Exact power (watts)

2

power.zone

targetValueOne=high, targetValueTwo=low

Power zone (FTP%)

2

power.zone

zoneNumber=1-7

HR zone (named)

4

heart.rate.zone

zoneNumber=1-5

HR custom (BPM)

4

heart.rate.zone

targetValueOne=low, targetValueTwo=high

Pace zone

6

pace.zone

targetValueOne=max m/s, targetValueTwo=min m/s

Always set BOTH workoutTargetTypeId AND workoutTargetTypeKey — the validation pipeline catches mismatches.


Upstream Features Preserved

171 tools total once the coaching modules are registered — counted from the @app.tool() registrations, with no duplicate names. The upstream surface is preserved in full:

Module

Tools

health_wellness

29

sleep, HRV, body battery, stress, respiration, steps

activity_management

23

list, get, GPS track, edit, rename, retype, manual entry, delete

training

15

CTL/ATL/TSB, HRV trend, VO2 max, FTP, lactate threshold

workouts

14

upload, schedule, unschedule, delete, list, download

nutrition

14

food log, custom foods, meals, hydration targets

challenges

9

badges, ad-hoc and virtual challenges

devices

6

device list, settings, solar data, alarms

weight_management

5

weigh-ins by day and range, add, delete

user_profile

4

profile, settings, personal records

activity_analysis

4

FIT parsing, power duration curve, Di2 shift summary

womens_health

3

menstrual cycle and pregnancy data

gear_management

3

gear list with stats, associate/dissociate per activity

data_management

4

body composition, blood pressure (add + delete), hydration

courses

3

list, upload GPX, delete

That is 135 tools, plus the 5 upstream workout builders kept inside workout_builders.py (create_walk_run_workout, create_run_workout, create_z2_walk_workout, create_strength_workout, schedule_week) — 140 upstream-derived. The remaining 31 are new: 18 triathlon builders and 13 coaching tools.

delete_activity and delete_blood_pressure were added so that create_manual_activity and set_blood_pressure are undoable through the server — garminconnect had both deletes and neither was registered.

Note that nutrition returns HTTP 403 on accounts without Garmin's nutrition feature, reads and writes alike, before any of this code runs. That is account-level, not a defect here.

Synced with upstream through a16f057, which adds search_foods, set_nutrition_daily_settings and Garmin Coach workout access, and carries upstream's DXT, stdio-corruption and nested-target-bounds fixes.

Two upstream defects were found here and submitted back:

  • get_device_solar_data read six fields that do not exist in the response, so it reported no data for solar watches that had a full day of readings. It now reads solarDailyDataDTOs[].localConnectDate and derives utilisation from solarInputReadings[]. Verified against an Instinct 2X Solar with 1254 readings. (upstream PR #247)

  • get_endurance_score crashed on Garmin's explicit null for a section with no data — .get("enduranceScoreDTO", {}) does not help when the key is present and the value is None. (upstream PR #246)

Both fixes are carried here regardless of whether upstream merges them.


Upstream Setup (Claude Desktop, Codex, Docker)

See upstream documentation for:

  • Claude Desktop configuration

  • Codex/opencode TOML config

  • Docker deployment

  • HTTP transport mode

  • Garmin Connect China


Credits

License

MIT (same as upstream)

Available Tools

172 tools
add_body_compositionC

Add body composition data

Args: date: Date in YYYY-MM-DD format weight: Weight in kg percent_fat: Body fat percentage percent_hydration: Hydration percentage visceral_fat_mass: Visceral fat mass bone_mass: Bone mass muscle_mass: Muscle mass basal_met: Basal metabolic rate active_met: Active metabolic rate physique_rating: Physique rating metabolic_age: Metabolic age visceral_fat_rating: Visceral fat rating bmi: Body Mass Index

ParametersJSON Schema
NameRequiredDescriptionDefault
bmiNo
dateYes
weightYes
basal_metNo
bone_massNo
active_metNo
muscle_massNo
percent_fatNo
metabolic_ageNo
physique_ratingNo
percent_hydrationNo
visceral_fat_massNo
visceral_fat_ratingNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

C2.7/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description omits any details about side effects, such as whether the operation overwrites existing entries, requires prior data, or is idempotent. With no annotations and only a bare action phrase, the agent has no insight into the write behavior or potential consequences.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise, consisting of a single action line followed by a parameter list. However, it is not structured as a narrative or front-loaded with a summary; it reads as a bare enumeration, which is efficient but could be more organized.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description lacks information about the return value or success/failure indications, despite the presence of an output schema. It also does not highlight which parameters are required (date and weight) or any constraints, leaving the agent with incomplete context for invoking the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The description adds moderate value by listing units for some parameters (e.g., weight in kg, percent_fat as percentage), but it does not provide units for all fields (e.g., basal_met) or clarify expected ranges. Given the schema has 0% coverage, this is helpful but incomplete, earning a middle score.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb 'Add' and the resource 'body composition data', making the tool's purpose unambiguous. It is also distinct from sibling tools like add_weigh_in and get_body_composition, so an agent can easily identify this as the write operation for body composition metrics.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives such as add_weigh_in or get_body_composition. There is no explanation of prerequisites, relationship to other data, or typical use cases, leaving the agent to infer context from the tool name alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

add_gear_to_activityB

Associate gear with an activity

Links a specific piece of gear (like shoes, bike, etc.) to an activity.

Args: activity_id: ID of the activity gear_uuid: UUID of the gear to add (get from get_gear)

ParametersJSON Schema
NameRequiredDescriptionDefault
gear_uuidYes
activity_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It only states the high-level association action and does not reveal potential side effects, whether existing gear associations are replaced or appended, error behavior for invalid IDs, or any authorization requirements.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and well-organized, with a clear one-line summary followed by a focused Args section. Every sentence contributes value, and the parameter documentation is front-loaded and easy for an agent to parse quickly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple two-parameter tool with an output schema present, the description covers the basics of what the tool does and what the parameters mean. However, it omits important behavioral context such as whether the operation is idempotent, what happens if gear is already associated, and whether the activity must exist, making it adequate but not fully complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 0% description coverage, so the description must compensate, and it largely does. The Args block provides semantic meaning for both parameters: activity_id is 'ID of the activity' and gear_uuid is 'UUID of the gear to add (get from get_gear)', adding useful sourcing context beyond the raw schema types.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states a specific verb and resource: 'Associate gear with an activity' and 'Links a specific piece of gear to an activity.' It distinguishes the action from unrelated siblings by defining exactly what is being linked, though it does not explicitly contrast itself with remove_gear_from_activity or get_activity_gear.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The only usage guidance is the hint that the gear UUID should be obtained from get_gear, which is a helpful prerequisite but not a tool-selection guideline. It does not explain when to choose this tool over remove_gear_from_activity or other gear-related tools, nor does it provide any exclusions or alternative conditions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

add_hydration_dataA

Add a hydration entry for a date.

Entries accumulate: logging 500 then 250 leaves the day at 750. A NEGATIVE value_in_ml subtracts, which is how an over-log is corrected — verified live.

A day cannot go below zero: Garmin answers a further subtraction with "Daily hydration is already at 0 and cannot be reduced." There is no delete for hydration in the Garmin API, so zeroing a day is as close to removing it as the platform allows — the entry stays, reading 0 mL.

Args: value_in_ml: Amount of liquid in millilitres. Negative subtracts from the day's running total. date: Date in YYYY-MM-DD format. (Named to match get_hydration_data; the argument was once called cdate.) timestamp: Timestamp in YYYY-MM-DDThh:mm:ss.sss format

ParametersJSON Schema
NameRequiredDescriptionDefault
dateYes
timestampYes
value_in_mlYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden and does so exceptionally well: it discloses accumulation semantics, negative-value subtraction, the zero floor, the exact API error message on further subtraction, and the fact that the Garmin API offers no delete so zeroing is the closest to removal. These are exactly the traits an agent needs before mutating hydration data.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The most decision-relevant facts (accumulation, negative correction, zero floor) are front-loaded in the first two paragraphs, and the Args block is a compact reference. Length is justified by genuinely non-obvious behavior rather than padding.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 3-param mutation tool with an output schema, the description covers everything the agent needs: mutation semantics, correction workflow, floor behavior, error text, and the absence of a delete path. Return values are legitimately omitted since an output schema exists.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, and the Args block does: value_in_ml is 'amount of liquid in millilitres' with negative-subtracts semantics, date is YYYY-MM-DD, and timestamp is YYYY-MM-DDThh:mm:ss.sss. It even explains why the argument is named `date` rather than the historical `cdate`.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The opening sentence gives a specific verb and resource ('Add a hydration entry for a date'), so the agent immediately knows this is a write operation on daily hydration. It references the sibling get_hydration_data only incidentally (for the date argument naming), not as a routing cue, so the sibling distinction is present but not fully exploited.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description effectively tells the agent when and how to use this tool: entries accumulate, and a negative value_in_ml is the documented way to correct an over-log. It does not, however, explicitly say to read with get_hydration_data first or name any exclusions, so it stops short of full when/when-not guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

add_weigh_inB

Add a new weight measurement

Args: weight: Weight value unit_key: Unit of weight — 'kg' or 'lbs'. Not 'lb': the underlying client validates against {'kg', 'lbs'} and rejects anything else with "unitKey must be one of {'kg', 'lbs'}".

ParametersJSON Schema
NameRequiredDescriptionDefault
weightYes
unit_keyNokg

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden, and it does disclose a genuine behavioral fact: the client validates unit_key against {'kg','lbs'} and rejects anything else with a specific error. That is real value beyond the schema. It still omits other mutation-critical behavior — whether a date/time is assumed, whether re-submitting overwrites, and any auth requirements.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The purpose is front-loaded and the docstring-style Args block is compact and easy to scan. The quoted error string is somewhat verbose but earns its place as a warning about a non-obvious validation trap.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained, and the critical unit_key footgun is covered. But for a mutation with zero annotations, key facts are missing: the time/date semantics that distinguish it from add_weigh_in_with_timestamps and any indication of write scope or reversibility.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate, and it largely does: unit_key is given its valid enum values ('kg' or 'lbs'), explicitly rules out the ambiguous 'lb', and quotes the validation error. weight is documented only as 'Weight value', which is thin but sufficient for a plain number.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Add a new weight measurement'), so the agent knows exactly what the tool does. However, it gives no differentiation from close siblings like add_weigh_in_with_timestamps or add_body_composition, which a reader cannot distinguish from the description alone.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance and no alternatives named. The sibling add_weigh_in_with_timestamps is an obvious fork (does this one use the current time instead of supplied timestamps?) and the description never addresses it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

add_weigh_in_with_timestampsB

Add a new weight measurement with specific timestamps

Args: weight: Weight value unit_key: Unit of weight — 'kg' or 'lbs'. Not 'lb'; see add_weigh_in. date_timestamp: Local timestamp in format YYYY-MM-DDThh:mm:ss gmt_timestamp: GMT timestamp in format YYYY-MM-DDThh:mm:ss

ParametersJSON Schema
NameRequiredDescriptionDefault
weightYes
unit_keyNokg
gmt_timestampNo
date_timestampNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden. It does not state what happens if the timestamps are omitted (schema defaults them to null), whether an entry for that date is overwritten, whether permissions are needed, or any reversibility/duplication behavior for this write operation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The one-line summary is front-loaded and the Args block is compact and scannable. Slight redundancy between the arg list and the schema, but no wasted prose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained, but for an unannotated mutation tool the description should clarify the relationship to add_weigh_in and whether the two timestamps are jointly required. Those gaps leave the agent guessing about when and how to use this variant.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate, and it largely does: it documents all four parameters, gives the accepted unit values plus a spelling warning, and specifies the exact timestamp format YYYY-MM-DDThh:mm:ss for both timestamp fields. It falls short on 'weight: Weight value' (tautological) and does not explain that both timestamps are optional or how they interact.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a clear verb+resource ('Add a new weight measurement') plus the distinguishing feature 'with specific timestamps', which separates it from the sibling add_weigh_in at a high level. However, it never explicitly states how it differs from add_weigh_in beyond the name, so the differentiation is implied rather than stated.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The only reference to an alternative is inside the unit_key arg note ('Not lb; see add_weigh_in'), which is guidance about unit spelling, not about tool selection. Nothing tells the agent when to prefer this timestamps variant over add_weigh_in, or whether timestamps are required.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

count_activitiesA

Get total count of activities in the user's Garmin account

Returns the total number of activities recorded.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It accurately portrays a read-only operation via 'Get' and 'Returns', and does not claim any side effects. However, it does not explicitly disclose that it does not modify data or mention error conditions, though these are largely implied by the verb.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is exceptionally concise, consisting of two short sentences that convey the purpose and output without any redundant or extraneous information. The structure is clean and front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's simplicity (no parameters, no annotations, no explicit output schema), the description fully explains what the tool does and what it returns. Agents can invoke it correctly without needing additional context.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has no parameters, so schema coverage is effectively 100%. The description adds no parameter-specific information, but none is needed. The baseline score of 3 applies because the schema fully covers the (empty) parameter space.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose with a specific verb ('Get') and resource ('total count of activities'), and implicitly distinguishes it from get_activities by focusing on the count instead of the list. The two sentences are concise and unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies the primary use case (obtaining a count) but does not explicitly state when to prefer this over get_activities or mention any alternative. No guidance on when not to use it is provided, leaving the decision to the agent.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_brick_bike_run_workoutA

Create a brick workout (bike → run) and upload it to Garmin Connect.

Multi-sport workout with two segments. Garmin auto-adds T1/T2 transitions.

Args: name: Workout name (e.g. "Brick 60m/20m") bike_duration_min: Bike segment duration in minutes (default 60) run_duration_min: Run segment duration in minutes (default 20) bike_hr_zone: HR zone for bike (default Z2) run_hr_zone: HR zone for run (default Z3) bike_warmup_min: Bike warmup in minutes (default 15) run_cooldown_min: Run cooldown in minutes (default 5)

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYes
run_hr_zoneNoZ3
bike_hr_zoneNoZ2
bike_warmup_minNo
run_cooldown_minNo
run_duration_minNo
bike_duration_minNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It usefully discloses that the workout is uploaded to Garmin Connect and that Garmin auto-adds T1/T2 transitions, but says nothing about authentication, whether the workout is also scheduled, or what happens on upload failure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the one-line purpose and side effect, then presents parameters as a scannable list. The Args formatting is slightly verbose but every line carries parameter meaning.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained, and all inputs are documented. Only the workflow context (scheduling, auth, idempotency of re-uploading the same name) is left unaddressed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, yet the Args block documents all seven parameters with units, defaults, and examples (e.g. HR zones defaulting to Z2/Z3), which is meaningfully more than the bare schema titles and defaults provide.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Create a brick workout (bike → run)') plus the side effect ('upload it to Garmin Connect'). The two-segment structure distinguishes it clearly from siblings like create_brick_swim_bike_workout and create_walk_run_workout.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied by the purpose and by the default durations/zones, which suggest a standard brick session, but the description never states when to choose this over create_brick_swim_bike_workout or the single-discipline workout creators, nor any prerequisites.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_brick_swim_bike_workoutA

Create a swim → bike brick workout and upload it to Garmin Connect.

Multi-sport workout: swim segment with pace target → bike segment with HR zone.

Args: name: Workout name (e.g. "Brick Swim/Bike") swim_distance_m: Swim distance in meters (default 1500) bike_duration_min: Bike segment duration in minutes (default 60) swim_pace: Target swim pace as "M:SS/100m" bike_hr_zone: HR zone for bike (default Z2) swim_warmup_m: Swim warmup distance in meters (default 400) bike_cooldown_min: Bike cooldown in minutes (default 10) pool_length: Pool length in meters (default 25)

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYes
swim_paceNo1:45/100m
pool_lengthNo
bike_hr_zoneNoZ2
swim_warmup_mNo
swim_distance_mNo
bike_cooldown_minNo
bike_duration_minNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries full burden. It usefully discloses the external side effect ('upload it to Garmin Connect'), which the schema does not convey, but it omits auth requirements, whether a duplicate is created on repeat calls, and any scheduling behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The opening sentence front-loads the action and the upload side effect, and the Args block is efficient per-line; it is slightly long but each line earns its place by defining an otherwise-undocumented parameter.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values need not be explained, and parameters plus the upload side effect are covered. The only gap for a mutation tool with zero annotations is the absence of prerequisite/auth and repeat-call behavior.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, and it does: every one of the 8 parameters is documented with units and defaults (e.g. 'swim_distance_m: Swim distance in meters (default 1500)', pace as "M:SS/100m"), giving the agent more than the bare schema titles.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource ('Create a swim → bike brick workout') and explicitly names the destination ('upload it to Garmin Connect'), which cleanly distinguishes it from siblings like create_brick_bike_run_workout and create_swim_endurance_workout.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The multi-sport structure ('swim segment with pace target → bike segment with HR zone') implies the use case, but there is no explicit statement of when to pick this over a plain swim or cycling workout, nor any prerequisite or exclusion guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_custom_foodA

Create a custom food in the user's Garmin nutrition library

Creates a new food item with nutritional information per serving. On success the response includes foodId and servingId needed for log_custom_food. If the API returns no data (204), use get_custom_foods(search=food_name) to retrieve those IDs.

All nutrient amounts are ABSOLUTE values per serving, not %DV. Nutrition labels often print %DV for calcium/iron/vitamin D — convert to absolute units before passing.

Args: food_name: Name of the custom food (e.g. "Homemade Chocolate Cookies") calories: Calories per serving serving_unit: Unit for serving size (e.g. "G", "ML", "OZ"). Default "G" number_of_units: Serving size in the specified unit. Default 100 brand_name: Brand or vendor name (e.g. "Three Bridges") carbs: Carbohydrates in grams per serving protein: Protein in grams per serving fat: Total fat in grams per serving fiber: Fiber in grams per serving sugar: Sugar in grams per serving saturated_fat: Saturated fat in grams per serving sodium: Sodium in mg per serving cholesterol: Cholesterol in mg per serving potassium: Potassium in mg per serving trans_fat: Trans fat in grams per serving calcium: Calcium in mg per serving (NOT %DV) iron: Iron in mg per serving (NOT %DV) vitamin_d: Vitamin D in mcg per serving (NOT %DV)

ParametersJSON Schema
NameRequiredDescriptionDefault
fatNo
ironNo
carbsNo
fiberNo
sugarNo
sodiumNo
calciumNo
proteinNo
caloriesYes
food_nameYes
potassiumNo
trans_fatNo
vitamin_dNo
brand_nameNo
cholesterolNo
serving_unitNoG
saturated_fatNo
number_of_unitsNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It discloses the success response (foodId and servingId) and the 204 fallback, but does not mention potential errors, authentication, or other side effects beyond creation. This is adequate but not fully transparent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-structured with an intro, behavioral notes, and an Args list. It is slightly long but every sentence adds value, including the fallback and unit conversion reminders. No redundant or filler content.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the schema covers all parameters and an output schema exists (per context), the description adds necessary context about the response and fallback. It tells the agent what to do with the output (use IDs for logging) and how to handle a 204. Missing minor details like uniqueness constraints, but overall complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema has no descriptions, but the Args section in the description adds meaningful detail to every parameter, including units (mg, g, mcg) and clarifications like 'NOT %DV' for micronutrients. It also specifies allowed examples for serving_unit and defaults. This fully compensates for the schema's lack of descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the specific verb 'Create' and the resource 'custom food in the user's Garmin nutrition library'. It also differentiates from siblings by noting that the returned foodId and servingId are needed for log_custom_food, making the purpose unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides actionable guidance: it explains the fallback to get_custom_foods when a 204 is returned, and instructs converting %DV values to absolute units. While it does not explicitly say when to use this vs update or delete, the purpose is clear and these guidelines are helpful.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_cycling_endurance_workoutA

Create a steady endurance cycling workout and upload it to Garmin Connect.

Builds a single continuous aerobic ride with HR zone target.

Args: name: Workout name (e.g. "Z2 Endurance 90m") duration_min: Duration of the main ride in minutes hr_zone: Target heart-rate zone (Z1-Z5, default Z2) warmup_min: Warmup duration in minutes (default 15) cooldown_min: Cooldown duration in minutes (default 15)

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYes
hr_zoneNoZ2
warmup_minNo
cooldown_minNo
duration_minYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It does disclose the important side effect that the workout is "upload[ed] to Garmin Connect," which an agent needs to know this is a write operation. However, it says nothing about auth requirements, whether re-creating a same-named workout duplicates or overwrites, or failure behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the purpose and upload side effect in the first sentence, followed by a compact structured Args block. Every line carries information; the only mild redundancy is restating the ride shape in sentence two after sentence one already established it.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained, and the parameter docs cover the 0%-coverage schema. The remaining gap is behavioral: for a mutation/upload tool with no annotations, the agent is not told about authentication or duplicate-handling. Still broadly sufficient to invoke correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, and it does: it documents all five parameters including the Z1-Z5 range for hr_zone and defaults for hr_zone/warmup_min/cooldown_min. It even clarifies that duration_min is the main ride length, implicitly distinct from warmup and cooldown. Minor gap: no explicit statement that name/duration_min are required.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ("Create a steady endurance cycling workout") and elaborates the structure ("a single continuous aerobic ride with HR zone target"). This clearly distinguishes it from the many cycling siblings (tempo, sweet spot, interval, over-under), which are all named after different workout shapes.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies when to use it by defining the workout type (steady endurance vs. tempo/intervals), but never explicitly states when to prefer this tool over create_cycling_tempo_workout or create_cycling_sweet_spot_workout, nor any prerequisites such as needing a connected Garmin account.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_cycling_ftp_test_workoutA

Create a 20-minute FTP test cycling workout and upload it to Garmin Connect.

No power target — free effort. The athlete rides at maximum sustainable pace and records the average watts from the test block.

Args: name: Workout name (default "FTP Test") warmup_min: Warmup duration in minutes (default 20) test_min: FTP test duration in minutes (default 20) cooldown_min: Cooldown duration in minutes (default 15)

ParametersJSON Schema
NameRequiredDescriptionDefault
nameNoFTP Test
test_minNo
warmup_minNo
cooldown_minNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the burden and does disclose the key side effect — the workout is uploaded to Garmin Connect — plus the deliberate absence of a power target and how the test is executed. It does not say whether the workout is scheduled, whether it requires authentication/device setup, or what happens on upload failure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the action and resource, followed by protocol rationale, then a compact Args block. Slightly redundant with the schema, which already carries the same defaults and types, but nothing is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values need not be described, and the description covers parameters, the upload side effect, and the workout's intent. The remaining gap is scheduling/authentication context and any constraints on duration parameters.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, and it does: all four parameters are named with their meaning and units (durations in minutes) and their defaults. It stops short of constraints such as minimum/maximum durations or whether test_min changes the workout structure.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The first sentence states a specific verb (Create), a specific resource (20-minute FTP test cycling workout), and the side effect (upload to Garmin Connect). It clearly distinguishes itself from the many sibling create_cycling_* tools (endurance, tempo, sweet spot, interval, over/under) by naming the FTP-test protocol.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explains the workout's nature (no power target, free effort, max sustainable pace, average watts recorded) which implies when it is appropriate, but never states when to choose this over get_cycling_ftp or the other cycling workout builders, nor any prerequisite (e.g., needing a power meter).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_cycling_interval_workoutA

Create a power-based VO2max cycling interval workout and upload it to Garmin Connect.

Uses exact watt targets (power.between, ID 6) — the device displays "250-270W" not "Zone X".

Args: name: Workout name (e.g. "VO2max 5x3m") reps: Number of intervals (default 5) work_sec: Duration of each interval in seconds (default 180) rest_sec: Recovery duration between intervals in seconds (default 180) power_low: Lower power target in watts (default 250) power_high: Upper power target in watts (default 270) warmup_min: Warmup duration in minutes (default 15) cooldown_min: Cooldown duration in minutes (default 10)

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYes
repsNo
rest_secNo
work_secNo
power_lowNo
power_highNo
warmup_minNo
cooldown_minNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It usefully discloses that the tool both creates AND uploads to Garmin Connect, and that targets are encoded as power.between (ID 6) so the device shows '250-270W'. It does not cover authentication requirements, rate limits, duplicate-workout behavior, or failure handling on upload — meaningful gaps for an unannotated write tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The action is front-loaded in sentence one, followed by a short implementation rationale and a compact Args list. The Args block duplicates defaults already in the schema, but given 0% schema coverage that redundancy is largely justified.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be described, and the parameter list is fully documented. For a mutation tool with no annotations, the only real gap is operational context (auth, upload failure, scheduling follow-up) rather than anything an agent needs to form a correct call.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% (properties carry only a title), so the description must compensate, and it does: the Args block defines all 8 parameters with units (watts, seconds, minutes), semantics, and defaults that match the schema. This is close to ideal compensation for a zero-coverage schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb, resource, and training modality: create a power-based VO2max cycling interval workout and upload it to Garmin Connect. This is clearly distinguishable from siblings like create_cycling_tempo_workout, create_cycling_sweet_spot_workout, and create_cycling_over_under_workout without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The 'VO2max' and 'power-based' framing implies the intended training context, but there is no explicit when-to-use statement, no exclusion criteria, and no named alternative among the many create_*_workout siblings. The note about exact watt targets implies it is the right choice when the athlete wants absolute watts rather than zones, but that inference is left to the agent.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_cycling_over_under_workoutA

Create an over/under threshold cycling workout and upload it to Garmin Connect.

Alternates between above-threshold (over) and sub-threshold (under) power targets based on FTP percentages.

Args: name: Workout name (e.g. "Over/Under 3x(1m/2m)") reps: Number of over/under pairs (default 3) over_sec: Duration of the over segment in seconds (default 60) under_sec: Duration of the under segment in seconds (default 120) over_pct: FTP percentage for over segment (default 105) under_pct: FTP percentage for under segment (default 90) warmup_min: Warmup duration in minutes (default 15) cooldown_min: Cooldown duration in minutes (default 10)

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYes
repsNo
over_pctNo
over_secNo
under_pctNo
under_secNo
warmup_minNo
cooldown_minNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden. It usefully discloses the two-step side effect (create then upload to Garmin Connect), but says nothing about the FTP prerequisite, authentication needs, error behavior, or whether the created workout is persisted/scheduled. Partial disclosure for a mutation tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded purpose sentence followed by a one-line workout explanation and a clean Args block. Every line earns its place given the 0% schema coverage, though the description is slightly longer than strictly necessary.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need no explanation, and the description adequately covers the create action plus the upload side effect. It is nearly complete for a create tool; the main gap is the undocumented FTP dependency and any failure/permission context, which matters more without annotations.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate, and it largely does: all 8 parameters are named with meaning, units (seconds, minutes, FTP percentage), and defaults. It stops short of giving valid ranges or interactions (e.g., reps vs segment durations), but adds substantial meaning beyond the bare schema titles.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb (create), a specific resource (over/under threshold cycling workout), and a second action (upload to Garmin Connect). The over/under alternation pattern clearly distinguishes it from the many sibling cycling-workout tools (tempo, sweet spot, interval, endurance, FTP test).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The workout concept (alternating over/under FTP targets) implies when an athlete would use it, but there is no explicit when-to-use, when-not-to-use, or routing to alternative workout tools. Usage must be inferred from the training description.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_cycling_sweet_spot_workoutA

Create a sweet spot interval cycling workout and upload it to Garmin Connect.

Repeated sweet spot blocks (Z4) with active recovery between.

Args: name: Workout name (e.g. "Sweet Spot 3x20") reps: Number of sweet spot repeats (default 3) work_min: Duration of each sweet spot block in minutes (default 20) rest_min: Recovery duration between blocks in minutes (default 5) warmup_min: Warmup duration in minutes (default 15) cooldown_min: Cooldown duration in minutes (default 10)

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYes
repsNo
rest_minNo
work_minNo
warmup_minNo
cooldown_minNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description must carry behavioral disclosure. It does state that the workout is uploaded to Garmin Connect, which is an important side effect, and it describes the workout structure and all default durations. However, it omits prerequisites such as authentication, error behavior, and whether an existing workout is overwritten.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The definition is front-loaded with the core action and then gives a compact workout structure summary followed by an Args list. It is appropriately sized, with only minor redundancy in repeating defaults already present in the schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has a moderate six-parameter schema with no description coverage, no annotations, and an output schema. The description compensates for the parameter gap and discloses the Garmin upload, but it lacks guidance about when to use this workout creator versus sibling cycling workout tools and does not mention upload prerequisites.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It documents all six parameters with meanings, units, and defaults, and gives an example format for the workout name. It does not mention accepted ranges or validation constraints, but it substantially closes the schema's semantic gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: create a sweet spot interval cycling workout and upload it to Garmin Connect. It distinguishes this tool from sibling workout creators such as create_cycling_interval_workout and create_cycling_tempo_workout by specifying the sweet spot (Z4) block structure with active recovery.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies the tool's usage through the workout type and structure, but it does not explicitly state when to choose this over alternative cycling workout creators. There are no when-not conditions or named alternatives provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_cycling_tempo_workoutA

Create a tempo/threshold cycling workout and upload it to Garmin Connect.

Same structure as endurance but with a higher HR zone default (Z3).

Args: name: Workout name (e.g. "Tempo 60m") duration_min: Duration of the main tempo block in minutes hr_zone: Target heart-rate zone (Z1-Z5, default Z3) warmup_min: Warmup duration in minutes (default 15) cooldown_min: Cooldown duration in minutes (default 15)

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYes
hr_zoneNoZ3
warmup_minNo
cooldown_minNo
duration_minYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It does disclose the key side effect that the workout is uploaded to Garmin Connect (a remote, non-read-only action), which is more than most siblings offer. It says nothing about authentication requirements, whether the upload is idempotent, whether it lands in the library vs. being scheduled, or failure behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core action and the most decision-relevant fact (upload to Garmin Connect), followed by a compact argument list. No filler sentences; only minor redundancy in restating defaults that also appear in the schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained, and parameter documentation is solid despite 0% schema coverage. What is missing is disambiguation among the five cycling workout variants and any note on the upload's effect on the athlete's Garmin account state.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, and it does: all five parameters are named with meaning, defaults for hr_zone/warmup_min/cooldown_min are restated, and an example value ('Tempo 60m') is given for name. Critically, it supplies the Z1-Z5 range for hr_zone, a constraint absent from the schema (no enums defined).

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Create a tempo/threshold cycling workout') and adds a notable second action, uploading to Garmin Connect. It partially differentiates from siblings by anchoring to the endurance variant ('Same structure as endurance but with a higher HR zone default'), though it does not distinguish itself from the sweet-spot, interval, or over-under cycling workouts.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'Same structure as endurance but with a higher HR zone default (Z3)' gives an implicit signal for choosing this over create_cycling_endurance_workout. However, there is no explicit when-to-use/when-not-to-use guidance and no mention of how it differs from the other cycling workout creators.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_manual_activityA

Log a manual activity in Garmin Connect — useful for activities done without a watch.

The type_key must match a Garmin activity type. Use get_activity_types to see the full list. Common values: yoga, strength_training, meditation, indoor_cycling, pilates, bouldering, fitness_equipment.

Args: type_key: Activity type key (e.g. "yoga", "strength_training") date: Date of the activity in YYYY-MM-DD format duration_minutes: Duration of the activity in minutes start_time: Start time as HH:MM (24-hour, default 09:00) activity_name: Optional title; defaults to the type_key if not provided distance_km: Distance in kilometres (default 0.0 for non-distance activities) time_zone: IANA time zone for the activity (default UTC)

ParametersJSON Schema
NameRequiredDescriptionDefault
dateYes
type_keyYes
time_zoneNoUTC
start_timeNo09:00
distance_kmNo
activity_nameNo
duration_minutesYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden of behavioral disclosure. It only says 'Log' and lists input constraints; it does not mention side effects, permissions, reversibility, validation failures, or whether it overwrites anything. For a write operation this is a significant gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with the purpose, followed by a useful type_key note and a compact Args list. It repeats a few defaults already in the schema, but the extra format details earn their place. No filler or fluff.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 7-parameter write tool with no annotations, the description covers the parameter space well and directs agents to the right reference for valid type_key values. It avoids explaining return values because an output schema exists. Missing error-handling details, but the essentials are present.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, but the description compensates thoroughly: it documents every parameter with formats (YYYY-MM-DD, HH:MM), defaults (09:00, UTC, 0.0, default to type_key), and a meaningful controlled vocabulary for type_key, including a pointer to get_activity_types. This adds exactly the meaning the schema lacks.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: 'Log a manual activity in Garmin Connect'. The phrase 'manual activity' and 'without a watch' distinguishes it from workout-creation and upload tools like create_strength_workout or upload_workout. Clear intent, no ambiguity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly gives the primary use case ('useful for activities done without a watch') and instructs the agent to use get_activity_types for valid type_key values. It does not explicitly name alternatives or exclusions, but the context is sufficient for tool routing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_run_easy_workoutA

Create an easy aerobic run workout and upload it to Garmin Connect.

Simpler than create_run_workout — defaults to Z2 with shorter warmup/cooldown.

Args: name: Workout name (e.g. "Easy Run 45m") duration_min: Duration of the main run in minutes hr_zone: Target heart-rate zone (Z1-Z5, default Z2) warmup_min: Warmup duration in minutes (default 10) cooldown_min: Cooldown duration in minutes (default 10)

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYes
hr_zoneNoZ2
warmup_minNo
cooldown_minNo
duration_minYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It usefully discloses that the workout is uploaded to Garmin Connect (an external write side effect), but says nothing about authentication requirements, whether re-running creates duplicates, or how failures behave. Adequate but thin for a mutating tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the action and the sibling comparison in the first two sentences, followed by a compact args list. No filler; every line adds information an agent needs.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained, and parameter defaults are covered. For a write tool with no annotations, the remaining gap is operational context such as auth requirements and duplicate-handling, but the core information needed to invoke it correctly is present.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, and it does: all five parameters are documented with meaning, units (minutes), and defaults (Z2, 10/10), including the Z1-Z5 range for hr_zone. Only the exact duration constraints and name formatting rules are left implicit.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Create an easy aerobic run workout') plus the side effect ('upload it to Garmin Connect'). It explicitly positions itself against the sibling create_run_workout, so an agent can distinguish the two without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Names the alternative (create_run_workout) and gives the selecting condition: this one defaults to Z2 with shorter warmup/cooldown, i.e. use it for simpler easy runs. It does not state exclusions or what to do if the user wants a specific zone other than Z2, so it falls short of a full 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_run_hills_workoutA

Create a hill repeat running workout and upload it to Garmin Connect.

Hill reps with jog-down recovery. No target — effort-based.

Args: name: Workout name (e.g. "Hills 8x1m") reps: Number of hill repeats (default 8) hill_sec: Duration of each hill effort in seconds (default 60) jog_down_sec: Jog-down recovery duration in seconds (default 90) warmup_min: Warmup duration in minutes (default 15) cooldown_min: Cooldown duration in minutes (default 10)

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYes
repsNo
hill_secNo
warmup_minNo
cooldown_minNo
jog_down_secNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It usefully discloses the key side effect ('upload it to Garmin Connect') and the effort-based (non-target) nature of the workout, but says nothing about authentication needs, whether an existing same-named workout is overwritten, error behavior, or whether uploading also schedules it.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the action and side effect, then uses a compact structured Args block. Every line earns its place; only minor trimming is possible.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a creation tool with an output schema (so return values need no explanation), the description covers purpose, workout structure, and all parameters. Gaps around upload side effects (overwrite/schedule behavior) and permissions keep it from being fully complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate — and it does, documenting all six parameters with meaning, unit (seconds vs minutes), and defaults. It falls short of full marks only because it omits valid ranges/constraints for values like reps or hill_sec.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Create a hill repeat running workout') and additionally specifies the structure ('Hill reps with jog-down recovery') and the target mode ('No target — effort-based'). This clearly distinguishes it from siblings like create_run_intervals_workout or create_run_tempo_workout, which are paced/target-based.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The 'No target — effort-based' phrasing implies when this workout is appropriate versus target-based siblings, but no explicit when-to-use, prerequisites, or named alternatives are given. Usage is inferable rather than stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_run_intervals_workoutA

Create a track-style running interval workout and upload it to Garmin Connect.

Uses distance-based end condition for the work intervals (e.g. 400m).

Args: name: Workout name (e.g. "Track 6x400m") reps: Number of intervals (default 6) distance_m: Distance per interval in meters (default 400) rest_sec: Recovery duration between intervals in seconds (default 120) hr_zone: Target heart-rate zone for intervals (Z1-Z5, default Z5) warmup_min: Warmup duration in minutes (default 10) cooldown_min: Cooldown duration in minutes (default 10)

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYes
repsNo
hr_zoneNoZ5
rest_secNo
distance_mNo
warmup_minNo
cooldown_minNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It does disclose that the tool creates a workout and uploads it to Garmin Connect, which is important side-effect information, but it omits permissions, reversibility, rate limits, and error behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded purpose statement followed by a clean Args list; no filler. Could be slightly tighter by merging defaults, but it is well-structured and easy to scan.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given seven parameters, no annotations, and an output schema that handles return values, the description covers purpose and all parameter semantics. It is missing explicit usage context and deeper behavioral details, but is otherwise complete enough to invoke correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, and the description fully compensates by documenting all seven parameters with meaning, units, defaults, and examples (e.g., distance_m in meters, hr_zone Z1-Z5). This adds substantial value beyond the bare schema titles.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Specific verb 'Create' and resource 'track-style running interval workout', with scope 'upload to Garmin Connect' and a distinguishing detail about distance-based work intervals. Clearly separates it from sibling run workout creators like create_run_easy_workout or create_run_tempo_workout.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The workout type implies when to use it (track-style intervals with distance-based ends), but there is no explicit when-to-use guidance, no mention of alternatives, and no exclusions. An agent must infer the context from the name and description.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_run_long_workoutA

Create a long run with custom BPM range and upload it to Garmin Connect.

Uses targetValueOne/Two for a custom HR band — the device displays "130-145 bpm" instead of "Zone X". This fixes the old encoder's silent-drop bug on heart.rate custom ranges.

Args: name: Workout name (e.g. "Long Run 90m") duration_min: Duration of the long run in minutes hr_min: Lower HR target in bpm (default 130) hr_max: Upper HR target in bpm (default 145) warmup_min: Warmup duration in minutes (default 10) cooldown_min: Cooldown duration in minutes (default 10)

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYes
hr_maxNo
hr_minNo
warmup_minNo
cooldown_minNo
duration_minYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden. It does add real value by disclosing the upload side effect and the custom HR-band encoding detail (targetValueOne/Two, device showing '130-145 bpm' instead of 'Zone X'). But it omits auth/permission requirements, whether the upload is immediate or reversible, and failure behavior for a mutating, remote-upload tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose is front-loaded in the first sentence, followed by the differentiating HR-band detail and a clean Args list. The encoder-bug sentence is implementation trivia but earns some place by explaining why the tool exists; overall it is tight and skimmable.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained, and the description covers all parameters plus the notable upload side effect. The main remaining gap is the absence of any sibling differentiation, which matters given how crowded the create_*_workout family is.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate — and it does: the Args section documents all six parameters with units and defaults (hr_min 130, hr_max 145, warmup_min 10, cooldown_min 10), and the HR-band paragraph explains what hr_min/hr_max actually produce on-device. It stops short of stating valid ranges or constraints.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource ('Create a long run ... upload it to Garmin Connect') and the opening sentence makes the outcome clear. However, it never distinguishes this from the many sibling workout creators (create_run_easy_workout, create_run_tempo_workout, create_run_intervals_workout, create_run_workout, etc.), so an agent must infer the distinction from the name alone.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is only implied by the word 'long run' — the description gives no explicit when-to-use versus when-not, and does not route to any of the ~10 similar run/workout creation siblings. There is no prerequisite or exclusion guidance, so the agent must infer context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_run_progression_workoutA

Create a progressive run workout and upload it to Garmin Connect.

Sequential blocks (e.g. Z2→Z3→Z4) without recovery between them. The athlete progresses through increasing intensity zones.

Args: name: Workout name (e.g. "Progression Z2-Z3-Z4") blocks: List of dicts with keys: - duration_sec (int): duration in seconds - hr_zone (str): target HR zone (Z1-Z5) Example: [{"duration_sec": 900, "hr_zone": "Z2"}, {"duration_sec": 900, "hr_zone": "Z3"}, {"duration_sec": 900, "hr_zone": "Z4"}] warmup_min: Warmup duration in minutes (default 15) cooldown_min: Cooldown duration in minutes (default 10)

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYes
blocksYes
warmup_minNo
cooldown_minNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the burden. It usefully discloses the side effect that the workout is uploaded to Garmin Connect, which is meaningful for a mutation. However it says nothing about permissions, whether an existing workout is overwritten, or upload failure behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose is front-loaded, then a compact conceptual explanation, then a well-organized Args section with a worked example. The two sentences about sequential blocks and increasing zones overlap slightly, but nothing is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained, and the description covers the create-and-upload behavior and all four parameters. The only gap is the absence of any mention of prerequisites or how the upload interacts with existing workouts.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must carry everything — and it does: it documents name, the nested block keys (duration_sec as int seconds, hr_zone as Z1-Z5), a concrete three-block example, and the defaults for warmup_min and cooldown_min. This fully compensates for the undocumented schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Create a progressive run workout and upload it to Garmin Connect') and defines the structure (sequential blocks with increasing zones, no recovery), which distinguishes it conceptually from interval-style siblings. It never names an alternative sibling (e.g. create_run_intervals_workout) to make the routing explicit, so it stops short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The definition of a progression (sequential blocks, no recovery, increasing intensity zones) implicitly tells the agent when this tool fits, but there is no explicit when/when-not guidance and no sibling named as an alternative. Usage is inferable but not stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_run_tempo_workoutA

Create a tempo/threshold run workout and upload it to Garmin Connect.

Same structure as easy run but with a higher HR zone default (Z4).

Args: name: Workout name (e.g. "Tempo Run 30m") duration_min: Duration of the tempo block in minutes hr_zone: Target heart-rate zone (Z1-Z5, default Z4) warmup_min: Warmup duration in minutes (default 10) cooldown_min: Cooldown duration in minutes (default 10)

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYes
hr_zoneNoZ4
warmup_minNo
cooldown_minNo
duration_minYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It usefully discloses that it uploads a new artifact to Garmin Connect and lists the defaults applied when args are omitted, but it says nothing about authentication/permission needs, idempotency, or what happens on a duplicate name.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose is front-loaded in the first sentence, followed by a compact, well-labeled Args block. Every line carries information; the only mild redundancy is restating the Z4 default already implied by the annotation-free schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need no explanation, and the description covers purpose, all parameters, and defaults for a creation/upload tool. The main gap is the absence of any behavioral caveats (permissions, upload failures) in the absence of annotations.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 0% schema description coverage the description must compensate, and it documents all five parameters individually, including an example value for name and the Z1-Z5 range for hr_zone (an enum the schema lacks). It adds real meaning beyond the bare parameter titles, though it doesn't clarify units or interaction between duration_min and the tempo block.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Specific verb ('Create') plus precise resource ('tempo/threshold run workout') and an explicit side effect ('upload it to Garmin Connect'). It even differentiates from the sibling create_run_easy_workout by noting the higher HR zone default, so an agent can separate it from the other run-workout creators.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'tempo/threshold run' implies when this is appropriate, and the comparison to 'easy run' hints at the alternative, but it never states when to choose this over create_run_easy_workout, create_run_intervals_workout, or the other run-workout siblings. Usage is left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_run_workoutA

Create a continuous run workout and upload it to Garmin Connect.

Builds a single uninterrupted run interval with warmup and cooldown walks.

Targets a named Garmin heart-rate zone by default. Named zones (Z1-Z5) don't line up with every real training target -- e.g. a 136-148 bpm Zone 2 goal straddles Garmin's Z2 (118-137) and Z3 (138-157). Pass hr_min and hr_max together to target that exact bpm range instead; the watch will then show "in range" only for the range you actually want, not a whole zone that over- or under-shoots it.

Args: name: Workout name (e.g. "Step 8 - 30min continuous") run_seconds: Duration of the run in seconds warmup_min: Warmup walk duration in minutes cooldown_min: Cooldown walk duration in minutes hr_zone: Target heart-rate zone (Z1-Z5, default Z3). Ignored if hr_min/hr_max are given. hr_min: Optional custom target heart rate range, minimum bpm (must be given with hr_max) hr_max: Optional custom target heart rate range, maximum bpm (must be given with hr_min)

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYes
hr_maxNo
hr_minNo
hr_zoneNoZ3
warmup_minYes
run_secondsYes
cooldown_minYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden of transparency. It discloses the primary side effect (uploading to Garmin Connect) and the behavior of heart rate targeting with both named zones and custom ranges. It does not mention failure modes or reversibility, but for a creation operation this level of transparency is reasonably complete.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-structured with a headline, an explanatory paragraph, and a parameter list. The heart rate zone example is detailed and valuable but adds a bit of length; overall, it is concise and every sentence serves a purpose without fluff.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity and the existing output schema (though not displayed), the description adequately covers the input semantics and the primary action. It does not discuss return values, but since an output schema is indicated, that is acceptable. No prerequisites or error conditions are mentioned, but they are not essential for this creation operation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The description explicitly explains every parameter in a bulleted list, including the nuanced behavior of hr_min and hr_max (must be given together, overrides hr_zone) and the default value of hr_zone. This adds substantial meaning beyond the raw schema, which only lists types and defaults, ensuring the agent understands exactly how to supply each argument.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool creates a continuous run workout and uploads it to Garmin Connect, with a detailed explanation of a single uninterrupted run interval. It does not explicitly contrast with sibling workout creation tools like create_walk_run_workout or create_strength_workout, but the 'continuous run' phrasing and the example parameter values make the purpose unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear guidance on when to use hr_zone versus hr_min/hr_max, explaining the pitfall of named zones and the requirement to pass both min and max together. However, it does not explicitly state when to choose this tool over alternative workout creation sibling tools, leaving some ambiguity about the broader selection criteria.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_strength_workoutA

Create a strength workout and upload it to Garmin Connect.

Each exercise becomes a reps-based step. The name is kept in the step description; it is also sent as exerciseName, which Garmin only retains when it matches one of its own exercise keys (e.g. "FARMERS_CARRY").

Args: name: Workout name exercises: List of dicts with keys: name, sets, reps, rest_seconds and an optional category. Category is omitted from the payload when not given; Garmin accepts that. When given it must be one of Garmin's exercise categories (e.g. SQUAT, DEADLIFT, PUSH_UP, CARRY, SLED) — anything else, including "UNASSIGNED" and "OTHER", is rejected with 400 Invalid category. Full list: https://connect.garmin.com/web-data/exercises/Exercises.json

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYes
exercisesYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral disclosure burden. It covers the create+upload side effect, the conversion of exercises to reps-based steps, how exerciseName is only retained when it matches Garmin's keys, and that invalid categories cause a 400 rejection. It also links to the authoritative category list, which is a strong transparency signal.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The purpose is front-loaded in the first sentence, followed by concise behavioral details and an Args block. Every sentence adds necessary information, and the long allowed-category list is linked rather than inlined. The length is justified by the number of constraints the agent must handle.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 2-parameter tool with no annotations and a sparse schema, the description is the only guide to correct invocation. It fully specifies the payload shape, category validation, and Garmin's exerciseName behavior. The presence of an output schema means return values need no description, and nothing critical appears missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% and the schema only defines 'name' and an opaque 'exercises' array. The description compensates fully by specifying the required dict keys (name, sets, reps, rest_seconds, optional category), explaining category omission and allowed values, and documenting the failure mode. This adds essential meaning beyond the raw schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource ('Create a strength workout') and immediately adds the distinctive behavior 'upload it to Garmin Connect' plus 'Each exercise becomes a reps-based step.' This clearly differentiates it from siblings like create_run_workout or create_walk_run_workout, even without naming them.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides detailed parameter constraints but never tells an agent when to prefer this tool over closely related siblings such as create_run_workout, create_walk_run_workout, or upload_workout. No alternatives or exclusions are mentioned, so an agent must infer tool selection entirely from the name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_swim_drills_workoutA

Create a swim drill set and upload it to Garmin Connect.

Each drill step includes name, distance, equipment, and stroke.

Args: name: Workout name (e.g. "Drills 1200m") drills: List of dicts with keys: - name (str): drill name (e.g. "Catch-up") - distance_m (int): drill distance - equipment (str): "none", "fins", "paddles", "pull_buoy", "kickboard" - stroke (str): stroke type warmup_m: Warmup distance in meters (default 300) cooldown_m: Cooldown distance in meters (default 200) pool_length: Pool length in meters (default 25)

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYes
drillsYes
warmup_mNo
cooldown_mNo
pool_lengthNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It does disclose a meaningful behavioral trait beyond the schema: the set is created and then uploaded to Garmin Connect, implying remote persistence. It omits authentication/scope requirements, whether the workout is auto-scheduled, and whether a failed upload leaves a partial state.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose is front-loaded in the first sentence and the Args block is compact and skimmable. The parameter documentation is longer than usual but justified because the schema supplies no field descriptions.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be described, and the description covers the creation-plus-upload behavior plus every input including nested drill keys. Nothing critical is missing for an agent to invoke this correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the nested 'drills' items schema is a bare object with additionalProperties, so the description is doing the heavy lifting. It documents all five parameters, the per-drill keys, defaults for warmup_m/cooldown_m/pool_length, and the accepted equipment vocabulary, compensating for the schema gap; only 'stroke (str): stroke type' remains vague.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Create a swim drill set') plus the side effect ('upload it to Garmin Connect'), which is more than a name restatement. However, it does not distinguish itself from the several sibling swim-workout creators (create_swim_endurance_workout, create_swim_intervals_workout, create_swim_threshold_workout), so an agent cannot choose among them from the description alone.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no statement of prerequisites, and no mention of the alternatives despite the crowded set of sibling swim/run/cycling workout creators. Usage is only implied by the tool name and the drill-oriented parameters.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_swim_endurance_workoutA

Create a continuous swim workout with pace target and upload it to Garmin Connect.

Args: name: Workout name (e.g. "Endurance Swim 1500m") distance_m: Main set distance in meters (default 1500) pace: Target pace as "M:SS/100m" (default "1:45/100m") stroke: Stroke type — freestyle, backstroke, breaststroke, butterfly warmup_m: Warmup distance in meters (default 400) cooldown_m: Cooldown distance in meters (default 200) pool_length: Pool length in meters — 25 or 50 (default 25)

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYes
paceNo1:45/100m
strokeNofreestyle
warmup_mNo
cooldown_mNo
distance_mNo
pool_lengthNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It usefully discloses a side effect beyond the schema — that the workout is uploaded to Garmin Connect — but says nothing about permissions, failure behavior, whether an existing workout is overwritten, or other operational constraints.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose is front-loaded in the first sentence and the per-parameter list is scannable. It is slightly verbose but every line carries needed information, and nothing appears to be padding.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given seven parameters, no annotations, and an existing output schema (so return values need not be explained), the definition is nearly complete: purpose, side effect, and all parameter semantics are covered. The only gap is the lack of any routing guidance relative to its many workout-creation siblings.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 0% schema description coverage, the description must compensate fully, and it does: it documents all seven parameters, gives the pace format ("M:SS/100m"), enumerates valid stroke types, and restricts pool_length to 25 or 50. This adds substantial meaning beyond the bare schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (create), resource (swim workout), and key attribute (continuous, with pace target), which implicitly separates it from the swim_intervals sibling. However, it does not explicitly distinguish itself from create_swim_threshold_workout or create_swim_drills_workout, so an agent still has to reason about the boundary.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit when-to-use or when-not-to-use statement and no named alternative among the several create_swim_* siblings. The word 'continuous' hints at the distinction from intervals, but this is only implied and left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_swim_intervals_workoutB

Create a CSS threshold swim interval workout and upload it to Garmin Connect.

Args: name: Workout name (e.g. "CSS 4x200m") reps: Number of intervals (default 4) distance_m: Distance per interval in meters (default 200) rest_sec: Rest between intervals in seconds (default 30) pace: Target pace as "M:SS/100m" (default "1:40/100m") stroke: Stroke type warmup_m: Warmup distance in meters (default 400) cooldown_m: Cooldown distance in meters (default 200) pool_length: Pool length in meters (default 25)

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYes
paceNo1:40/100m
repsNo
strokeNofreestyle
rest_secNo
warmup_mNo
cooldown_mNo
distance_mNo
pool_lengthNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It does disclose the key side effect — the workout is uploaded to Garmin Connect — which is genuinely useful behavioral information. However, it says nothing about auth requirements, whether the workout is only uploaded or also scheduled, failure behavior, or whether it overwrites anything.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The purpose sentence is front-loaded and followed by an ordered arg list that reads quickly. Mild redundancy in restating defaults already present in the schema, but the added units and pace format justify most of the length.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained, and the description covers the creation-plus-upload behavior and all parameter meanings. It falls short only on the missing differentiation from sibling swim-workout creators and on failure/permission behavior for an unannotated mutation tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% (properties have titles only), so the Args block does the heavy lifting and does it reasonably well: it supplies units for every distance parameter, the 'M:SS/100m' format for pace, and an example value for name. Stroke type is named but its allowed values are not enumerated, which is the main gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The first sentence gives a concrete verb+resource ('Create a CSS threshold swim interval workout and upload it to Garmin Connect'), which is more specific than the bare name. However, it never differentiates itself from the very similar sibling 'create_swim_threshold_workout' (nor 'create_swim_endurance_workout' / 'create_swim_drills_workout'), leaving the agent to guess which swim workout creator to pick.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use, when-not-to-use, or alternative routing guidance. Given four sibling swim-workout creators, the description should say what makes this interval/CSS variant the right choice, but it offers nothing beyond the parameter list.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_swim_threshold_workoutA

Create a single threshold swim set and upload it to Garmin Connect.

Good for CSS testing — one continuous threshold block after warmup.

Args: name: Workout name (e.g. "CSS Test 800m") distance_m: Threshold set distance in meters (default 800) pace: Target pace as "M:SS/100m" stroke: Stroke type warmup_m: Warmup distance in meters (default 400) cooldown_m: Cooldown distance in meters (default 200) pool_length: Pool length in meters (default 25)

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYes
paceNo1:42/100m
strokeNofreestyle
warmup_mNo
cooldown_mNo
distance_mNo
pool_lengthNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden, and it does disclose the key side effect: the workout is uploaded to Garmin Connect, i.e. a remote resource is created. However, it says nothing about auth requirements, whether re-uploading duplicates, or failure behavior for a mutation tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the action and the use case in two short sentences, then a structured Args list. The Args block largely restates schema fields but is justified by the 0% schema description coverage; no filler sentences.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need no explanation, and all 7 parameters are covered. For an unannotated mutation tool the remaining gap is error/auth/idempotency behavior, but the creation-and-upload semantics are clear enough to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, and it documents all 7 parameters with meaning plus format guidance ('pace as "M:SS/100m"', name example 'CSS Test 800m'). This substantially exceeds what the bare schema titles provide, though it doesn't explain unit/validity constraints for pool_length or stroke values.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource: 'Create a single threshold swim set and upload it to Garmin Connect,' with the scope narrowed by 'single' and 'one continuous threshold block after warmup.' This implicitly separates it from create_swim_endurance_workout, create_swim_intervals_workout, and create_swim_drills_workout, though it never names those siblings explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Good for CSS testing' gives an intended use case, which is more than nothing, but there is no explicit when-to-use/when-not guidance and no routing to the endurance, intervals, or drills swim-workout siblings. Usage is implied rather than stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_walk_run_workoutA

Create a walk/run interval workout and upload it to Garmin Connect.

Builds the internal Garmin JSON automatically and returns the new workout ID.

Args: name: Workout name (e.g. "W3 Mié 2:2") run_seconds: Duration of each run interval in seconds walk_seconds: Duration of each walk/recovery interval in seconds repeats: Number of run/walk repetitions warmup_min: Warmup duration in minutes cooldown_min: Cooldown duration in minutes hr_zone: Target heart-rate zone (Z1-Z5, default Z3)

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYes
hr_zoneNoZ3
repeatsYes
warmup_minYes
run_secondsYes
cooldown_minYes
walk_secondsYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of disclosing behavior. It clearly states that the tool uploads to Garmin Connect and returns a new workout ID, which are the primary side effects and return value. It does not mention potential failure modes or authorization requirements, but the main behavioral traits are transparent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise, consisting of two sentences that convey the purpose and key outputs, followed by a parameter list that adds necessary semantics. No redundant or filler content is present. The structure is clear and easy to scan.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's moderate complexity (7 parameters) and the existence of a return value (workout ID), the description covers the essential purpose and parameters. It does not describe the output schema, but the returned ID is sufficient for many use cases. It lacks details on error handling or prerequisites, but within the context of sibling tools, it is reasonably complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema provides only titles and types with zero descriptions, so the parameter explanations in the description are essential. Each parameter is given a brief semantic definition (e.g., 'Duration of each run interval in seconds', 'Target heart-rate zone (Z1-Z5, default Z3)'), covering all seven parameters. It lacks constraints like minimum values or exact allowed enum values for hr_zone, but the core meaning is clear.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb 'Create' and the resource 'walk/run interval workout', and specifies that it uploads to Garmin Connect. The mention of 'walk/run interval' distinguishes it from sibling tools like create_run_workout and create_z2_walk_workout, making the purpose unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description does not provide any explicit guidance on when to use this tool versus the alternative creation tools (e.g., create_run_workout, create_z2_walk_workout) or upload tools. There is no mention of 'use this when...' or 'not for...', leaving the selection criteria implicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_weekly_planA

Create and schedule a full week of workouts from a plan file.

Parses a JSON/YAML plan file, maps each entry to the appropriate workout builder, uploads all workouts, schedules each on its own date, and returns a summary.

Plan file format (JSON): [ { "date": "2026-07-13", "sport": "cycling", "type": "endurance", "name": "Z2 Endurance 90m", "duration_min": 90, "hr_zone": "Z2" } ]

Every entry needs its own date; an entry without one is uploaded but reported with schedule_error rather than scheduled.

Args: plan_yaml_path: Absolute path to the plan JSON/YAML file

ParametersJSON Schema
NameRequiredDescriptionDefault
plan_yaml_pathYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden and does well: it discloses the multi-step side effects (uploads AND schedules all workouts), the required per-entry date, and the non-obvious failure mode where a missing date still uploads but reports schedule_error. It lacks any auth/permission or idempotency note, keeping it short of a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core action, then structured into behavior, format example, and edge case. The embedded JSON sample consumes space but earns it by pinning down the file format for a 0%-coverage parameter. Slightly long, but no filler sentences.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so the description needn't detail return values, and it only briefly notes it 'returns a summary.' Combined with thorough format and edge-case documentation, it is complete enough for an agent to call correctly; only the absence of sibling routing and auth details keeps it from 5.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% and the lone parameter has no schema description, so the description must compensate — and it does, defining plan_yaml_path as an absolute path to a JSON/YAML file and specifying the exact expected entry structure. This is far more than the bare schema conveys.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb+resource: 'Create and schedule a full week of workouts from a plan file,' and even enumerates the internal steps (parse, map, upload, schedule). The purpose is unmistakable and reads as a high-level orchestrator. It does not, however, explicitly differentiate itself from close siblings like schedule_week or upload_workouts, so it stops short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied by the plan-file input: an agent can infer this is the bulk/week-batch entry point rather than the single-workout creators. But there is no explicit when-to-use/when-not guidance and no alternative named (e.g. schedule_week or the individual create_*_workout tools), leaving the routing decision to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_z2_walk_workoutA

Create a steady Z2 walking workout and upload it to Garmin Connect.

Args: name: Workout name duration_min: Main walking block duration in minutes hr_min: Minimum heart rate in bpm (used for description; target is Z2) hr_max: Maximum heart rate in bpm (used for description; target is Z2)

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYes
hr_maxYes
hr_minYes
duration_minYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden, and it does disclose the key external side effect: upload to Garmin Connect. It also clarifies that hr_min and hr_max are used only for the description rather than setting the target, which prevents an incorrect assumption. However, it does not address duplicate names, overwrite behavior, or required permissions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two compact sections: a front-loaded purpose statement followed by a terse argument list. Every sentence contributes information, with no filler or repetition of the schema titles.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers the operation, all required parameters, and the key side effect, and an output schema exists so return values need not be specified. It is only slightly incomplete in not addressing when to prefer this over sibling workout-creation or upload tools and in not documenting edge cases such as duplicate workout names.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the parameter semantics fall entirely on the description, and the description documents all four required arguments with units and purpose. It usefully explains that hr_min and hr_max are descriptive only while the target is Z2, adding meaning the schema does not provide.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: create a steady Z2 walking workout and upload it to Garmin Connect. The Z2 walking scope differentiates it from sibling tools like create_run_workout, create_strength_workout, and the more generic create_walk_run_workout.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The use case (steady Z2 walking workout plus upload) is implied by the first sentence, but there is no explicit when-to-use guidance or exclusions. An agent is not told when to choose this tool over upload_workout, create_walk_run_workout, or create_run_workout.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_activityA

Permanently delete an activity from Garmin Connect.

This cannot be undone — the activity and its recorded data are gone. Intended for removing a manual activity logged in error; think twice before pointing it at a recorded session.

Exists because create_manual_activity could write a record the server had no way to remove (§D9k).

Args: activity_id: ID of the activity to delete

ParametersJSON Schema
NameRequiredDescriptionDefault
activity_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full behavioral burden and does well: it states the deletion is permanent, cannot be undone, and takes the activity and its recorded data with it. It does not mention authentication requirements or any server-side restrictions, but the core destructive profile is clearly disclosed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The permanent-deletion warning is front-loaded and the guidance is efficient. The paragraph explaining why the tool exists ('Exists because create_manual_activity could write a record the server had no way to remove (§D9k)') is backstory that does not help an agent select or invoke the tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive single-parameter tool with no annotations, the description covers purpose, irreversibility, data loss, and intended use. An output schema exists so return values need not be explained, but permission prerequisites and whether recorded activities are technically deletable are left implicit.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate. It adds 'ID of the activity to delete,' which gives the parameter's meaning, but the schema allows either integer or string and the description does not clarify which form is expected or provide examples.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: 'Permanently delete an activity from Garmin Connect.' It clearly distinguishes itself from sibling readers like get_activity and sibling mutators like set_activity_name by naming the destructive action and target resource.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly says it is 'Intended for removing a manual activity logged in error' and warns to 'think twice before pointing it at a recorded session.' This gives clear when-to-use and when-to-be-cautious guidance, though it does not name an alternative tool for any edge case.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_blood_pressureA

Delete a single blood pressure measurement.

set_blood_pressure returns the date and version of the reading it wrote; pass both back here to remove it. Without this the server could create a record it had no way to delete (§D9k).

Args: date: Date of the measurement in YYYY-MM-DD format version: The measurement's version, as returned by set_blood_pressure or get_blood_pressure

ParametersJSON Schema
NameRequiredDescriptionDefault
dateYes
versionYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden, and it does disclose the destructive nature plus the version-matching behavior (the version must correspond to the reading being removed, an optimistic-concurrency constraint not visible in the schema). It stops short of stating irreversibility, permission requirements, or failure modes on a version mismatch.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The core instruction is front-loaded in the first two sentences, and the Args block is tight and useful. The parenthetical rationale about the server needing a delete path (§D9k) is meta-commentary that adds little for an agent choosing a tool, slightly diluting the otherwise efficient text.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Two required parameters are both documented with formats and provenance, an output schema exists so return values need no explanation, and the workflow dependency on set_blood_pressure is spelled out. Nothing needed to invoke this correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, and it does: 'date' is specified as YYYY-MM-DD and 'version' is defined as the value returned by set_blood_pressure or get_blood_pressure, which is essential given the schema's loose integer-or-string anyOf.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Delete') and resource ('a single blood pressure measurement'), and explicitly scopes to one reading rather than a bulk operation. It also names the sibling that produces the required identifiers (set_blood_pressure), making it distinguishable from get_blood_pressure and set_blood_pressure without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear workflow context: call after set_blood_pressure and pass back the date and version it returned. It does not state explicit exclusions (e.g. what happens if the version is stale or whether deletion is permanent), but the 'when' is unambiguous for the primary path.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_courseA

Delete a course from Garmin Connect.

Args: course_id: ID of the course to delete (get IDs from get_courses).

ParametersJSON Schema
NameRequiredDescriptionDefault
course_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description carries the full behavioral burden because no annotations are provided. It identifies the action as destructive, but does not disclose whether the deletion is permanent or irreversible, what dependent data may be affected, whether authorization is required, or how invalid IDs are handled. For a destructive tool with zero annotation coverage, this is a significant gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is very compact and well-structured: the first sentence states the operation, and the second line defines the sole parameter and how to source it. Every sentence earns its place; there is no redundancy or filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter destructive tool with an output schema present, the description supplies the essential call information: what to delete and how to get a valid ID. It lacks explicit irreversibility or edge-case caveats, but the overall operation is simple and the output schema covers return shape.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, and the description meaningfully compensates by explaining that `course_id` is the ID of the course to delete and by directing the agent to `get_courses` as the reliable source of valid IDs. This adds real value beyond the bare integer field.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb-resource pair ('Delete a course') and identifies the exact domain ('Garmin Connect'), making the intended tool unambiguous even among siblings like `delete_workout`, `delete_weigh_ins`, and `upload_course`. No other plausible interpretation exists.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives a genuine usage hint by telling the agent to obtain course IDs from `get_courses`, which is a useful prerequisite. However, it does not explicitly say when to use this tool versus alternatives or state any exclusions. The usage context remains mostly implied rather than explicitly spelled out.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_custom_foodA

Delete a custom food from the user's Garmin nutrition library

Permanently removes a custom food entry. The food must not be actively referenced in a logged meal to be deleted. Use get_custom_foods to find the foodId.

Args: food_id: ID of the custom food to delete — a 32-char hex string (from get_custom_foods or create_custom_food)

ParametersJSON Schema
NameRequiredDescriptionDefault
food_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations present, the description carries the full burden and does disclose the key behavioral facts: the deletion is permanent, and there is a referential-integrity guard against foods used in logged meals. It does not say what happens on failure (error vs silent no-op) or whether the caller needs any special permission, which leaves a modest gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose and precondition are front-loaded, and the prose is tight. The trailing 'Args:' block is partly redundant with the schema but earns its place by supplying the hex-string format and provenance that the schema lacks.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need no explanation, and for a single-parameter mutation the description covers the destructive nature and the blocking precondition. Remaining omissions are error semantics and auth expectations, which are secondary for an agent deciding whether and how to call it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% and the schema itself only offers a bare anyOf(integer|string) with no description, so the description must compensate and largely does: it defines food_id as a 32-char hex string and names the two tools that produce one. Note a minor tension — the schema also permits an integer, while the description implies a hex string is expected.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The first sentence states a specific verb and resource ('Delete a custom food from the user's Garmin nutrition library'), scoping it to the custom-food library rather than food logs or workouts. This is enough to distinguish it from delete_food_log, delete_workout, and update_custom_food without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives a clear workflow pointer ('Use get_custom_foods to find the foodId') and an explicit precondition ('must not be actively referenced in a logged meal'), which functions as a when-not rule. It stops short of naming a sibling alternative for removing logged entries (delete_food_log), so routing between the two delete tools still requires inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_food_logA

Delete a food log entry

Permanently removes a logged food item from the nutrition log. Works for both QUICK_ADD and REGULAR_LOG entry types. Use get_nutrition_daily_food_log to find the logId and date.

Args: log_id: Log entry ID to delete — a 32-char hex UUID (from get_nutrition_daily_food_log) meal_date: Date of the log entry in YYYY-MM-DD format

ParametersJSON Schema
NameRequiredDescriptionDefault
log_idYes
meal_dateYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden, and it delivers the most safety-critical fact: the removal is permanent (irreversible). It also notes the entry-type scope, which helps interpret effects. It omits auth/permission needs and error behavior, keeping it below 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the destructive action, then scope, then sourcing guidance, then per-argument detail. Every sentence adds information; the Args section cleanly mirrors the two parameters without waste.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 2-param destructive tool with an output schema (so return values needn't be explained), the description covers purpose, irreversibility, entry-type scope, ID sourcing, and both parameter formats. It's complete enough to call correctly, only missing edge/error handling nuance.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% and the schema only offers a bare anyOf(integer,string) for log_id, so the description must compensate. It does: log_id is a 32-char hex UUID sourced from get_nutrition_daily_food_log, and meal_date is YYYY-MM-DD. This adds substantial meaning the schema lacks.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Delete) and resource (food log entry / logged food item from the nutrition log), and clarifies it covers both QUICK_ADD and REGULAR_LOG entry types. This reasonably distinguishes it from the sibling delete_food (custom foods) and delete_activity, though it doesn't name those siblings explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly directs the agent to call get_nutrition_daily_food_log first to obtain the logId and date, which is real prerequisite guidance. It stops short of stating when-not-to-use or what to do if the entry is missing, so it doesn't reach a full 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_weigh_insB

Delete weight measurements for a specific date

Args: date: Date in YYYY-MM-DD format delete_all: Whether to delete all measurements for the day

ParametersJSON Schema
NameRequiredDescriptionDefault
dateYes
delete_allNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are present, so the description carries the burden. It states the destructive action, but does not disclose whether deletion is permanent, what happens if delete_all is false, or any return/error behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Short, direct, front-loaded with the action and argument list; no unnecessary words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool lacks output description and leaves delete_all false ambiguous; for a destructive operation, more context is needed to be complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

It adds date format and describes delete_all as controlling whether all measurements for the day are deleted. It does not explain the false behavior or default, and there are no property-level descriptions in the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Clearly identifies the action (delete), the target (weight measurements), and the scope (specific date). It is distinct from sibling add/get weigh-in tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The purpose statement implies the use case: call this when deleting weight measurements for a date. It does not explicitly mention alternatives or when not to use it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_workoutA

Delete a workout from Garmin Connect

Permanently removes a workout from your Garmin Connect workout library.

Args: workout_id: ID of the workout to delete (get IDs from get_workouts)

ParametersJSON Schema
NameRequiredDescriptionDefault
workout_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It explicitly discloses that deletion is permanent, which is the critical behavioral warning for a destructive operation. It does not mention side effects on scheduled workouts, but the permanence statement is valuable and non-obvious.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and front-loaded: the main action, the permanent effect, and the only parameter are stated in three short sections. There is no filler or redundant information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter tool, the description is mostly adequate for invoking it correctly. However, it omits important context about the sibling delete_workouts, whether deleting also unschedules workouts, and any authorization requirements. The output schema presence makes missing return-value details acceptable.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema only defines workout_id as an integer with 0% description coverage. The description compensates by explaining the parameter's role and pointing to get_workouts as the source of valid IDs, which is actionable guidance beyond the schema field name.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Clearly identifies the action as deleting a single workout from Garmin Connect and emphasizes that the removal is permanent. It does not explicitly contrast with the sibling delete_workouts, but the singular wording and single workout_id make the scope reasonably clear.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides a useful prerequisite by directing the agent to get IDs from get_workouts. However, it does not mention when to prefer this tool over delete_workouts or unschedule_workout, nor does it explain scenarios where deletion may be inappropriate.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_workoutsA

Delete multiple workouts from Garmin Connect in a single call

Permanently removes multiple workouts from your Garmin Connect workout library.

Args: workout_ids: List of workout IDs to delete (get IDs from get_workouts)

ParametersJSON Schema
NameRequiredDescriptionDefault
workout_idsYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the burden of behavioral disclosure. It clearly labels the operation as permanent removal, which is essential for a destructive tool, and notes the batch nature of the call. It does not discuss failure behavior or permissions, but the core destructive consequence is disclosed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with the key purpose, then adds the permanence warning and a compact Args explanation. No extraneous content.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple one-parameter deletion tool, the description plus schema covers the action, the permanent effect, and the ID source. An output schema exists, so return-value details are not required. It could still mention the singular delete_workout alternative or what happens with invalid IDs, but nothing critical is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description compensates by explaining that workout_ids is a list of workout IDs and pointing to get_workouts as the source. This is sufficient for the only parameter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific action and resource: deleting multiple workouts from Garmin Connect and permanently removing them from the workout library. The word 'multiple' and 'single call' distinguish it from the sibling delete_workout, which appears to handle a single workout.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description makes clear this is for deleting multiple workout IDs at once and tells the agent to obtain IDs from get_workouts. It does not explicitly mention when to use the singular delete_workout instead, so there is a small exclusion gap.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

download_activity_fileA

Download an activity and save it to disk as a file.

Saves the activity in the requested format. Defaults to the original .fit file; Garmin also supports gpx, tcx, and csv.

Directory resolution (first match wins):

  1. output_dir argument (one-off; not persisted)

  2. GARMIN_FIT_DOWNLOAD_DIR environment variable

  3. persisted config (set via set_fit_download_dir) If none is configured, returns status "needs_setup" with a suggested default (the server's current working directory). In that case, ask the user where to save, call set_fit_download_dir(path), then call this tool again.

Files are named "{activity_id}.{ext}" and overwrite any existing file.

Args: activity_id: Garmin activity ID format: One of fit, gpx, tcx, csv (default fit) output_dir: Optional one-off directory override (not persisted)

ParametersJSON Schema
NameRequiredDescriptionDefault
formatNofit
output_dirNo
activity_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, but the description fully discloses behavioral details: it saves a file to disk, overwrites existing files, supports a one-off output directory override that is not persisted, and returns a needs_setup status when no directory is configured.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-structured and front-loaded, but the directory resolution and needs_setup behavior are described twice in nearly identical terms, adding minor redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers the key operational context and user prompting flow, but it does not describe the actual return value on success or the response schema, leaving some ambiguity about what the caller receives.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Although the schema description coverage is 0%, the prose explains all three parameters: activity_id is the Garmin activity ID, format has default 'fit' and supported values, and output_dir is an optional non-persisted override.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool downloads an activity and saves it to disk as a file. It identifies the resource (activity), the action (download/save), and the output (file), which differentiates it from related tools like download_course_gpx and download_workout.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides explicit guidance on when to use the tool: it explains supported formats (fit, gpx, tcx, csv), the directory resolution precedence, and the fallback needs_setup behavior with instructions to ask the user and call set_fit_download_dir.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

download_workoutA

Download a workout as a FIT file

Downloads the workout in FIT format. The binary data cannot be returned directly through the MCP interface, but this confirms the workout is available.

Args: workout_id: ID of the workout to download

ParametersJSON Schema
NameRequiredDescriptionDefault
workout_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the burden of explaining behavior. It discloses that binary data is not returned directly but confirms availability, which is a meaningful behavioral trait. However, it does not mention permissions, side effects, or any other behavioral details.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise and front-loaded with the primary purpose. It uses a clear first line followed by a brief elaboration and an explicit Args section, making it easy to scan and understand.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is simple, and the description covers the core action and an important limitation. While it does not specify the exact return value, the statement 'confirms the workout is available' implies a confirmation response, which is sufficient for this simple tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema provides only the parameter name and type (integer) with no description. The description adds the explanation 'ID of the workout to download', which clarifies the parameter's purpose but provides no additional detail about constraints, format, or default behavior.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool downloads a workout as a FIT file, naming the resource (workout) and action (download). It is distinct from sibling tools like download_course_gpx and download_activity_file, which target different resources.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description does not indicate when to use this tool versus alternatives such as download_activity_file or upload_workout. It mentions a limitation (binary data cannot be returned) but offers no explicit guidance on selection criteria or preconditions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_activitiesA

Get activities with pagination support.

Retrieves a paginated list of activities ordered newest-first. Use this for browsing through large activity lists when you do not need to filter by date range, or as a complement to get_activities_by_date.

Each activity includes an event_type field. Common values: "race", "training", "uncategorized" (no event type set by the user — common for Peloton imports and untagged runs). Filter for races with event_type == "race" rather than excluding "training", as many non-race activities appear as "uncategorized" rather than "training".

Args: start: Starting index (default 0) limit: Maximum number of activities to return (default 20, max 100)

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
startNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Because no annotations are provided, the description carries the full burden, and it delivers: pagination behavior, newest-first ordering, event_type semantics, common values, and the important caveat about 'uncategorized' activities and Peloton imports. This goes well beyond the minimal 'retrieves a list' statement.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-structured and front-loaded, with purposeful paragraphs for purpose, usage, and parameter details. It is slightly redundant because the opening sentence and the second sentence both convey pagination support, but the overall length is justified by the valuable event_type guidance.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the output schema exists, the description doesn't need to spell out return fields—it covers the important contextual points: pagination defaults, ordering, event_type filtering, and relationship to the date-filtered sibling. An agent has everything needed to call this correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, but the description fully compensates with an Args section: start is the starting index with default 0, limit is max activities with default20 and max100. This is exactly the meaning an agent needs beyond the bare integer types.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb and resource ('Get activities'), and adds critical distinguishing detail: paginated list, ordered newest-first. It also explicitly contrasts itself with get_activities_by_date, so an agent can select it without opening schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It tells the agent exactly when to use the tool: 'when you do not need to filter by date range' or as a complement to get_activities_by_date. It also gives concrete filtering guidance for event_type, which is rare and highly actionable.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_activities_by_dateA

Get activities between specified dates with pagination support.

For accounts with large activity histories, broad date ranges can return thousands of activities in a single response. Use page and page_size to retrieve activities in manageable chunks and avoid "result too large" errors. Activities are ordered newest-first.

Pagination: when has_more is true the response includes next_page — pass that value as page on the next call to retrieve the following page. Repeat until has_more is false.

Note: total_count for a date range is not available from the Garmin API without fetching all results. Use has_more / next_page to walk pages.

Each activity includes an event_type field with values such as:

  • "race" — explicitly tagged as a race by the user

  • "training" — explicitly tagged as a training activity

  • "uncategorized" — no event type set; common for Peloton imports and untagged outdoor runs. Distinct from "training": filter for races with event_type == "race" rather than excluding "training", since many non-race activities appear as "uncategorized" not "training"

  • field absent — API returned no eventType for this activity; not observed in practice in any activity back to 2012 (oldest activities sampled on this account)

Args: start_date: Start date in YYYY-MM-DD format end_date: End date in YYYY-MM-DD format activity_type: Optional activity type filter (e.g., cycling, running, swimming) page: Zero-based page number (default 0) page_size: Number of activities per page, max 200 (default 100)

ParametersJSON Schema
NameRequiredDescriptionDefault
pageNo
end_dateYes
page_sizeNo
start_dateYes
activity_typeNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are present, so the description carries the full behavioral burden. It clearly discloses newest-first ordering, the pagination contract, the total_count API limitation, and nuanced event_type semantics including the absent-field edge case.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is substantial but every paragraph adds necessary operational detail: pagination behavior, API constraints, event_type filtering guidance, and parameter formats. It is front-loaded with a clear purpose before caveats.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 5-parameter tool with no annotations, the description covers invocation details, pagination mechanics, filtering semantics, and known API limitations. The output schema handles return-value documentation, so nothing essential is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, and the description fully compensates. It documents date format for start_date/end_date, examples for activity_type, zero-based indexing for page, and the max/default for page_size. Every parameter receives actionable meaning.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description states a specific operation: 'Get activities between specified dates,' with pagination support. The verb, resource, and date-range scope are clear, distinguishing it from sibling tools like get_activity or count_activities.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides strong practical guidance: explains when pagination is needed, how to walk pages with has_more/next_page, and how to interpret event_type for filtering. However, it never explicitly contrasts this tool with sibling alternatives or states when to prefer another tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_activities_fordateA

Get activities for a specific date

Args: date: Date in YYYY-MM-DD format

ParametersJSON Schema
NameRequiredDescriptionDefault
dateYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.5/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. It states the operation is a read ('Get') but does not describe the return shape, pagination, date inclusivity, or any other runtime behavior. This is a minimal disclosure for a tool with zero annotation support.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is short, front-loaded with the core purpose, and includes a clean Args block for the parameter. Every element earns its place without redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple one-parameter getter with an output schema present, the description provides enough to invoke the tool correctly: the operation, the resource, and the date format. It lacks sibling differentiation and behavioral detail, but the low complexity and output schema reduce the burden.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, but the description compensates by specifying the exact date format (YYYY-MM-DD), which is the key semantic information for the single required parameter. It adds meaning beyond the schema's bare 'Date' title.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states a specific verb ('Get') and resource ('activities') scoped to a date. However, it does not differentiate from the sibling tool get_activities_by_date, which appears to serve a nearly identical purpose.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'for a specific date' implies this tool is for date-scoped activity retrieval, but there is no explicit guidance on when to choose this over get_activities, get_activities_by_date, or count_activities. Usage context is implied rather than stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_activityA

Get detailed information for a single activity.

Returns a comprehensive summary including timing, distance, heart rate, elevation, training effect, and an event_type field. Common event_type values: "race", "training", "uncategorized" (no event type set by the user). The field is omitted for very old activities that pre-date event type support in the Garmin API.

Args: activity_id: ID of the activity to retrieve

ParametersJSON Schema
NameRequiredDescriptionDefault
activity_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the behavioral disclosure burden. It goes beyond a simple 'get' by explaining the returned fields and the event_type behavior, including that the field is omitted for old activities. It does not cover errors, permissions, or rate limits, but for a read-only single-resource tool it provides meaningful behavioral context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and well-structured, front-loading the purpose and then adding return details and the argument definition. There is slight redundancy between 'Get detailed information' and 'Returns a comprehensive summary,' but no filler or unnecessary content.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the simple single-parameter schema and the presence of an output schema, the description is largely complete for invoking the tool correctly. It covers the core return fields and a notable edge case (old activities lacking event_type). Explicit guidance on when to use this over sibling activity tools would make it complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, but the description compensates by explicitly documenting activity_id as 'ID of the activity to retrieve.' This adds meaning beyond the schema's bare type/title, though it could clarify accepted ID formats or whether integer/string forms are interchangeable.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with 'Get detailed information for a single activity,' naming a specific verb, resource, and singular scope. This clearly distinguishes it from sibling list tools like get_activities and activity-specific tools like get_activity_splits.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'single activity' implies use when you have an activity_id and need one activity's details, but no explicit guidance is given about when to choose this over alternatives like get_activities or get_activity_splits. Usage context is implied rather than stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_activity_detailsA

Get the GPS track and per-point route data for an activity.

Returns the recorded GPS polyline (latitude, longitude, altitude, time, distance, speed per point) plus the track's bounding box and start/end coordinates. Garmin downsamples the track server-side to at most max_points points, evenly spread over the activity, so the route shape survives at any size.

Use get_activity for the summary metrics and get_activity_fit_data for full per-second sensor series; this tool is for the route itself (mapping, course analysis, where a climb or interval happened).

Indoor activities carry no GPS track; the response then says so and still lists which per-point metric keys Garmin recorded.

Args: activity_id: ID of the activity to retrieve the track for max_points: Maximum number of track points to return (default 500). Raise it (e.g. 4000) for a finer trace on long routes.

ParametersJSON Schema
NameRequiredDescriptionDefault
max_pointsNo
activity_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and delivers: it discloses the return contents (lat/lon/altitude/time/distance/speed, bounding box, start/end), the server-side downsampling behavior ('at most max_points, evenly spread ... shape survives at any size'), and the indoor-activity edge case where no track exists but metric keys are still listed. This is exactly the kind of behavioral context annotations would otherwise provide.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core action, then returns, then alternatives, then args — a logical order with no filler sentences. It runs longer than strictly necessary (returns could be leaned on the output schema), but each section adds distinct value, so it stays efficient rather than padded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given a mutation-free read tool with an output schema and a 2-param schema lacking descriptions, the definition covers everything an agent needs: purpose, alternatives, return shape, downsampling semantics, param defaults, and the no-GPS edge case. Nothing material is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate, and it does: activity_id is defined as 'ID of the activity to retrieve the track for,' and max_points gets its default (500), its meaning (max track points returned), and actionable guidance to raise it (~4000) for finer traces on long routes. No parameter is left undocumented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Opens with a specific verb+resource ('Get the GPS track and per-point route data for an activity') and immediately names the siblings it is not (get_activity, get_activity_fit_data), stating its distinct niche: 'this tool is for the route itself.' An agent can route to it without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly names two alternatives and what each is for ('get_activity for the summary metrics', 'get_activity_fit_data for full per-second sensor series'), then lists concrete use cases (mapping, course analysis, locating a climb or interval). When-to-use and when-not-to-use are both covered.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_activity_exercise_setsB

Get exercise sets for strength training activities

Args: activity_id: ID of the activity to retrieve exercise sets for

ParametersJSON Schema
NameRequiredDescriptionDefault
activity_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the burden of behavioral disclosure. 'Get' implies a read-only operation, which is appropriate, but it does not specify what happens for non-strength activities, whether an empty result is possible, or any error behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise and front-loaded with the main purpose. The Args section is slightly redundant with the schema, but it is not bloated or confusing.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-parameter getter with an output schema, the basics are covered: what the tool retrieves and which parameter to pass. However, it lacks any context about activity type expectations, empty results, or relationship to similar tools, leaving some ambiguity for an agent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It does explain that activity_id is the ID of the activity to retrieve sets for, but this adds only modest meaning beyond the parameter name and schema type.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states a specific verb and resource: 'Get exercise sets for strength training activities.' It does not explicitly compare against siblings, but the resource is distinct enough that an agent can infer what this tool does.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'for strength training activities' implies the intended use case, but the description gives no explicit guidance on when to choose this over related activity data tools like get_activity_splits or get_activity_fit_data, and no exclusions are provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_activity_fit_dataA

Download and parse FIT file for an activity to expose advanced cycling data.

Returns data not available through the standard REST API, including:

  • DI2 / electronic shifting events with cadence at time of shift, grade at shift, gear combinations, shift quality classification, and terrain-grouped shift analysis

  • Cycling dynamics per session and lap: platform center offset (PCO), left/right power balance, torque effectiveness, pedal smoothness

  • Variability Index (NP / avg_power) per session and lap

  • Climb detection with VAM (vertical ascent rate), avg power/cadence/HR per climb, and W/kg per climb (using auto-fetched body weight from Garmin)

  • Grade-correlated stats: avg power, cadence, HR broken down by terrain steepness

  • HR drift / cardiac drift coefficient (aerobic decoupling for rides ≥60 min)

  • Temperature correlation: avg HR/power in hottest vs. coolest portions of ride

  • Power Duration Curve: best mean maximal power at 5s, 30s, 1min, 5min, 10min, 20min, 60min

  • Optional full per-second time series when include_records=True

Shift quality:

  • proactive: shifted at 70-100 rpm (ideal cadence range)

  • reactive: shifted below 70 rpm (already grinding before shifting)

  • coasting: shifted at 0 rpm (mid-stop or freewheeling)

  • spun_out: shifted above 100 rpm (waited too long in easy gear)

Note: DI2 data requires Shimano Di2 / SRAM eTap. Cycling dynamics require a compatible power meter (e.g., Garmin Rally, Favero Assioma, PowerTap P1 pedals).

Args: activity_id: Garmin activity ID include_records: Include full per-second time series (default False). Warning: adds significant data volume for long rides.

ParametersJSON Schema
NameRequiredDescriptionDefault
activity_idYes
include_recordsNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of transparency. It discloses a side effect—'Warning: adds significant data volume for long rides'—and hardware dependencies that affect success. It does not explicitly state that the operation is read-only or non-destructive, but given the lack of annotations and the 'download and parse' wording, this is a reasonable transparency level.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Despite being lengthy, the description is well-organized with bullet points and sections (shift quality, notes, args). It packs substantive information without redundancy or filler. Each sentence contributes to understanding the tool's capabilities, prerequisites, or parameters, making the length appropriate for the complexity of the data returned.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description is comprehensive for an agent: it details the output categories (DI2 events, cycling dynamics, climb detection, etc.), defines shift quality terms, lists hardware requirements, and explains arguments. No output schema is provided, but the enumerated data types give sufficient context for expected results. The inclusion of a dedicated note about include_records ensures the agent understands the trade-off. Overall, nothing essential is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema has zero descriptions for its properties, so the description fully compensates. It explicitly explains activity_id as 'Garmin activity ID' and include_records as 'Include full per-second time series (default False)' with a data volume warning. This covers both parameters completely, leaving no ambiguity about their meanings or defaults.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description begins with a clear verb+resource: 'Download and parse FIT file for an activity to expose advanced cycling data.' It distinguishes itself from the standard REST API by stating it returns data 'not available through the standard REST API,' and enumerates specific advanced metrics such as DI2 shifting, cycling dynamics, and power curves. This makes the tool's unique purpose obvious relative to the many sibling tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides useful usage guidance by noting hardware prerequisites ('DI2 data requires Shimano Di2 / SRAM eTap') and a warning about include_records adding significant data volume. It also states that it returns data not available through the standard REST API, which tells the agent when to prefer this tool. However, it does not explicitly name alternative sibling tools or give conditional 'use this vs. that' instructions, so it falls short of fully explicit when/when-not guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_activity_gearA

Get gear data used for an activity

Args: activity_id: ID of the activity to retrieve gear data for

ParametersJSON Schema
NameRequiredDescriptionDefault
activity_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The name and description imply a read-only operation, but no explicit statement about side effects or lack thereof is made. Since annotations are absent, the description carries the full burden; it does not clarify whether the returned gear data includes all fields or just summaries, nor does it mention any rate limits or authentication requirements.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely concise, consisting of one clear sentence plus the parameter explanation. No unnecessary words or redundant information is present.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple retrieval tool, the description is adequate. It does not cover return value details or error conditions, but since an output schema exists (as indicated by context), the lack of return-value explanation is acceptable. The description covers the essential purpose and parameter.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The description explicitly explains the sole parameter: 'ID of the activity to retrieve gear data for'. This provides clear meaning for activity_id. However, it does not specify whether the ID is numeric or string-based, which could cause minor ambiguity given the anyOf type in the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action: 'Get gear data used for an activity' and identifies the specific resource (gear for an activity). This verb+resource combination is unambiguous and distinct from sibling tools like get_gear or add_gear_to_activity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives such as get_gear (list all gear) or add/remove gear tools. The description does not mention any conditions, prerequisites, or scenarios that would make this tool the preferred choice.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_activity_hr_in_timezonesB

Get heart rate data in different time zones for an activity

Args: activity_id: ID of the activity to retrieve heart rate time zone data for

ParametersJSON Schema
NameRequiredDescriptionDefault
activity_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It only restates the get operation and data scope; it does not clarify read-only behavior, timezone conversion semantics, whether data is available for all time zones, or any permissions or rate limits. The output schema may describe the return shape, but the description itself adds little behavioral transparency beyond the tool name.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and front-loaded: the core purpose appears in the first sentence, followed by a short Args block. There is no fluff or repetition of schema details, so every part of the description earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Complexity is low with one required parameter, and an output schema exists, so the description does not need to explain return values. However, with no annotations and no usage guidance, the description leaves timezone semantics and when-to-use questions unanswered. It is minimally complete for a simple getter but not richly contextual.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It does add some semantic value with 'activity_id: ID of the activity to retrieve heart rate time zone data for,' which explains the parameter's role. However, it does not clarify accepted formats beyond the schema's anyOf, how to find or interpret the ID, or behavior when the activity does not exist, so the compensation is minimal though adequate for a single obvious parameter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names the specific operation: 'Get heart rate data in different time zones for an activity.' This gives a clear verb, resource, and scope. It distinguishes the tool from the closely named sibling get_activity_power_in_timezones by specifying heart-rate data rather than power data.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance about when to use this tool versus alternatives such as get_heart_rates, get_activity_splits, or get_activity_power_in_timezones. The only contextual clue is 'for an activity,' which implies an activity_id is needed but does not explain when this endpoint is the right choice.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_activity_power_in_timezonesA

Get power distribution across training zones for an activity.

Returns time spent in each power zone with watt thresholds and duration. Requires a power meter. Zones are based on the athlete's FTP configured in Garmin Connect.

Args: activity_id: ID of the activity to retrieve power zone data for

ParametersJSON Schema
NameRequiredDescriptionDefault
activity_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries full burden. It discloses output details (time spent, watt thresholds, duration) and prerequisites (power meter, FTP). It does not state side effects, but the read-only nature is implicit from the name and description, and there is no contradiction.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise and well-structured: a one-sentence summary, output details, prerequisites, and parameter explanation. No redundant or vague wording.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Provides essential context: what it returns, prerequisites, and parameter meaning. It does not mention error cases or what happens when no power data exists, but these are minor gaps given the tool's simplicity and the presence of related siblings with similar patterns.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The single parameter activity_id is clearly explained in the Args section as 'ID of the activity to retrieve power zone data for'. Although the schema lacks description, the description fully compensates, making the parameter's meaning unambiguous.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states a specific action ('Get power distribution across training zones') for a specific resource ('an activity'). This distinguishes it from sibling tools like get_activity_hr_in_timezones or get_activity_splits.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides condition 'Requires a power meter' and explains FTP basis, which tells the agent when this tool is appropriate. It does not explicitly mention when to choose an alternative, but the unique purpose is clear enough.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_activity_seriesA

Get raw per-activity measurements for a date range, with no interpretation.

One record per activity: date, discipline, duration, distance, average and max HR, power, elevation and Garmin's training effect. Pace is deliberately not computed — what counts as a comparable effort is a coaching decision.

Activities are returned oldest-first (Garmin serves them newest-first, which silently reverses anything treating list position as time).

Args: start_date: Start date in YYYY-MM-DD format (inclusive) end_date: End date in YYYY-MM-DD format (inclusive) sports: Optional filter, any of ["running", "cycling", "swimming"]. Activities outside the three triathlon disciplines have sport: null and are excluded when this filter is set. include_hr_zones: Attach per-zone seconds. Costs one extra Garmin request per activity — HR zones are not in the activity payload. Off by default.

ParametersJSON Schema
NameRequiredDescriptionDefault
sportsNo
end_dateYes
start_dateYes
include_hr_zonesNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does so well: it discloses the ordering quirk (oldest-first, because Garmin serves newest-first and this 'silently reverses anything treating list position as time'), that include_hr_zones costs one extra Garmin request per activity and is off by default, and that non-triathlon activities come back with sport: null and are dropped when the filter is set.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with purpose, then behaviour, then a clean Args block. The two-sentence pace explanation is slightly editorial but it justifies a deliberate omission, so it earns its place; no filler remains.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only series tool with 4 parameters, no annotations and an output schema present, the description covers purpose, filtering, ordering, field contents and request cost. Return-value formatting is correctly left to the output schema, so nothing material is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must supply all parameter meaning, and it does: date formats with inclusive bounds, the allowed sport values, null-sport exclusion behaviour, and the cost/default of include_hr_zones. Nothing an agent needs to construct the call is left undocumented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The opening sentence names a specific verb (get) and resource (raw per-activity measurements) and pins the scope with 'for a date range, with no interpretation', then enumerates the returned fields. This implicitly separates it from analytic siblings such as get_performance_trend or get_workout_compliance, and from get_activity_details by stating the one-record-per-activity grain.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives clear context for when the tool is appropriate ('raw', 'no interpretation', pace deliberately omitted because that's a coaching decision) and explains the sports filter's exclusion semantics. However, it never names an alternative tool or an explicit when-not-to-use condition, so routing still requires inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_activity_splitsC

Get splits for an activity

Args: activity_id: ID of the activity to retrieve splits for

ParametersJSON Schema
NameRequiredDescriptionDefault
activity_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden of behavioral disclosure. It only states the action and parameter, giving no information about side effects, permissions, rate limits, or whether the operation is read-only. The verb 'get' implies a read, but this is not explicit, and there is no mention of error conditions or return behavior. This is a minimal disclosure that does not go beyond the literal action.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely concise and front-loaded, with the core purpose stated in the first sentence. The parameter explanation is redundant with the schema but does not add fluff. It is well-structured as a docstring, though the brevity limits the information conveyed. For a one-parameter tool, this length is appropriate, but it could be more informative without losing efficiency.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description is incomplete in the context of the tool ecosystem. It does not clarify what type of splits are returned (e.g., lap splits, typed splits, summaries), nor does it address the existence of sibling tools that might be more appropriate for specific use cases. The presence of an output schema mitigates the need to describe return values, but the tool's role among the many split-related tools is ambiguous. An agent cannot confidently choose this tool over its siblings without additional information.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The description says 'ID of the activity to retrieve splits for', which adds only slightly to the schema's 'Activity Id' field. It explains the parameter's role but provides no additional context about format, constraints, or typical usage. With 0% schema description coverage, the description should compensate with richer semantics, but it essentially repeats what the parameter name already conveys. This is insufficient.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a clear verb-resource pairing: 'Get splits for an activity'. It identifies the resource (activity splits) and the action (get). However, it does not differentiate from closely related siblings like get_activity_typed_splits or get_activity_split_summaries, which likely also retrieve split data. This is a clear purpose but lacks the specificity to distinguish among similar tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives. It does not mention any conditions, prerequisites, or exclusions, nor does it reference sibling tools. An agent must infer the appropriate usage, which is risky given the number of split-related tools available in the sibling list.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_activity_split_summariesC

Get split summaries for an activity

Args: activity_id: ID of the activity to retrieve split summaries for

ParametersJSON Schema
NameRequiredDescriptionDefault
activity_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

There are no annotations provided, so the description carries the full burden of behavior disclosure. It only says 'get' without mentioning return format, volume, pagination, error cases, or any read-only cautions, leaving the agent to infer from the schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is appropriately short and front-loaded with the purpose, but the Args block largely duplicates the input schema, adding no real value. Both sentences together are concise, but not every line strongly earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

This is a simple one-parameter read tool, so the minimum is present, but the description does not clarify the meaning or format of 'split summaries' relative to sibling functions like get_activity_splits. Since an output schema exists, return documentation is not strictly needed, yet more context about what the summaries represent would make the description complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With schema_description_coverage at 0%, the description should compensate for missing parameter documentation. It does say 'activity_id: ID of the activity to retrieve split summaries for', which adds a little semantic context over the schema name. However, it does not explain the accepted types or where the activity_id comes from.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a clear verb and resource, 'Get split summaries for an activity', so the agent knows what the tool does at a basic level. However, it does not distinguish this from the closely related siblings get_activity_splits and get_activity_typed_splits, leaving a subtle terminology gap.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description contains no guidance on when to use this tool instead of get_activity_splits or get_activity_typed_splits. It only repeats the activity_id parameter in the Args block, which is not a usage guideline.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_activity_typed_splitsB

Get typed splits for an activity

Args: activity_id: ID of the activity to retrieve typed splits for

ParametersJSON Schema
NameRequiredDescriptionDefault
activity_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.3/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description provides no information about side effects, read-only status, data scope, or any behavioral characteristics. With no annotations present, the description itself should disclose such details, but it is entirely silent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, direct sentence with no extraneous words. It is concise and well-structured, conveying the essential information without padding.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple GET operation with one parameter, the description is adequately complete. The lack of output schema or annotations lowers the burden, and the description covers the core functionality. A brief mention of return type would be beneficial but is not strictly necessary given the low complexity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The single parameter activity_id is clearly described as 'ID of the activity to retrieve typed splits for'. The meaning and purpose are unambiguous, fully satisfying the parameter explanation requirement.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Get') and a specific resource ('typed splits'), making the primary purpose clear. However, it does not explicitly differentiate from sibling tools like get_activity_splits or get_activity_split_summaries, so some ambiguity remains about what 'typed' adds.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives such as get_activity_splits or get_activity_split_summaries. The description offers no context for selecting this tool over others.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_activity_typesA

Get all available activity types

Returns a list of all activity types supported by Garmin Connect, useful for filtering activities by type.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description must carry the behavioral disclosure. It does indicate a read-only retrieval ('Get', 'Returns a list') and the data scope ('supported by Garmin Connect'). However, it leaves out details such as whether the list includes custom or localized activity types, which could matter depending on the agent's task.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is only two sentences: the first is a direct, front-loaded statement of the tool's function, and the second adds a relevant use case. Every word earns its place with no filler or repetition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has no parameters and an output schema is available, the description adequately covers what the tool returns and why it is useful. No critical information needed for correct invocation is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, and schema description coverage is 100%. The description correctly adds no parameter-specific information because none is needed, matching the baseline for parameterless tools.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific action and resource: 'Get all available activity types' and clarifies it returns 'a list of all activity types supported by Garmin Connect'. This clearly distinguishes it from sibling tools that set activity types or retrieve activity records.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly identifies a practical context—'useful for filtering activities by type'—which gives the agent a clear cue for when to invoke it. It does not mention exclusions or alternatives, but the intended use case is sufficiently clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_activity_weatherA

Get weather data for an activity.

Garmin's weather endpoint returns temperatures in Fahrenheit (from the weather-station source) with no unit indicator, regardless of account settings. This tool converts them to the account's display unit: metric accounts get Celsius, statute_us accounts keep Fahrenheit. The temperature_unit field ("F" or "C") states which unit was returned.

Wind speed, unlike temperature, is already returned in the account's display unit (km/h for metric, mph for statute_us), so it is passed through unconverted and labeled via the wind_speed_unit field.

Args: activity_id: ID of the activity to retrieve weather data for

ParametersJSON Schema
NameRequiredDescriptionDefault
activity_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden, and it does so thoroughly. It explains that the raw endpoint returns Fahrenheit without a unit indicator, that this tool converts temperatures to the account's display unit, and that wind speed is passed through unconverted with its own unit label. This gives the agent essential behavioral context beyond the tool name.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description opens with a clear one-sentence purpose and then uses short, organized paragraphs for unit-conversion behavior, followed by an Args section. Every sentence adds useful information, and the structure makes the content easy to scan.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has only one required parameter and an output schema is present, the description covers all essential context: what the tool does, what the activity_id means, and the key behavioral quirk around unit conversion. Nothing critical is missing for an agent to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has no parameter descriptions, so the description must compensate. The Args section says 'activity_id: ID of the activity to retrieve weather data for,' which clarifies which activity the ID refers to beyond the schema title 'Activity Id.' For a single, self-explanatory identifier parameter, this is sufficient semantic guidance.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The opening sentence states a specific verb and resource: 'Get weather data for an activity.' This clearly distinguishes the tool from the many activity-related siblings, none of which are about weather. The resource and scope are exactly what an agent needs to select it.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description makes the usage context clear: use this when weather data for a particular activity is needed. It does not name exclusions or alternative tools, but the purpose is specific enough that an agent can identify this as the weather-specific activity getter among the sibling tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_adhoc_challengesA

Get user-created social/group challenges (e.g., step competitions with friends)

Returns challenges created by users to compete with connections. These are different from official Garmin badge challenges.

Args: start: Starting index for pagination (default 0) limit: Maximum number of challenges to return (default 20, max 100)

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
startNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden. It states that the tool 'returns' challenges, indicating a read operation, but does not disclose potential errors, rate limits, or any additional behavior beyond the basic retrieval. The differentiation from badge challenges adds some context but lacks thorough behavioral disclosure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise, consisting of two sentences that front-load the purpose and then list the parameters. There is no fluff or redundancy; every sentence adds value.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is positioned within the context of challenges by explicitly contrasting with official badge challenges, and the parameters are fully explained. However, the description does not mention the response structure or any related usage notes, leaving a minor gap in completeness for an agent encountering this tool for the first time.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The description provides meaningful semantics for both parameters: 'start' is the starting index for pagination, and 'limit' is the maximum number to return (with a max of 100). This goes beyond the bare schema types and defaults, giving the agent clear guidance on how to use the parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states that the tool retrieves user-created social/group challenges, distinguishing them from official Garmin badge challenges. The verb 'Get' and specific resource type make the purpose unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly notes that these challenges are different from official badge challenges, which implies a contrast with sibling tools like get_badge_challenges. However, it does not explicitly state 'use this when you need user-created challenges, use the other for official ones,' leaving some inference to the agent.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_all_day_eventsC

Get daily wellness events data

Args: date: Date in YYYY-MM-DD format

ParametersJSON Schema
NameRequiredDescriptionDefault
dateYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

C2.5/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

There are no annotations, and the description does not disclose any behavioral traits such as read-only nature, data source, or potential side effects. Since nothing is disclosed beyond the basic action, the description fails to inform the agent about the tool's behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely concise and to the point, containing only the essential information about the parameter format. There is no redundancy or irrelevant detail, making it efficient for the agent to parse.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description lacks essential context about what 'wellness events' actually represents, what the response structure will be, or how this data is used. Without output schema details or examples, the agent cannot fully anticipate the tool's behavior or results.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The description provides the date format 'YYYY-MM-DD' in the Args section, which adds meaning beyond the schema's bare type 'string'. This helps the agent correctly format the parameter, but it is the only parameter and the description is otherwise minimal.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a verb 'Get' and a resource 'daily wellness events data', but the term 'wellness events' is ambiguous and could refer to various types of data (e.g., stress, body battery, all-day stress). It does not clearly differentiate from sibling tools like get_all_day_stress or get_body_battery_events.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool over alternatives. The description gives no conditions, examples, or hints about appropriate usage scenarios, leaving the agent to guess based on the name alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_all_day_stressB

Get all-day stress data

Args: date: Date in YYYY-MM-DD format

ParametersJSON Schema
NameRequiredDescriptionDefault
dateYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are present, so the description carries the full burden. It states only the action 'Get' without clarifying whether it is read-only, if there are side effects, or how it behaves on errors or missing data. The name implies a read operation but this is not explicit.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely concise, consisting of one sentence and a parameter format line. No unnecessary information is included, and every word adds value.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description is complete for a simple getter in terms of its action and parameter, but it does not describe the output format, data structure, or possible error conditions. Given the simplicity, this is acceptable but leaves some gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has a single required parameter 'date' with no description. The description adds the format 'YYYY-MM-DD' but does not explain the semantic meaning (e.g., the specific day for which stress is requested). This is minimal but adequate for a straightforward date parameter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Get all-day stress data' clearly states the action (get) and the resource (all-day stress data). It is specific enough to distinguish from generic stress queries, though it does not explicitly differentiate from sibling tools like get_stress_data.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternative stress-related tools such as get_stress_data, get_stress_summary, or get_weekly_stress. The description does not mention conditions, fallbacks, or selection criteria.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_athlete_contextA

Get the athlete's own thresholds, HR zones and physical profile.

Returns LTHR, cycling and running FTP, per-sport HR zone floors, VO2max, weight, height and unit preferences — the numbers that make an absolute heart rate or power target meaningful for this athlete rather than for a generic one.

Thresholds carry as_of and is_stale. Garmin marks a threshold stale once it has stopped reflecting the athlete; prescribing against a stale FTP means prescribing against a test nobody has repeated.

not_available names what Garmin does not expose, so a missing field reads as a platform limitation rather than a gap in this profile.

No arguments.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does well: it discloses that thresholds carry as_of and is_stale flags, explains the operational consequence of staleness (prescribing against a stale FTP), and notes that not_available marks platform limitations rather than profile gaps. It stops short of covering auth requirements or rate limits.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core purpose, then progressively adds return contents and caveats in short paragraphs. Every sentence carries information, though the multi-paragraph format is slightly longer than strictly necessary for a no-argument read tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained, yet the description usefully annotates the tricky fields (as_of, is_stale, not_available). Combined with the zero-parameter schema, the definition gives the agent everything needed to call it correctly; only explicit sibling routing is absent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero arguments, so the baseline is 4. The description confirms 'No arguments,' which is consistent with the empty input schema and leaves no ambiguity about invocation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource and then enumerates exactly what it returns (LTHR, cycling/running FTP, HR zone floors, VO2max, weight, height, unit preferences). This clearly separates it from sibling tools like get_cycling_ftp, get_lactate_threshold, get_user_profile, and get_unit_system, which each cover a subset of this data.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives a clear rationale for use — 'the numbers that make an absolute heart rate or power target meaningful for this athlete rather than for a generic one' — which tells the agent to fetch this before prescribing absolute targets. It does not, however, name an explicit alternative or state when not to use it (e.g., versus get_user_profile or get_lactate_threshold).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_athlete_status_snapshotA

Get a point-in-time athlete snapshot with baseline deviations.

Current HRV, body battery, resting HR, readiness and sleep, each alongside Garmin's own seven-day baseline and the deviation from it. Deviation against a published baseline is arithmetic; what a −22% HRV deviation means for today's session is not, and no gate is applied.

Args: date: Date in YYYY-MM-DD format. Defaults to today — but before the watch syncs, today holds nothing, so pass yesterday when a morning call comes back empty.

ParametersJSON Schema
NameRequiredDescriptionDefault
dateNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the burden and does it well: it discloses that deviations are arithmetic against Garmin's baseline, that no interpretation or gating of the session is performed, and that today's data may be empty until sync. It omits auth/permission requirements, which keeps it just under 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose is front-loaded in the first sentence with the metric list immediately after, and the argument note sits at the end where it belongs. The aside ('what a −22% HRV deviation means... is not') is slightly discursive but justifies the important 'no gate is applied' caveat rather than padding.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need no explanation. Given the tool's complexity (multi-metric aggregate with derived deviations), the description covers what is returned, how deviations are derived, what it will not do, and the one real invocation pitfall (empty pre-sync day). An agent has everything needed to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% (a bare nullable string with no description), so the description must compensate and largely does: it gives the YYYY-MM-DD format, the today default, and the non-obvious workaround of passing yesterday when a morning call is empty. It does not address timezone handling, which is the only remaining ambiguity for a date parameter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Get a point-in-time athlete snapshot with baseline deviations') and enumerates exactly what is inside: HRV, body battery, resting HR, readiness and sleep, each against a seven-day baseline. This distinguishes it from the single-metric siblings like get_hrv_data, get_body_battery and get_sleep_data by making clear it is a combined snapshot rather than a raw series.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear contextual guidance tied to the call site: it is a point-in-time snapshot, no gate is applied, and if a morning call returns empty you should pass yesterday because the watch has not synced. It does not explicitly name alternative tools (e.g. use get_hrv_trend for trends), so it stops short of the 5 tier.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_available_badge_challengesA

Get official Garmin badge challenges available to join

Returns monthly/seasonal challenges from Garmin that the user can join. These challenges award badges and points upon completion.

Args: start: Starting index for pagination (starts at 1) limit: Maximum number of challenges to return (default 20, max 100)

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
startNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description is a simple retrieval operation with no side effects, but it does not explicitly state that it is read-only or mention any permissions, rate limits, or error conditions. Since annotations are absent, the description carries the burden and provides basic transparency but lacks depth.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise, with a clear one-sentence purpose, a brief explanation of returns, and a simple 'Args:' section. No unnecessary words or redundant information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Since the tool has an output schema (as indicated by 'Has output schema: true'), the description need not enumerate output fields. However, it could briefly mention that the result is a paginated list of challenges, but the pagination parameters already imply that. Overall, the description is sufficient given the schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Both parameters are explained in the description: 'start' is the pagination index starting at 1, and 'limit' is the maximum number of results with a default of 20 and a max of 100. This adds meaningful detail that the schema does not provide (the schema only includes type and default).

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states it retrieves official Garmin badge challenges that are available to join, and distinguishes them from earned badges, ad-hoc challenges, and non-completed challenges by explicitly mentioning 'available to join' and 'monthly/seasonal'.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No explicit guidance is given about when to use this tool versus other badge-related tools like get_badge_challenges or get_earned_badges. The description implies its use case but does not state conditions or alternatives, leaving the agent to infer.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_badge_challengesA

Get all badge challenges the user has joined (completed and in-progress)

Returns the user's history of badge challenges including progress, completion status, and earned dates.

Args: start: Starting index for pagination (starts at 1) limit: Maximum number of challenges to return (default 20, max 100)

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
startNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden of explaining behavior. It clearly indicates the scope (user's joined challenges) and the return content, but it does not explicitly state that the operation is read-only or mention any side effects, authentication requirements, or rate limits. However, the verb 'Get' makes the read-only nature implicit.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise and well-structured: a single sentence for the main purpose, one for the return details, and a clear breakdown of parameters. There is no extraneous information, making it easy for an agent to parse quickly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description provides enough context for an agent to know exactly what the tool does and what parameters mean. Since an output schema is indicated as existing, the description does not need to enumerate the return structure. It covers the essential information for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema only provides types and defaults for 'start' and 'limit', but the description adds meaningful semantics: 'start' is the starting index (1-based) and 'limit' is the maximum number to return with a default of 20 and max of 100. This fully explains the parameters beyond the raw schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool retrieves all badge challenges the user has joined, including both completed and in-progress ones. It also lists the specific data returned (progress, completion status, earned dates), making the purpose unambiguous and distinct from sibling tools like get_earned_badges or get_available_badge_challenges.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies when to use this tool—when you need the user's full history of joined badge challenges—but it does not explicitly contrast it with alternatives such as get_non_completed_badge_challenges or get_available_badge_challenges. The guidance is inferred rather than stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_blood_pressureC

Get blood pressure data

Args: start_date: Start date in YYYY-MM-DD format end_date: End date in YYYY-MM-DD format

ParametersJSON Schema
NameRequiredDescriptionDefault
end_dateYes
start_dateYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description must carry the behavioral disclosure burden. It only says 'Get blood pressure data' and lists parameter formats; it does not disclose read-only behavior, date-range inclusivity, timezone handling, or response shape. The agent cannot infer side effects or edge-case behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is short and free of irrelevant content. The core purpose is front-loaded and the argument details are listed cleanly. It could use a sentence or two more of behavioral context, but as written it is economical.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With only two required date parameters and an output schema present, the tool does not need an explanation of return fields. Yet the complete lack of usage guidance and behavioral transparency leaves gaps: the agent is not told how to interpret the date range, whether results are aggregated or raw, or whether there are any constraints. This is a minimal viable description, not a complete one.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema provides only parameter names and string types, so the description adds the YYYY-MM-DD format for both date arguments. This is helpful but minimal: it still omits whether the date range is inclusive, what time of day applies, or how invalid dates are handled. For a tool with 0% schema coverage, the compensation is partial.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource ('Get blood pressure data'), and the resource name distinguishes it from sibling tools like set_blood_pressure. However, it does not go beyond the tool name to describe what the returned blood pressure data contains or how it is scoped, so it is clear but minimally informative.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives, nor any exclusions or prerequisites. Sibling tools include several health-data getters and set_blood_pressure, but the description does not help the agent choose this one over those or clarify whether it is the only blood-pressure read query.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_body_batteryC

Get body battery data with events

Args: start_date: Start date in YYYY-MM-DD format end_date: End date in YYYY-MM-DD format

ParametersJSON Schema
NameRequiredDescriptionDefault
end_dateYes
start_dateYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It implies a read operation but does not explain event semantics, data availability, date-range behavior, or any other runtime characteristics beyond taking two dates.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is short and well-structured, front-loading the tool's purpose and then cleanly listing arguments. It avoids filler, though the argument section mostly restates schema property names.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple two-parameter read tool with an output schema present, this is minimally viable. The description is incomplete in that it never differentiates from get_body_battery_events, and the meaning of 'events' is unexplained, creating a real selection risk.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It does add meaning by exposing the YYYY-MM-DD format and naming start/end as a date range, but it stops short of clarifying inclusivity, timezone handling, or max range limits.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific action and resource: 'Get body battery data with events'. However, the phrase 'with events' is ambiguous, and the close sibling get_body_battery_events exists, so the description does not clearly distinguish this tool from that one.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is given for when to use this tool versus get_body_battery_events or other health-data getters. The description only lists arguments, leaving the agent to infer the appropriate use case.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_body_battery_eventsC

Get body battery events data

Args: date: Date in YYYY-MM-DD format

ParametersJSON Schema
NameRequiredDescriptionDefault
dateYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

C2.9/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description gives no information about side effects, read-only nature, return format, or any behavioral details. Since annotations are absent, the description bears full responsibility and fails to disclose anything beyond the action name.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely concise, with no redundant words or fluff. It directly states the purpose and parameter format, making it highly efficient.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given many similar tools (e.g., get_body_battery, get_body_composition), the description lacks context about what 'events' specifically means, what the output looks like, or how it differs from siblings. No output schema is provided, leaving the user without a complete picture.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The single parameter 'date' is explained with a format specification 'YYYY-MM-DD' in the docstring, which is sufficient for usage. However, no further semantic details (e.g., timezone handling) are provided, but the format note meets the basic need.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states a 'Get' action on 'body battery events data', identifying the resource and verb. However, it does not differentiate from the sibling tool 'get_body_battery', so purpose is clear but not fully distinct.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives like 'get_body_battery' or which conditions apply. The description just restates the function name without any context for selection.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_body_compositionA

Get body composition data for a single date or date range

Args: start_date: Date in YYYY-MM-DD format or start date if end_date provided end_date: Optional end date in YYYY-MM-DD format for date range

ParametersJSON Schema
NameRequiredDescriptionDefault
end_dateNo
start_dateYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.7/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description alone must convey behavior, but it only states that data is retrieved. It does not disclose read-only status, potential errors, or any side effects, leaving the agent without information about consequences of invoking the tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is brief and direct, leading with the purpose and followed by parameter details. It avoids redundant information and is easy to parse.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers the tool's function and all parameters, and since an output schema exists, it need not describe return values. However, it lacks any examples or edge-case notes, which could be useful but are not essential.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The description fully explains both parameters: start_date is required and can be a single date or the start of a range, and end_date is optional and completes the range. It also specifies the YYYY-MM-DD format, which is not present in the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('Get') and resource ('body composition data') in the first sentence, making the tool's purpose unambiguous. It also specifies the date-based scope, which distinguishes it from other data retrieval tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies when to use this tool (when body composition data is needed) but does not explicitly compare it to alternatives like get_weigh_ins or get_stats_and_body. No conditions or exclusions are provided, so guidance is minimal.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_cardiac_drift_analysisA

Measure cardiac drift (aerobic decoupling) for a specific activity.

Returns the percentage change in the power:HR ratio from the first half of the activity to the second. Where the boundary between "coupled" and "decoupled" sits is a coaching decision, so no label is attached — a common reading is that beyond about 5% the athlete was decoupling, but that number belongs in the coaching layer.

TWO PRECONDITIONS, both checked before you waste a call:

  1. The activity must carry POWER. Drift is a power:HR ratio, so heart rate alone cannot produce it. Cycling needs a power meter. Running usually does not — most modern Garmin watches estimate running power natively and record it on every sample.

  2. At least 3600 records with both power and heart rate, which is 60 minutes at 1-second sampling. Shorter sessions do not qualify, and a device set to smart recording rather than 1-second sampling may not reach it even on a long one.

When either fails the response is an error with a reason naming which, plus records_total, records_with_power and records_usable so the gap is visible. That is a limitation of the recorded data, not a fault in the activity — do not report it as a training finding.

Args: activity_id: Garmin activity ID (numeric or string)

ParametersJSON Schema
NameRequiredDescriptionDefault
activity_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the burden and does well: it discloses the failure response shape (error with reason plus records_total/records_with_power/records_usable) and warns not to report data limitations as training findings. It omits any explicit statement of read-only safety or auth/permission needs.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the metric definition, then the two preconditions, then error behavior. The coaching-layer aside about the 5% threshold is slightly tangential but justifies the absence of a label, so it earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter analysis tool with an output schema present, the description covers the metric meaning, preconditions, and failure diagnostics. Nothing an agent needs to call it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% for the single activity_id parameter, but the description only restates it as 'Garmin activity ID (numeric or string)', which the schema's anyOf already conveys. It adds marginal meaning over the structured field.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Measure cardiac drift (aerobic decoupling) for a specific activity') and immediately defines the metric as the percentage change in the power:HR ratio between halves. This distinguishes it from sibling analysis tools like get_activity_power_in_timezones or get_power_duration_curve.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives two explicit preconditions (activity must carry power; at least 3600 records with both power and HR) and explains when each fails, which is strong when-to-use guidance. It does not name an alternative tool to use instead when drift cannot be computed, so it stops short of a full 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_coursesA

List all courses saved on Garmin Connect.

Returns a curated list of courses with id, name, distance, activity type and creation date.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are present, so the description carries the full burden. It indicates a read-only operation returning a curated list of fields, which is transparent about the result. It does not mention error handling or rate limits, but the behavior is clear enough for a simple list tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise, using two sentences to convey purpose and output fields. It is well-structured and free of unnecessary detail.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description names the specific fields returned, which covers the output structure. However, it does not mention pagination or limits, which could be relevant for a list endpoint. Still, for a simple listing tool, it is reasonably complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There are no parameters, so schema coverage is 100% (empty). Per the rubric, the baseline is 3 since the description does not need to elaborate on absent parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: to list all courses saved on Garmin Connect. It also specifies the returned fields, distinguishing it from related tools like get_course_details or delete_course.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description does not explicitly mention when to use this tool versus alternatives. It implies a simple 'list all' operation, but lacks guidance on scenarios like needing a specific course or filtering options, so usage guidance is minimal.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_custom_foodsA

Search or list user's custom foods

Returns custom foods the user has created. Use the search parameter to find existing foods by name before creating duplicates — the response includes foodId and servingId needed for log_custom_food.

For branded catalog foods (FatSecret), use search_foods instead.

Args: search: Search term to filter foods by name (default: list all) start: Starting index for pagination (default 0) limit: Maximum number of results (default 20)

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
startNo
searchNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries full responsibility. It clearly indicates this is a read-only lookup and mentions returned fields, but it does not explicitly state that no data is modified or describe pagination edge behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is succinct and well-structured, leading with the main purpose, followed by usage guidance and parameter details. It avoids unnecessary fluff.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple list/search tool with an output schema, the description provides sufficient context: it identifies user-specific custom foods, explains the purpose of returned IDs, and directs users to the correct sibling tool for branded foods. No essential information is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema has no parameter descriptions, but the Args section explains all three parameters: search filters by name, start is the pagination index, and limit is the maximum result count. Defaults are also included.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: search or list the user's custom foods. It also distinguishes this tool from search_foods, which handles branded catalog foods.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly instructs to use the search parameter before creating duplicates and notes that the response includes foodId and servingId for log_custom_food. It also names search_foods as the alternative for branded foods, so usage guidance is concrete.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_custom_food_serving_unitsA

Get available serving units for custom foods

Returns the list of valid serving units (e.g. G, ML, OZ) that can be used when creating custom foods.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description must carry behavioral disclosure. It correctly states the tool returns a list of valid units, implying a read-only, side-effect-free operation. However, it does not detail pagination, auth requirements, or whether the list is exhaustive, although for a simple lookups these are minor gaps.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and front-loaded with the main statement, followed by a clarifying detail with examples. The opening sentence closely mirrors the tool name, causing slight redundancy, but the rest earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

This is a simple, parameterless lookup with an output schema present, so the description is nearly complete. It conveys the return concept and its use case, which is enough for an agent to call it correctly; nothing important is missing for this level of complexity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, and the schema coverage trivially covers them. The description adds value by explaining what the returned units are used for, even though no parameter semantics are needed. Baseline 4 for 0 params is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Get') and resource ('available serving units for custom foods'), and clarifies with examples (G, ML, OZ). It is clearly distinct from sibling tools like get_custom_foods, so an agent can tell the difference without opening schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly frames the tool as a lookup to be used 'when creating custom foods,' giving clear context for when to call it. It does not name alternative tools or exclusion conditions, but no obvious alternative exists for this resource, so the guidance is sufficient.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_cycling_ftpA

Get the latest cycling Functional Threshold Power (FTP) data.

Returns the most recent cycling FTP estimate available from Garmin.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the behavioral disclosure burden. It clearly indicates a read-only, no-argument operation returning the latest cycling FTP estimate from Garmin, which is sufficient for a simple getter; it does not detail edge cases like missing data or sync delay, but the output schema covers return structure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is very short and front-loaded, but the first sentence and second sentence partly restate the same idea ('Get the latest' vs. 'Returns the most recent'). Still, it is compact and easy to scan.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter read with an output schema, the description is complete: it names the exact data (cycling FTP) and the recency behavior (latest estimate). No additional usage steps or preconditions are needed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters and 100% schema description coverage, so there are no parameter semantics to document. The baseline of 4 applies because the description correctly implies no arguments are needed.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Get'), a resource ('cycling Functional Threshold Power data'), and a qualifier ('latest', 'from Garmin'). This clearly distinguishes the tool from sibling metric getters such as get_lactate_threshold or get_vo2max_trend.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies the tool is appropriate when the agent needs the most recent cycling FTP estimate, but it does not explicitly state when to prefer it over alternatives or mention exclusions. Usage context is inferred from the tool name and wording rather than directly guided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_daily_stepsB

Get steps data for a date range

Args: start_date: Start date in YYYY-MM-DD format end_date: End date in YYYY-MM-DD format

ParametersJSON Schema
NameRequiredDescriptionDefault
end_dateYes
start_dateYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden of behavioral disclosure. It states that data is retrieved but does not mention read-only behavior, authentication requirements, rate limits, pagination, or whether the returned data is aggregated or raw.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise and front-loaded with the action and object. The Args section is directly useful and adds no unnecessary detail.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is simple and the description covers the core operation, but it omits return-value details such as the unit of measurement, whether summaries or per-day records are returned, and any constraints on the date range. These are useful for an agent selecting and using the tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The description explains the date parameters and their YYYY-MM-DD format, adding value over the bare schema. However, it does not clarify whether the date range is inclusive, whether a maximum span applies, or how timezones are handled.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific action ('Get steps data') and a clear scope ('for a date range'). It is distinct enough from most siblings, though the name 'get_daily_steps' is needed to clarify that the result is daily step values rather than generic step data.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided about when to use this tool instead of similar siblings such as get_steps_data or get_weekly_steps. The description gives no context about the expected use case or when this tool is preferred.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_daily_weigh_insA

Get weight measurements for a specific date

Args: date: Date in YYYY-MM-DD format

ParametersJSON Schema
NameRequiredDescriptionDefault
dateYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It only states the operation and the date format, without mentioning what is returned, whether multiple measurements can exist for one date, or any limitation or error behavior. This is minimal beyond the core purpose.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely brief and front-loaded with the core purpose, followed by a single argument specification. There is no fluff, and the only additional detail is the required date format, which is necessary.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's low complexity (one required parameter, no nested objects, and an output schema), the description covers the basic calling needs. However, it does not distinguish this from get_weigh_ins or mention any behavioral nuances, leaving some ambiguity given the large sibling set.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema provides only a string property named 'date' with no description (0% schema description coverage), so the description must compensate. It does so by specifying the exact YYYY-MM-DD format, which is essential for correct invocation and is not present in the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with the verb 'Get' and clearly identifies the resource 'weight measurements' scoped to 'a specific date.' This directly differentiates it from sibling tools like get_weigh_ins (likely broader) and add/delete weigh-in tools, giving an agent enough to select it correctly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage for a specific date, but it does not explicitly state when to choose this over get_weigh_ins or other weigh-in tools. No exclusions or alternative recommendations are included, so the agent must infer the distinction from the phrase 'for a specific date.'

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_device_alarmsA

Get alarms from all Garmin devices

Returns all configured alarms with their schedules, sounds, and enabled status.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description uses 'Get' and 'Returns', indicating a read-only operation, but it does not explicitly disclose side effects, rate limits, or pagination behavior. Since no annotations are provided, the description carries the burden and is somewhat minimal but not misleading.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is short and to the point, consisting of two concise sentences. It avoids unnecessary detail and is well-structured for quick comprehension.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, the description does not need to explain return values in depth, but it does provide a helpful summary of the alarm attributes (schedules, sounds, enabled status). It lacks any mention of authentication or potential limitations, but these are not critical for a straightforward getter.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, and the schema coverage is 100% (vacuously). The description does not add parameter information because none exist, which is consistent with the baseline for high schema coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose with a specific verb ('Get') and resource ('alarms from all Garmin devices'), and it specifies the returned data (schedules, sounds, enabled status). No sibling tool appears to handle alarms, so there is no ambiguity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage for retrieving alarms but does not explicitly state when to use it versus alternative tools. However, since there is no other alarm tool among the siblings, the necessity is implicit. It lacks explicit when-to-use/when-not-to-use guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_device_last_usedB

Get information about the last used Garmin device

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are present, and the description does not disclose whether the tool is read-only, has side effects, or requires special permissions. As a getter, it is likely safe, but this is not explicitly stated.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, succinct sentence with no unnecessary words. It is well-structured and immediately conveys the tool's purpose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description is minimal and does not elaborate on what 'information' includes or the return format. Given the existence of many similar device-related tools, additional context about the specific data returned would improve completeness, but the core purpose is clear.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has no parameters, so schema coverage is 100% vacuously. The description adds no extra meaning about parameters, but none are needed. Baseline score of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action (get) and the target (information about the last used Garmin device). It is specific enough to differentiate from many sibling tools, though it could be more explicit about what 'information' entails (e.g., model, firmware, battery).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives like get_devices or get_primary_training_device. The description does not mention any conditions or exclusions, leaving the agent to infer appropriateness.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_devicesA

Get all Garmin devices associated with the user account

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description uses 'Get' which clearly implies a read-only operation. With no annotations provided, the description carries the full burden, and it sufficiently communicates that this tool has no side effects.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, concise sentence with no redundant words. It front-loads the action and clearly defines the scope.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple getter with no parameters, the description is sufficient for an agent to understand what the tool does. It does not elaborate on the return format, but that is not strictly necessary for this type of operation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, so there is nothing to describe. The schema coverage is complete (100%), and the description does not need to add parameter context. A baseline of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('Get') and the exact resource ('all Garmin devices associated with the user account'). It is unambiguous and distinguishes from tools that fetch individual device details.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage by indicating 'all' devices, but it does not explicitly mention when to use this tool instead of related device-specific tools like get_device_settings or get_primary_training_device. The context is clear but not formally guided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_device_settingsA

Get settings for a specific Garmin device

Returns device configuration including time/date format, units, activity tracking settings, and alarm information.

Args: device_id: Device ID (optional; defaults to the most recently used device when omitted; can be obtained from get_devices or get_device_last_used)

ParametersJSON Schema
NameRequiredDescriptionDefault
device_idNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden of disclosing behavior. It states that it returns device configuration and lists the included settings, but it does not mention potential errors (e.g., device not found) or the exact output structure. This provides some transparency but not comprehensive detail.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise and well-structured. It uses two short sentences for the purpose and a clear parameter explanation, with no wasted words or redundant information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple read operation with one optional parameter, the description provides enough context for an agent to invoke the tool correctly. It explains the return content, and the parameter guidance covers how to obtain the device ID. The output schema is marked as present, so the absence of explicit return fields in the description is acceptable.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The description fully explains the only parameter (device_id): it is optional, defaults to the most recently used device, and can be obtained from specific sibling tools. This adds significant meaning beyond the schema, which only specifies type and default with no description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb 'Get' and the resource 'settings for a specific Garmin device', and it lists the types of settings returned (time/date format, units, activity tracking, alarms). This distinguishes it from sibling tools like get_devices and get_device_last_used, which return device lists or IDs.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It provides guidance on how to obtain the device_id parameter via get_devices or get_device_last_used and explains the default behavior when omitted. It does not explicitly state alternatives for the same purpose (since there are none), but the guidance is sufficient for typical use.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_device_solar_dataA

Get solar data for a specific device

Returns solar charging data for devices with solar panels (e.g., Instinct Solar, Fenix Solar). Only applicable to solar-capable devices.

Args: device_id: Device ID (can be obtained from get_devices) date: Date in YYYY-MM-DD format

ParametersJSON Schema
NameRequiredDescriptionDefault
dateYes
device_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description must carry the behavioral burden. 'Get' and 'Returns' imply a read-only operation, and the solar-capable-device constraint is useful behavioral context. However, it does not disclose authentication needs, rate limits, or what happens if a non-solar device is queried, leaving notable gaps.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with the core action, then provides return context, applicability scope, and parameter documentation. The Args block is clean and useful. Slight redundancy between 'Get solar data' and 'Returns solar charging data' keeps it from being maximally concise.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has an output schema, so return values need not be explained in depth. For a two-parameter read tool with 0% schema description coverage, the description adequately covers both parameters and the key solar-capable-device constraint. Minor gaps remain around error behavior and non-solar device handling.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It documents both parameters meaningfully: device_id's source ('can be obtained from get_devices') and date's format ('YYYY-MM-DD'). It does not explain the anyOf integer/string device_id type, but it covers the essential parameter meaning.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'Get solar data for a specific device,' then clarifies the returned data as 'solar charging data.' It also scopes the tool to solar-capable devices, distinguishing it from the many sibling device and activity tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It clearly states the applicability condition: 'Only applicable to solar-capable devices.' It also tells the agent where device_id comes from ('get_devices'). However, it does not name an alternative tool for non-solar devices or explain when not to use it beyond the solar-capability constraint.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_earned_badgesB

Get earned badges for user

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are present, so the description carries full burden. 'Get' implies a read-only operation, but the description does not disclose whether it returns the current user's badges, whether authentication is required, or any filtering/ordering/pagination behavior. Minimal beyond the operation itself.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single concise sentence with no filler, front-loads the operation and resource. It earns its place despite lacking detail.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is simple, has no parameters, and an output schema exists, so full return documentation is not needed. However, the description is thin: it fails to clarify the implicit user scope (current user?) and gives no routing signal among the many badge-related siblings. Adequate for a trivial getter but with clear gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There are zero parameters, so the schema provides no target information. The description adds the concept of 'user' as the scope of the query, which is meaningful; however, it leaves unspecified whether this means the authenticated user, which is a minor ambiguity. Baseline for 0 params is 4.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a clear verb ('Get') and resource ('earned badges') and implies a user scope, so an agent can identify the operation. It doesn't explicitly differentiate from sibling badge tools like get_badge_challenges or get_available_badge_challenges, but the 'earned' qualifier distinguishes it from challenge/available lists.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to call this vs alternatives; no exclusions or context. The tool name and siblings imply a decision between earned badges, available challenges, and non-completed challenges, but the description does not state it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_endurance_scoreA

Get endurance score data between dates

Args: start_date: Start date in YYYY-MM-DD format end_date: End date in YYYY-MM-DD format

ParametersJSON Schema
NameRequiredDescriptionDefault
end_dateYes
start_dateYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the burden of specifying behavior. It indicates a read operation (getting data) but does not mention potential side effects, rate limits, or the structure of the returned data. It is adequate but not richly transparent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise, using a single sentence plus a brief arg list. Every word is purposeful and there is no fluff.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple date-range retrieval, the description is minimally sufficient but leaves out details about the output format, units, or any filtering options. Given its simplicity, it does not mislead, but it lacks the completeness that would fully prepare an agent for the response.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The description adds format details (YYYY-MM-DD) for both start_date and end_date that are not in the schema. This clarifies the expected input, though it does not specify whether dates are inclusive or how timezones are handled.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb 'Get' and the resource 'endurance score data' with a date range scope. It is unambiguous and distinguishes itself from sibling tools that focus on other metrics like steps, heart rate, or training readiness.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explains what the tool does but provides no guidance on when to use it versus alternatives, nor any context about date inclusivity, timezone handling, or typical use cases. It is not misleading but lacks explicit usage instructions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_fitnessage_dataB

Get fitness age data

Args: date: Date in YYYY-MM-DD format details: If True, include component breakdown (BMI, RHR, vigorous activity) with targets and improvement suggestions

ParametersJSON Schema
NameRequiredDescriptionDefault
dateYes
detailsNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of disclosing behavior. It only says 'get' and does not disclose whether the operation is read-only, what happens on missing data, or any rate-limit or authentication concerns. This is a meaningful gap for an API tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and front-loads the core purpose before the Args list. Each parameter description adds value beyond the schema and there is no extraneous text, though a single-line alternative could be slightly tighter.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

It has an output schema, and the description covers the essential usage information: required date format and the behavior of optional details. It does not mention error handling or rate limits, but for a read operation these are less critical given the existing structured output.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema has 0% description coverage, so the description must compensate. It adds crucial meaning by specifying the exact date format (YYYY-MM-DD) and clarifying that 'details' controls whether BMI, RHR, and vigorous-activity breakdowns with targets and suggestions are returned.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states the resource ('fitness age data') and a clear action ('get'), so the purpose is easily understood. It doesn't explicitly distinguish this from sibling getters, but the resource name is specific enough to avoid confusion in the agent's selection.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is given about when to use this tool versus alternatives, nor are any exclusions or alternative tool names mentioned. The only hints are embedded in parameter descriptions rather than a clear use-case statement.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_floorsA

Get floors climbed data

Args: date: Date in YYYY-MM-DD format

ParametersJSON Schema
NameRequiredDescriptionDefault
dateYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden, but it only states the action and does not disclose any side effects, permissions, or limitations beyond getting data.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely concise and front-loaded with the core purpose, with no unnecessary words or repetition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple read-only query tool, the description is sufficient to understand what it does and what input is required, though it does not describe the return format or content.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The single parameter 'date' is described with a specific format (YYYY-MM-DD) in the description, adding useful detail beyond the schema's generic 'Date' title.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Get floors climbed data' uses a specific verb and resource, clearly distinguishing this tool from the many other get_* tools in the sibling list.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is given on when to use this tool versus alternatives like get_stats or get_user_summary, which might also provide similar data.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_full_nameA

Get user's full name from profile

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. It only says 'get', which implies read-only, but gives no explicit statements about characteristics such as authentication requirements, whether the data is always available, or any side effects. For a getter, this is minimal; no extra behavioral concerns are revealed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is one short sentence with no filler. It is fully front-loaded and every word earns its place. For a zero-parameter getter, this is appropriately concise.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's simplicity (0 params) and the existence of an output schema, the description is basically complete for invoking the tool. The only gap is the lack of usage guidance and differentiation from similar siblings, but as a simple getter, the description sufficiently conveys the operation. It does not need to explain return values because an output schema exists.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There are zero parameters in the schema, so the baseline is 4. The description adds minor context by noting the source 'from profile', but otherwise the action is fully defined by the name. No additional parameter semantics are needed because there are none.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states exactly what it does with a specific verb and resource: 'Get user's full name from profile'. It clearly distinguishes itself from sibling tools like get_user_profile, which would return the full profile, not just the full name. This is a precise, unambiguous getter.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance about when to use this tool versus alternatives. The description does not mention on the many sibling getters that could provide the full name (e.g., get_user_profile, get_user_summary) or when this specific tool is preferable. The agent receives no context for selection.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_garmin_coach_workoutsA

Get Garmin Coach workouts around the given date

Returns workouts from the active Garmin Coach/training plan, including plan metadata, workout identifiers, dates, sport, duration, completion status, rest days, race days, and workout intent when Garmin provides them. Adaptive plans expose only Garmin's currently generated window, typically the current week; future dates may return no workouts even while a plan is active. The count includes rest-day entries.

Garmin's standalone Daily Suggested Workouts are generated on compatible devices. As of July 31, 2026, no supported or known Garmin Connect web/API endpoint, including those exposed by this project's python-garminconnect dependency, returns the device's upcoming DSW schedule. This tool returns Garmin Coach/training-plan workouts and does not synthesize device-generated suggestions.

This is the preferred tool for Garmin Coach requests. The legacy get_training_plan_workouts tool returns the same data; do not call both.

Adaptive Coach plans typically expose workout_uuid; other plan families may expose numeric workout_id. Pass whichever identifier is present to get_workout_by_id. Rest-day UUIDs may return minimal detail without workout segments.

Args: calendar_date: Reference date in YYYY-MM-DD format (returns week's workouts)

ParametersJSON Schema
NameRequiredDescriptionDefault
calendar_dateYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full transparency burden. It does a good job disclosing that future dates may return no workouts, that count includes rest days, that rest-day UUIDs may return minimal detail, and that no synthesis of Daily Suggested Workouts occurs. It does not explicitly state read-only behavior, but the getter semantics are clear from the wording.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is fairly long with multiple explanatory paragraphs, but each section adds meaningful context: return fields, plan limitations, sibling comparison, and identifier guidance. The structure is readable and useful, though slightly more verbose than strictly necessary.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers the primary input semantics, output content, plan-family differences, rest-day behavior, and relationship to sibling tools. It does not specify the exact response schema or possible error conditions, but the provided context is sufficient for typical use.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema has 0% description coverage, so the description must compensate. It explains calendar_date format and that the returned data covers the week around that date. However, it does not clarify week boundaries, timezone handling, or any edge cases like invalid dates, which leaves some ambiguity for a bare string parameter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states that the tool returns Garmin Coach/training-plan workouts for a given date. It also distinguishes the tool from the legacy get_training_plan_workouts sibling, making its purpose unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly names the preferred use case for Garmin Coach requests, warns against calling both this tool and get_training_plan_workouts, and clarifies that Daily Suggested Workouts are not available through this endpoint. This gives concrete when-to-use guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_gearA

Get all gear registered with the user account

Returns complete gear inventory including usage statistics and default activity associations. No parameters required - user profile is fetched automatically.

Args: include_stats: Include usage statistics for each gear item (default True). Set to False for faster response with large gear collections.

ParametersJSON Schema
NameRequiredDescriptionDefault
include_statsNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It discloses what is returned, that the user profile is fetched automatically, and the performance tradeoff of include_stats. It does not explicitly declare read-only status or rate-limit behavior, but none are implied.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the main action and resource, followed by a tight Args block. Each sentence adds useful information, though the first sentence and 'Returns complete gear inventory' are slightly redundant.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with one optional parameter and an output schema, the description is largely complete: it covers purpose, return contents, parameter semantics, defaults, and performance behavior. The only notable gap is the lack of explicit alternative guidance.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, but the description fully compensates. It explains that include_stats controls usage statistics inclusion, states the default True, and gives a concrete reason to set it to False for large gear collections.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: 'Get all gear registered with the user account' and reinforces scope with 'complete gear inventory'. This clearly distinguishes it from activity-specific gear tools like get_activity_gear.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description makes the intended use clear: retrieving the user's full gear inventory. However, it never explicitly names alternatives such as get_activity_gear or states when not to use this tool, so routing decisions are mostly left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_goalsA

Get Garmin Connect goals (active, future, or past)

Args: goal_type: Type of goals to retrieve. Options: "active", "future", or "past"

ParametersJSON Schema
NameRequiredDescriptionDefault
goal_typeNoactive

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The word 'Get' clearly indicates a read-only operation, and there are no annotations to contradict this. It does not mention side effects, permissions, or error behavior, but for a simple getter this level of transparency is adequate.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise and well-structured, with a clear one-line purpose followed by an Args section. It contains no filler and every sentence adds useful information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description provides enough context for an agent to call the tool correctly, including the parameter options. While it does not describe the return format, an output schema exists, so that omission is acceptable.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema only provides a type and default for goal_type, but the description explicitly lists the allowed options ('active', 'future', 'past'), adding meaningful semantic detail beyond the schema. Parameter semantics are fully covered.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('Get') and the resource ('Garmin Connect goals'), and further specifies the three goal types (active, future, past). This is specific enough to distinguish the tool from the many other get_* sibling tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explains the purpose and the acceptable goal_type values, making it clear when to use this tool. It does not explicitly contrast with sibling tools, but since no other sibling tool retrieves 'goals', the usage guidance is sufficient.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_health_seriesA

Get raw multi-day health measurements in one call, with no interpretation.

Returns one record per day carrying only what Garmin measured. No thresholds are applied and no verdict is rendered — this is the input a coaching layer reasons over.

Days with no data are absent from days. Days whose fetch FAILED are listed in errors; an empty errors array is what makes a short series trustworthy. Never read a gap as a negative finding when errors is non-empty.

Body Battery returns all four values Garmin records — at_wake is the overnight recharge, most_recent is the end-of-day drain. They mean different things; pick deliberately.

Args: start_date: Start date in YYYY-MM-DD format (inclusive) end_date: End date in YYYY-MM-DD format (inclusive, max 60 days) metrics: Subset of ["body_battery", "hrv", "resting_hr", "sleep", "stress", "training_load", "readiness"]. Defaults to all seven. Fewer metrics means fewer Garmin requests; body_battery and stress share one endpoint.

ParametersJSON Schema
NameRequiredDescriptionDefault
metricsNo
end_dateYes
start_dateYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden and does well: it discloses that no thresholds or verdicts are applied, that empty days are simply absent, that failed fetches surface in `errors`, and that `at_wake` vs `most_recent` mean different things. It does not cover auth/permission requirements or explicit rate limits, leaving some gaps for a data-fetch tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the core purpose in the first sentence, then layers error semantics and field meanings in short, scannable paragraphs. It is longer than minimal, but nearly every sentence carries actionable information rather than filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Despite no annotations, the definition covers purpose, all parameters, error/missing-data handling, and field interpretation, and an output schema exists so return values need not be fully re-explained. An agent has everything needed to select and invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate and it does: it documents date formats, inclusivity, the 60-day max on end_date, the full metric enum with a stated default, and the endpoint-sharing nuance between body_battery and stress. All three parameters gain meaning beyond the bare schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource — 'Get raw multi-day health measurements in one call' — and pins down scope ('multi-day', 'no interpretation'), which distinguishes it from the per-domain siblings like get_sleep_data or get_body_battery. It never names an alternative explicitly, so the differentiation is inferred from scope rather than stated.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear context: this is the raw input a coaching layer reasons over, and 'Fewer metrics means fewer Garmin requests' steers efficient usage. It also tells the agent how to interpret results (empty `errors` makes a short series trustworthy; don't read a gap as negative when `errors` is non-empty). It stops short of naming sibling tools as alternatives for single-metric queries.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_heart_ratesA

Get full heart rate time-series data

Note: This returns detailed 2-minute interval data (~25KB). For a compact summary, use get_heart_rates_summary().

Args: date: Date in YYYY-MM-DD format

ParametersJSON Schema
NameRequiredDescriptionDefault
dateYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It mentions the return type (detailed time-series) and size (~25KB), which gives some behavioral expectations. It does not explicitly state side effects, but 'get' implies read-only, which is acceptable.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is succinct, with the purpose, a distinguishing note, and parameter format all in a few lines. No unnecessary details.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple get operation with one parameter, the description is complete. It provides the tool's purpose, output characteristics, and the parameter format. The contrast with the summary tool adds context without excessive length.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema has one parameter (date) with no description, but the tool description adds 'Date in YYYY-MM-DD format', providing the required format and meaning. This fully covers the single parameter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool retrieves full heart rate time-series data for a given date, distinguishing it from the summary variant via 'full' and 'detailed'. The sibling list includes get_heart_rates_summary, which reinforces the distinction.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly names the alternative tool get_heart_rates_summary() for compact summaries, implying this tool is for detailed data needs. The note about 2-minute intervals and ~25KB size further guides when to use it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_heart_rates_summaryA

Get heart rate summary with essential metrics (lightweight version)

Returns a compact summary (~500 bytes) instead of full time-series data (~25KB). Ideal for daily health checkups and LLM integrations.

Args: date: Date in YYYY-MM-DD format

ParametersJSON Schema
NameRequiredDescriptionDefault
dateYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It discloses the output size (~500 bytes) and that it returns a compact summary. It does not mention error handling or missing data, but overall it is reasonably transparent for a read-only summary tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise, with a front-loaded purpose and use case. It avoids unnecessary verbosity while still conveying the key differences from the full data tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

It provides enough context for an agent to decide when to use this tool versus get_heart_rates. It could specify what 'essential metrics' include, but that is not required for invocation. Given the simplicity of the tool, the description is nearly complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The only parameter 'date' is given a clear format ('YYYY-MM-DD'), which is essential for correct invocation. Since the schema has no description for this parameter, the description adds necessary meaning beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific action: 'Get heart rate summary with essential metrics' and clearly distinguishes it from the full time-series version with 'lightweight version' and 'compact summary'. The resource and verb are unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly says 'instead of full time-series data', which identifies when to use this tool over the fuller alternative. It also mentions 'Ideal for daily health checkups and LLM integrations', providing contextual use cases.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_hill_scoreA

Get hill score data between dates

Args: start_date: Start date in YYYY-MM-DD format end_date: End date in YYYY-MM-DD format

ParametersJSON Schema
NameRequiredDescriptionDefault
end_dateYes
start_dateYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are present, and the description does not disclose whether the operation is read-only, requires authentication, or has any side effects. The simplicity of a 'get' implies read-only, but this is not explicitly stated.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely concise, with the key action and parameters presented in two short sentences. It is well-structured and front-loaded with the purpose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple getter with a date range, the description is complete enough. It does not elaborate on the return value, but given the existence of an output schema (not shown), this is acceptable. No critical context is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The description explicitly provides the format (YYYY-MM-DD) for both start_date and end_date, which adds meaning beyond the basic schema types. It does not specify inclusivity/exclusivity or additional constraints, but the format guidance is valuable.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb 'get' and the resource 'hill score data' with a date range scope, which is specific and distinguishes it from other getter tools in the sibling list.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool over similar getter tools (e.g., get_heart_rates, get_sleep_data). The description only states the action without contextual usage conditions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_hrv_dataB

Get Heart Rate Variability (HRV) data

Args: date: Date in YYYY-MM-DD format return_timeseries: If True, include detailed 5-minute HRV readings (can be large)

ParametersJSON Schema
NameRequiredDescriptionDefault
dateYes
return_timeseriesNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description must convey side effects or output characteristics. It only notes that time series data 'can be large', but fails to disclose what the default response contains (e.g., daily summary fields) or any performance implications beyond size. The read-only nature is not explicitly stated.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is succinct, using a brief leading phrase and two bullet points for parameters. Every sentence adds value, with no redundant or verbose content.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

While the tool is simple and an output schema exists (not shown here), the description lacks details about the structure of the returned data (e.g., HRV metrics like RMSSD or stress score). It also does not clarify the relationship with related trend tools, leaving minor gaps for an agent deciding whether this is the right call.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Both parameters in the schema are fully explained: the date format is given as YYYY-MM-DD, and return_timeseries is described as including detailed 5-minute readings. This exceeds the simple schema by providing practical meaning and usage details.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the function is to retrieve HRV data for a given date, with an option to include time series. However, it does not differentiate from sibling tools like get_hrv_trend or get_heart_rates, which weakens precision.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives. The description only explains parameters and does not mention scenarios where this tool is preferred or avoided, leaving the agent to infer usage from the name alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_hrv_trendA

Get HRV (Heart Rate Variability) trend over a date range.

Returns daily HRV values and weekly rolling averages. Single-day HRV is too noisy to act on — use this tool to identify baseline shifts that signal accumulated fatigue or recovery. A drop of >10ms from the 7-day baseline warrants reducing training load.

Recommended range: 7-21 days. Maximum: 30 days.

Args: start_date: Start date in YYYY-MM-DD format end_date: End date in YYYY-MM-DD format

ParametersJSON Schema
NameRequiredDescriptionDefault
end_dateYes
start_dateYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description must disclose behavior. It states that it 'Returns daily HRV values and weekly rolling averages', revealing the output shape. However, it does not explicitly confirm that the operation is read-only or describe any potential side effects, though 'Get' implies no mutation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-structured with a purpose statement, return details, usage guidance, and parameter descriptions. It is slightly longer than necessary due to the clinical interpretation note ('A drop of >10ms...'), but this information is relevant and not redundant, so it earns a high score.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple two-parameter getter, the description fully covers the tool's purpose, usage, return content, and parameter formats. The existence of an output schema further fills in structural details, making the description complete for a caller to use the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has no descriptions for either parameter, so the description's statement of 'Start date in YYYY-MM-DD format' and 'End date in YYYY-MM-DD format' adds essential format information. It does not specify inclusivity of dates or ordering, but the phrase 'over a date range' implies inclusive bounds, and ordering is conventionally understood.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb 'Get' and the specific resource 'HRV trend over a date range'. It further distinguishes this from a simple data fetch by mentioning weekly rolling averages, making the purpose unambiguous even without comparing to sibling tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly explains when to use this tool: 'Single-day HRV is too noisy to act on — use this tool to identify baseline shifts'. It also provides practical usage constraints with 'Recommended range: 7-21 days. Maximum: 30 days.', guiding the caller on appropriate date ranges.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_hydration_dataB

Get hydration data

Args: date: Date in YYYY-MM-DD format

ParametersJSON Schema
NameRequiredDescriptionDefault
dateYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

There are no annotations and the description does not disclose any behavioral aspects such as read-only nature, potential side effects, or data format. The name implies a simple retrieval, but the lack of explicit statements leaves behavior under-transparent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely concise, using only two short sentences to convey the purpose and parameter. There is no unnecessary verbosity or redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description does not specify what the response contains (e.g., structure, units, possible values) and no output schema is provided. This leaves the agent without critical context about the return data, making the tool incomplete for practical use.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The description explicitly defines the only parameter 'date' with its format 'YYYY-MM-DD', fully covering the input schema. The schema itself lacks a description, but the prose compensates comprehensively.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action 'Get hydration data' and identifies the resource (hydration data). It is unambiguous about the tool's core purpose, though it does not elaborate on what specific hydration metrics are included.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives. It does not mention any conditions, prerequisites, or compare with other getter functions in the sibling list, leaving the agent to infer usage from the name alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_inprogress_virtual_challengesA

Get in-progress virtual challenges/expeditions

Returns virtual challenges (like walking expeditions on famous trails) that the user is currently participating in.

Args: start: Starting index for pagination (default 1, must be >= 1; garminconnect 0.3.2 rejects 0 for this endpoint) limit: Maximum number of challenges to return (default 20, max 100)

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
startNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden of behavioral disclosure. It provides important runtime details: the pagination start index must be >= 1, the limit max is 100, and a version-specific note that garminconnect 0.3.2 rejects 0. The read-only nature is implied by 'Get/Returns' but not explicitly declared, and rate limits or error behavior are not discussed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and well-structured: a one-sentence purpose statement followed by a concise Args list. Every sentence provides actionable information, including a practical version-specific warning, with no fluff or repetition of schema defaults.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple read-only endpoint with two optional parameters and an output schema present, the description covers the essential usage: what the tool returns, what param constraints exist, and how pagination works. The only missing element is explicit guidance on when to use this over sibling challenge tools, which is already acknowledged in the usage dimension.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema only provides types and defaults for the two parameters. The description adds essential meaning: start is a pagination index with a lower bound, limit is the maximum number of challenges with a max cap, and a real-world compatibility quirk about start=0. This fully compensates for the 0% schema description coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific action ('Get') on a specific resource ('in-progress virtual challenges/expeditions') and defines what qualifies: challenges the user is currently participating in. This clearly sets it apart from sibling tools like get_badge_challenges or get_adhoc_challenges, which target different challenge types.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description clearly conveys the intended use: retrieve virtual challenges the user is currently participating in. However, it does not explicitly mention alternatives or provide when-not-to-use guidance, so it stops short of full routing instructions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_lactate_thresholdA

Get lactate threshold data

Returns lactate threshold information, which is the exercise intensity at which lactate starts to accumulate in the blood. This is a key metric for endurance training.

Args: start_date: Start date in YYYY-MM-DD format (optional, omit for latest) end_date: End date in YYYY-MM-DD format (optional, omit for latest)

ParametersJSON Schema
NameRequiredDescriptionDefault
end_dateNo
start_dateNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the disclosure burden. It does explain that omitting dates returns the latest data and that date args are optional, which is useful. However, it does not describe output granularity, units, possible missing-data behavior, or any API-specific constraints, leaving the behavioral picture incomplete.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and front-loaded with the core purpose, followed by a one-sentence domain definition and a clear Args section. It has a small redundancy in repeating 'lactate threshold' and separate near-identical wording for start_date and end_date, but every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple, optional-parameter retrieval tool with an output schema, the description covers the essential invocation details: what the metric is, which parameters exist, their format/optionality, and the 'latest' fallback. It does not address date-range inclusiveness or sibling-tool selection, but those are minor gaps given the low complexity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema provides only names, types, and defaults with no descriptions. The description compensates fully by documenting both parameters: 'Start date in YYYY-MM-DD format (optional, omit for latest)' and the equivalent for end_date. This gives the agent the exact format and the meaningful default behavior.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Get lactate threshold data', then defines what lactate threshold means. Among a large set of sibling get_* tools, the metric is clearly identified and not easily confused with related metrics like get_cycling_ftp or get_heart_rates.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance about when to choose this tool versus related metric or trend tools. The only context is that lactate threshold is a key endurance metric, but the description never states when an agent should call this instead of a sibling like get_vo2max_trend, get_training_status, or get_cycling_ftp.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_lifestyle_logging_dataB

Get lifestyle logging data for a specific date

Returns lifestyle logging data which allows users to track behaviors and their impact on health metrics.

Args: date: Date in YYYY-MM-DD format

ParametersJSON Schema
NameRequiredDescriptionDefault
dateYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are present, and the description only states that it returns data. It does not disclose potential side effects, read-only nature, rate limits, or authentication requirements, placing full burden on the description which remains minimal.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise, using two short sentences plus an Args section. No redundant information or filler is present; it is straight to the point.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The term 'lifestyle logging data' is vague and does not specify what types of metrics are included (e.g., stress, sleep, steps). Given the many closely related sibling tools, more context is needed to fully understand what this tool returns and how it differs.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The only parameter (date) is described with a specific format (YYYY-MM-DD), which adds meaningful detail beyond the schema's basic type/title. This helps the agent pass a correctly formatted value.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('Get') and the resource ('lifestyle logging data') with a date parameter. It is distinguishable from siblings like get_steps_data or get_sleep_data, though it does not explicitly name an alternative.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus other get_* tools. It only explains the date parameter but does not mention conditions or scenarios where this tool is appropriate.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_menstrual_calendar_dataA

Get menstrual calendar data between specified dates

Automatically chunks requests longer than 92 days, Garmin's server-side limit, and stitches the responses together.

Args: start_date: Start date in YYYY-MM-DD format end_date: End date in YYYY-MM-DD format

ParametersJSON Schema
NameRequiredDescriptionDefault
end_dateYes
start_dateYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description discloses a key behavioral trait: it chunks long requests and stitches responses, indicating internal handling of pagination or limits. It does not mention return format, potential errors, or whether it is read-only (though 'Get' implies so), leaving some behavioral aspects unspecified.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely concise, consisting of two sentences that deliver the primary purpose and a relevant behavioral note. No unnecessary words or redundant details, making it highly efficient.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given that an output schema exists (as noted in context), the description does not need to explain return values. It adequately covers the essential usage context, including the chunking behavior for long date ranges. Minor details like date inclusivity are omitted but not critical for basic use.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The description adds crucial meaning to both parameters by specifying 'YYYY-MM-DD format' for start_date and end_date, which the schema lacks. This clarifies the expected input format and resolves potential ambiguity.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('Get') and resource ('menstrual calendar data') with a specific date range. It does not explicitly contrast with the sibling tool 'get_menstrual_data_for_date', but the range vs. single-date distinction is implicit, making the purpose understandable without ambiguity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides a guideline about automatic chunking for requests over 92 days, which helps with long ranges. However, it does not mention when to prefer this tool over the single-date alternative or any other usage context, relying on the user to infer the appropriate scenario.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_menstrual_data_for_dateB

Get menstrual data for a specific date

Args: date: Date in YYYY-MM-DD format

ParametersJSON Schema
NameRequiredDescriptionDefault
dateYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the responsibility for indicating side effects. The verb 'get' implies read-only behavior, but the description does not explicitly state that it has no side effects, nor does it mention any data-return limitations or error conditions. This is adequate but not thorough.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely concise and front-loaded. It states the purpose in a single sentence and then documents the only parameter. There is no unnecessary verbiage or redundant information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description tells what the tool does but not what the returned menstrual data will look like, what fields are included, or how it might compare to related menstrual endpoints. Since the output schema is not included in the visible definition, the lack of any expected-response detail leaves the agent somewhat under-informed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema only indicates that 'date' is a required string, but the description adds the exact expected format ('YYYY-MM-DD'), which is valuable. No additional context such as timezone or date-range constraints is provided, but for a single parameter this is reasonably well covered.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('Get') and the resource ('menstrual data') for a specific date. However, it does not explicitly distinguish itself from the sibling tool 'get_menstrual_calendar_data' beyond the date-specific wording, so it is not fully self-contained in differentiating alternatives.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives such as 'get_menstrual_calendar_data' or other health-data getters. The description simply states what it does without explaining the appropriate context or selection criteria.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_morning_briefA

Get the morning's measurements in a single call.

Aggregates sleep, recovery (body battery, HRV, resting HR), Garmin's training readiness and today's scheduled workout — five endpoints in one round trip. Body battery is reported both at wake and most-recent, because they answer different questions.

No alerts and no recommendation: those thresholds live in the coaching skill, where they can be read and changed in one place.

Args: date: Date in YYYY-MM-DD format

ParametersJSON Schema
NameRequiredDescriptionDefault
dateYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the burden and does meaningful work: it discloses scope exclusions (no alerts, no recommendations, and why), the aggregation/latency behavior (one round trip vs five), and that body battery is returned at two distinct points (wake and most-recent) because they answer different questions. It does not address failure/partial-data behavior or whether a future date is rejected, which are the remaining gaps.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The purpose and aggregation scope are front-loaded in the first sentence, and each subsequent sentence carries distinct information (dual body-battery reporting, deliberate absence of thresholds). The rationale about thresholds living in the coaching skill is slightly editorial but defensible as it explains an omission an agent might otherwise assume was a bug.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return structure need not be re-described, and the description correctly focuses on what is aggregated and what is deliberately excluded. For a single-parameter read tool this is nearly complete; only edge behavior (missing source data, timezone handling for the date) is unaddressed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, and it does supply the one thing missing: the required 'date' format (YYYY-MM-DD). It adds no timezone or date-range semantics, which matter for a 'morning' aggregate tied to wake time, so it falls short of fully compensating.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('get the morning's measurements') and then enumerates exactly what is aggregated: sleep, recovery (body battery, HRV, resting HR), training readiness, and today's scheduled workout. This clearly distinguishes it from the five sibling tools it wraps (get_sleep_data, get_body_battery, get_training_readiness, get_scheduled_workouts), so an agent can tell it apart without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The 'five endpoints in one round trip' framing strongly implies when to prefer this over the individual siblings, and it explicitly bounds scope by noting no alerts and no recommendations are returned ('those thresholds live in the coaching skill'). It stops short of naming the alternative tools or stating 'use this instead of calling get_sleep_data + get_body_battery + ...', so usage is clear but not fully spelled out.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_morning_training_readinessA

Get morning training readiness score

Returns the morning training readiness assessment, which evaluates recovery status and readiness to train based on overnight metrics.

Args: date: Date in YYYY-MM-DD format

ParametersJSON Schema
NameRequiredDescriptionDefault
dateYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the burden of conveying the operation. It clearly states this is a retrieval operation by saying 'Get' and 'Returns', which implies read-only behavior with no side effects. It also explains what the returned assessment evaluates, giving the agent useful behavioral context beyond the tool name.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact, starts with the core purpose, adds one useful clarifying sentence about what the score evaluates, and includes the essential parameter format. There is no redundant or filler content; every line contributes to correct invocation.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

This is a simple one-parameter retrieval tool, and an output schema exists to define the return shape. The description covers the tool's purpose, the meaning of the score, and the exact date format. No critical information needed to call the tool correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It fully documents the single required parameter, 'date', and provides the crucial format constraint 'YYYY-MM-DD'. This goes beyond the schema's generic title 'Date' and gives the agent everything needed to supply a valid value.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb 'Get' with a specific resource 'morning training readiness score', and explains what the score represents: recovery status and readiness to train based on overnight metrics. The 'morning' qualifier and the reference to overnight metrics help distinguish this from sibling tools like get_training_readiness and get_training_status.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies the tool is for retrieving the morning readiness assessment, and the overnight-metrics wording hints at when it is appropriate. However, it does not explicitly state when to prefer this over related tools such as get_training_readiness or get_body_battery, nor does it mention any exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_non_completed_badge_challengesA

Get badge challenges currently in progress (not yet completed)

Returns active challenges the user has joined but hasn't completed yet. Useful for tracking current progress toward badge goals.

Args: start: Starting index for pagination (starts at 1) limit: Maximum number of challenges to return (default 20, max 100)

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
startNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of transparency. It explains the core functionality and pagination parameters, but it does not mention whether the operation is read-only, what happens when no challenges are found, or any potential errors. It also lacks detail on the order of results, which is important for pagination.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is succinct and well-organized: a one-line summary, a brief explanatory paragraph, and a clear Args section. No unnecessary words or filler, and the structure makes the tool's purpose and parameters easy to scan.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has a simple getter function with an output schema, so the description does not need to explain the return format. It covers the essential context: what the tool does, which challenges it returns, and how pagination works. The description is complete for an agent to decide when and how to use it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema only defines 'start' and 'limit' as integers with defaults, but the description adds critical semantics: start is a 1-based index and limit has a max of 100. This goes beyond the schema and provides the information needed to use the parameters correctly.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description explicitly states that the tool retrieves badge challenges that are currently in progress (not yet completed), distinguishing it from sibling tools like get_badge_challenges (all challenges) and get_available_badge_challenges (joinable ones). It clearly identifies the resource and the specific subset returned.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description mentions it is 'useful for tracking current progress toward badge goals,' which gives context for when to use it. It does not explicitly name alternatives or state when not to use it, but the purpose is clear enough that an agent can infer appropriate usage from the description and sibling tool names.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_nutrition_daily_food_logA

Get daily food consumption records for a date

Returns food items logged throughout the day including calories, macronutrients, and meal associations.

Args: date: Date in YYYY-MM-DD format

ParametersJSON Schema
NameRequiredDescriptionDefault
dateYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It discloses that the tool returns food items logged throughout the day including calories, macronutrients, and meal associations, which sets expectations for a read operation. It does not discuss side effects or prerequisites, but the 'get' semantics and simple single-date scope make those gaps minor.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and well-structured: a purpose sentence, a return-content sentence, and a single Args entry. There is no redundant filler or repetition of schema fields.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple one-parameter daily read with an output schema available, the description is complete. It states the resource scope, the return content categories, and the required date format, leaving no significant gap for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It provides the necessary date format (YYYY-MM-DD) that the schema does not include, and it directly documents the only required parameter. This is sufficient for an agent to invoke the tool correctly.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The opening sentence states a specific verb and resource: 'Get daily food consumption records for a date.' The second sentence clarifies what is returned — food items, calories, macronutrients, and meal associations. This clearly distinguishes it from nutrition siblings such as get_nutrition_daily_meals and get_nutrition_daily_settings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description clearly indicates this tool is for retrieving a single day's food consumption records. It provides a clear context for use, but it does not explicitly name alternatives or state when not to use this tool, so it falls just short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_nutrition_daily_mealsA

Get daily meal summaries for a date

Returns meal-level summaries (breakfast, lunch, dinner, snacks) with nutritional totals for each meal. Each meal includes a mealId needed for logging food items to that meal.

Args: date: Date in YYYY-MM-DD format

ParametersJSON Schema
NameRequiredDescriptionDefault
dateYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It discloses that the tool returns meal summaries and each meal includes a mealId for logging, which is useful. However, it does not explicitly state that this is a read-only operation, nor mention date handling, timezone, or error behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise, well-structured, and front-loaded with the primary purpose. The returns detail and parameter documentation are directly relevant with no filler content.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is simple with one required parameter and an output schema already present. The description explains what the tool returns and why mealId matters, making it sufficient for correct invocation. It could be slightly more complete with explicit sibling differentiation or date edge-case behavior, but these are minor gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It does by documenting the date parameter as YYYY-MM-DD format, which is essential for correct invocation. It does not add additional context like valid date ranges or timezone behavior, but for a single required parameter the format guidance is significant.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool gets daily meal summaries and specifies the resource (breakfast, lunch, dinner, snacks) with nutritional totals. It does not explicitly differentiate from related siblings like get_nutrition_daily_food_log or get_nutrition_daily_settings, but the meal-summary focus is reasonably distinct.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear context: use it to retrieve meal-level summaries and obtain mealId values needed for logging food items. It does not explicitly state when not to use it or name alternative tools for food logs, but the intended use case is clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_nutrition_daily_settingsA

Get nutrition plan/settings for a date

Returns the user's nutrition goals and targets including calorie targets, macronutrient goals, and plan configuration.

Args: date: Date in YYYY-MM-DD format

ParametersJSON Schema
NameRequiredDescriptionDefault
dateYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are present, so the description carries the full burden. It implies a read-only operation via 'Get' and 'Returns' but does not explicitly mention side effects, permissions, rate limits, or error behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is brief and focused, with no redundant wording. It front-loads the purpose and immediately explains the return content and parameter format.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple read-only getter with one parameter, the description is complete: it states what is retrieved, what is returned, and the date format. No output schema is provided, but the description sufficiently describes the expected response.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema provides only 'date' as a string with no description. The tool description adds the crucial format 'YYYY-MM-DD', which fully clarifies the parameter's expected value.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Clearly states the tool 'Get nutrition plan/settings for a date' and describes the returned data: calorie targets, macronutrient goals, and plan configuration. This distinguishes it from sibling tools like set_nutrition_daily_settings and get_nutrition_daily_meals.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No explicit guidance on when to use this tool versus alternatives such as get_nutrition_daily_meals or get_nutrition_daily_food_log. The read-only intent is implied by 'Get' but not stated as a direct instruction.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_performance_trendA

Get per-activity pace or power over time, with the regression slope.

Returns each activity's raw value and average HR, plus the slope across them. No HR normalisation is applied — normalise against the athlete's own thresholds from get_athlete_context if you want that.

Args: metric: Metric to track — "pace" (default) or "power" sport: Sport type — "running" or "cycling" days: Number of days to analyze (default 90)

ParametersJSON Schema
NameRequiredDescriptionDefault
daysNo
sportNorunning
metricNopace

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does a reasonable job: it discloses that the result contains each activity's raw value plus average HR and a regression slope, and warns that no HR normalisation is applied. It omits any note on permissions, limits, or pagination, but for a read-only analytics call this is solid disclosure beyond structured fields.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core capability, then the return shape and the normalisation caveat, ending in an Args block. The Args section largely duplicates the schema, making it slightly longer than necessary, but it remains readable and efficient.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values need not be spelled out, yet the description still clarifies them (raw value, average HR, slope). Combined with the enum values and the normalisation caveat, an agent has everything needed to call this read-only tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate — and it does, describing all three parameters including accepted value sets ('pace'/'power', 'running'/'cycling') that are absent from the schema and would otherwise be guesswork. Defaults are also restated, which is redundant with the schema but harmless.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb and resource ('Get per-activity pace or power over time, with the regression slope') and details the analyzed unit (per activity) and outputs (raw value, average HR, slope). It distinguishes itself from sibling trend tools by being about pace/power movement across activities, though it never names those siblings explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied by the metric/sport framing, and it correctly points to get_athlete_context for HR normalisation — a useful cross-reference. However, it gives no guidance on when to choose this tool versus the many other *_trend siblings (vo2max, hrv, respiration, training_load).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_personal_recordD

Get personal records for user

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

D1.8/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of disclosing behavior. It only says 'get', implying a read operation, but provides no detail on side effects, required permissions, data scope, or output characteristics. This is insufficient for a tool that likely returns sensitive personal data.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single short sentence, but it is under-specified rather than concise. It lacks essential detail, so the brevity is not a virtue. A concise description would still convey purpose and usage in a few words without losing meaning.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the lack of annotations, no parameter explanations, and no description of the output schema, this description is grossly incomplete. The tool likely returns a set of records, but the agent has no idea what to expect, how to interpret the results, or when to call it. This is far below the minimum viable standard.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, so the schema is fully covered. The baseline for 0 params is 4. The description adds no parameter-specific meaning because there are no parameters, but this is acceptable given the absence of parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Get personal records for user' is essentially a restatement of the tool name with minimal added context. It names a verb and resource but is vague about what 'personal records' means (e.g., fitness records, personal bests?) and does not distinguish this from many similar get_* siblings like get_user_summary or get_stats. It barely clarifies the tool's specific function.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives. No context, no exclusions, no mention of what distinguishes it from the dozens of other get_* tools. An agent has no way to know which scenario calls for this tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_power_duration_curveA

Get season-best Power Duration Curve across recent activities.

Downloads FIT files for recent cycling activities and computes best mean maximal power at each standard duration. Returns season bests with which activity and date each best came from.

Durations: 5s (sprint), 30s, 1min, 5min (VO2 max proxy), 10min, 20min (FTP proxy), 60min

Use the 20-minute best × 0.95 as a strong FTP estimate without a formal test.

Warning: downloads multiple FIT files — may take 30-60 seconds for 20 activities.

Args: num_activities: Number of recent activities to analyze (default 20, max 50) activity_type: Activity type to filter (default "cycling")

ParametersJSON Schema
NameRequiredDescriptionDefault
activity_typeNocycling
num_activitiesNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It transparently reveals that it downloads multiple FIT files and may take 30-60 seconds. It also mentions it computes and returns season bests. However, it does not explicitly state it is a read-only operation or discuss any side effects beyond downloads, lacking full transparency.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is slightly verbose with repeated mentions of downloading and durations, but it is well-structured and front-loaded with the core purpose. It earns a 4 because the extra details like the FTP formula and warning are valuable, though some redundancy exists.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description provides a comprehensive context: it explains the output ('Returns season bests with which activity and date each best came from'), includes the standard durations, and gives a practical usage tip for FTP estimation. It is complete for the tool's complexity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The description adds meaning to both parameters beyond the schema: 'num_activities' is explained as number of recent activities with a max of 50, and 'activity_type' is explained as a filter with default 'cycling'. This fully covers the parameter semantics.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states 'Get season-best Power Duration Curve across recent activities' with a specific verb and resource. It distinguishes itself from sibling tools like get_cycling_ftp and get_lactate_threshold by focusing on the power duration curve derived from multiple activities.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explains when to use the tool, especially for estimating FTP via the 20-minute best × 0.95 formula, and includes a warning about download time. However, it does not explicitly mention alternatives or when not to use it, though the FTP estimation context is clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_pregnancy_summaryA

Get pregnancy summary data

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, and the description only says 'Get', which implies a read-only operation. It does not explicitly confirm side effects or lack thereof, but for a simple getter the inference is reasonable.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely concise, consisting of a single sentence with no unnecessary words. It is well-structured and easily digestible.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the simplicity of the tool (no parameters, likely a simple summary return), the description is adequately complete. The output schema is stated to exist, so return values are covered externally.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has no parameters, so schema coverage is technically 100%. The baseline for high coverage is 3, and the description adds no additional parameter-related information, which is acceptable given the absence of parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action (Get) and the resource (pregnancy summary data), which is distinct from all sibling tools. However, it lacks detail on what the summary includes, but it is sufficient for basic identification.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit guidance on when to use this tool versus alternatives. While it is likely the only tool for pregnancy summary, the description does not state this or provide any context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_primary_training_deviceB

Get information about the primary training device

Returns details about the device designated as primary for training metrics, along with other wearable devices on the account.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of disclosing behavioral traits. It states that the tool returns details, but does not explicitly mention whether it is read-only, what data format is returned, or any side effects, permissions, or rate limits.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise and well-structured, consisting of two clear sentences. It avoids unnecessary details and gets straight to the point, making it easy for an agent to parse quickly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple, parameterless getter, the description is sufficiently complete. It clearly states what information is returned and the concept of a primary training device. It could slightly benefit from noting how this differs from get_devices, but the current level is adequate for the tool's simplicity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has no parameters, so schema coverage is effectively 100% and there is no parameter information to add. The description does not need to explain parameters, but it also does not provide any extra semantic value beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool gets information about the primary training device and explains what is returned, including details about the designated primary device and other wearable devices. This distinguishes it from the more generic get_devices sibling in the provided context.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no explicit guidance on when to use this tool versus alternatives such as get_devices or get_device_last_used. There is no mention of appropriate use cases, exclusions, or conditions that would help an agent choose this tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_progress_summary_between_datesB

Get progress summary for a metric between dates

Args: start_date: Start date in YYYY-MM-DD format end_date: End date in YYYY-MM-DD format metric: Metric to get progress for (e.g., "elevationGain", "duration", "distance", "movingDuration")

ParametersJSON Schema
NameRequiredDescriptionDefault
metricYes
end_dateYes
start_dateYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.2/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

There are no annotations provided, and the description does not disclose any side effects, read-only nature, or required permissions. The user is left without information about the tool's behavior beyond its apparent purpose.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise and well-structured, leading with the primary purpose in the first sentence and clarifying parameters in the second. Every sentence adds value without redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple getter, the description covers the essential purpose and parameters. Since an output schema exists, the return value does not need to be explained. Minor caveats like possible empty results are not mentioned but are not critical.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The description adds meaning to all three parameters (start_date, end_date, metric) by specifying formats and giving an example for metric. This compensates for the schema's lack of descriptions, though it could be more explicit about valid metric values.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('Get') and the resource ('progress summary') with a specific scope ('between dates' for a given metric). This unambiguously distinguishes it from the many other getter tools in the sibling list.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is given on when to use this tool versus alternatives. It does not mention any prerequisites, limitations, or contrast with similar getters like 'get_stats' or 'get_goals'.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_race_predictionsA

Get predicted race times based on current fitness level

Returns Garmin's predictions for 5K, 10K, half marathon, and marathon finish times based on the user's recent training data and VO2 max.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

There are no annotations to indicate side effects or read-only behavior. The description states that it 'Returns' predictions, which implies a read-only operation, and the tool name begins with 'get', reinforcing this. However, it does not explicitly mention that no data is modified, so it partially relies on convention. Given the absence of annotations, this score is appropriate.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise and well-structured, consisting of two sentences. The first sentence is a clear summary, and the second provides essential details about the output (distances and basis). No unnecessary information is present.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple getter with no parameters, the description is complete. It specifies what is returned (predictions for four race distances) and the underlying data (recent training data and VO2 max). It does not describe an output schema, but the presence of an output schema is indicated in the context, and the description is sufficient for an agent to invoke the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, and the schema coverage is complete (100%). The description does not need to explain parameters. According to the baseline rule for 0 parameters, this scores a 4; the description adds no parameter-specific details because none exist.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states that the tool retrieves predicted race times for specific distances (5K, 10K, half marathon, marathon) based on fitness data. It uses the specific verb 'Get' and identifies the resource ('predicted race times'), making its purpose unambiguous and distinct from sibling tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implicitly indicates when to use the tool (when the user wants race time predictions based on training data and VO2 max). It does not explicitly differentiate from alternatives, but given the tool's unique function among many getters, the usage context is clear enough. No explicit exclusions are needed.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_respiration_dataA

Get full respiration time-series data

Note: This returns detailed interval data (~20KB). For a compact summary, use get_respiration_summary().

Args: date: Date in YYYY-MM-DD format

ParametersJSON Schema
NameRequiredDescriptionDefault
dateYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided (no readOnlyHint or destructiveHint), so the description carries the burden. It mentions the response size (~20KB) and implies a read-only 'Get' operation, but it doesn't explicitly state that there are no side effects or that the request may be large/slow beyond the size note.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is short and well-structured: a clear purpose, a note about data size, an alternative reference, and an Args section. Every sentence adds value and there is no redundant content.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the single parameter and no output schema shown, the description covers the key context: what data is returned, its approximate size, the alternative for summaries, and the expected date format. It's sufficient for an agent to call it correctly, though a brief note on the response structure would fully round it out.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema provides only 'date' with type string and no description (0% schema coverage). The description compensates by specifying the format 'YYYY-MM-DD', which adds meaning beyond the schema. However, it doesn't explain what date range or timezone assumptions apply, so a small gap remains.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Get') and resource ('full respiration time-series data'). It explicitly distinguishes itself from the sibling tool get_respiration_summary, so an agent can identify which tool to use without opening other schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly names the alternative get_respiration_summary() and gives the condition for using it ('For a compact summary'). It also notes the data size (~20KB), providing clear guidance on when to choose this tool versus the summary.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_respiration_summaryA

Get respiration summary with essential metrics (lightweight version)

Returns a compact summary (~300 bytes) instead of full time-series data (~20KB).

Args: date: Date in YYYY-MM-DD format

ParametersJSON Schema
NameRequiredDescriptionDefault
dateYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It usefully discloses that the response is a compact summary (~300 bytes) rather than full time-series data (~20KB), but it does not describe exact included metrics, possible date restrictions, or any other behavioral nuances beyond the size contrast.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and front-loaded: it states the purpose, immediately gives the key behavioral distinction (summary vs full data), then documents the single parameter. Every sentence contributes useful information with no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple one-parameter read operation with a provided output schema, the description is largely complete: it explains the tool's purpose, the compact nature of the response, and the required date format. It falls short only by not explicitly naming sibling tools for differentiation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema provides only the parameter name 'date' with type string and no description, so the description's explicit 'Date in YYYY-MM-DD format' adds the necessary semantic detail. For a single-parameter tool, this fully compensates for the 0% schema description coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool retrieves a respiration summary and frames it as the lightweight version that returns compact data rather than full time-series. It distinguishes the tool from the full-data sibling implicitly, though it does not name get_respiration_data or get_respiration_trend explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'instead of full time-series data' gives an implied usage context: use this when a compact summary is sufficient and full data is not needed. However, it does not explicitly name the alternative tool or state when to prefer a different respiration endpoint such as get_respiration_trend.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_respiration_trendA

Get overnight respiration rate trend over a date range.

Elevated resting respiration rate (compared to personal baseline) is an early warning sign for overreaching, illness, or poor recovery. Use this alongside HRV trend for a complete recovery picture.

Recommended range: 7-21 days. Maximum: 30 days.

Args: start_date: Start date in YYYY-MM-DD format end_date: End date in YYYY-MM-DD format

ParametersJSON Schema
NameRequiredDescriptionDefault
end_dateYes
start_dateYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

There are no annotations, so the description carries the burden. It adds domain context (elevated respiration as a warning sign) and a range constraint, but it does not disclose edge-case behavior such as what happens beyond 30 days, timezone handling, or the exact definition of 'overnight.' The 'Get' verb implies read-only behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well organized: purpose, domain context, usage guidance, then args. Each sentence adds value, though the clinical context sentence is helpful but not strictly necessary for invoking the tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The output schema covers return values, and the description covers the two required parameters with format and range guidance. It is complete enough for a simple read-only trend tool, but it could be stronger by explicitly distinguishing this from get_respiration_data and get_respiration_summary.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It documents both parameters with format ('YYYY-MM-DD') and adds a recommended/maximum range, which is meaningful beyond the bare schema that only declares them as strings.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'Get overnight respiration rate trend over a date range.' It clearly conveys what the tool returns and the 'trend' wording distinguishes it from sibling tools like get_respiration_data and get_respiration_summary, though it does not explicitly name those alternatives.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It provides clear usage context: 'Use this alongside HRV trend for a complete recovery picture' and gives a 'Recommended range: 7-21 days. Maximum: 30 days.' This tells an agent when the tool is appropriate, but it does not state when to prefer get_respiration_data or get_respiration_summary instead.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_rhr_dayC

Get resting heart rate data

Args: date: Date in YYYY-MM-DD format

ParametersJSON Schema
NameRequiredDescriptionDefault
dateYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The verb 'Get' implies a read-only operation, but with no annotations the description carries the full burden of explaining behavior. It does not state what is returned, whether any side effects occur, or how the response is structured, leaving the actual behavior under-specified.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is short and free of irrelevant content, with the Args block clearly listing the parameter. It is not padded, though the first sentence is somewhat generic and could have been more specific without adding length.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema and a large sibling list, the description leaves important context missing: it does not state what the returned resting heart rate data looks like, what units are used, or how this endpoint differs from the many other heart-rate and summary tools. The single date argument helps, but the overall contract is incomplete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The description adds the date format 'YYYY-MM-DD' beyond the raw schema, which is helpful for the single date parameter. However, it does not clarify the meaning of the date, timezone considerations, or any constraints such as valid ranges or historical limits.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states the verb and resource, 'Get resting heart rate data,' but uses the vague term 'data' and does not explicitly mention the single-day scope that the tool name and date argument imply. It also lacks any differentiation from sibling tools like get_heart_rates or get_heart_rates_summary.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no guidance on when to use this tool instead of related heart-rate tools. There is no mention of scenarios, exclusions, or alternatives, so an agent cannot distinguish this endpoint from get_heart_rates, get_heart_rates_summary, or other cardiovascular endpoints.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_scheduled_workoutsA

Get scheduled workouts between two dates with curated summary list

Returns workouts that have been scheduled on the Garmin Connect calendar, including their scheduled dates and completion status.

Args: start_date: Start date in YYYY-MM-DD format end_date: End date in YYYY-MM-DD format

ParametersJSON Schema
NameRequiredDescriptionDefault
end_dateYes
start_dateYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the responsibility of conveying behavior. It states the tool 'Returns workouts' and is phrased as a read-only operation, which clearly implies no side effects. It does not explicitly mention that data is unchanged, but the 'Get' verb and return wording are sufficient for a user to infer read-only behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is composed of two brief, focused sentences: a one-line summary and a slightly more detailed explanation. There is no redundant information, jargon, or filler. It efficiently conveys the essential purpose and output without unnecessary length.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description confirms the output is a summary list including scheduled dates and completion status, giving users a preview of the return structure. Since an output schema exists but is not shown here, the description cannot fully describe the response format. However, it provides enough context for basic usage and expectations.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The description specifies that start_date and end_date are in YYYY-MM-DD format and that the query is 'between two dates'. This gives clear semantics for the two required parameters, though it does not state whether the range is inclusive or exclusive. The schema only lists string types, so the description adds important format details.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool retrieves scheduled workouts for a date range and returns a summary list with dates and completion status. It uses the verb 'Get' and specifies the resource ('scheduled workouts'), making the purpose obvious. It does not explicitly contrast with sibling tools like schedule_workout or get_workouts, but the meaning is still unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explains what the tool does but does not provide explicit guidance on when to use it versus other workout-related tools (e.g., get_workouts for saved templates, schedule_workout for creating schedules). It implies usage for viewing the calendar of scheduled workouts, but the absence of alternative references leaves some ambiguity.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_sleep_dataA

Get full sleep data with all details

Note: This returns detailed sleep data (~50KB). For a compact summary, use get_sleep_summary().

Args: date: Date in YYYY-MM-DD format

ParametersJSON Schema
NameRequiredDescriptionDefault
dateYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It adds one useful trait—the response is detailed and roughly 50KB—which is relevant for an agent deciding whether to call it. However, it does not mention authentication, rate limits, or other behavioral constraints, though as a read-only getter the risk is modest.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and well-structured with the purpose first, followed by the critical payload-size note and then the argument format. Every sentence earns its place, and there is no redundant or vague wording.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the single required parameter, an output schema, the explicit date format, and the pointer to a lighter-weight alternative, the description provides enough context for correct invocation. The presence of an output schema means return values do not need to be explained in the description.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema only defines 'date' as a string with no format information, and schema description coverage is 0%. The description compensates by specifying the exact expected format (YYYY-MM-DD). For a single-parameter tool, this fully clarifies the input requirement.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource ('Get full sleep data') and indicates it provides all details. It also distinguishes the tool from the sibling get_sleep_summary by framing this as the detailed alternative, so an agent can differentiate between them without examining schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly names the sibling tool get_sleep_summary and provides the selection criterion: use that tool when a compact summary is desired. This clearly signals when to prefer the alternative, effectively telling the agent when this tool is the right choice.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_sleep_summaryA

Get sleep summary with only essential metrics (lightweight version)

This endpoint returns a compact summary of sleep data (~350 bytes) instead of the full granular data (~50KB). Ideal for daily health checkups and LLM integrations where the full time-series data would overwhelm the context window.

Args: date: Date in YYYY-MM-DD format

ParametersJSON Schema
NameRequiredDescriptionDefault
dateYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the burden of behavioral disclosure. It states that the endpoint returns a compact summary rather than full time-series data, which is a meaningful behavioral trait. It also explains the payload size tradeoff. It does not mention error behavior or auth, but for a simple read-only getter this is reasonably transparent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and front-loaded: the first line states the purpose and the lightweight nature, followed by a clear use-case sentence and a minimal Args section. Every sentence adds useful information without redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is simple with one required parameter and an output schema present, so return-value documentation is already covered. The description covers the key tradeoff and parameter format. It could optionally point to get_sleep_summary_range for multi-day needs, but that is not essential for this endpoint's daily-summary purpose.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema provides no description for the date parameter (0% coverage), but the description compensates by specifying 'Date in YYYY-MM-DD format.' This adds concrete format guidance beyond the schema's bare string type, though it does not elaborate on timezone or edge-case semantics.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb 'Get', the resource 'sleep summary', and the key differentiator: 'only essential metrics (lightweight version)' versus full granular data. It also quantifies the difference (~350 bytes vs ~50KB), making it easy for an agent to distinguish this from get_sleep_data.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives a clear usage context: 'Ideal for daily health checkups and LLM integrations where the full time-series data would overwhelm the context window.' It implies the alternative is the full-granularity endpoint, though it does not explicitly name get_sleep_data or get_sleep_summary_range as alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_spo2_dataB

Get SpO2 (blood oxygen) data

Args: date: Date in YYYY-MM-DD format

ParametersJSON Schema
NameRequiredDescriptionDefault
dateYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden of behavioral disclosure, but it only states that SpO2 data is retrieved. It does not mention what happens when no data exists for the date, whether the returned value has units or a specific structure, or any other behavioral traits. The name implies a read operation, but the description adds little beyond that.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two short sentences and front-loads the tool's purpose before the argument documentation. There is no filler or repetition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple one-parameter getter with an output schema available, the description covers the essential input contract. It is slightly thin on behavior/no-data semantics, but that is partially mitigated by the output schema and the low tool complexity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The only parameter, date, is given meaning beyond the schema by specifying 'YYYY-MM-DD format'. Since schema description coverage is 0%, this is essential and adequate for a single-parameter tool, though it does not clarify timezone or what date the metric refers to.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific action ('Get') and resource ('SpO2 (blood oxygen) data'), making the core purpose immediately clear. It does not explicitly differentiate itself from sibling data getters, but SpO2 is a distinct health metric so confusion is unlikely.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is given about when to use this tool versus alternatives like get_respiration_data or get_heart_rates. The description simply states what it does and provides no conditions, prerequisites, or exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_statsA

Get daily activity stats with curated essential metrics

Returns a summary of daily health and activity data including steps, calories, heart rate, stress, body battery, and sleep metrics.

Args: date: Date in YYYY-MM-DD format

ParametersJSON Schema
NameRequiredDescriptionDefault
dateYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden of behavioral transparency. It states that the tool returns a summary of daily health metrics, which is clear, but it does not disclose any potential side effects, data source nuances, or response formatting details. This is sufficient for a simple read operation, though not richly informative.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact, using two sentences to state the purpose and list the included metrics, followed by a clear Args section. It is front-loaded with the primary action and avoids any redundant or unnecessary wording.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple read-only tool with one parameter and no output schema, the description is nearly complete. It specifies the input format and the type of data returned, leaving no critical gaps for the agent to make a correct call. Minor details like timezone handling are absent but not essential for the tool's basic use.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema only defines 'date' as a required string with no description. The tool description adds crucial format information ('YYYY-MM-DD'), which goes beyond the schema and helps the agent provide a valid input. This increases the semantic clarity of the parameter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb 'Get' and the resource 'daily activity stats', and it enumerates the specific metrics included (steps, calories, heart rate, stress, body battery, sleep). This distinguishes it from sibling tools that retrieve only individual metrics (e.g., get_steps_data, get_body_battery), though it does not explicitly name or contrast with those alternatives.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description does not provide explicit guidance on when to use this tool versus the many sibling tools. It implies a 'curated summary' but gives no direct instruction such as 'use this for a quick overview instead of calling multiple endpoints', leaving the agent to infer the intended use case.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_stats_and_bodyB

Get stats and body composition data

Args: date: Date in YYYY-MM-DD format

ParametersJSON Schema
NameRequiredDescriptionDefault
dateYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are present, and the description does not disclose whether the operation is read-only, whether it has side effects, or any other behavioral characteristics. Since the name implies a GET, it is likely read-only, but this is not explicitly stated, and the lack of annotations leaves the burden on the description.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise and to the point, using a single sentence that directly states the purpose. There is no extraneous information or verbose wording.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description explains what data is retrieved (stats and body composition) but does not specify the structure or content of the response. Given the tool has no output schema, this could lead to uncertainty about the exact fields returned. However, for a simple getter, this level of detail may be acceptable.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The only parameter, date, is clearly described with the expected format (YYYY-MM-DD). This provides sufficient semantic meaning for the parameter, though no additional constraints (e.g., required range) are mentioned.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action (get) and the resource (stats and body composition data), distinguishing it from sibling tools like get_stats or get_body_composition by combining both. However, 'stats' is somewhat vague and could be interpreted as multiple data types.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus the many sibling tools that retrieve similar data (e.g., get_stats, get_body_composition, get_user_summary). The description does not mention any conditions or preferences for choosing this tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_steps_dataA

Get detailed steps data with 15-minute intervals

Note: This returns full interval data (~14KB). For a compact summary, use get_stats() which includes total_steps.

Args: date: Date in YYYY-MM-DD format

ParametersJSON Schema
NameRequiredDescriptionDefault
dateYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description discloses a key behavioral trait: the response is large (~14KB) and contains full interval data, which is useful for expectations and resource planning. It could also mention read-only or timezone handling, but for a simple read operation this is sufficient.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is brief and front-loaded with the main purpose, followed by a practical note about payload size and an alternative. Every sentence adds value with no wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is simple with one parameter, and the description covers the data granularity, payload size, and a sibling alternative. Since an output schema exists, return structure is already available, so nothing critical is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The only parameter `date` has zero description coverage in the schema, but the description compensates fully by specifying the exact YYYY-MM-DD format. This is exactly the meaning an agent needs to call the tool correctly.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool returns detailed steps data at 15-minute intervals, which is a specific verb, resource, and granularity. It also distinguishes this tool from the compact get_stats alternative.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly tells agents to use get_stats() when a compact summary with total_steps is needed, providing a clear condition and alternative. This guides tool selection without ambiguity.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_stress_dataA

Get full stress time-series data

Note: This returns detailed interval data (~35KB) including body battery. For a compact summary, use get_stress_summary().

Args: date: Date in YYYY-MM-DD format

ParametersJSON Schema
NameRequiredDescriptionDefault
dateYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It discloses the response size (~35KB), that it returns interval-level detail, and that body battery is included. These are meaningful behavioral traits beyond the schema. It does not discuss errors or availability, but for a simple read operation with an output schema this is reasonably transparent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact, front-loaded with the core purpose, and includes only useful additions: data size, contents, and a pointer to the summary alternative. The Args section is minimal and directly relevant.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-parameter read tool with an output schema, this description covers the purpose, the parameter format, the payload size, and the main alternative. Nothing essential for an agent to select and call it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It does so by specifying the date format ('YYYY-MM-DD') and implying the date identifies the day for which stress data is returned. This adds real meaning beyond the schema's bare 'date' string property.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states a specific verb and resource: 'Get full stress time-series data.' It also distinguishes itself from get_stress_summary by noting it returns detailed interval data including body battery, so an agent can tell it apart from the closest sibling without ambiguity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly provides an alternative and the condition for choosing it: 'For a compact summary, use get_stress_summary().' This gives clear when-to-use guidance relative to the primary sibling and implies this tool should be used when full detail is needed.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_stress_summaryA

Get stress summary with essential metrics (lightweight version)

Returns a compact summary (~400 bytes) instead of full time-series data (~35KB). Ideal for daily health checkups and LLM integrations.

Args: date: Date in YYYY-MM-DD format

ParametersJSON Schema
NameRequiredDescriptionDefault
dateYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description must carry behavioral disclosure. It clearly notes the compact output size (~400 bytes vs ~35KB), making the response behavior predictable, but it omits potential side effects, errors, or read-only semantics.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Description is brief and to the point, with no redundant information. The byte-size comparison and ideal-use note add value without bloating the text.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple one-parameter tool, the description provides the necessary context: what it returns (compact summary), why it exists (lightweight), and the input format. It does not enumerate the specific metrics in the summary, but the core use is clear.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema only defines 'date' as a required string with no property description, so the docstring's explicit 'YYYY-MM-DD' format is essential and compensates for the schema gap. Additional semantics like allowed date ranges or timezone are not covered.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Clearly identifies the resource (stress summary), the action (get), and distinguishes itself as a lightweight version returning compact data instead of full time-series. This differentiates it from sibling tools like get_stress_data.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

States ideal use cases (daily health checkups, LLM integrations) and explicitly contrasts with full time-series data, guiding when to prefer this tool. It does not name sibling alternatives explicitly, but the usage context is clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_training_effectB

Get training effect data for a specific activity

Args: activity_id: ID of the activity to retrieve training effect for

ParametersJSON Schema
NameRequiredDescriptionDefault
activity_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, and the description does not mention side effects, permissions, or read-only nature; it only states the action.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise and well-structured, with the main purpose first followed by the parameter explanation.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers the essential calling requirement (activity_id), and with an output schema present, no return details are needed, so it is complete for its simplicity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The description explains the single activity_id parameter meaning, though it lacks format or example details.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool retrieves training effect data for a specific activity, but it does not explicitly differentiate it from sibling getter tools, so a 4 is appropriate.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool compared to other activity-related getters; the description lacks alternative references.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_training_load_balanceA

Get Garmin's Load Focus — the distribution of the trailing-month training load across Aerobic Low, Aerobic High, and Anaerobic intensity bands, plus the system's feedback phrase (e.g. AEROBIC_HIGH_SHORTAGE, BALANCED, ANAEROBIC_SHORTAGE).

Use this to assess whether the athlete's training mix is balanced or deficient in a particular intensity band. Each band reports its load alongside Garmin's target range; a status of "below", "within", or "above" is computed from the load relative to that range.

Args: date: Date in YYYY-MM-DD format

ParametersJSON Schema
NameRequiredDescriptionDefault
dateYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden. It discloses meaningful logic: each band reports load alongside a target range, and a status of below/within/above is computed relative to that range. It also explains the feedback phrase with examples. It does not explicitly state read-only behavior, but the 'get' verb and absence of mutation language make it reasonably inferable.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-structured and front-loaded: a one-sentence definition, a one-sentence use case, a compact explanation of the output status, and a single parameter listing. Every sentence contributes necessary information without fluff or repetition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter getter with an output schema available, the description is complete enough. It covers what the tool returns, how status is derived, what the feedback phrase means, and the expected date format. The output schema presumably handles the full return structure, so the description does not need to re-document it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It provides the date format ('YYYY-MM-DD') and gives context by describing the trailing-month load distribution, implying that the date anchors the trailing month. This adds meaning beyond the bare schema field, though it could further clarify valid date ranges or what happens if no data exists.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Get Garmin's Load Focus,' and immediately defines the exact content: the distribution of training load across three intensity bands plus a feedback phrase. It is clearly distinct from siblings like get_training_load_trend because it names the specific metric and its components.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description includes an explicit use statement: 'Use this to assess whether the athlete's training mix is balanced or deficient in a particular intensity band.' It gives clear context for when to invoke the tool, though it does not mention alternatives or when not to use it, which prevents a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_training_load_breakdownB

Get training load breakdown by sport for a date range.

Returns total minutes and percentage distribution across running, cycling, swimming and other sports, plus Garmin's acute/chronic load, ACWR and TSB as of the end date.

Args: start_date: Start date in YYYY-MM-DD format end_date: End date in YYYY-MM-DD format

ParametersJSON Schema
NameRequiredDescriptionDefault
end_dateYes
start_dateYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden, and it does disclose the composition of the response (minutes, percentage distribution, acute/chronic load, ACWR, TSB as of the end date). However it says nothing about permissions, data availability limits, or whether the computation is expensive/slow — gaps for a tool with zero annotation coverage.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The payload sentence is front-loaded and the Args block is compact and purposeful given the empty schema descriptions. Slightly redundant to enumerate return values when an output schema exists, but it is not verbose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values need not be explained (though they are, harmlessly). With no annotations, the description covers date formats and the analytical content well; only the lack of sibling differentiation and edge-case behavior keeps it from being fully complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the schema contributes nothing about the two date parameters. The description compensates by documenting both start_date and end_date with the exact YYYY-MM-DD format, and clarifies that end_date anchors the ACWR/TSB computation. It does not explain behavior for inverted or very long ranges, but the core semantics are covered.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Get') and resource ('training load breakdown by sport') with scope ('for a date range'), which is far more informative than the bare name. It partially distinguishes itself from near-siblings like get_training_load_balance and get_training_load_trend by saying 'breakdown by sport', but never explicitly contrasts with them.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no prerequisites, and no mention of the closely related training-load siblings (trend, balance, weekly progression). The agent must infer usage entirely from the description of returned data.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_training_load_trendA

Get the Performance Management Chart (CTL/ATL/TSB) over a date range.

Returns Chronic Training Load (CTL, 42-day fitness), Acute Training Load (ATL, 7-day fatigue), Training Stress Balance (TSB = CTL - ATL, form/freshness), and Acute:Chronic Workload Ratio (ACWR) per day. Use this to assess whether the athlete is building fitness, peaking, or accumulating too much fatigue.

Recommended range: 4-8 weeks. Maximum: 90 days.

Args: start_date: Start date in YYYY-MM-DD format end_date: End date in YYYY-MM-DD format

ParametersJSON Schema
NameRequiredDescriptionDefault
end_dateYes
start_dateYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the behavioral burden and handles it well by explaining each metric's time window (42-day fitness, 7-day fatigue), the TSB formula, and the ACWR. It adds date-range limits (recommended 4-8 weeks, max 90 days) and per-day output semantics. It does not cover potential error conditions or data-availability behavior, but those are minor for a read-only trend query and the output schema covers return structure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well organized, front-loaded with the core purpose, followed by return details, usage guidance, and parameter documentation. Every sentence adds value and is skimmable, with formulas and ranges clearly highlighted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given a two-parameter tool, a provided output schema, and no annotations, this description gives an agent everything needed to select and call the tool: purpose, metric definitions, use case, limits, and date format. The only minor omissions (error handling, inclusivity of end date) are unlikely to hinder correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, and the description fully compensates by documenting both parameters with their required YYYY-MM-DD format. It also adds practical range constraints (recommended 4-8 weeks, maximum 90 days). This is exactly the semantic context the bare schema lacks.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description uses a specific verb 'Get' plus resource 'Performance Management Chart (CTL/ATL/TSB)' and enumerates the returned metrics (CTL, ATL, TSB, ACWR), clearly distinguishing it from sibling training-metric tools. It also defines the metrics in parenthetical terms, leaving no ambiguity about what the tool does.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly states when to use the tool: 'Use this to assess whether the athlete is building fitness, peaking, or accumulating too much fatigue.' It also gives a recommended range (4-8 weeks) and a maximum (90 days). It does not name alternative sibling tools or explicitly state when not to use it, so it stops just short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_training_plan_workoutsA

Compatibility alias for get_garmin_coach_workouts

Prefer get_garmin_coach_workouts for new requests. This legacy tool returns the same Garmin Coach/training-plan data; do not call both for one request. Adaptive plans expose only Garmin's currently generated window, typically the current week; future dates may return no workouts even while a plan is active.

Adaptive training plans typically expose workout_uuid; other plan families may expose numeric workout_id. Pass whichever identifier is present to get_workout_by_id. The returned count includes rest days.

Args: calendar_date: Reference date in YYYY-MM-DD format (returns week's workouts)

ParametersJSON Schema
NameRequiredDescriptionDefault
calendar_dateYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden, and it delivers. It discloses the alias behavior, the fact that adaptive plans only expose the current generated week, the possibility that future dates return no workouts, and that the returned count includes rest days. These non-obvious behavioral details are highly useful to an agent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact yet information-dense, with a clear alias statement, usage guidance, edge-case caveats, and an Args section. Every sentence contributes actionable information with no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given an output schema exists, return-value details are not needed. The description covers the legacy status, preferred alternative, date format, plan-window limitation, rest-day behavior, and how to route identifiers to get_workout_by_id. This is complete for correct selection and invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema only provides a parameter name and type with 0% description coverage. The description compensates fully by specifying the YYYY-MM-DD format, the reference-date semantics, and that it returns that week's workouts.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly identifies the tool as a compatibility alias for get_garmin_coach_workouts and states it returns the same Garmin Coach/training-plan data. This precise resource and explicit sibling relationship distinguish it from the many sibling tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly instructs agents to prefer get_garmin_coach_workouts for new requests, labels this tool as legacy, and warns against calling both for one request. This gives unambiguous when-to-use and when-not-to-use guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_training_readinessB

Get training readiness data with curated metrics

Returns training readiness score and contributing factors.

Args: date: Date in YYYY-MM-DD format

ParametersJSON Schema
NameRequiredDescriptionDefault
dateYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, and the description does not disclose whether the operation is read-only, requires authentication, or has any side effects. The verb 'Get' implies non-destructive behavior, but this is not explicitly stated.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is very concise, consisting of two short sentences. It immediately states the purpose and mentions the return value, with no extraneous information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description mentions the return value (score and contributing factors) but does not elaborate on the structure or meaning. Given the single parameter and straightforward nature of the data, this is adequate but leaves some ambiguity about what 'training readiness' specifically includes.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The description explicitly states the date format (YYYY-MM-DD) and implies that the parameter represents the date for which readiness data is requested. This provides sufficient meaning beyond the bare schema definition.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb 'Get' and the resource 'training readiness data', and mentions that it returns a score and contributing factors. However, it does not differentiate from sibling tools like 'get_morning_training_readiness' or 'get_training_status', which could lead to ambiguity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives. The description does not mention any specific use cases or scenarios where this endpoint is preferred over similar readiness or training tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_training_readiness_compositeB

Get Garmin's training readiness score with its factor breakdown.

Returns Garmin's own score and the percentage each of its six factors contributed — sleep, recovery time, HRV, acute load, sleep history, stress history. No interpretation: what the score means for today's session is a coaching decision.

Args: date: Date in YYYY-MM-DD format

ParametersJSON Schema
NameRequiredDescriptionDefault
dateYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It does disclose the factor breakdown and explicitly bounds the tool's scope (raw score, no interpretation), which is genuine added context. It is silent on data-availability behavior (e.g. what is returned when no readiness data exists for the date) and on whether the read requires a synced device.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded: the purpose and the return composition come first, followed by the scope disclaimer, then the argument. Every sentence earns its place; the 'Args:' block is slightly formal/redundant against the schema but conveys the date format.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be re-explained (though they usefully are). The remaining gap is sibling routing in a family dense with readiness/status tools, plus no mention of data-availability failure modes — the description is adequate but incomplete for this context.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

One parameter with 0% schema description coverage — the schema only says type string, title 'Date'. The description compensates by specifying the expected format, 'Date in YYYY-MM-DD format', which is exactly the information the schema omits. It does not cover edge cases like date ranges or future dates.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Specific verb + resource ('Get Garmin's training readiness score') with the return composition spelled out: the score plus the percentage contribution of six named factors. However, it never distinguishes itself from close siblings like get_training_readiness or get_morning_training_readiness, so the agent cannot tell from the description alone which of the three to call.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance is given. The line 'No interpretation: what the score means for today's session is a coaching decision' scopes the tool's output but does not tell the agent when to pick this over get_training_readiness or get_morning_training_readiness, nor does it state any prerequisite such as needing recent device data.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_training_statusB

Get training status with curated metrics

Returns comprehensive training status including load, VO2 max, recovery, and training readiness indicators.

Args: date: Date in YYYY-MM-DD format

ParametersJSON Schema
NameRequiredDescriptionDefault
dateYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. It states that the tool 'returns' training status but does not explain any operational behavior, such as required permissions, whether the date is a single day or a range, how metrics are computed, or how the tool behaves with invalid or missing data. This is a real gap for an unannotated tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is short and front-loaded with the core purpose. The list of metrics is useful, but the opening phrase 'Get training status with curated metrics' is somewhat redundant with the 'Returns comprehensive training status' line, so it is not perfectly tight.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The output schema exists, so return structures are documented elsewhere. However, given the large sibling list and the lack of usage guidance, the description does not fully orient an agent on when this tool is the right choice or what distinguishes it from adjacent training metrics tools. It is adequate but has clear contextual gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema description coverage is 0%, so the description must compensate. It does add meaningful detail by specifying the date format as YYYY-MM-DD and naming the parameter. With only one parameter, this is nearly sufficient, though it could clarify whether the date represents a single day or a range endpoint.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the resource ('training status') and lists the included metrics: load, VO2 max, recovery, and training readiness. However, it does not explicitly differentiate from closely named siblings like get_training_readiness or get_training_load_trend, so the agent may struggle to pick this over similar tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided about when to use this tool versus the many training-related siblings. The description is purely declarative and gives no exclusions, alternatives, or context to help an agent choose correctly among get_training_readiness, get_training_load_trend, and similar tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_unit_systemA

Get user's preferred unit system from profile

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden of behavioral disclosure. 'Get ... from profile' clearly signals a read-only, non-mutating operation, but it does not provide additional behavioral context such as default behavior, error cases, or authentication requirements. For a zero-parameter getter, this is acceptable but not rich.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence with no filler, front-loading the action and the resource. Every word contributes meaning.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given that the tool has zero parameters, and an output schema is present, the description is largely complete for invoking the tool correctly. It could add a note about which sibling tools might overlap, but that is not essential for this simple lookup.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has no parameters, so there is nothing for the description to clarify beyond the empty input schema. The baseline score of 4 applies because parameter semantics are trivially complete.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly identifies the action (get), the resource (user's preferred unit system), and the source (profile). This distinguishes it from broader profile-related siblings like get_user_profile and get_userprofile_settings by focusing specifically on unit system.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies the tool should be used when the agent needs the user's preferred unit system, which is adequate for a simple getter. However, it does not explicitly state when not to use it or mention alternative tools that might also expose unit system data.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_user_profileB

Get user profile information

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden of behavioral disclosure. It only restates the 'get' action and gives no detail about response contents, permission requirements, or any other behavioral nuances, which is insufficient for meaningful transparency.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single efficient sentence with no filler or redundancy. It is appropriately brief for a parameterless tool, though it could have added a clarifying clause about scope without much cost.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is simple and has an output schema, which reduces the need to explain return values. However, the description leaves the exact meaning of 'user profile information' ambiguous, especially given closely named siblings, so it is only minimally complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters and the schema is empty, so there is nothing the description needs to add about parameter meaning. The baseline for no-parameter tools is 4.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Get user profile information' states a clear verb and resource, so an agent knows this is a read operation for a user profile. However, it does not differentiate from similar siblings like get_userprofile_settings or get_full_name, so it stops short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance about when to use this tool versus the many sibling tools that also fetch profile-like data. There is no mention of alternatives, exclusions, or the intended context for this specific read call.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_userprofile_settingsC

Get user profile settings

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The verb 'Get' implies a read-only operation, but with no annotations the description carries the full burden and does not explicitly disclose side effects, rate limits, or auth requirements.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single direct sentence, concise and front-loaded with the operation.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple zero-parameter tool with an output schema, the one-line description is mostly adequate, but it does not clarify what 'profile settings' includes or how it differs from the sibling get_user_profile.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There are no parameters, so schema coverage is trivially 100%; the baseline of 3 applies and no parameter documentation is needed.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Get') and resource ('user profile settings'), but 'settings' is somewhat vague and it is not explicitly differentiated from sibling get_user_profile or get_unit_system.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives like get_user_profile; there is no context, exclusions, or examples.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_user_summaryB

Get user summary data (compatible with garminconnect-ha)

Args: date: Date in YYYY-MM-DD format

ParametersJSON Schema
NameRequiredDescriptionDefault
dateYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.1/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, and the description does not disclose whether the operation is read-only, what side effects exist, or any error behavior. It only states the function name and parameter.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely concise, containing only the essential information without any redundant or verbose text. It is well-structured and easy to parse.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple getter, the description is adequate: it states the purpose and documents the parameter format. While it does not specify output details or the exact content of the summary, the presence of an output schema and the straightforward nature of the tool make this acceptable.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The date parameter is described with a specific format (YYYY-MM-DD), which adds useful meaning beyond the bare string type. However, it does not clarify what the date represents (e.g., activity date, log date) or whether a range is expected.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states a specific verb ('Get') and object ('user summary data'), and mentions compatibility with garminconnect-ha, which gives context. It is distinct from nearby getters like get_stats or get_daily_steps, but does not elaborate on what 'user summary' includes.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description offers no guidance on when to use this tool versus the many alternative getters. It lacks context on preferred scenarios, limitations, or how it differs from similar tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_vo2max_trendA

Get VO2 max trend over a date range.

Returns daily VO2 max estimates from Garmin's FirstBeat algorithm. Use this to track whether training is producing fitness gains over weeks or months. Flat or declining VO2 max over 4+ weeks suggests insufficient training stimulus or overreaching.

Note: VO2 max estimates are smoothed and update gradually — daily changes of <0.5 are within normal noise. Focus on the 4-6 week trend direction.

Garmin records a new VO2 max value only on days with a recompute (after an activity). Days in between carry the last known value forward and are marked with "carried_forward": true, matching the trend chart in Garmin Connect.

If historical values are unavailable, the current profile estimate is returned separately and is not represented as a historical trend point.

Recommended range: 4-12 weeks. Maximum: 90 days.

Args: start_date: Start date in YYYY-MM-DD format end_date: End date in YYYY-MM-DD format

ParametersJSON Schema
NameRequiredDescriptionDefault
end_dateYes
start_dateYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description fully carries the transparency burden. It discloses key behaviors: estimates are smoothed, daily changes <0.5 are noise, days without recompute carry forward the last value (with a 'carried_forward' flag), and if historical data is unavailable, the current profile estimate is returned separately. This is thorough and honest.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is moderately verbose but each sentence adds meaningful information (purpose, interpretation, behavior, format, range). There is a slight redundancy with the recommended range stated twice (4-12 weeks and 4-6 week trend), but it does not detract significantly from clarity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given that an output schema is indicated (though not shown), the description need not detail return fields. It sufficiently covers what the tool does, when to use it, and the important behavioral edge cases. It could mention error conditions or rate limits, but those are not central to invoking it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema provides no descriptions for start_date or end_date (0% coverage), but the description adds the YYYY-MM-DD format and recommends a 4–12 week range with a 90-day maximum. It does not explicitly define each parameter individually, but the format and range guidance substantially compensate for the missing schema descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the specific resource ('VO2 max trend') and the action ('Get') over a date range, which distinguishes it from the many other get_* trend tools. The first sentence alone provides a precise purpose.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It provides explicit guidance on when to use the tool (to track fitness gains over weeks) and how to interpret results (flat/declining over 4+ weeks suggests insufficient stimulus). It lacks explicit comparison to alternatives like get_training_load_trend or get_hrv_trend, but the focus on VO2 max is clear enough.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_weekly_intensity_minutesA

Get weekly intensity minutes data aggregates

Returns weekly intensity minutes (moderate and vigorous) for the specified number of weeks ending at end_date.

Args: end_date: End date in YYYY-MM-DD format weeks: Number of weeks to fetch (default 4, max 52)

ParametersJSON Schema
NameRequiredDescriptionDefault
weeksNo
end_dateYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the transparency burden. It does state that the tool returns aggregated data and includes moderate/vigorous intensity, but it does not disclose read-only status, edge cases, or behavior such as how end_date is handled at boundaries. It is adequate but not rich.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and front-loaded: it states the purpose first, then gives return details and parameter explanations. The Args section is directly useful and there is little wasted text.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple two-parameter read-only getter with an output schema present, the description covers the essential semantics: what is returned, the date anchor, and the valid weeks range. Nothing critical is missing for an agent to invoke this correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 0% description coverage, but the description compensates fully by specifying end_date format (YYYY-MM-DD) and weeks semantics including default (4) and max (52). This adds meaning well beyond the bare schema fields.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific resource ('weekly intensity minutes'), distinguishes moderate and vigorous data, and specifies the time window ('number of weeks ending at end_date'). This clearly differentiates it from the many sibling getter tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies use when weekly intensity minutes are needed, but it does not explicitly state when to choose this tool over alternative getters or mention any exclusions. Context is clear but no dedicated usage guidance or alternative routing is provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_weekly_load_progressionA

Get training minutes per ISO week and the week-over-week change.

Arithmetic only. Whether a given ramp rate is too fast is a coaching decision.

Args: weeks: Number of weeks to analyze (default 12, max 52)

ParametersJSON Schema
NameRequiredDescriptionDefault
weeksNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It usefully discloses the non-obvious caveat that output is 'arithmetic only' with no coaching interpretation, which is genuinely valuable behavioral context. It stops short of stating units, computation source, or rounding, so gaps remain despite a strong caveat.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the core purpose in the first sentence, then the scope caveat, then args. Appropriately sized with little waste. The trailing 'Args:' block is slightly format-heavy but earns its place by adding the max-52 bound.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return-format explanation is unnecessary. For a one-param, read-only tool the description covers purpose, scope caveat, and parameter bounds. It is essentially complete, with only the ambiguity against load/trend siblings left unaddressed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the single param's schema only carries a default. The description compensates by explaining 'weeks: Number of weeks to analyze' and adding the 'default 12, max 52' constraint, which is meaningfully beyond the schema. This is nearly complete param semantics for a one-param tool.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: 'Get training minutes per ISO week and the week-over-week change.' This is clear and concrete, telling the agent exactly what data comes back. However, it does not distinguish itself from nearby siblings like get_weekly_intensity_minutes or get_training_load_trend, which an agent would need to disambiguate.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The line 'Whether a given ramp rate is too fast is a coaching decision' implicitly scopes the tool as informational rather than interpretive, which is useful guidance. But it names no alternatives and gives no explicit when-to-use vs when-not conditions relative to the many load/trend siblings. Usage is only implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_weekly_stepsA

Get weekly step data aggregates

Returns weekly step totals for the specified number of weeks ending at end_date.

Args: end_date: End date in YYYY-MM-DD format weeks: Number of weeks to fetch (default 4, max 52)

ParametersJSON Schema
NameRequiredDescriptionDefault
weeksNo
end_dateYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden of behavioral transparency. It clearly indicates a read-only aggregation operation by saying it 'Returns weekly step totals'. It does not mention edge cases like week-boundary handling or invalid inputs, but these are not critical for a simple read query.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise and well-structured: a one-line summary, a clarifying sentence, and a short args list. No unnecessary words or redundant information are present.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description provides enough context for typical use, covering the query target and all parameters. It does not describe the response structure, but an output schema is indicated as present, and the tool is a simple aggregate query.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The description fully compensates for the missing schema descriptions by explaining end_date as a YYYY-MM-DD date and weeks as the number of weeks with default 4 and max 52. Both parameters are given meaningful, actionable semantics.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the specific verb 'Get' and the resource 'weekly step data aggregates', with a precise scope of weekly step totals for a specified number of weeks ending at end_date. This makes it readily distinguishable from daily step or raw step data siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description clearly implies use when weekly aggregate step totals are needed rather than daily step data, especially through the 'weekly step totals' and 'weeks' parameters. However, it does not explicitly contrast itself with get_daily_steps or get_steps_data, leaving the differentiation implicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_weekly_stressC

Get weekly stress data aggregates

Returns weekly stress values for the specified number of weeks ending at end_date.

Args: end_date: End date in YYYY-MM-DD format weeks: Number of weeks to fetch (default 4, max 52)

ParametersJSON Schema
NameRequiredDescriptionDefault
weeksNo
end_dateYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

C2.9/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

There are no annotations provided, so the description must carry the burden of disclosing behavioral traits. The description only states what the tool returns ('Returns weekly stress values') and does not explicitly mention whether it is read-only, has side effects, or requires specific permissions. Since it lacks any explicit behavioral disclosure, the transparency is minimal.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely concise, with a one-line summary followed by a clearly organized argument list. It avoids unnecessary detail and front-loads the core purpose. The structure is easy to parse and directly addresses the tool's functionality.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given that the tool has an output schema (as indicated by 'Has output schema: true'), the description does not need to explain the return structure. It provides the essential information: what the tool does, the parameters, and their meanings. It is complete for the typical use case, though it lacks examples or edge-case handling, which are not strictly necessary here.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The description enriches both parameters: 'end_date' is explained as 'End date in YYYY-MM-DD format' and 'weeks' as 'Number of weeks to fetch (default 4, max 52)'. This goes beyond the bare schema by providing format, defaults, and constraints, giving the agent meaningful semantic information. The coverage is complete for the two parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's function: 'Get weekly stress data aggregates' and elaborates that it 'Returns weekly stress values for the specified number of weeks ending at end_date.' This specifies the resource (stress data) and the scope (weekly aggregates), making the purpose obvious. It does not explicitly differentiate from sibling stress tools like get_stress_summary or get_stress_data, but the wording is specific enough to suggest a distinct aggregate view.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives. It does not mention any conditions, preferences, or contrasting scenarios with other stress-related tools. Without any usage direction, an agent cannot determine when this tool is the best choice.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_weigh_insA

Get weight measurements between specified dates

Args: start_date: Start date in YYYY-MM-DD format end_date: End date in YYYY-MM-DD format

ParametersJSON Schema
NameRequiredDescriptionDefault
end_dateYes
start_dateYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the burden of behavioral disclosure. The verb 'Get' implies a read-only operation and the date range is clearly scoped, but the description does not explicitly state that no data is modified, what granularity of data is returned, or whether both dates are inclusive.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The purpose is stated in a single front-loaded sentence, followed by a compact Args block with no filler. Every sentence earns its place, and the structure is easy to scan.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the presence of an output schema and only two simple required parameters, the description provides the essential calling contract. It could be more complete by clarifying date-boundary behavior or differentiating from get_daily_weigh_ins, but these are minor gaps for a simple retrieval tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, but the description compensates by explicitly defining both parameters and their exact YYYY-MM-DD format. It does not address inclusivity or ordering, but the parameter names and required flags already convey the basic contract.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Get'), a clear resource ('weight measurements'), and a date-range scope ('between specified dates'). It does not explicitly distinguish itself from sibling tools like get_daily_weigh_ins, but the core purpose is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no guidance on when to use this tool versus alternatives such as get_daily_weigh_ins, get_stats_and_body, or get_progress_summary_between_dates. There is no mention of exclusions, prerequisites, or preferred contexts.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_workout_by_idA

Get detailed information for a specific workout

Returns workout details including segments and step structure.

Accepts either:

  • Numeric workout ID (from get_workouts, get_scheduled_workouts, or training-plan families that expose workout_id)

  • Workout UUID (from adaptive Garmin Coach/training-plan workouts)

Rest-day UUIDs can resolve to a minimal record without a workout name or segments.

Args: workout_id: Workout ID (numeric) or UUID (for training plan workouts)

ParametersJSON Schema
NameRequiredDescriptionDefault
workout_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description discloses an important behavioral nuance: 'Rest-day UUIDs can resolve to a minimal record without a workout name or segments.' This goes beyond a simple 'get' and informs the caller of an edge case. Since annotations are absent, this added detail improves transparency, though it doesn't cover error responses or other potential side effects.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise and well-structured: a one-sentence purpose, a clear breakdown of the two ID types, and a brief note on rest-day behavior. No redundant or extraneous information is present, and the key details are front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the presence of an output schema and the tool's relative simplicity, the description is complete. It explains what data is returned (segments, step structure), the accepted parameter formats, and the rest-day exception. The output schema handles detailed return types, so no further elaboration is needed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema only defines workout_id as anyOf integer/string with no description, so the description carries the full burden. It adds significant meaning by explaining that numeric IDs come from specific sources (e.g., get_workouts, scheduled workouts) and UUIDs from adaptive coach/training-plan workouts, and it clarifies the special behavior for rest-day UUIDs. This fully compensates for the lack of schema descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: 'Get detailed information for a specific workout' and specifies it returns segments and step structure. It distinguishes from sibling tools by focusing on retrieval by ID rather than listing or scheduling workouts. The mention of both numeric IDs and UUIDs further clarifies its scope.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides useful context on when to use this tool by explaining the origin of accepted IDs: numeric IDs from get_workouts, get_scheduled_workouts, or training-plan families, and UUIDs from adaptive Garmin Coach/training-plan workouts. While it doesn't explicitly state 'use this when you have a workout ID' or contrast with download_workout, the guidance is clear enough for most cases.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_workout_complianceA

List scheduled workouts alongside activities recorded the same day.

Returns the pairing, not a compliance score. Matching is on date only — Garmin's schedule query returns no sport to match against — so whether a given activity completes a given workout is left to the caller.

Args: start_date: Start date in YYYY-MM-DD format end_date: End date in YYYY-MM-DD format

ParametersJSON Schema
NameRequiredDescriptionDefault
end_dateYes
start_dateYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and delivers: it discloses that no score is computed, that matching is on date only, and the upstream reason (Garmin's schedule query returns no sport). It omits permissions, pagination, and ordering behavior, keeping it short of a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core behavior and the important caveat before the args list. The Args section partially duplicates the schema, but it adds format detail, so little is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return structure need not be re-explained, and the description covers the key interpretive caveat an agent needs to use the result correctly. Given no annotations, it is nearly complete, missing only operational details like limits.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, and it documents both parameters with the required YYYY-MM-DD format. This adds real meaning the bare schema lacks, though it gives no range or boundary semantics.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('List scheduled workouts alongside activities recorded the same day') and immediately corrects the misleading name by clarifying it returns a pairing, not a compliance score. An agent can distinguish this from siblings like get_scheduled_workouts or get_activities_by_date without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explains the crucial usage constraint (matching is date-only and the compliance determination is left to the caller), which implies when the tool is appropriate. However, it never names an alternative sibling or states explicit when-not-to-use conditions, so routing guidance is inferred rather than stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_workoutsA

Get all workouts with curated summary list

Returns a count and list of workout summaries with essential metadata only. For detailed workout information including segments, use get_workout_by_id.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Description indicates a read-only retrieval operation returning a summary list; no side effects are mentioned. While not explicitly stating it is non-destructive, the behavior is clear from the 'get' action and return description.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise, with two short sentences that convey purpose, return type, and the alternative for detailed data.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Description gives enough context for a summary list tool and points to the detailed variant. It does not enumerate the exact summary fields, but 'essential metadata only' is sufficient given the existence of a detailed counterpart.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, so there is no parameter semantics to document. The description is complete in this regard.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description clearly states the tool gets all workouts and returns a curated summary list with count and essential metadata. It distinguishes itself from get_workout_by_id, which provides detailed workout information.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly directs users to get_workout_by_id when detailed workout information including segments is needed, making the appropriate use case clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_zone_distributionB

Get heart rate zone distribution across sports for a date range.

Calculates the percentage of time spent in each HR zone (Z1-Z5) for running, cycling, and swimming activities.

Args: start_date: Start date in YYYY-MM-DD format end_date: End date in YYYY-MM-DD format

ParametersJSON Schema
NameRequiredDescriptionDefault
end_dateYes
start_dateYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It usefully discloses the computed output (percentage of time in Z1-Z5 across three sports), but says nothing about permissions, rate limits, or how missing data is handled. The 'get' verb reasonably implies a read operation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose is front-loaded in the first sentence, with the calculation detail and parameter formats following. The Args block is slightly mechanical but earns its place by supplying the missing date format.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values needn't be explained. For a simple two-parameter read-only aggregate, the description covers purpose, scope, computation, and parameter formats adequately; only usage routing to alternatives is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It does document both required parameters and, importantly, adds the YYYY-MM-DD format the schema omits. Minor gap: no mention of range limits or timezone handling.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (get) and resource (heart rate zone distribution) with clear scope: across running/cycling/swimming for a date range. This implicitly distinguishes it from per-activity siblings like get_activity_hr_in_timezones, but never names them explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description says what the tool returns but gives no when-to-use guidance, prerequisites, or named alternatives (e.g., vs get_activity_hr_in_timezones for a single activity). Usage must be inferred from scope alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

log_custom_foodA

Log a food item to a meal on a date

Adds a food entry to the nutrition log. The meal is determined automatically by matching meal_time against each meal's startTime/endTime window; falls back to SNACKS if no window matches.

Food sources:

  • "GARMIN" (default): user's custom food library. Use get_custom_foods to find food_id and serving_id.

  • "FATSECRET": branded/catalog food from FatSecret. Use search_foods to find food_id and serving_id. Pass the source value from the search_foods result (e.g. "FATSECRET").

Garmin custom food IDs are 32-char hex UUIDs; FatSecret IDs are numeric strings (e.g. "4132350"). Passing the wrong source for a given food_id returns a 400 from Garmin.

Args: meal_date: Date in YYYY-MM-DD format meal_time: Time in HH:MM:SS format (e.g. "12:30:00", account timezone) food_id: Food ID from get_custom_foods (GARMIN) or search_foods (FATSECRET) serving_id: Serving ID from get_custom_foods or search_foods serving_qty: Number of servings (default 1) source: Food namespace — "GARMIN" (default) or "FATSECRET"

ParametersJSON Schema
NameRequiredDescriptionDefault
sourceNoGARMIN
food_idYes
meal_dateYes
meal_timeYes
serving_idYes
serving_qtyNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does well: it explains automatic meal-window matching, the SNACKS fallback, ID format differences, and the 400 error on source mismatch. It omits duplicate/idempotency behavior and whether an existing entry is overwritten, which matters for a write tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Well structured with source and lookup cues front-loaded, but there is mild redundancy between the title line and 'Adds a food entry to the nutrition log', and the Args block duplicates parameter info. Still, every section serves a purpose and it is easy to scan.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained, and the description covers meal resolution, source namespaces, ID formats, and a failure mode. The notable gap is how this tool relates to log_food and upsert_and_log, which an agent needs to choose correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, and it does: it documents all six parameters with concrete formats (YYYY-MM-DD, HH:MM:SS with account timezone), default values, and enum semantics for source. It even gives example ID shapes for both namespaces, adding meaning well beyond the bare schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Log a food item to a meal on a date') and elaborates what it does. However, it never distinguishes itself from its close siblings log_food and upsert_and_log, which appear to do overlapping things, so an agent cannot tell them apart from the description alone.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear routing guidance: use get_custom_foods for GARMIN IDs and search_foods for FATSECRET IDs, and warns that a wrong source yields a 400. It lacks any when/when-not guidance relative to log_food or upsert_and_log, so the alternative-selection story is incomplete.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

log_foodA

Quick-add a food entry with macro values to the nutrition log

Logs food directly by name and macros without requiring a food ID. Uses Garmin's Quick Add feature. The meal is determined automatically by matching meal_time against each meal's startTime/endTime window; falls back to SNACKS if no window matches.

Args: meal_date: Date in YYYY-MM-DD format name: Display name for the food entry calories: Calories (kcal) carbs: Carbohydrates in grams protein: Protein in grams fat: Fat in grams meal_time: Time in HH:MM:SS format (account timezone)

ParametersJSON Schema
NameRequiredDescriptionDefault
fatYes
nameYes
carbsYes
proteinYes
caloriesYes
meal_dateYes
meal_timeYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full disclosure burden, and it delivers meaningful context: it reveals the non-obvious meal-determination logic (matching meal_time against meal startTime/endTime windows, falling back to SNACKS) and the timezone caveat for meal_time. For an unannotated mutation tool it could still disclose more (e.g., whether repeated calls duplicate entries, or reversibility), but the critical hidden behavior is surfaced.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with the purpose, followed by one mechanism sentence, then a compact args list. Each line in the Args block adds format or unit information the schema lacks, so it earns its place. Slightly longer than strictly necessary, but nothing is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 7-required-parameter tool with 0% schema coverage and zero annotations, the description is notably complete: it covers purpose, mechanism, all parameter formats/units, and the non-obvious meal-window fallback. The output schema covers return-value expectations, so that gap is not the description's fault. The only real absence is explicit guidance on when to prefer sibling tools, which is also reflected in the usage_guidelines score.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description fully compensates by documenting all 7 parameters with useful semantics: exact formats ('YYYY-MM-DD', 'HH:MM:SS'), units (kcal, grams), and the timezone qualifier on meal_time. It also explains how meal_time behaviorally maps to meal selection, which the bare schema cannot convey. No parameter is left underspecified.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Quick-add a food entry with macro values to the nutrition log.' It further differentiates from siblings by stating it logs 'directly by name and macros without requiring a food ID,' which clearly separates it from log_custom_food and search_foods. The purpose is unmistakable.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies when to use this tool ('without requiring a food ID' and 'Uses Garmin's Quick Add feature') and explains the automatic meal assignment behavior. However, it never explicitly names alternatives (e.g., log_custom_food for ID-based logging) or states when-not-to-use conditions. The routing guidance is implicit rather than explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

remove_gear_from_activityA

Remove gear association from an activity

Unlinks a specific piece of gear from an activity.

Args: activity_id: ID of the activity gear_uuid: UUID of the gear to remove

ParametersJSON Schema
NameRequiredDescriptionDefault
gear_uuidYes
activity_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It clarifies that the operation unlinks rather than deletes gear, which is useful, but it does not disclose side effects, idempotency, error behavior, or whether the gear must already be associated with the activity.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is short and front-loaded, but the first two sentences are largely redundant ('Remove gear association' vs 'Unlinks a specific piece of gear'). The Args section is useful and compact.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple two-parameter tool with an output schema, the description covers the core action and parameters. However, it lacks usage guidance and behavioral details such as prerequisites or failure modes, leaving some gaps for an agent to infer.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate. The Args section explains activity_id as the activity's ID and gear_uuid as the UUID of the gear to remove, adding meaning beyond the bare schema titles. It could further specify where to obtain these values, but it is adequate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Remove'/'Unlinks') and a clear resource ('gear association from an activity'), making the operation unambiguous. It is naturally distinguished from sibling tools like add_gear_to_activity and get_activity_gear.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The intended use is implied by the action ('unlink gear from an activity'), but there is no explicit guidance about when to choose this over alternatives such as add_gear_to_activity or get_activity_gear. No exclusions or prerequisites are stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

request_reloadA

Ask Garmin to re-process the epoch data for a date.

Use when a day's metrics look incomplete after a sync. The response says whether the reload was accepted, not whether new data appeared — re-read the day afterwards to see that.

Args: date: Date in YYYY-MM-DD format

ParametersJSON Schema
NameRequiredDescriptionDefault
dateYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does well: it discloses that the response only reports acceptance, not whether new data appeared, and that the caller must re-read the day to confirm. It omits auth requirements, rate limits, and whether reloads are idempotent, so it stops short of a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences, front-loaded with the action, then the trigger, then the return caveat, followed by a one-line arg format note. Every sentence earns its place with no padding.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values needn't be enumerated, yet the description still clarifies the meaning of the response. Combined with the when-to-use and post-call guidance, an agent has everything needed to invoke and interpret it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the schema only titles the field 'Date', so the description's 'Date in YYYY-MM-DD format' adds real value. One parameter is well covered; the low baseline expectation for a single-param tool is met.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource: it asks Garmin to 're-process the epoch data for a date'. No sibling tool performs a reload/reprocess action, so the agent can immediately distinguish it from the many get_*/set_* siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives a clear triggering condition ('when a day's metrics look incomplete after a sync') and a follow-up action ('re-read the day afterwards'). It doesn't name explicit alternatives, but no sibling provides equivalent reprocessing, so there is little to disambiguate.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

schedule_weekA

Schedule a list of workouts for the week in a single call.

Idempotent: if a workout is already scheduled for that date, it is reported as already scheduled and the POST is skipped (avoids duplicating calendar entries).

A bad entry is reported and the rest of the batch continues; nothing aborts the whole week.

Args: week: List of dicts with keys: calendar_date (YYYY-MM-DD) and workout_id (int). The key is calendar_date, matching schedule_workout and schedule_workouts.

ParametersJSON Schema
NameRequiredDescriptionDefault
weekYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does well: it discloses idempotency (existing dates reported and POST skipped to avoid duplicates) and partial-failure semantics (a bad entry is reported while the batch continues). It omits auth/permission requirements and any rate-limit context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core action, then error/idempotency behavior, then argument detail. Each sentence earns its place, with only mild redundancy in the Args restatement.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter batch mutation with an output schema, the description covers what an agent needs: action, idempotent behavior, and parameter shape. No annotations exist, so it could add permission/response framing, but the output schema absorbs the return-value burden.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, and it does: it documents the week array's inner dict keys (calendar_date as YYYY-MM-DD, workout_id as int) and even anchors the key convention to sibling tools. It doesn't clarify whether additional keys are accepted, though the schema sets additionalProperties: true.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

It states a specific verb and resource ('Schedule a list of workouts for the week in a single call'), clearly distinguishing batch-of-one-week intent from the singular schedule_workout. It stops short of differentiating from the sibling schedule_workouts, which is also a batch scheduling tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The 'in a single call' phrasing plus the idempotency note implies when this batch tool is preferred, and the partial-failure behavior is stated. However, it never names schedule_workouts or explains which of the two batch tools to choose, so routing between siblings is left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

schedule_workoutA

Schedule a workout to a specific calendar date

This adds an existing workout from your Garmin workout library to your Garmin Connect calendar on the specified date.

Idempotent: if the workout is already scheduled for that date, this is a no-op that reports success without creating a duplicate entry.

Args: workout_id: ID of the workout to schedule (get IDs from get_workouts) calendar_date: Date to schedule the workout in YYYY-MM-DD format

ParametersJSON Schema
NameRequiredDescriptionDefault
workout_idYes
calendar_dateYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It discloses that the operation is idempotent (no-op on duplicate) and that it 'reports success'. It also implies the action modifies the calendar. However, it does not mention potential failure modes (e.g., invalid workout ID or date format issues) or side effects beyond scheduling, so some behavioral traits remain undisclosed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise and well-structured: a one-sentence summary, a brief explanatory line, a note about idempotency, and a clear list of arguments. Every sentence adds value without redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the straightforward nature of the tool, the description is complete: it states the purpose, required parameters (with formats), and idempotency. The output schema exists, so return details are not required. No critical context (e.g., prerequisites beyond having a valid workout ID) is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has no per-parameter descriptions (coverage 0%), but the tool description compensates with an explicit 'Args' section that explains 'workout_id: ID of the workout to schedule (get IDs from get_workouts)' and 'calendar_date: Date to schedule the workout in YYYY-MM-DD format'. This fully covers the meaning and format of both parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action: 'Schedule a workout to a specific calendar date' and explains that it 'adds an existing workout from your Garmin workout library to your Garmin Connect calendar on the specified date.' This is a specific verb and resource, and it distinguishes itself from sibling tools like schedule_workouts (plural) by focusing on a single workout scheduling.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides context for use: it tells the agent that it takes an existing workout ID and a date, and it even directs to get IDs from get_workouts. It also mentions idempotency, which is helpful for deciding when to call. However, it does not explicitly state when to prefer this over schedule_workouts or other scheduling tools, though the singular vs plural distinction is implicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

schedule_workoutsA

Schedule multiple workouts to specific calendar dates

This adds workouts to your Garmin Connect calendar in a single call. Each item can either reference an existing workout by ID, or provide inline workout_data to upload-and-schedule in one step.

Args: schedules: List of workout schedules, each with: - calendar_date (str): Date to schedule the workout in YYYY-MM-DD format (required) - workout_id (int): ID of an existing workout to schedule (required unless workout_data is provided) - workout_data (dict): Inline workout JSON to upload first, then schedule (optional). When provided, workout_id is not required. Uses the same structure and target-value rules as upload_workout.

Examples: Schedule existing workouts by ID: [{"workout_id": 123456, "calendar_date": "2024-01-15"}, {"workout_id": 789012, "calendar_date": "2024-01-17"}]

Upload and schedule inline:
[{"calendar_date": "2024-01-15", "workout_data": {"workoutName": "Easy Run", ...}},
 {"workout_id": 789012, "calendar_date": "2024-01-17"}]
ParametersJSON Schema
NameRequiredDescriptionDefault
schedulesYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It explains the high-level behavior (adds to calendar, supports inline upload), but it does not mention potential side effects like partial failures, duplicate handling, whether inline uploads create permanent workouts, or any permissions. This is decent but lacks depth.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-structured with a summary, a compact Args section using bullet points, and clear examples. It is succinct without unnecessary jargon, and the key information is front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers the core use cases, parameter constraints, and provides examples. However, it leaves some edge cases ambiguous, such as what happens if both workout_id and workout_data are provided in the same object, and it does not address error handling or return values (though output schema exists, which mitigates the latter). Overall it is fairly complete for basic usage.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema is minimal (just an array of objects with additionalProperties), but the description thoroughly explains each parameter: calendar_date required, workout_id required unless workout_data is provided, workout_data optional with structure and rules referencing upload_workout. The mutual exclusivity is clear and examples support the semantics.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool 'Schedule multiple workouts to specific calendar dates' and explains it adds workouts to the Garmin Connect calendar in a single call. It distinguishes itself from sibling tools like schedule_workout by explicitly being plural and batch-oriented, and it also clarifies two modes (reference existing or inline upload).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies when to use this tool (for multiple workouts, single call) and provides examples, but it does not explicitly contrast with schedule_workout or state when not to use it. The alternative is only implied by the plural name and 'single call' wording, not directly named with a condition.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_foodsA

Search Garmin's general food catalog (FatSecret + Garmin custom foods)

Searches across the entire food catalog including FatSecret-sourced branded and generic foods, not just the user's Garmin custom foods. Use this to find branded packaged foods by name before logging them.

Returns food_id, source, name, brand, and all available servings with macros. The source field ("FATSECRET" or "GARMIN") and food_id together identify the right routing for log_custom_food — pass both to log_custom_food's food_id and source parameters respectively.

For the user's own custom foods only, use get_custom_foods instead.

Args: query: Food name or brand to search for (e.g. "Cheerios", "Greek yogurt") start: Starting index for pagination (default 0) limit: Maximum number of results per page (default 20)

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
queryYes
startNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It discloses the search scope, the source field values ('FATSECRET' or 'GARMIN'), and the returned fields. It doesn't mention rate limits or error behavior, but for a search operation the core behavioral traits are well covered.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with a one-line summary, then expands into scope, return semantics, routing, and parameter details. Every sentence adds value, and the Args section is clearly structured. No redundant or filler content.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers selection (when to use vs get_custom_foods), invocation (all parameters), output semantics (returned fields and source values), and downstream usage (passing food_id and source to log_custom_food). It is complete for an agent to select and call the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must fully compensate. It does: query is explained as a food name or brand with examples, start is described as the pagination starting index, and limit as the maximum results per page. Defaults are also stated.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Search Garmin's general food catalog (FatSecret + Garmin custom foods)'. It clearly distinguishes itself from get_custom_foods by stating it searches the entire catalog, not just the user's custom foods.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly says to use this tool to find branded packaged foods by name before logging them, and explicitly names get_custom_foods as the alternative for the user's own custom foods. It also explains how results route into log_custom_food, giving clear when-to-use guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

set_activity_descriptionA

Set or update the free-text description (notes) of an activity.

This is the notes field shown on the activity page — useful for recording how a session felt, kit used, conditions, niggles, etc. Pass an empty string to clear an existing description.

Args: activity_id: ID of the activity to update description: New description text (empty string clears it)

ParametersJSON Schema
NameRequiredDescriptionDefault
activity_idYes
descriptionYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the behavioral disclosure burden. It clearly states this is a mutation ('Set or update'), identifies the target field, and documents the important empty-string-clearing behavior. It does not mention auth or validation, but for a simple field setter this is adequate.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The first sentence front-loads the core purpose, followed by useful usage context and the critical clearing behavior. The Args section is necessary given the schema's lack of descriptions. It is slightly redundant in places but remains well-organized and appropriately sized.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple two-parameter setter, the description covers purpose, usage context, parameter semantics, and clearing behavior. An output schema exists, so return values are covered elsewhere. Missing auth/error details are minor for this tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. The Args section explains both parameters: activity_id is the ID of the activity to update, and description is the new text with empty-string clearing semantics. This adds real meaning beyond the bare schema, though it omits constraints like max length.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb and resource: 'Set or update the free-text description (notes) of an activity.' It clearly identifies this as the notes field, distinguishing it from sibling tools like set_activity_name and set_activity_type without ambiguity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear context for when to use the tool, such as recording how a session felt, kit used, conditions, and niggles. It does not explicitly name alternatives or exclusions, but the purpose is distinct enough that an agent can select it correctly.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

set_activity_event_typeA

Set the event type of an activity.

Event type categorises the activity's purpose. Valid keys: race, recreation, specialEvent, training, transportation, touring, geocaching, fitness, uncategorized.

Args: activity_id: ID of the activity to update event_type: Target event type key (e.g. 'race', 'training')

ParametersJSON Schema
NameRequiredDescriptionDefault
event_typeYes
activity_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.5/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations present, the description carries the full burden of explaining behavior. It only says 'Set', implying a mutation, but does not mention side effects, return values, or error handling. The valid values list is helpful but insufficient to fully understand the effect.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise and well-structured, with a clear opening line, a brief explanation of valid keys, and a straightforward argument list. There is no unnecessary fluff or repetition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is a simple setter with two parameters, and the description covers the essentials. However, it does not mention the output or success/failure behavior, which would be valuable given the tool's mutation nature. The presence of an output schema helps, but the description leaves some context to be inferred.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The description adds meaning beyond the schema by explaining activity_id as 'ID of the activity to update' and event_type as 'Target event type key' with a list of valid values. This compensates for the schema's lack of parameter descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's function with the verb 'Set' and the resource 'event type of an activity.' It also lists all valid values for the event_type parameter, leaving no ambiguity about its purpose.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no explicit guidance on when to use this tool versus alternatives like set_activity_type or set_activity_name. It does not mention conditions, prerequisites, or scenarios where this setter is preferred.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

set_activity_feelA

Set how an activity felt ('How did you feel?').

Mirrors Garmin Connect's 5-point feel rating, stored as one of: 0 = very tired / poor 25 = tired 50 = normal 75 = good 100 = strong Higher is better.

Args: activity_id: ID of the activity to update feel: One of 0, 25, 50, 75, 100

ParametersJSON Schema
NameRequiredDescriptionDefault
feelYes
activity_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description must carry the behavioral burden. It discloses that this is a mutation ('set'), and it explains the feel value semantics with labels and 'higher is better.' However, it does not mention whether existing feel values are overwritten, whether the activity must exist first, or any side effects. The presence of an output schema helps cover the return-value aspect.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is tight and well-structured: a one-line primary action, a compact scale table, and a short Args section. Every sentence earns its place, and the most important information is front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple two-parameter setter, the description is essentially complete: it identifies the resource, the accepted values, and the parameter meanings. The output schema covers return details. The only notable omission is guidance on how this tool relates to perceived effort and other activity-editing tools, but that is not required for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and feel has no enum, so the description fully compensates. It defines activity_id as 'ID of the activity to update' and specifies exactly the valid feel values (0, 25, 50, 75, 100) with human-readable meanings. This gives an agent everything needed to construct valid arguments.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states a specific verb and resource: 'Set how an activity felt' with an explicit 5-point scale. The Garmin Connect feel-rating mapping distinguishes it from other activity setters, though it does not explicitly contrast it with set_perceived_effort or other sibling tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied: use this tool to record how an activity felt via the Garmin feel scale. There is no explicit when-to-use or when-not-to-use guidance, and no alternatives are named, but the purpose is clear enough for an agent to infer the intended context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

set_activity_nameA

Set or update the name of an activity.

Args: activity_id: ID of the activity to update activity_name: New activity name

ParametersJSON Schema
NameRequiredDescriptionDefault
activity_idYes
activity_nameYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are present, so the description carries the full burden of behavioral disclosure. It only states that the name is set/updated, without mentioning side effects, whether the activity must already exist, or what happens on failure. This is minimal transparency for a mutating operation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and well-structured: a one-line purpose statement followed by an Args block. There is no filler, and the core action is front-loaded. Every sentence serves a clear function.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple two-parameter setter, the description covers the operation and both arguments sufficiently. An output schema exists, so return values need not be described. The main gap is the lack of usage context relative to sibling tools, but basic invocation is fully supported.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, but the description's Args section provides meaning for both parameters: activity_id is the activity to update and activity_name is the new name. This compensates for the bare schema titles, though it does not elaborate on validation rules or accepted formats.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Set or update') and a clear resource ('the name of an activity'), which is distinct from sibling tools like set_activity_type, set_activity_description, or set_activity_event_type. An agent can immediately identify what attribute this tool modifies without ambiguity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided about when to use this tool versus the many sibling set_* tools. The description does not mention selection criteria, prerequisites, or exclusions. Usage must be inferred solely from the tool name and parameter names.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

set_activity_typeA

Change the activity type (sport) of an activity.

Useful for reclassifying a mislabelled activity, e.g. flipping a run logged as 'trail_running' to 'running', or a 'treadmill_running' walk to 'treadmill_walking'. Call get_activity_types to see all valid type keys.

Args: activity_id: ID of the activity to update type_key: Target activity type key (e.g. 'running', 'trail_running', 'treadmill_running', 'cycling', 'lap_swimming')

ParametersJSON Schema
NameRequiredDescriptionDefault
type_keyYes
activity_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

There are no annotations, so the description carries the burden of disclosing behavior. It clearly identifies the operation as a mutation and adds a useful constraint that type_key must come from get_activity_types. However, it does not address permissions, reversibility, failure behavior, or side effects on related activity data.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-structured: a short action statement, a purposeful rationale with examples, a cross-tool reference, and a compact Args block. Every sentence adds value and the explanation is front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple two-parameter setter with an output schema present, the description covers the core requirements: what the tool does, when to use it, and how to find valid values. It leaves some edge-case behavior implicit, but an agent has enough information to call the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It does: activity_id is explained as the ID of the activity to update, and type_key is explained with concrete examples. It could be stronger by noting how to obtain activity_id, but the provided semantics are sufficient.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a clear verb and resource: 'Change the activity type (sport) of an activity.' It also distinguishes itself from sibling setters by focusing specifically on the sport/type field, with examples that make the operation unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear context: it is useful for reclassifying mislabelled activities, and it directs the agent to call get_activity_types for valid keys. It does not explicitly name alternative tools or state when not to use it, but the usage context is clear enough.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

set_blood_pressureC

Set blood pressure values

Args: systolic: Systolic pressure (top number) diastolic: Diastolic pressure (bottom number) pulse: Pulse rate notes: Optional notes

ParametersJSON Schema
NameRequiredDescriptionDefault
notesNo
pulseYes
systolicYes
diastolicYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It identifies the operation as 'Set' but does not clarify whether this appends a new reading or overwrites an existing one, what units are expected, or what happens on invalid input. These are material unknowns for a health-data write operation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact with a clear front-loaded purpose line followed by a structured args list. Every sentence earns its place, and the arg documentation is not redundant given the schema's 0% description coverage. Minor room for improvement: the arg list could be folded into prose or trimmed.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is simple (4 scalar params, no nesting) and an output schema exists, so return-value documentation is not needed. The main gaps are units, the append-vs-replace write semantics, and expected value ranges—each of which an agent would want to know before recording health data. Adequate but not complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It adds useful disambiguation ('top number' vs 'bottom number') and marks notes as optional, which the bare schema titles do not convey. However, it omits units (mmHg, bpm), valid ranges, and any context on how values are interpreted, so the compensation is only partial.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Set') and resource ('blood pressure values'), clearly identifying the operation. It is not a tautology and the set/get contrast with the sibling get_blood_pressure makes the intent unmistakable, though the description does not explicitly name that sibling.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is given on when to use this tool versus alternatives. There is no mention that this tool records a new blood pressure reading as opposed to fetching existing data via get_blood_pressure, nor any note about prerequisites (e.g., user authorization). The intended use is implied only by the tool's name, not stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

set_fit_download_dirA

Set and persist the default directory for downloaded activity files.

Stores the absolute path in a small JSON config file (~/.garminconnect_fit_config.json, overridable via GARMIN_FIT_CONFIG) so download_activity_file can save files without asking again.

Args: path: Directory where activity files (.fit/.gpx/.tcx/.csv) are saved. Pass the current working directory to keep files where the server runs.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden of explaining side effects. It clearly states that the tool stores the path in a JSON config file and affects future downloads, which is transparent about persistence behavior. It does not mention potential overwrite or error behavior, but those are not critical for this simple setter.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise and well-structured, with a clear purpose statement followed by implementation detail and parameter guidance. Every sentence adds useful information without redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple state-setting tool, the description provides all necessary context: what it does, where it persists, and how it affects another tool. No return schema details are needed for this action, and the description is sufficient for an agent to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Although the schema provides no field description, the tool description fully covers the only parameter ('path') by explaining that it is a directory and offering recommended usage. This gives the agent complete semantic understanding of the argument.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('set and persist') and clearly identifies the resource ('default directory for downloaded activity files'). It is distinct from sibling tools, which mostly retrieve data, and it explicitly names the dependent tool download_activity_file.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explains that this tool persists the directory so download_activity_file can save files without asking, giving clear context for when to use it. It also provides practical guidance to pass the current working directory, though it could be more explicit about sequencing relative to download_activity_file.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

set_nutrition_daily_settingsA

Update daily nutrition goals (calorie target and macronutrient targets).

Reads the current settings for the date, applies the supplied overrides, and writes the merged result back. Only the fields you provide are changed; omitted fields keep their existing values.

Garmin stores macros as grams. The calorie goal should match 4carbs + 4protein + 9*fat to within a small rounding margin — Garmin accepts minor mismatches but will silently correct large discrepancies.

Args: date: Date in YYYY-MM-DD format (settings are typically set once and inherited across days, but Garmin accepts per-day overrides) calorie_goal: Daily calorie target in kcal carbs_grams: Daily carbohydrate target in grams fat_grams: Daily fat target in grams protein_grams: Daily protein target in grams

ParametersJSON Schema
NameRequiredDescriptionDefault
dateYes
fat_gramsNo
carbs_gramsNo
calorie_goalNo
protein_gramsNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description fully discloses side effects: it reads current settings, writes merged results, and notes that Garmin may silently correct large calorie-macro discrepancies. This gives the agent a realistic expectation of behavior beyond a simple set.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-structured with a clear intro, behavioral notes, and an Args list. It is slightly verbose due to the calorie-macro consistency explanation, but that information is essential for correct usage, so the length is justified.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Provides enough context to call the tool correctly: explains merge behavior, parameter semantics, and a critical consistency rule. Does not mention return value or error conditions, but an output schema exists (not shown) and the operation is straightforward; still, a brief note on expected response would improve completeness.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, but the description's Args section adds meaningful semantics for all 5 parameters: explains date format and inheritance behavior, units for calorie_goal (kcal), and units for macros (grams). This fully compensates for the schema's lack of descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb 'update' and resource 'daily nutrition goals' (calorie and macronutrient targets). Distinguished from sibling getter tools like get_nutrition_daily_settings by its setter nature and explicit description of the update behavior.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides clear usage instructions: mentions the merge behavior (reads current settings, applies overrides, writes merged result) and explains that only provided fields are changed, omitted fields keep existing values. Also includes important consistency note about calorie and macro relationships, guiding correct usage.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

set_perceived_effortA

Set the perceived effort (RPE) for an activity.

Mirrors Garmin Connect's 'Perceived Effort' rating on a 0-10 scale, where 0 clears the rating. Internally Garmin stores this multiplied by 10 (so RPE 7 is stored as 70); this tool handles the conversion.

Args: activity_id: ID of the activity to update rpe: Perceived effort from 0 to 10 (0 clears the rating)

ParametersJSON Schema
NameRequiredDescriptionDefault
rpeYes
activity_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden of behavioral disclosure. It goes beyond a bare 'sets RPE' by explaining the 0-10 scale, that 0 clears the rating, and that Garmin stores the value multiplied by 10 while the tool handles conversion. It does not mention mutation side effects or permissions, but the setter behavior and clearing semantics are well covered.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and front-loaded with the core action. Every sentence earns its place: the scale, the clear-rating behavior, the internal conversion detail, and the parameter list are all relevant and free of filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is simple and the two required parameters are fully described. An output schema exists, so return-value details are not required. Minor gaps such as explicit alternative-routing guidance and potential validation behavior prevent a perfect score, but the description is otherwise complete for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, and it does: activity_id is identified as the target activity, and rpe is explained with its 0-10 valid range and the special meaning of 0. The schema provides only type information, making the description's parameter explanations essential and largely sufficient.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly names the verb ('Set') and the resource ('perceived effort (RPE) for an activity'), which is unambiguous among the many set_activity_* sibling tools. It does not explicitly differentiate from siblings like set_activity_feel, but the RPE target and 0-10 scale make the purpose clear.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no explicit guidance on when to use this tool versus alternative setters such as set_activity_feel or set_activity_name. It implies usage by naming what it does, but does not state conditions, exclusions, or when another sibling tool would be preferred.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

unschedule_workoutA

Remove a scheduled workout from the Garmin Connect calendar

Deletes a calendar entry without deleting the underlying workout template — the workout stays in your library and can be re-scheduled.

IMPORTANT: scheduled_workout_id is the calendar-entry id, which is different from the workout's id. Get it from get_scheduled_workouts (the "scheduled_workout_id" field), not from get_workouts.

Note: the scheduled-workouts listing is an eventually-consistent index. If you just scheduled this workout, allow a moment before unscheduling so the id is available.

Args: scheduled_workout_id: Calendar-entry id from get_scheduled_workouts

ParametersJSON Schema
NameRequiredDescriptionDefault
scheduled_workout_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It discloses the side effect of not deleting the template and warns about the eventual-consistency delay. It does not mention error behavior, idempotency, or permissions, which would be useful but are not essential for basic operation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is somewhat lengthy due to the important notes, but each sentence adds value. The use of 'IMPORTANT' and 'Note' labels makes it well-structured and easy to parse.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given that an output schema exists, the description need not explain return values. The tool is simple with one parameter, and the description covers purpose, parameter semantics, side effects, and timing, providing all necessary context for an agent to decide when and how to invoke it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema only provides type integer with zero description coverage. The description fully compensates by explaining that the parameter is the calendar-entry ID, distinct from the workout ID, and precisely specifies where to obtain it (from get_scheduled_workouts' scheduled_workout_id field).

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('Remove a scheduled workout from the Garmin Connect calendar') and distinguishes it from deleting the underlying template, making the tool's purpose unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly describes the use case and clarifies that this operation preserves the workout template, implying when to use this over a full deletion. The note about eventual consistency provides timing guidance. However, it does not name alternative sibling tools explicitly, but the context is clear enough.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

unschedule_workoutsA

Remove multiple scheduled workouts from the Garmin Connect calendar

Deletes multiple calendar entries in a single call. The underlying workout templates are left intact in your library.

IMPORTANT: each id is a calendar-entry id (the "scheduled_workout_id" field from get_scheduled_workouts), not a workout id.

Args: scheduled_workout_ids: List of calendar-entry ids from get_scheduled_workouts

ParametersJSON Schema
NameRequiredDescriptionDefault
scheduled_workout_idsYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are present. The description discloses the primary side effect (deletes calendar entries) and a non-side effect (templates remain), which is helpful. However, it does not mention irreversibility, permissions, or other potential side effects beyond the delete operation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is brief and well-organized, with a clear action, a side-effect clarification, and an important ID warning. Every sentence adds value and there is no redundant filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the output schema exists and the sibling tools provide surrounding context, the description gives enough information to call the tool correctly. It could mention return behavior, but that is not essential because the output schema is provided.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema only provides an array of integers with no description, but the tool description supplies critical parameter semantics: the IDs are scheduled_workout_id values from get_scheduled_workouts, not workout IDs. This is highly useful and compensates for the schema's lack of detail.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb 'Remove' and the target 'scheduled workouts' from the Garmin Connect calendar, and notes the underlying templates remain intact. It is distinct from the singular 'unschedule_workout' sibling and other workout-related tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly says this is for multiple scheduled workouts in a single call and gives the crucial distinction that IDs are calendar-entry IDs, not workout IDs. It does not explicitly name the singular alternative, but the 'multiple' qualifier plus sibling names make the usage clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update_custom_foodA

Update an existing custom food in the user's Garmin nutrition library

Fetches the food's current record before writing so that omitted optional fields (brand, carbs, protein, fat, micros, etc.) preserve their existing values rather than being cleared. Only the fields you explicitly pass are changed; everything else is carried forward from the current record.

All nutrient amounts are ABSOLUTE values per serving, not %DV. Nutrition labels often print %DV for calcium/iron/vitamin D — convert to absolute units before passing.

Use get_custom_foods first to find the foodId and servingId.

Args: food_id: ID of the custom food to update (from get_custom_foods) serving_id: Serving ID of the food (from get_custom_foods) food_name: Name of the custom food calories: Calories per serving serving_unit: Unit for serving size (e.g. "G", "ML", "OZ"). Default "G" number_of_units: Serving size in the specified unit. Default 100 brand_name: Brand or vendor name; omit to preserve the existing value carbs: Carbohydrates in grams per serving protein: Protein in grams per serving fat: Total fat in grams per serving fiber: Fiber in grams per serving sugar: Sugar in grams per serving saturated_fat: Saturated fat in grams per serving sodium: Sodium in mg per serving cholesterol: Cholesterol in mg per serving potassium: Potassium in mg per serving trans_fat: Trans fat in grams per serving calcium: Calcium in mg per serving (NOT %DV) iron: Iron in mg per serving (NOT %DV) vitamin_d: Vitamin D in mcg per serving (NOT %DV)

ParametersJSON Schema
NameRequiredDescriptionDefault
fatNo
ironNo
carbsNo
fiberNo
sugarNo
sodiumNo
calciumNo
food_idYes
proteinNo
caloriesYes
food_nameYes
potassiumNo
trans_fatNo
vitamin_dNo
brand_nameNo
serving_idYes
cholesterolNo
serving_unitNoG
saturated_fatNo
number_of_unitsNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does well: it discloses the fetch-before-write behavior, that omitted optional fields preserve existing values rather than being cleared, and that nutrients are absolute values (not %DV). It omits error/auth behavior and permission requirements, keeping it short of a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Well front-loaded: purpose, then the critical preserve-omitted-fields semantics, then the %DV warning, then the arg list. Slightly redundant (the preserve-existing-values point is restated twice) but every block is relevant and earn its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 20-parameter mutation tool with no annotations, the description covers behavior and parameter semantics thoroughly, and an output schema exists so return values need not be explained. Missing only edge-case/error handling, which keeps it from a 5.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate — and it does, annotating all 20 parameters with meaning and units (grams, mg, mcg) and explicitly flagging that calcium/iron/vitamin_d are NOT %DV. This adds substantial value the bare schema lacks.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource: 'Update an existing custom food in the user's Garmin nutrition library.' This clearly distinguishes it from create_custom_food, delete_custom_food, and get_custom_foods without needing to open any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly instructs 'Use get_custom_foods first to find the foodId and servingId,' establishing a clear prerequisite workflow. It gives strong usage context but does not name alternate tools (create_custom_food) or state when-not-to-use.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

upload_courseA

Upload a GPX file as a Garmin Connect Course.

The course can then be loaded onto the watch (sync or "Send to Device") and used as a navigation course or to build a PacePro strategy.

Args: gpx_path: Absolute path to the .gpx file on disk. course_name: Override the course name. Defaults to the name parsed from the GPX file. activity_type: One of running, cycling, hiking, walking, trail_running, mountain_biking, road_biking, gravel_cycling. Defaults to running. description: Optional description shown on the course detail page.

ParametersJSON Schema
NameRequiredDescriptionDefault
gpx_pathYes
course_nameNo
descriptionNo
activity_typeNorunning

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the behavioral disclosure burden. It clearly states that this uploads a GPX file into a Garmin Connect Course and describes post-upload usage, but it does not disclose side effects such as duplicate handling, overwrite behavior, required authentication, or validation rules.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The purpose is front-loaded in the first sentence, followed by a compact and useful operational context paragraph, then a clearly structured Args block. Every line adds information; there is no filler or repetition of the schema titles.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given that an output schema exists and all input parameters are thoroughly documented, the description is complete enough for normal invocation. The only gaps are explicit failure-mode behavior and routing guidance relative to sibling tools, which are minor in the context of the other rich details provided.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema has 0% description coverage, but the Args section fully compensates: gpx_path is defined as an absolute path, course_name's override behavior and default are explained, the full allowed set of activity_type values is listed, and description's purpose is clarified. This goes well beyond what the schema provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Upload a GPX file as a Garmin Connect Course.' It also explains the purpose of the resulting course (navigation, PacePro), which clearly differentiates it from sibling tools like download_course_gpx and delete_course.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides practical context about loading the course onto a watch, which implies when this tool is useful, but it does not explicitly mention alternatives or when not to use it. An agent can infer the use case but is not directly routed away from related tools like upload_workout.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

upload_workoutA

Upload a workout from JSON data

Creates a new workout in Garmin Connect from structured workout data.

IMPORTANT: Step types must use Garmin's DTO format:

  • Use "ExecutableStepDTO" for regular steps (warmup, interval, cooldown, recovery)

  • Use "RepeatGroupDTO" for repeat/interval groups with numberOfIterations. Always include endCondition with conditionTypeId 7 and conditionTypeKey "iterations"; omitting conditionTypeId causes the API to silently corrupt the repeat count.

IMPORTANT: Heart rate targets come in two forms:

  • Named zone (e.g. Zone 2): set targetType to "heart.rate.zone" and use "zoneNumber" (1-5). Do NOT put the zone number in targetValueOne.

  • Custom HR range (e.g. 105-143 bpm): set targetType to "heart.rate.zone" and use "targetValueOne" (low bpm) / "targetValueTwo" (high bpm). Do NOT set "zoneNumber". This matches Garmin Connect's "Custom" heart rate target. For non-HR targets (pace, power, cadence), use targetValueOne/targetValueTwo directly. Target values are fields on the workout step, alongside targetType; do not put targetValueOne, targetValueTwo, or zoneNumber inside the targetType object. Use either zoneNumber or targetValueOne/targetValueTwo, not both. Garmin silently discards a custom range when a named zone is also present.

Note: a safety check converts targetValueOne 1-5 to zoneNumber when zoneNumber is missing, to catch the common mistake of putting a zone index in targetValueOne. Typical bpm values (e.g. 105, 143) are not affected.

IMPORTANT: Target type IDs and keys must match Garmin's canonical mapping. Garmin treats workoutTargetTypeId as authoritative, so mismatches are rejected before upload. Known mappings:

  • workoutTargetTypeId 1 -> "no.target"

  • workoutTargetTypeId 2 -> "power.zone" (cycling power zone 1-7, use zoneNumber)

  • workoutTargetTypeId 4 -> "heart.rate.zone"

  • workoutTargetTypeId 6 -> "pace.zone" (running/swim). NOT power: a cycling upload with "power.between" is stored as pace.zone.

IMPORTANT: For cycling power targets use the correct target type:

  • Power zone (zone 1-7 based on FTP %): use workoutTargetTypeId 2, key "power.zone", and "zoneNumber" (1-7).

  • Absolute watt range (e.g. 200-250 W): use workoutTargetTypeId 2, key "power.zone", and "targetValueOne" (low watts) / "targetValueTwo" (high watts). Using workoutTargetTypeId 2 with key "power.between" is a silent Garmin bug: the workout uploads but Garmin stores it as "power.zone" and the intent is lost.

Use {"workoutTargetTypeId": 4, "workoutTargetTypeKey": "heart.rate.zone"} with targetValueOne/targetValueTwo for custom heart-rate ranges.

IMPORTANT: Sport type IDs for workouts (different from activity API!):

  • 1 = running, 2 = cycling, 5 = strength_training, 6 = cardio, 11 = walking

IMPORTANT: End condition IDs and keys must match Garmin's canonical mapping. Garmin treats conditionTypeId as authoritative, so mismatches such as {"conditionTypeId": 4, "conditionTypeKey": "heart.rate"} are rejected before upload because Garmin would interpret them as "calories". Use {"conditionTypeId": 6, "conditionTypeKey": "heart.rate"} for heart-rate end conditions.

Available Templates: Instead of building workout JSON from scratch, you can use these MCP resources as starting points:

  • workout://templates/simple-run - Basic warmup/run/cooldown structure

  • workout://templates/interval-running - Interval training with repeat groups

  • workout://templates/tempo-run - Tempo run with heart rate zone targets

  • workout://templates/strength-circuit - Strength training with exercises, reps, rest

  • workout://reference/structure - Complete JSON structure reference with all fields

Access these resources using your MCP client's resource reading capability, modify the template as needed, and pass the resulting JSON as the workout_data parameter.

Strength training workouts require these additional fields on each exercise step:

  • "category": exercise category (e.g. "BENCH_PRESS", "PULL_UP", "CURL", "SHOULDER_PRESS", "ROW", "SQUAT", "DEADLIFT", "TRICEPS_EXTENSION", "PLANK", "LUNGE", "CARDIO")

  • "exerciseName": specific exercise (e.g. "BARBELL_BENCH_PRESS", "PULL_UP", "DUMBBELL_BICEPS_CURL", "DUMBBELL_SHOULDER_PRESS", "BENT_OVER_ROW_WITH_DUMBELL", "BODY_WEIGHT_DIP", "BARBELL_SQUAT", "BARBELL_DEADLIFT")

  • "weightValue" (optional): weight as number (e.g. 24.0)

  • "weightUnit" (optional): {"unitId": 8, "unitKey": "kilogram", "factor": 1000.0} Use endCondition reps (conditionTypeId: 10) for exercises, rest (stepTypeId: 5) between sets.

Example strength exercise step: { "type": "ExecutableStepDTO", "stepOrder": 1, "stepType": {"stepTypeId": 3, "stepTypeKey": "interval"}, "endCondition": {"conditionTypeId": 10, "conditionTypeKey": "reps"}, "endConditionValue": 10.0, "targetType": {"workoutTargetTypeId": 1, "workoutTargetTypeKey": "no.target"}, "category": "BENCH_PRESS", "exerciseName": "BARBELL_BENCH_PRESS", "weightValue": 60.0, "weightUnit": {"unitId": 8, "unitKey": "kilogram", "factor": 1000.0} }

Example running workout with HR zone target: { "workoutName": "My Workout", "sportType": {"sportTypeId": 1, "sportTypeKey": "running"}, "workoutSegments": [{ "segmentOrder": 1, "sportType": {"sportTypeId": 1, "sportTypeKey": "running"}, "workoutSteps": [{ "type": "ExecutableStepDTO", "stepOrder": 1, "stepType": {"stepTypeId": 3, "stepTypeKey": "interval"}, "endCondition": {"conditionTypeId": 2, "conditionTypeKey": "time"}, "endConditionValue": 1200.0, "targetType": {"workoutTargetTypeId": 4, "workoutTargetTypeKey": "heart.rate.zone"}, "zoneNumber": 3 }] }] }

Args: workout_data: Dictionary containing workout structure (name, sport type, segments, etc.)

ParametersJSON Schema
NameRequiredDescriptionDefault
workout_dataYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and delivers rich behavioral detail: silent corruption of repeat counts, silent discarding of custom HR ranges, silent API bugs for power targets, and a safety check that converts targetValueOne. It does not cover authentication requirements or what happens on success, but the API-specific failure modes are thoroughly disclosed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is very long but largely justified by the complex nested JSON payload. It is front-loaded with the basic purpose and then organized into focused IMPORTANT sections. Some repetition across target-type and heart-rate sections slightly reduces efficiency.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given a single opaque JSON parameter and no annotations, the description provides everything needed to construct a valid payload, including templates and canonical mappings. An output schema exists, so return-value details are correctly omitted.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% and the single object parameter has no documented properties, yet the description compensates with extensive detail on the expected JSON structure, step DTO types, target formats, sport type IDs, and strength exercise fields. This is exceptional semantic enrichment for an otherwise opaque parameter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'Upload a workout from JSON data' and 'Creates a new workout in Garmin Connect from structured workout data.' It is clear about what the tool does, but it does not distinguish itself from siblings like upload_workouts (plural) or the many create_* workout builders.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit guidance on when to use this tool versus alternatives such as create_run_workout or upload_workouts. The mention of MCP template resources explains how to prepare input data, not when this tool is the correct choice.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

upload_workoutsA

Upload multiple workouts from JSON data in a single call

Creates multiple new workouts in Garmin Connect. Each item in the list uses the same structure as upload_workout.

IMPORTANT: Step types must use Garmin's DTO format:

  • Use "ExecutableStepDTO" for regular steps (warmup, interval, cooldown, recovery)

  • Use "RepeatGroupDTO" for repeat/interval groups with numberOfIterations. Always include endCondition with conditionTypeId 7 and conditionTypeKey "iterations"; omitting conditionTypeId causes the API to silently corrupt the repeat count.

IMPORTANT: For named heart rate zone targets, use "zoneNumber" (1-5), NOT targetValueOne/targetValueTwo. For custom heart-rate ranges, use targetType {"workoutTargetTypeId": 4, "workoutTargetTypeKey": "heart.rate.zone"} with targetValueOne/targetValueTwo. Target values belong on the workout step, alongside targetType, not inside it. For cycling power zone targets (zone-based), use workoutTargetTypeId 2, key "power.zone". For cycling absolute watt range targets, use workoutTargetTypeId 2, key "power.zone", with targetValueOne (low watts) and targetValueTwo (high watts). Target type IDs and keys must match Garmin's canonical mapping.

IMPORTANT: End condition IDs and keys must match Garmin's canonical mapping. Garmin treats conditionTypeId as authoritative, so mismatches are rejected before upload.

Args: workouts: List of workout dictionaries, each containing workout structure (name, sport type, segments, etc.) — same format as upload_workout.

ParametersJSON Schema
NameRequiredDescriptionDefault
workoutsYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does so well: it warns that omitting conditionTypeId silently corrupts the repeat count, that Garmin rejects mismatched mappings before upload, and that conditionTypeId is authoritative. These are non-obvious behavioral traits; it stops short of covering permissions, idempotency, or error semantics.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose is front-loaded and the technical IMPORTANT blocks are dense with actionable detail rather than filler. It is long, and the repeated 'IMPORTANT:' headers plus some redundant zone/end-condition restatements add mild bloat, but the length is justified by Garmin DTO complexity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tricky schema-less payload into a strict external API, the description supplies the structural and error-handling context an agent needs, and the output schema covers return values. It is nearly complete, missing only permission/auth and failure-response specifics.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% and the schema is only an open array of objects, yet the description compensates heavily by enumerating the step DTO types, repeat-group structure, heart-rate and power zone target fields, and end-condition rules. It adds substantial meaning, though the 'same format as upload_workout' reference punts some detail to a sibling.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Upload multiple workouts', 'Creates multiple new workouts in Garmin Connect') and distinguishes itself from the singular sibling by emphasizing 'multiple ... in a single call'. An agent can tell it apart from upload_workout without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The 'single call' framing and 'Each item in the list uses the same structure as upload_workout' make the batch-vs-single choice clear by implication. However, there is no explicit when-to-use/when-not statement or routing to an alternative beyond the structural reference to upload_workout.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

upsert_and_logA

Find-or-create a custom food then log it in one step

Searches the user's custom food library for food_name. If found, logs it immediately. If not found, creates it with the provided nutrition data and then logs it. This avoids duplicate food entries and removes the need for separate search → create → log round-trips.

Args: meal_date: Date in YYYY-MM-DD format meal_time: Time in HH:MM:SS format (account timezone); used to determine the meal automatically food_name: Name of the food to find or create calories: Calories per serving carbs: Carbohydrates in grams per serving protein: Protein in grams per serving fat: Total fat in grams per serving serving_unit: Unit for serving size (e.g. "G", "ML", "OZ"). Default "G" number_of_units: Serving size in the specified unit. Default 100 serving_qty: Number of servings to log (default 1)

ParametersJSON Schema
NameRequiredDescriptionDefault
fatNo
carbsNo
proteinNo
caloriesYes
food_nameYes
meal_dateYes
meal_timeYes
serving_qtyNo
serving_unitNoG
number_of_unitsNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries full responsibility for conveying behavior. It explicitly describes the find-or-create-log workflow, including the conditional logic (search, then either log or create and log). It also clarifies that it uses the provided nutrition data only when creating a new food, implying that existing foods are logged as-is. No side effects are hidden or contradicted.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-structured and appropriately concise. It opens with a one-sentence summary, adds a brief elaboration of the workflow, and then presents a clear, labeled parameter list. There is minimal redundancy; the opening sentence and the following paragraph complement each other rather than repeating information unnecessarily.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers all necessary contextual information: it specifies the food source (custom food library), the condition for creating vs. logging, how meal_time is used, and the defaults for optional parameters. Since an output schema exists, the absence of return value details is acceptable. No critical usage details are missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The description includes an 'Args:' section that provides semantic meaning for all 10 parameters, including the purpose of meal_time ('used to determine the meal automatically'), the units for nutrient values, and defaults for serving_unit and number_of_units. This fully compensates for the lack of schema property descriptions (0% schema coverage).

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: it finds or creates a custom food and logs it in one step. It explicitly names the resource (custom food library and nutrition log) and the actions (search, create, log). It also distinguishes itself from sibling tools by mentioning it avoids duplicate food entries and removes the need for separate search, create, and log round-trips.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear usage guidance by stating the conditional behavior: 'If found, logs it immediately. If not found, creates it...' It also explains when this tool is appropriate by highlighting the benefit of avoiding duplicate entries and multiple round-trips, which indirectly contrasts with using separate search, create, and log tools. This gives an agent sufficient context to decide when to invoke this tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 172 tool updatesv1.0.0
    • First observedadd_body_composition
    • First observedadd_gear_to_activity
    • First observedadd_hydration_data
    • First observedadd_weigh_in
    • First observedadd_weigh_in_with_timestamps
    • First observedcount_activities
    • First observedcreate_brick_bike_run_workout
    • First observedcreate_brick_swim_bike_workout
    • First observedcreate_custom_food
    • First observedcreate_cycling_endurance_workout
    • First observedcreate_cycling_ftp_test_workout
    • First observedcreate_cycling_interval_workout
    • First observedcreate_cycling_over_under_workout
    • First observedcreate_cycling_sweet_spot_workout
    • First observedcreate_cycling_tempo_workout
    • First observedcreate_manual_activity
    • First observedcreate_run_easy_workout
    • First observedcreate_run_hills_workout
    • First observedcreate_run_intervals_workout
    • First observedcreate_run_long_workout
    • First observedcreate_run_progression_workout
    • First observedcreate_run_tempo_workout
    • First observedcreate_run_workout
    • First observedcreate_strength_workout
    • First observedcreate_swim_drills_workout
    • First observedcreate_swim_endurance_workout
    • First observedcreate_swim_intervals_workout
    • First observedcreate_swim_threshold_workout
    • First observedcreate_walk_run_workout
    • First observedcreate_weekly_plan
    • First observedcreate_z2_walk_workout
    • First observeddelete_activity
    • First observeddelete_blood_pressure
    • First observeddelete_course
    • First observeddelete_custom_food
    • First observeddelete_food_log
    • First observeddelete_weigh_ins
    • First observeddelete_workout
    • First observeddelete_workouts
    • First observeddownload_activity_file
    • First observeddownload_workout
    • First observedget_activities
    • First observedget_activities_by_date
    • First observedget_activities_fordate
    • First observedget_activity
    • First observedget_activity_details
    • First observedget_activity_exercise_sets
    • First observedget_activity_fit_data
    • First observedget_activity_gear
    • First observedget_activity_hr_in_timezones
    • First observedget_activity_power_in_timezones
    • First observedget_activity_series
    • First observedget_activity_split_summaries
    • First observedget_activity_splits
    • First observedget_activity_typed_splits
    • First observedget_activity_types
    • First observedget_activity_weather
    • First observedget_adhoc_challenges
    • First observedget_all_day_events
    • First observedget_all_day_stress
    • First observedget_athlete_context
    • First observedget_athlete_status_snapshot
    • First observedget_available_badge_challenges
    • First observedget_badge_challenges
    • First observedget_blood_pressure
    • First observedget_body_battery
    • First observedget_body_battery_events
    • First observedget_body_composition
    • First observedget_cardiac_drift_analysis
    • First observedget_courses
    • First observedget_custom_food_serving_units
    • First observedget_custom_foods
    • First observedget_cycling_ftp
    • First observedget_daily_steps
    • First observedget_daily_weigh_ins
    • First observedget_device_alarms
    • First observedget_device_last_used
    • First observedget_device_settings
    • First observedget_device_solar_data
    • First observedget_devices
    • First observedget_earned_badges
    • First observedget_endurance_score
    • First observedget_fitnessage_data
    • First observedget_floors
    • First observedget_full_name
    • First observedget_garmin_coach_workouts
    • First observedget_gear
    • First observedget_goals
    • First observedget_health_series
    • First observedget_heart_rates
    • First observedget_heart_rates_summary
    • First observedget_hill_score
    • First observedget_hrv_data
    • First observedget_hrv_trend
    • First observedget_hydration_data
    • First observedget_inprogress_virtual_challenges
    • First observedget_lactate_threshold
    • First observedget_lifestyle_logging_data
    • First observedget_menstrual_calendar_data
    • First observedget_menstrual_data_for_date
    • First observedget_morning_brief
    • First observedget_morning_training_readiness
    • First observedget_non_completed_badge_challenges
    • First observedget_nutrition_daily_food_log
    • First observedget_nutrition_daily_meals
    • First observedget_nutrition_daily_settings
    • First observedget_performance_trend
    • First observedget_personal_record
    • First observedget_power_duration_curve
    • First observedget_pregnancy_summary
    • First observedget_primary_training_device
    • First observedget_progress_summary_between_dates
    • First observedget_race_predictions
    • First observedget_respiration_data
    • First observedget_respiration_summary
    • First observedget_respiration_trend
    • First observedget_rhr_day
    • First observedget_scheduled_workouts
    • First observedget_sleep_data
    • First observedget_sleep_summary
    • First observedget_spo2_data
    • First observedget_stats
    • First observedget_stats_and_body
    • First observedget_steps_data
    • First observedget_stress_data
    • First observedget_stress_summary
    • First observedget_training_effect
    • First observedget_training_load_balance
    • First observedget_training_load_breakdown
    • First observedget_training_load_trend
    • First observedget_training_plan_workouts
    • First observedget_training_readiness
    • First observedget_training_readiness_composite
    • First observedget_training_status
    • First observedget_unit_system
    • First observedget_user_profile
    • First observedget_user_summary
    • First observedget_userprofile_settings
    • First observedget_vo2max_trend
    • First observedget_weekly_intensity_minutes
    • First observedget_weekly_load_progression
    • First observedget_weekly_steps
    • First observedget_weekly_stress
    • First observedget_weigh_ins
    • First observedget_workout_by_id
    • First observedget_workout_compliance
    • First observedget_workouts
    • First observedget_zone_distribution
    • First observedlog_custom_food
    • First observedlog_food
    • First observedremove_gear_from_activity
    • First observedrequest_reload
    • First observedschedule_week
    • First observedschedule_workout
    • First observedschedule_workouts
    • First observedsearch_foods
    • First observedset_activity_description
    • First observedset_activity_event_type
    • First observedset_activity_feel
    • First observedset_activity_name
    • First observedset_activity_type
    • First observedset_blood_pressure
    • First observedset_fit_download_dir
    • First observedset_nutrition_daily_settings
    • First observedset_perceived_effort
    • First observedunschedule_workout
    • First observedunschedule_workouts
    • First observedupdate_custom_food
    • First observedupload_course
    • First observedupload_workout
    • First observedupload_workouts
    • First observedupsert_and_log

TDQS

B3.1/5.0

Scored across 172 tools

Disambiguation3/5

With 172 tools there is heavy overlap: get_activities_fordate / get_activities_by_date / get_activities / count_activities, get_stats / get_user_summary / get_stats_and_body, and a large family of run-workout builders (create_run_workout vs create_run_easy_workout/tempo/long/intervals/hills/progression) that differ only by defaults. The very detailed descriptions and explicit 'use X instead' notes (e.g. summary vs full variants, the get_garmin_coach_workouts alias) mitigate much of the confusion, but several boundaries remain fuzzy.

Naming Consistency4/5

Almost everything follows a snake_case verb_noun pattern (get_*, set_*, create_*, delete_*, upload_*, schedule_*, log_*, search_*), which is highly predictable. Minor deviations exist: get_activities_fordate lacks the underscore used in get_activities_by_date, singular/plural varies (get_personal_record vs get_race_predictions), and upsert_and_log uses a non-standard verb.

Tool Count1/5

172 tools is far beyond any reasonable agent-facing surface and reflects extreme specialization, redundant aliases (schedule_workout/schedule_workouts/schedule_week, upload_workout/upload_workouts), and near-duplicate summary/full pairs. This is the extreme-mismatch end of the scale even for a broad platform like Garmin Connect.

Completeness4/5

Coverage is remarkably broad: activity read/update/delete/create, health and sleep metrics, nutrition CRUD plus logging, workout build/upload/schedule/unschedule/delete, courses, gear, challenges and rich training analytics. Only minor gaps remain (no workout edit, no true hydration delete, no update for scheduled entries), and several are documented platform limitations the agent can work around.

Maintenance

ActivitySlowing
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • F
    license
    B
    quality
    D
    maintenance
    Integrates with Garmin Connect to retrieve activity data, health metrics, and provide AI-powered training insights and personalized coaching recommendations based on your fitness activities and performance trends.
    14
    -
  • A
    license
    Not graded
    quality
    B
    maintenance
    Enables AI assistants to read Garmin activities and create/schedule structured workouts and multi-week training plans on Garmin Connect, syncing to the user's watch.
    1
    MIT
  • F
    license
    Not graded
    quality
    B
    maintenance
    Provides read access to Garmin Connect health and training data (sleep, heart rate, stress, HRV, activities) and supports writes like creating workouts, editing activities, and exporting files.
    -